# ORB Learn · Thinking in Probabilities

Edition 04 · Nine chapters · Unrecorded. Old recordings removed at the owner’s request.

## Begin with a claim you can lose

Come closer to the crystal ball and choose a gold light. Before asking whether its story feels plausible, ask what would make it wrong. A folding iPhone is a useful starting point because you can separate the invitation to an event from the product you hope to see. The invitation establishes a stage. It does not establish what will appear on it. Our test asks for an official introduction of a display that physically folds. It does not count a patent or a clever accessory.

Now turn toward the AI helper. Its test concerns the published purchase-approval requirement, not a claim that every action is safe. You have moved from a product milestone to a permission boundary, while keeping the same discipline: name the observable result before it happens. The source tells you what the company says. Independent testing would be needed to show how reliably the system behaves.

## Read a calendar with room for delay

Imagine an orbital sunset and a small cargo vehicle approaching the station. The image makes the destination feel close; the schedule may tell another story. NASA’s current Progress 96 page marks its timing under review. An earlier plan named dates, but those dates must now be qualified. This is how a careful forecast changes its confidence without inventing a technical explanation.

Move across to Crew-13. A no-earlier-than window is a boundary, not a reservation. Our guess is that the crew mission has not lifted off by September 22. Notice how the two tests differ. The cargo branch asks for completed docking; the crew branch asks whether liftoff has occurred. You cannot score either one from an impressive rehearsal or a new target date. Each needs the actual milestone, with its own timestamp.

## Let the envelope surprise you

The light changes from orbital blue to the warmth of an awards stage. The Pitt and Hacks have broad nomination recognition, which makes them reasonable favorites for this exercise. But recognition across many categories is not a hidden count of final votes. Sixty percent leaves plenty of room for another nominee.

Before the ceremony, give the strongest rival a fair hearing. What could concentrate support around a different show? You do not need to know the answer to admit that the missing ballot information matters. After the ceremony, use the exact series category. An acting award cannot quietly become a win for our drama or comedy forecast. Both predictions draw on a similar signal, so two correct picks would still be a small, partly related result.

## Match the evidence to the sporting question

Look now at the sports lights. Historical strength can support an expectation about a medal table, but the first few days may favor a different team because of the event schedule. A full-tournament record and an early-table snapshot are not interchangeable. Our Asian Games call stops at September 22 and uses the stated medal tie-break rule.

The US Open branch asks a smaller, more playful question: will the men’s final finish in at least four completed sets? The format allows it, but that fact is not a probability model. Last year’s four-set final is only one example. A fair record therefore preserves the modest estimate and its limitations. Retirement is left unresolved under this rule, because a partial match does not answer the intended question about a completed final.

## Keep the measurement and the warning separate

Return to Earth and two numbers with practical consequences. A monthly core-price reading describes a national statistical basket. It does not tell you exactly how your own bills changed. The original prediction specifies the initial rounded release, so a later revision cannot move the scoring target.

The Atlantic branch has an equally important boundary. It asks whether a new cyclone designation appears within a short interval. A quiet formation outlook does not promise safe weather at a particular place. The retrieved outlook has its own issue time, and this saved orb does not fetch new warnings. For real decisions, turn to current official guidance. For this learning experiment, keep the original estimate and record the operational advisory evidence that resolves it.

## Listen for the signal, then demand the reveal

Turn toward the lights where imagined worlds begin. StarCraft carries the pull of a familiar name, but an unusual website is still an ambiguous signal. Our new-game test refuses to turn a crossover, remaster or unexplained teaser into a win. You can enjoy the suspense while keeping the category firm.

Project ZIRCON begins from a different kind of evidence: an announced demonstration. Here the uncertainty is whether the presentation delivers actual play. Pause with that distinction. One forecast anticipates an unconfirmed product; the other anticipates fulfillment of a published plan. Their probabilities should not be mistaken for equally daring glimpses of the future.

## Watch an idea acquire hands

Let the imagined world dissolve into a busy exhibition hall. A robot arm waits beside a person; a small aircraft waits beside a production cell. The intriguing moment is the connection between steps. Did the lesson shape the robot’s action? Did the manufacturing cell produce the drone that flew? Our tests require those connections to be documented.

Now look for the people around the demonstration. Someone may arrange the objects, correct a failed attempt or finish an integration step. Their contribution belongs in the account. A successful demonstration can still be valuable with human assistance, provided you describe it accurately. What remains unknown is performance across many attempts, unfamiliar situations and a full working day.

## Follow the handoff, not just the conversation

Return to the screen with the same attention to hidden work. Two agents in a diagram do not yet show cooperation. Trace a task from one identifiable role to another, then look for its result. Our Dreamforce test asks for a visible handoff, while allowing a prerecorded product demonstration if the resolution labels it honestly.

Connect this to the older permission branch. Giving software more ways to cooperate also raises questions about who can authorize an action and who can stop a mistake. A polished example does not settle those questions. The useful next experiment would introduce an error and observe recovery. For now, keep the demonstration claim small enough to check, and its implications open enough to investigate.

## Return with a complete record

Step back until the whole sphere fits in your view. Themes sit above, guesses cross the middle, and evidence with open questions rests below. The lower points are useful precisely because they show where the research stops. You can continue a branch without pretending the missing knowledge was already present.

When the window closes, keep every result: yes, no, provisional or unresolved, each with a source. One simple score is the squared distance between your probability and a binary outcome. Lower is better, but fifteen selected cases cannot establish calibration, AGI or access to the future. Separate the ten retained calls from the five newest additions, while preserving each earlier batch and cutoff. The change of scope is part of the record, not something to hide.

The crystal ball remains a playful invitation. Its useful gift is a habit: define the claim, look for the strongest contrary evidence, mark the uncertainty and return honestly to what happened. Leave the orb with a sharper question, and leave room for the world to surprise you.