polyAether
Textbook · Chapter 3
Chapter 3

Why weather is predictable — but never certain

~11 min read

A forecast is not a guess and it is not a promise. It is something in between: a careful calculation that says here is what will probably happen, and here is how sure I am. Understanding that middle ground is the whole point of this chapter.

In Chapter 2 we met prediction markets — places where a price acts like a probability. In this book we care about markets that ask weather questions: Will tomorrow's high temperature in New York be above 90°F? To have any hope of trading such a market well, you first need to understand where a weather forecast comes from, and — just as importantly — why it can never be perfectly certain. That uncertainty is not a flaw. It turns out to be exactly where the opportunity lives.

01 — What a forecast actually is

The atmosphere is physics, not magic

Weather feels chaotic and moody, but the air above your head obeys a small set of physical laws. Warm air rises. Pressure differences push air from high to low, and we feel that push as wind. Water evaporates, cools, and condenses into clouds and rain. None of this is mysterious — it is the same physics that makes a kettle boil, just spread over a whole planet.

Because these laws are known, they can be written as equations — mathematical rules that say given the air's temperature, pressure, humidity, and motion right now, here is what they will be a few seconds from now. A numerical weather model (a computer program that simulates the atmosphere using those physics equations) does exactly this, over and over.

Here is the mental picture. Imagine dividing the whole atmosphere into a giant 3-D grid of boxes, like a colossal stack of sugar cubes wrapped around the Earth. For each box the computer stores a few numbers: how warm the air is, how much moisture it holds, which way and how fast it is moving. Then it applies the physics equations to nudge every box forward one small time-step, using its neighbors' values. Repeat that step millions of times and you have marched the simulated weather hours or days into the future.

To make that concrete, picture one box over Central Park at 8 a.m. It knows the box to its west is warmer and the box above it is drier. The equations say: warm air is flowing in from the west, so this box will heat up over the next few minutes; the dry air aloft is sinking, so humidity here will fall. The model applies those nudges, advances the clock a few seconds, and then asks the same question again — now with slightly updated neighbors. Nobody hand-writes tomorrow's weather; it emerges, one tiny bookkeeping step at a time, from millions of boxes each minding their own local physics. That is why a good model can produce a realistic thunderstorm it was never explicitly told to make: the storm is just what the equations do when warm, moist air gets pushed hard enough.

Key idea

A forecast is a physics simulation. Start from the atmosphere's current state, apply known physical laws step by step, and watch where the numbers go. It is closer to predicting the path of a thrown ball than to reading tea leaves.

02 — Why it can never be certain

Two unavoidable sources of doubt

If the physics is known, why isn't the forecast perfect? Two reasons, and both are baked in permanently — no amount of better computers fully removes them.

1. We never know the starting point exactly

The simulation needs to begin from the atmosphere's current state. But we measure that state from a scattered patchwork — weather stations, balloons, aircraft, satellites, buoys. There is no thermometer in every one of those millions of grid boxes. So the starting numbers are always a little bit wrong: a fraction of a degree off here, a slightly mis-measured wind there.

Think about the scale of the gap. A modern global model may carry tens of millions of grid boxes, but the planet only has a few thousand upper-air balloon launches a day and a scatter of surface stations. Everything in between has to be inferred — filled in from satellites that measure the atmosphere indirectly, and from the model's own prior forecast. The result is a genuinely good estimate of right now, but an estimate all the same. It is like trying to reconstruct the exact surface of a lake from a few dozen buoys: you will get the big waves right and miss the ripples, and some of those ripples matter.

2. Small errors grow

You might hope a tiny starting error stays tiny. In weather, it does the opposite. The atmosphere is a chaotic system — meaning small differences in the starting conditions can grow into large differences later. This is the famous butterfly effect: not that a butterfly literally causes a storm, but that an error too small to measure can, days later, be the difference between rain and sun.

A good analogy is a marble balanced at the top of a smooth hill. Give it the faintest nudge left and it rolls down the left side; nudge it right and it goes right. The nudges are almost identical; the outcomes are opposite. The atmosphere is full of moments like that marble.

There is an important pattern in how fast the error grows, and it is very good news for the kind of trading we do. The growth is roughly exponential: a small error stays small for a while, then accelerates. Practically, that means a forecast for tomorrow's high temperature is far tighter than a forecast for the same city ten days out. Skill decays with lead time. Since the weather markets we care about resolve on a one- or two-day horizon — not two weeks — we are living in the part of the curve where the physics is still firmly in control and the spread is narrow. The butterfly has not had time to flap yet.

Key idea

Forecast uncertainty is not laziness or bad equipment. It comes from two permanent facts: we can only measure today's weather approximately, and tiny approximations grow over time. A forecast that claims total certainty is lying.

03 — The clever fix: run it many times

What an ensemble is

Here is the beautiful trick that makes uncertainty measurable instead of just admitted.

If we cannot know the exact starting point, we do not run the simulation once and pretend we do. Instead we run it many times. Each run starts from a slightly different, equally plausible version of right now — we deliberately jiggle the starting numbers within the range of our measurement uncertainty, and we also nudge the model's own internal settings a little. This whole collection of runs is called an ensemble (a group of forecasts run together, each from a slightly tweaked starting point).

1
Take today's best guess of the atmosphere
Assembled from stations, satellites, balloons, and more.
2
Make dozens of tweaked copies
Each nudged within the range of what we honestly can't rule out.
3
Simulate all of them forward
Every copy marches into the future under the same physics.
4
Look at the spread of answers
Agreement means confidence; disagreement means doubt.

Now watch what the ensemble tells you — for free. Suppose you are forecasting tomorrow's high in a city and you run 100 tweaked simulations. If 90 of them land above 90°F and 10 land below, you have a natural, honest number: roughly a 90% chance of clearing 90°F. The fraction of runs that produce an outcome is the probability of that outcome. Counting how many copies of the future agree turns physics into odds.

Let's do a fuller worked example, because this counting step is the heart of everything that follows. Imagine a market asking Will tomorrow's high in New York land in the 88–89°F bucket? We run an ensemble and read off each member's predicted high temperature. Sorting them into 1-degree buckets, the tally comes out like this:

6
members land at 86–87°F
The cool tail — a few runs kept more cloud cover.
31
members land at 88–89°F
The middle of the pack — the single most likely bucket.
44
members land at 90–91°F
The fattest cluster — clear skies, full afternoon heating.
19
members land at 92°F or higher
The warm tail — a few runs dried the air out completely.

Out of 100 members, 31 fell in the 88–89°F bucket. That hands us a probability directly: about a 31% chance the answer lands in that exact band. Notice that the same ensemble simultaneously answers every related question. The chance of 90°F or hotter? Add the top two buckets: 44 + 19 = 63, so about 63%. The chance of below 88°F? Just the 6 cool runs, so about 6%. One ensemble, one pass of counting, and every temperature threshold the market might ask about gets a number at once. That is the machinery that lets us look at a whole ladder of market prices and check each rung against physics.

The spread — how much the runs disagree — is itself the message:

confident — runs agree uncertain — runs disagree
tight spread = high confidencewide spread = low confidence
Same idea, two moods. The horizontal axis is the forecast outcome (say, tomorrow's high temperature); the height shows how many of the tweaked runs landed there. A tall, narrow peak is a sure thing; a low, wide hump is a genuine coin-toss.

The same picture explains why the timing of a bet matters so much. Run the ensemble three days out and the runs scatter widely — a broad, low hump, because the marbles have had time to roll different ways. Run it the morning before resolution and the same market usually collapses to a tall, narrow spike: almost every member now agrees, because the atmosphere has settled and there is little future left for errors to grow into. A market that was a genuine coin-toss on Monday can be a near-certainty by Wednesday. As we'll see, that collapse is not just a curiosity — it shapes when a weather market is actually worth trading and when it isn't.

Key idea

An ensemble runs the same forecast many times with tiny tweaks to the starting point. The fraction of runs that agree on an outcome is its probability, and how tightly the runs cluster tells you how confident to be. This is how a physics simulation produces an honest number like 70% chance of rain.

04 — What polyAether actually leans on

More opinions, better odds

No single model is perfect, so the strongest approach is to blend several. Different national weather centers have built different models, each with its own strengths — a different grid resolution, different assumptions about clouds and turbulence, different histories of what they get right. Where several independent systems agree, you can trust the answer more; where they diverge, that disagreement is itself a real signal that the atmosphere is in a hard-to-call state.

polyAether pulls together an ensemble of about 122 members drawn from three of the world's leading systems — the American GFS, the German ICON, and the European ECMWF (each is just the name of a major weather model). Combining many members from several independent systems gives a broader, more trustworthy spread than any one model alone: it captures not only the we can't measure the start exactly uncertainty within a single model, but also the deeper reasonable experts built the physics slightly differently uncertainty across models. That second kind is invisible to any single system, which is exactly why a lone model tends to sound more confident than it has earned.

Those forecasts are checked against reality at roughly 80 curated weather stations — specific, reliable measurement sites — because in the end a bet on the high in New York is settled by what one real thermometer records. This is a critical and easily-missed point: the ensemble predicts a temperature, but the market pays out on a specific instrument's reading, rounded and reported by a specific rule. Translating the model's 90.4°F into what the official station will report is its own careful step. (How a market decides who won turns out to be surprisingly fiddly; that's the subject of Chapter 7.)

Why this matters for trading

An ensemble hands us a probability the market has to compete with. If our 122 runs say 72% chance, and the market price implies 60%, we may have found a disagreement worth acting on. But — as the next chapters show — a raw model number is not yet a trustworthy one. It has to be calibrated first: history has to confirm that when this ensemble says 72%, the thing really does happen about 72% of the time. Until then, 72% is a hypothesis, not an edge.

An honest note on where the edge actually is

It is tempting to imagine the ensemble as a crystal ball we simply point at any market to print money. It is not. In practice most apparent edges are mirages — they sit against a penny or two of leftover orders on unlikely tail outcomes that no serious trader would fill, and our own guardrails correctly refuse to touch them. Where prices are genuine and liquid, the market is usually already efficient and there is no disagreement to exploit. The real, tradeable window is narrow: often only about a day before a market resolves, when the ensemble has sharpened but the crowd hasn't fully caught up. The disciplined result is frequently no trade at all — and that is a feature, not a failure. The edge, when it exists, comes from having a better-calibrated probability than the crowd, not from forecasting genius and not from raw speed.

05 — Where this is heading

From weather to edge

You now have the core machinery. A forecast is a physics simulation; it is uncertain because we can't measure the start perfectly and errors grow; and an ensemble converts that uncertainty into an actual probability by running the simulation many times and counting the outcomes. Blending several independent systems widens that count into something more trustworthy than any single model can claim.

The next question is sharper: when the model says 70%, does it really happen 70% of the time? A number that looks like a probability is not yet one you can bet real money on. Turning a model's raw fraction into a trustworthy probability is the work of Chapter 4, on probability, calibration, and edge — where we'll see why markets systematically overpay for uncertainty (by a factor of roughly 1.3), and how a well-behaved, honestly-calibrated forecast can quietly exploit that gap.