Of everything called 70 percent, about 70 percent should land

A probability is a promise about frequency. If a system's "70% likely" events do not happen about 70% of the time, its numbers are decoration. Calibration is the test, the reliability curve is the receipt, and it should be published every cycle.

3 min read

Every intelligence system now speaks in probabilities. Risk scores of 0.7, forecasts with 80% confidence, alerts ranked high, medium and low. The grammar of uncertainty has been universally adopted. The discipline behind it, mostly, has not.

A probability is a promise about frequency. When a system says 70%, it is promising that across all the occasions it says 70%, roughly seven in ten will happen. This promise has a name, calibration, and a simple test, and the test can be run by anyone who keeps records. Which is why it is remarkable how rarely it is run.

The reliability curve, and how to read one

The test works like this. Collect every probability the system issued over a period, together with what actually happened. Group the predictions into bands: everything called 10 to 20%, everything called 20 to 30%, and so on. For each band, compute the fraction that actually occurred. Plot predicted against observed.

A trustworthy system hugs the diagonal. Its 30% band lands about 30% of the time, its 80% band about 80%. Deviations are diagnoses, and each has a distinct operational meaning:

  • The curve sags below the diagonal: overconfidence. The system's "80%" events land at 60%. Every decision threshold built on its numbers is too aggressive, and the misses will cluster exactly where confidence was highest.
  • The curve bows above: underconfidence. The system hedges, calling 60% on things that happen 80% of the time. Its warnings are being discounted when they should be acted on, and value leaks through excessive caution.
  • The curve is flat: the probabilities carry almost no information. Everything lands at the base rate regardless of what was predicted. The number dressing on the alerts is decoration.

Why calibration outranks accuracy for decisions

Accuracy asks whether the system was right on average. Calibration asks whether its stated uncertainty can be used. For decision-making, the second property is the load-bearing one, because decisions are sized by the odds: how much to hedge, when to escalate, whether to intervene now or wait a week. Mis-stated odds mis-size every one of those choices even when the headline accuracy looks respectable.

Calibration also fails silently in precisely the situations that matter. A model can hold a flattering accuracy score while its high-confidence band, the band that triggers action, drifts badly. Only the reliability curve, recomputed on live outcomes, exposes that drift while there is still time to correct for it.

The system that says "I do not know" at the right moments is worth more than the system that is often right and always certain.

Published, every cycle, or it does not count

One reliability curve at purchase time proves the vendor once assembled a good chart. Calibration is a maintenance property: models drift, data shifts, and last quarter's honest 70% becomes this quarter's optimistic one. The standard that means something is publication every cycle, on live decisions, with the drift visible when it happens and the correction on the record.

That standard also changes vendor conversations. "How accurate is it?" invites a rehearsed answer. "Show me last quarter's reliability curve" invites either a document or a revealing silence.

Trust the diagonal, not the demo.

Reliability curves are a standing feature of Prophesee's Foresight engine, recomputed as outcomes land; of everything called 70%, about 70% should land, and the receipt is on the table. See your own curve. Start here.

New essays land on LinkedIn first. Follow 3RDi to catch them, or get a demo to see Prophesee on your own data.