[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"article-body-of-everything-called-70-percent":3},"\nEvery intelligence system now speaks in probabilities. Risk scores of\n0.7, forecasts with 80% confidence, alerts ranked high, medium and low.\nThe grammar of uncertainty has been universally adopted. The discipline\nbehind it, mostly, has not.\n\nA probability is a promise about frequency. When a system says 70%, it\nis promising that across all the occasions it says 70%, roughly seven in\nten will happen. This promise has a name, calibration, and a simple\ntest, and the test can be run by anyone who keeps records. Which is why\nit is remarkable how rarely it is run.\n\n## The reliability curve, and how to read one\n\nThe test works like this. Collect every probability the system issued\nover a period, together with what actually happened. Group the\npredictions into bands: everything called 10 to 20%, everything called\n20 to 30%, and so on. For each band, compute the fraction that actually\noccurred. Plot predicted against observed.\n\nA trustworthy system hugs the diagonal. Its 30% band lands about 30% of\nthe time, its 80% band about 80%. Deviations are diagnoses, and each\nhas a distinct operational meaning:\n\n- **The curve sags below the diagonal**: overconfidence. The system's\n  \"80%\" events land at 60%. Every decision threshold built on its\n  numbers is too aggressive, and the misses will cluster exactly where\n  confidence was highest.\n- **The curve bows above**: underconfidence. The system hedges, calling\n  60% on things that happen 80% of the time. Its warnings are being\n  discounted when they should be acted on, and value leaks through\n  excessive caution.\n- **The curve is flat**: the probabilities carry almost no information.\n  Everything lands at the base rate regardless of what was predicted.\n  The number dressing on the alerts is decoration.\n\n## Why calibration outranks accuracy for decisions\n\nAccuracy asks whether the system was right on average. Calibration asks\nwhether its stated uncertainty can be used. For decision-making, the\nsecond property is the load-bearing one, because decisions are sized by\nthe odds: how much to hedge, when to escalate, whether to intervene now\nor wait a week. Mis-stated odds mis-size every one of those choices\neven when the headline accuracy looks respectable.\n\nCalibration also fails silently in precisely the situations that\nmatter. A model can hold a flattering accuracy score while its\nhigh-confidence band, the band that triggers action, drifts badly. Only\nthe reliability curve, recomputed on live outcomes, exposes that drift\nwhile there is still time to correct for it.\n\n> The system that says \"I do not know\" at the right moments is worth\n> more than the system that is often right and always certain.\n\n## Published, every cycle, or it does not count\n\nOne reliability curve at purchase time proves the vendor once assembled\na good chart. Calibration is a maintenance property: models drift, data\nshifts, and last quarter's honest 70% becomes this quarter's optimistic\none. The standard that means something is publication every cycle, on\nlive decisions, with the drift visible when it happens and the\ncorrection on the record.\n\nThat standard also changes vendor conversations. \"How accurate is it?\"\ninvites a rehearsed answer. \"Show me last quarter's reliability curve\"\ninvites either a document or a revealing silence.\n\n*Trust the diagonal, not the demo.*\n\nReliability curves are a standing feature of Prophesee's Foresight\nengine, recomputed as outcomes land; of everything called 70%, about\n70% should land, and the receipt is on the table. See your own curve.\n[Start here](/contact).\n",1786799034545]