[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"article-body-the-placebo-test":3},"\nMedicine settled a hard epistemological question a century ago: how do\nyou know a treatment works, rather than the patient merely improving?\nThe answer was the control group. A drug is approved not because people\ngot better, but because they got better than the people who received\nnothing dressed up as something.\n\nEnterprise AI, an industry built on claims of predictive power, has\nmostly not adopted the equivalent standard. Backtests are presented\nwithout controls, measured on histories the vendor selected, tuned until\nthey look strong. A good-looking backtest is easy to produce and easy to\nbelieve, and that combination is exactly what an evidential standard\nexists to defend against.\n\n## How the placebo test works\n\nThe mechanism is simple to state. When a candidate model finishes its\nbacktest on real history, the same test is run again, many times, on\nnoise: histories where the relationship between the inputs and the\ntarget has been deliberately destroyed, by shuffling the target series\nor randomising its alignment with the drivers. The structure, the\nvolumes and the statistical texture of the data stay realistic. The\nsignal, by construction, is gone.\n\nOur standard run is 20 tests, one real and 19 placebo. Three\noutcomes are possible, and each is informative:\n\n- **The real result stands clearly apart from all 19 fakes.** The\n  model found signal that does not exist in noise. It proceeds to the\n  next gate.\n- **The real result sits inside the spread of the fakes.** Whatever the\n  backtest chart looks like, the model has found luck. It does not\n  ship, whatever the demo would have looked like.\n- **Several placebo runs look strong.** This is the quiet, valuable\n  finding: the evaluation itself is too easy to fool, usually through\n  leakage or an over-flexible fit, and the testing harness gets fixed\n  before any model is judged with it.\n\n> The placebo arm does not just test the model. It tests the test.\n\n## Why this catches what ordinary validation misses\n\nStandard practice, holdout sets and cross-validation, guards against a\nmodel memorising its training data. It does not guard against the\nsubtler failure: a search over enough models, features and\nconfigurations will eventually produce something that performs well on\nany fixed holdout, by chance. The more diligent the experimentation,\nthe more certain this becomes. Diligence without a control group\nmanufactures false confidence at scale.\n\nThe placebo distribution gives every reported result a denominator.\n\"The model scored 0.81\" means nothing on its own. \"The model scored\n0.81, and 19 noise runs scored between 0.48 and 0.60\" is a\nstatement with evidential content: the gap between the real and the\nfake is the finding.\n\n## Questions to ask any vendor\n\nThe test generalises into procurement. Any claimed result, from any\nintelligence vendor, supports four questions:\n\n1. Was this measured on data the model never saw, selected by someone\n   other than the modeller?\n2. What did the same evaluation produce on noise, and how far apart are\n   the two?\n3. Are the misses in the record, or was the record assembled from the\n   survivors?\n4. Does the model beat the naive baseline, last year plus ten percent,\n   or only the strawman it was benchmarked against?\n\nA vendor with real results can engage with all four, because the answers\nare the product. A vendor without them will redirect the conversation to\nthe demo, which is its own answer.\n\n*Confidence is manufactured in slides. Evidence is manufactured in\ncontrols.*\n\nThe placebo test is how Prophesee's Foresight engine separates signal\nfrom luck. Models are held against their fakes before they are trusted,\nand the record travels with the model. See the approach on your own\ndata. [Start here](/contact).\n",1786799035592]