The placebo test

Before a model reaches production, we run its backtest 20 times, once on the real signal and 19 on noise. If the real result does not stand clearly apart from the fakes, it is not a result. Every vendor benchmark you have ever seen should be held to the same standard.

3 min read

Medicine settled a hard epistemological question a century ago: how do you know a treatment works, rather than the patient merely improving? The answer was the control group. A drug is approved not because people got better, but because they got better than the people who received nothing dressed up as something.

Enterprise AI, an industry built on claims of predictive power, has mostly not adopted the equivalent standard. Backtests are presented without controls, measured on histories the vendor selected, tuned until they look strong. A good-looking backtest is easy to produce and easy to believe, and that combination is exactly what an evidential standard exists to defend against.

How the placebo test works

The mechanism is simple to state. When a candidate model finishes its backtest on real history, the same test is run again, many times, on noise: histories where the relationship between the inputs and the target has been deliberately destroyed, by shuffling the target series or randomising its alignment with the drivers. The structure, the volumes and the statistical texture of the data stay realistic. The signal, by construction, is gone.

Our standard run is 20 tests, one real and 19 placebo. Three outcomes are possible, and each is informative:

  • The real result stands clearly apart from all 19 fakes. The model found signal that does not exist in noise. It proceeds to the next gate.
  • The real result sits inside the spread of the fakes. Whatever the backtest chart looks like, the model has found luck. It does not ship, whatever the demo would have looked like.
  • Several placebo runs look strong. This is the quiet, valuable finding: the evaluation itself is too easy to fool, usually through leakage or an over-flexible fit, and the testing harness gets fixed before any model is judged with it.

The placebo arm does not just test the model. It tests the test.

Why this catches what ordinary validation misses

Standard practice, holdout sets and cross-validation, guards against a model memorising its training data. It does not guard against the subtler failure: a search over enough models, features and configurations will eventually produce something that performs well on any fixed holdout, by chance. The more diligent the experimentation, the more certain this becomes. Diligence without a control group manufactures false confidence at scale.

The placebo distribution gives every reported result a denominator. "The model scored 0.81" means nothing on its own. "The model scored 0.81, and 19 noise runs scored between 0.48 and 0.60" is a statement with evidential content: the gap between the real and the fake is the finding.

Questions to ask any vendor

The test generalises into procurement. Any claimed result, from any intelligence vendor, supports four questions:

  1. Was this measured on data the model never saw, selected by someone other than the modeller?
  2. What did the same evaluation produce on noise, and how far apart are the two?
  3. Are the misses in the record, or was the record assembled from the survivors?
  4. Does the model beat the naive baseline, last year plus ten percent, or only the strawman it was benchmarked against?

A vendor with real results can engage with all four, because the answers are the product. A vendor without them will redirect the conversation to the demo, which is its own answer.

Confidence is manufactured in slides. Evidence is manufactured in controls.

The placebo test is how Prophesee's Foresight engine separates signal from luck. Models are held against their fakes before they are trusted, and the record travels with the model. See the approach on your own data. Start here.

New essays land on LinkedIn first. Follow 3RDi to catch them, or get a demo to see Prophesee on your own data.