Ask a vendor why their model is right for your problem and you will receive an architecture story: transformer this, foundation that, proprietary the other. Architecture stories are opinions. Which model best predicts a given target, on given data, at a given horizon, is an empirical fact, and an unstable one. It changes as the data changes, sometimes within a quarter.
The forecasting literature has made this point for decades, most famously through the M-competitions, where simple methods repeatedly embarrassed sophisticated ones on real series, and where no single method dominated across question types. A vendor whose product is one model, however capable, has answered an empirical question with a commitment made before your data was seen.
How the bake-off works
For every prediction question, Prophesee's Foresight engine runs a standing competition rather than a coronation:
- 62 models from 13 families compete: statistical methods, tree ensembles, neural approaches, and hybrids, spanning the five question types an operating business actually asks. How many? Which ones? How risky? When? Why?
- Scoring happens on unseen data, in rolling backtests, with the placebo harness behind it so a lucky fit cannot take a seat it did not earn.
- The table is re-scored daily as fresh outcomes arrive. A champion that drifts loses its seat to the contender that has not, without a meeting, without a migration project, without anyone having to defend last year's architecture choice.
- Every question gets its own champion. The model that wins weekly demand for fast movers is routinely the wrong model for intermittent spares, for regulatory case timing, or for attributing a variance. One question, one table, one current champion.
Nobody has to believe in a model family. The table settles it, and keeps settling it.
The contender that keeps everyone honest
Every table contains one entrant that does not care about elegance: the naive baseline. Last year plus ten percent. Same as last period. Whatever a sensible person would guess without a model.
The rule is absolute. A champion must beat naive on the question's agreed error metric, or there is no champion. This rule does real work. Research on forecast value repeatedly finds that a large share of sophisticated forecasts fail to beat naive methods, a result most organisations have never tested against their own numbers. When nothing beats naive, the honest output is not the least-bad model. It is the sentence: this target is not predictable yet with this data, and here is what would change that.
That sentence prevents the quiet catastrophe of enterprise forecasting, which is not bad models but confident automation of guesswork: infrastructure, review meetings and decisions built on numbers that a copied-forward spreadsheet cell would have matched.
What this replaces
The league table replaces two familiar failure modes. The first is the data science queue, where each new question waits months for a hand-built model, which then ossifies because nobody has time to revisit it. The second is the platform monoculture, where every question is answered by the vendor's one architecture, at whatever quality that architecture happens to achieve on it.
Against both, the competitive mechanism is boring, continuous and auditable: every model's history is on the table, every substitution has a scoring reason, and the answer to "why this model?" is never a belief. It is a row.
The league table runs as standard inside Prophesee's Foresight engine. 62 contenders, re-scored daily, naive baseline enforced. To see who wins on your data, start here.