Accuracy that does not change a decision is a vanity metric
Module · Demand Planning and Forecasting
Most accuracy work improves a number nobody acts on.
Prophesee Demand forecasts at the level the decision is taken, borrows a curve where there is no history, and labels the events your history never recorded.
Half the forecast is a person. Nobody scores that half.
The statistical forecast is adjusted by hand almost everywhere. Whether the adjustment helped is measured almost nowhere.
Sources: Fildes, Goodwin and De Baets, Forecast value added in demand planning, International Journal of Forecasting, 2024 (approximately 147,000 forecasts across six studies) · Fildes, Goodwin, Lawrence and Nikolopoulos, International Journal of Forecasting, 2009 (60,000+ forecasts, four companies; fieldwork more than fifteen years old) · Morlidge, Foresight, 2014 (eight companies, 300,000+ forecasts), reported via the SAS forecast value added white paper · Single specialty retailer override result reported in the same white paper, company not named.
From negotiating the number to explaining it
One number per item per period, with no band and no stated driver, so a planner cannot tell a confident forecast from a guess.
Quantile and hierarchical models per item, location and horizon, reconciled by arithmetic so the item, the family and the total agree. Every line carries a confidence band and a ranked list of the drivers behind it.
Sales adds a number, the planner adds another, and no record survives of who moved what or whether the move helped.
Every adjustment captured with its owner and scored against the outcome. Forecast value added by person, product and period, so the adjustments worth keeping are the ones that get kept.
The uplift is argued from the last promotion that felt similar. Cannibalisation and the post promotion dip are not modelled at all.
Model the uplift, the cannibalisation and the dip on the projected curve, then track actuals against the modelled plan so the promotion is separated from the trend.
The same product carries three codes across three systems, so history breaks at every rename and every launch looks like a brand new item.
Products, customers and locations resolved into one record, with launches, transitions and supersessions carried through history so the model sees a continuous series.
Turning demand challenges into decisions
One number per item per period, with no band and no driver, produced at a level nobody actually decides at.
A new item has no history, so the curve is borrowed from whichever past launch somebody in the room remembers.
Accuracy is reported after the period it describes, at a level too high to change any decision.
The spike in week eleven is recorded as noise, because nothing wrote down that a competitor was out of stock.
The uplift is claimed, the baseline is disputed, and nothing separates the promotion from the underlying trend.
14 AI applications that could be relevant
A sample of what becomes possible on the decision layer, not a fixed list: each application draws on the same data foundation and audit trail, and new ones are configured on the engines, not built from scratch.
A distribution per item and location, with a confidence band and the drivers behind every line.
Intermittent and lumpy items handled with quantile and Croston class methods rather than smoothing.
A launch curve built from weighted comparables, with the comparable set and its confidence attached.
What a detected event does to the plan, estimated from precedent as a range with a confidence bound.
Error weighted by the stock and the service it actually moved, not by unit volume.
Candidate events detected from feeds, documents and your own transactions, classified and ranked.
Every manual adjustment scored against what it did to accuracy, by owner and by product.
A change point in the series raised the day it happens rather than at the next cycle.
Model the uplift, the cannibalisation and the post promotion dip before you commit.
Ramp, peak, plateau, decline and floor adjusted inside guardrails, with every override recorded.
Which event classes have earned automatic application, and what each one needs to graduate.
Why this item moved, in plain language, with the evidence and the version behind it.
Your own history labelled with the events that were live, and what was knowable at the time.
Every event, what it was estimated to cost, and what it actually cost.
A day when the planner gets to decide
Today: A forecast built from six functions, each carrying an adjustment nobody can see.
Every item carries a distribution and its drivers. Fourteen have bands wide enough to change the stock decision.
Today: An item that broke trend in week two, noticed at the month end review.
Last cycle's adjustments scored against outcome. Two owners improved the forecast, one did not.
Today: A promotion argued from the last one that felt similar.
The promotion tested for uplift, cannibalisation and dip. She funds the one that moves the curve.
Today: Accuracy reported after the period it was supposed to change.
Error weighted by the cover it changed, not by units. That is the number that goes to the review.
Tomorrow's demand becomes today's decision.
Raise accuracy that changes stock
We agree the metric and the baseline in week one, and measure the result on your data.