Half the forecast is a person. Nobody scores that half.

Almost every enterprise forecast is adjusted by hand before it is used. The evidence says those adjustments help only about half the time. Yet organisations that measure model error to two decimal places do not measure the human half at all.

2 min read

Every enterprise forecast that matters has two authors. The first is a model, whose error is measured to two decimal places, tracked in reviews, and litigated in software selections. The second is a person, who adjusts the model's number before anyone uses it, and whose contribution is measured, in most organisations, never.

The evidence on that unmeasured half is not flattering. Fildes, Goodwin and De Baets, analysing roughly 147,000 forecasts across six studies, found judgemental adjustments improved accuracy for only just over half the items examined, with upward adjustments notably more likely to make the forecast worse (International Journal of Forecasting, 2024). A coin toss, purchased at the price of your most experienced people's time.

We measure the model obsessively and the person not at all, and the person is half the forecast.

Why the overrides exist, and why they should

The wrong conclusion is that adjustments should be banned. The planner adjusts because she knows things the model cannot: the promotion was cancelled yesterday, the customer is switching, the port strike will bite in week three. None of it is in the master data, and all of it belongs in the forecast. Some overrides are the most valuable information in the whole planning process.

The problem is that the valuable overrides and the destructive ones are currently indistinguishable, because nothing records which was which. The optimism nudge that pads the number to match the target travels through the same unmeasured channel as the genuine intelligence about the cancelled promotion. Both are labelled "experience".

Scoring the human half

The remedy is neither trust nor prohibition. It is measurement, the same courtesy extended to the models:

  • Record every adjustment as a decision: who, when, direction, size, and the stated reason, captured at the moment of the override rather than reconstructed later.
  • Score it against what the model alone would have done. Forecast value added is an old, simple idea: did the human step improve on the input it received? Run item by item, it separates the planner whose market knowledge consistently beats the model from the adjustment ritual that consistently subtracts value.
  • Feed the verdicts back. The planner with a strong record gets her overrides weighted up and her reasons mined for signals the model is missing. The adjustment class that reliably fails, typically the small, frequent, upward nudge, gets retired without ceremony.

Nobody's judgement is confiscated. Its results simply become visible, which is exactly what happened to the models.

One cultural objection deserves an answer: scoring people feels punitive. In practice the effect runs the other way. Unmeasured judgement is cheap to dismiss, which is why planners spend review meetings defending their numbers. Measured judgement with a track record is authority. The strongest argument a planner can bring to a forecast review is not seniority but a scored history of beating the model.

Judgement is not the enemy of the forecast. Unmeasured judgement is.

Prophesee records overrides as decisions and scores the human half of every forecast alongside the machine half. See what your overrides have been worth. Start here.

New essays land on LinkedIn first. Follow 3RDi to catch them, or get a demo to see Prophesee on your own data.