An evaluation is not a pilot

A pilot proves software can run. An evaluation proves it changes a number you care about, on your data, against a baseline agreed before anyone saw a result. Eight weeks is enough, if the metric comes first.

3 min read

The enterprise pilot has a design flaw so standard it is invisible: it is scoped to succeed. The success criteria are vague enough to be unfailable, "validate the use case", "explore the potential". The measurement is chosen after the results exist, which guarantees a favourable measurement exists to be chosen. And the vendor's best people hover over the deployment in a way they never will again.

Then the pilot succeeds, the purchase follows, the rollout meets reality, and the value fails to materialise at scale. The famous statistic about enterprise AI pilots delivering no measurable impact (and how it gets misread) is manufactured in large part by this exact design: not by technology failing, but by nothing being set up that could measure whether it worked.

A pilot proves the software runs. It was never designed to prove the software matters.

The one discipline that changes everything

The alternative is an evaluation, and nearly all of the difference is a single discipline applied at the start: the metric and the baseline are agreed in week one, in writing, before anyone has seen a result.

Which number should move: forecast error on these items, alert false-positive rate in this queue, days earlier that this risk is detected, hours out of this close process? What is that number today, measured honestly? What movement, by week eight, would justify a purchase, and what movement would not?

Everything else in the eight weeks is execution:

  • Weeks two to seven: run on your data, with your users, inside real decisions. Not a sandbox, not curated extracts, not the vendor driving. The system takes the inputs your operation actually produces, and its outputs land with the people who would own them in production.
  • Week eight: measure against the week-one baseline. The result is a number against a number agreed before the exercise could bias it. Either the metric moved past the threshold, or it did not.

Both outcomes are wins. A clean "yes" is a purchase justified by evidence. A clean "no" is a purchase avoided for the price of eight weeks, which is the cheapest failed AI initiative your organisation will ever run.

Why vendors resist, and what resistance tells you

An evaluation on these terms transfers risk from buyer to vendor, which is why the structure is rare. A vendor whose product creates measurable value inside eight weeks can accept the terms routinely; the evaluation is simply their product doing what it does, observed. A vendor whose value story depends on multi-year transformation narratives, integration programmes before any measurable effect, or metrics chosen in retrospect, cannot accept them, and the ways they decline are informative. "Value takes longer to emerge" sometimes means the value is real but slow. It reliably means the value is unmeasured.

Two honest caveats keep the standard credible. First, eight weeks suits decisions with fast feedback: forecasts, alerts, triage, process cycle times. Outcomes that resolve over years need staged proxies, agreed with the same week-one discipline. Second, an evaluation needs enough historical data to establish the baseline honestly; where the baseline itself is unmeasurable, that discovery is the first finding, and it is worth having before any purchase rather than after.

The deeper point is symmetry. The seller should be as exposed to the measurement as the buyer is to the purchase. Week-one metrics are how a buyer finds out, cheaply, which vendors have been exposing themselves to measurement all along.

Agree the number before anyone sees a result. Everything after that is just finding out.

The eight-week value evaluation is how Prophesee starts every engagement. Your data, your users, a week-one baseline, and a week eight measured against it, or the evidence that you should not buy. Agree your metric. Start here.

New essays land on LinkedIn first. Follow 3RDi to catch them, or get a demo to see Prophesee on your own data.