Your copilot is answering the wrong question

A copilot summarises what already happened, faster. But the questions that move a business are about what happens next, who acts, and whether the answer can be audited. Chat on top of records does not touch them.

3 min read

Every enterprise software vendor now ships a copilot. Every module, from every incumbent: a chat window resting on the system of record, summarising what the system already holds. Adoption has been effortless, because summarising records is genuinely useful and nearly risk-free.

It is also beside the point. MIT NANDA's State of AI in Business research found roughly 95% of enterprise generative AI pilots delivered no measurable profit and loss impact, and the failures stalled overwhelmingly at the workflow boundary: fluent output that did not change how work got decided. The copilot answers the question it can answer. It is the wrong question.

What the business actually asks

Sit in any operating review and listen to the questions that carry money:

  • Will we make the quarter, and what are the odds?
  • Which three of this week's thousand alerts actually need a human?
  • If we intervene now, what happens to the year-end number?
  • Who owns this decision, and what happened the last time we made it?

None of these are retrieval questions. They are prediction, triage, simulation and accountability questions. A chat interface over records can answer "what did we sell in Q2?" beautifully. It has nothing to offer "what will we sell in Q4, and should the plan change?", because nothing behind the window is computing odds, testing interventions, or carrying ownership.

The underlying audit problem

There is a second, harder problem. Ask a generative copilot the same question twice and you can receive two different answers. For drafting an email, that is harmless. For a number that feeds a decision, it is disqualifying: a figure that changes on refresh cannot be reconciled, cannot be explained to an auditor, and cannot be traced when the decision it fed goes wrong.

A probability you cannot audit is only a louder opinion.

Enterprises are noticing. Workiva's 2026 midyear benchmark of 2,272 finance, risk and sustainability professionals found one in four executives reporting that internal audits had detected AI errors which reached external audiences or board members (Workiva, The Verification Gap). The answer engine got deployed before the answer discipline.

What a decision layer needs that a chat window lacks

The copilot is not wrong, it is incomplete in a specific, structural way. Language models are the right tool for language (drafting, summarising, translating a question into a query). They are the wrong tool to be the number. A layer that supports decisions needs four things beneath the conversation:

  1. Predictions with published odds, backtested and calibrated, so "78% likely" is a tested statement rather than a tone of voice.
  2. Deterministic numbers. Same inputs, same figure, every time, so the answer can be audited and two viewers can reconcile what they saw.
  3. A named owner and a threshold, so an answer becomes an action rather than an interesting paragraph.
  4. A record of outcomes, so the organisation learns which answers deserved trust.

Where language helps, use it at the edges: asking in plain language, explaining a driver in plain language. The computation in the middle has to be a different kind of machine.

Prophesee is built in that order. Predictions, exceptions, plans and permission-aware answers that stay the same on refresh, with the copilot conveniences at the edges rather than the centre. See the difference on your own data. Start here.

New essays land on LinkedIn first. Follow 3RDi to catch them, or get a demo to see Prophesee on your own data.