[{"data":1,"prerenderedAt":4},["ShallowReactive",2],{"article-body-your-copilot-is-answering-the-wrong-question":3},"\nEvery enterprise software vendor now ships a copilot. Every module, from\nevery incumbent: a chat window resting on the system of record, summarising\nwhat the system already holds. Adoption has been effortless, because\nsummarising records is genuinely useful and nearly risk-free.\n\nIt is also beside the point. MIT NANDA's State of AI in Business research\nfound roughly 95% of enterprise generative AI pilots delivered no measurable\nprofit and loss impact, and the failures stalled overwhelmingly at the\nworkflow boundary: fluent output that did not change how work got decided.\nThe copilot answers the question it can answer. It is the wrong question.\n\n## What the business actually asks\n\nSit in any operating review and listen to the questions that carry money:\n\n- Will we make the quarter, and what are the odds?\n- Which three of this week's thousand alerts actually need a human?\n- If we intervene now, what happens to the year-end number?\n- Who owns this decision, and what happened the last time we made it?\n\nNone of these are retrieval questions. They are prediction, triage,\nsimulation and accountability questions. A chat interface over records can\nanswer \"what did we sell in Q2?\" beautifully. It has nothing to offer \"what\nwill we sell in Q4, and should the plan change?\", because nothing behind the\nwindow is computing odds, testing interventions, or carrying ownership.\n\n## The underlying audit problem\n\nThere is a second, harder problem. Ask a generative copilot the same\nquestion twice and you can receive two different answers. For drafting an\nemail, that is harmless. For a number that feeds a decision, it is\ndisqualifying: a figure that changes on refresh cannot be reconciled,\ncannot be explained to an auditor, and cannot be traced when the decision\nit fed goes wrong.\n\n> A probability you cannot audit is only a louder opinion.\n\nEnterprises are noticing. Workiva's 2026 midyear benchmark of 2,272\nfinance, risk and sustainability professionals found one in four\nexecutives reporting that internal audits had detected AI errors which\nreached external audiences or board members ([Workiva, The Verification\nGap](https://www.workiva.com/resources/executive-benchmark-survey-verification-gap)).\nThe answer engine got deployed before the answer discipline.\n\n## What a decision layer needs that a chat window lacks\n\nThe copilot is not wrong, it is incomplete in a specific, structural way.\nLanguage models are the right tool for language (drafting, summarising,\ntranslating a question into a query). They are the wrong tool to *be* the\nnumber. A layer that supports decisions needs four things beneath the\nconversation:\n\n1. **Predictions with published odds**, backtested and calibrated, so\n   \"78% likely\" is a tested statement rather than a tone of voice.\n2. **Deterministic numbers.** Same inputs, same figure, every time, so the\n   answer can be audited and two viewers can reconcile what they saw.\n3. **A named owner and a threshold**, so an answer becomes an action rather\n   than an interesting paragraph.\n4. **A record of outcomes**, so the organisation learns which answers\n   deserved trust.\n\nWhere language helps, use it at the edges: asking in plain language,\nexplaining a driver in plain language. The computation in the middle has to\nbe a different kind of machine.\n\nProphesee is built in that order. Predictions, exceptions, plans and\npermission-aware answers that stay the same on refresh, with the copilot\nconveniences at the edges rather than the centre. See the difference on\nyour own data. [Start here](/contact).\n",1786786820782]