One in four executives says an audit caught an AI error that got out

A quarter of executives say internal audits have caught AI-produced errors that reached external audiences or the board, and almost nobody trusts their data enough to feed the machines. Assurance over AI outputs is now internal audit's job, arriving faster than its methodology.

3 min read

Workiva's 2026 midyear benchmark surveyed 2,272 finance, risk and sustainability professionals, including 847 C-level executives, and produced the statistic this profession should sit with: one in four executives say internal audits have detected AI-generated errors that reached external audiences or board members. In the same study, only 11% said their organisation's data quality is sufficient for AI use.

Put the two findings together and the situation states itself. Machine-produced numbers are flowing into the most consequential communications an organisation makes, drawn from data the organisation itself does not trust, and the error rate is no longer hypothetical. The verification gap, in Workiva's phrase, is not coming. It is reporting to the audit committee.

Assurance demand is outrunning assurance capacity

Internal audit has seen this movie once before, with cyber. Gartner's 2026 audit planning research, surveying 160 chief audit executives, found 96% with cyber assurance activities planned, while only 48% expressed high confidence in their ability to actually provide that assurance. It took a decade for cyber to travel from emerging topic to universal audit plan item, and the confidence still has not caught up with the coverage.

AI assurance is starting the same journey with less runway. Boards that have watched an AI error reach them will ask, at the next committee, a version of: who is assuring these outputs? The chair does not care that the methodology is immature. And the function fielding the question is already stretched. Audit budgets have tightened while the profession absorbs new global standards and most audit leaders carry duties beyond audit itself.

Why the traditional toolkit cannot cover this

The deeper problem is not capacity but method. Internal audit's core instrument is periodic, sampled testing. Select a quarter, pull 25 items, examine them, conclude. Against an AI estate this instrument fails on every axis:

  • Volume. A model producing thousands of outputs daily makes a quarterly sample of a few dozen items statistically decorative.
  • Drift. A model that was accurate in March can be quietly wrong by June. Point-in-time conclusions expire faster than the audit cycle that produced them.
  • Provenance. Auditing an output requires knowing which model version produced it, trained on what data, approved by whom. In most estates that record does not exist, so the audit stalls at the first question.

You cannot sample your way to assurance over a system that never stops producing. The assurance has to be as continuous as the thing assured.

What assurable AI actually requires

The encouraging news is that the requirements are knowable, and they are infrastructure, not heroics:

  1. Provenance as a record, not a recollection. Every production model carries an immutable passport: version, owner, frozen training data, backtest with misses, approvals. Audit's first three questions become lookups.
  2. Continuous monitoring of the outputs themselves. Calibration tracked against outcomes every cycle, drift detected when it happens, exceptions routed to named owners; the control operates over the whole population, always, and audit assures the monitoring rather than re-performing it quarterly.
  3. Determinism where numbers are made. An output that cannot be regenerated from its inputs cannot be audited at all. Same inputs, same number is the precondition for everything above.

Audit leadership should also note the upside. The function that arrives at the committee with a working answer to "who assures the AI?" is having the trusted-advisor conversation the profession keeps saying it wants. The one that arrives with a sampling plan is having the other conversation.

The committee question is coming either way. The difference is whether audit meets it with a sampling plan or with standing infrastructure. Continuous output monitoring, model passports, deterministic recomputation. The Prophesee Audit module is the second answer. Make the AI estate auditable.

New essays land on LinkedIn first. Follow 3RDi to catch them, or get a demo to see Prophesee on your own data.