There is a sentence that ends more AI initiatives than any budget committee, and it is spoken in a tone of great responsibility. Our data is not ready. The room nods, a data programme is commissioned, and the AI conversation is adjourned for two years. Nobody has ever been fired for saying it, which is part of the problem.
The sentence is half right. Garbage in, garbage out is not a myth, and executives know how their estate would perform under examination. Only around one in nine say their data quality is sufficient for AI use (the verification gap), and audit-detected AI errors reaching boards suggest the other eight are correct to worry.
The half that is wrong is the sequencing, and the sequencing is where the years go.
The half that is right, stated precisely
Start with what the readiness instinct gets right, because the argument only works if it is honest. Models trained on drifted, duplicated, mislabelled data produce confident nonsense, and confident nonsense in front of a decision maker is worse than no model at all. The caution is earned.
But look at what the caution is actually about. It is almost never about volume. Two decades of ERP, CRM, incident registers, contracts and transaction history sit in the systems of record; the registers alone are training data that most functions have never used as such. Enterprises are not short of data. What they are short of is governance where the data gets used: one resolved identity per real thing, lineage that survives an export, policies that are enforced rather than published, and permissions that hold at the point of consumption.
You do not have a data shortage. You have enough data, ungoverned at the point of decision.
That distinction is the first link in a chain. Ungoverned data makes the AI built on it untrustworthy, and untrusted AI leaves decisions exactly where they were. What looks like three problems, data, AI and decisions, is one chain, and it breaks in the same place it can be fixed. Every year spent waiting adds to the decision debt: the compounding cost of decisions made late, blind or not at all while the estate was being made ready.
Why the clean-up programme never converges
The readiness instinct produces a specific project shape, and every large organisation has run one. Scope the whole estate. Define quality in the abstract, "fit for purpose" with no purpose named. Cleanse, deduplicate, catalogue. Declare partial victory at the budget review and quietly relapse the following year.
The shape fails for a structural reason, not a competence one. A data programme with no consuming decision has no forcing function. Nothing pulls on the data, so nothing establishes which defects matter, which of the four customer records is canonical, or what "done" would even mean. Quality is unmeasurable in the abstract because quality is a relationship between data and a use, and the use was postponed to phase two. The programme cannot converge on a target that has not been named.
Meanwhile the estate keeps moving. Systems are added, fields drift, and yesterday's cleansed table starts decaying the day the project team leaves. A one-time clean of a continuously drifting estate has the same half-life as any one-time fix of a living system, which is to say, one budget cycle.
Invert the sequence
The alternative is not to skip governance. It is to reverse the order and let the decision drive it.
Pick the decision first. The decision names its data, which is always a fraction of the estate; a demand forecast does not care that the marketing taxonomy is a mess. Wire the decision up, and every defect that matters now surfaces at the exact point it bites, in a prediction that misses, an exception that misfires, an audit trail that will not reconcile, with a named owner already attached because the decision has one. The forcing function the clean-up programme never had is simply use.
Run this way, governance stops being a programme and becomes a consequence of use. Entity resolution happens because the decision needs one supplier, not four spellings. Lineage exists because every figure must show its account of itself. Policy conformance is watched continuously because breaches now have somewhere to land. The data gets clean in the order the business actually needs it clean, and it stays clean because something is pulling on it every day.
The decision layer is the forcing function. Data quality is its by-product.
This is also where the readiness excuse quietly inverts. Waiting for clean data before wiring decisions guarantees neither. Wiring the decision first delivers both, and the organisations doing it are not braver, they are sequenced correctly.
Prophesee is built around that inversion, and Prophesee Data Governance is where the by-product becomes visible. An ontology resolves the corpus into entities, owners and applicable policies. The policy is watched continuously, with every breach routed to a named owner. Governed agents execute the approved clean-up, inheriting permissions, budgets and audit by construction. Governance produced this way is hard to copy for the same reason it is hard to fake; it is generated by operating decisions, not written beside them. Start here.