The Data Readiness Gap in AI Value Creation Plans

The AI line in a value creation plan is usually written as a model problem and executed as a data problem. That gap is where the timeline goes.

I have yet to see a portfolio company AI initiative fail because the model was not good enough. I have repeatedly seen them fail, or slip by a year, because the data the model needed was scattered across systems nobody owned, described inconsistently, incomplete in the fields that mattered, or unusable for the intended purpose on terms the company had already agreed to. The model layer is a commodity you can rent this afternoon. The data layer is the work.

This matters for underwriting because the two have completely different timelines. A plan that assumes a two-quarter path to an AI capability is implicitly assuming the data is ready. It very often is not, and that assumption is rarely tested during diligence.

Five preconditions

Access. Can the data be reached programmatically, in one place, without a person exporting it? A remarkable number of companies hold the data they need in systems with no usable interface, in a vendor platform that does not permit bulk export, or in a warehouse fed by a pipeline that breaks quietly. If getting the data requires a human every time, nothing built on top of it will run reliably.

Rights. Is the intended use permitted? This is the precondition that produces the ugliest surprises, because it is often discovered after the work is done. Customer data collected under a privacy policy that describes service delivery is not automatically available for model training. Data licensed from a third party for internal analytics is not automatically available for a product feature. And rights granted to a named entity do not automatically survive a change of control, which makes this a diligence question and not just an operating one.

Coverage. Does the data span the cases the system will meet in production? Models trained or evaluated on the common cases fail on the tail, and in most businesses the tail is where the cost and the risk concentrate. A dataset that represents the past two years of a business that has changed its mix in the last six months has a coverage problem that no amount of model quality will fix.

Labels. For anything requiring supervision or evaluation, are the outcomes recorded? This is the single most common gap. Companies have abundant records of what happened and almost no record of whether it was right. Without labeled outcomes you cannot fine-tune meaningfully and, more importantly, you cannot build an evaluation set — which means you cannot tell whether any change you make is an improvement.

Consistency. Does the same real-world thing have the same representation across systems? Two customer identifiers, three product taxonomies, and a date field that is local time in one system and UTC in another are ordinary conditions in a company that grew by acquisition. Each one becomes a correctness bug in anything built on top.

What the gap costs

For a mid-market portfolio company with a typical estate, closing these gaps is a two-to-four-quarter program before the AI work begins in earnest — and it is largely unglamorous engineering: pipelines, identity resolution, a warehouse that can be trusted, an annotation process, and someone accountable for the definitions.

That work is not wasted. It is the same foundation that makes reporting reliable, and it typically improves several things unrelated to AI. But it needs to be in the plan, with its own timeline and cost, rather than assumed away inside an AI initiative that will otherwise appear to be failing for nine months while the actual work happens.

The failure mode when it is not in the plan is predictable and I have watched it several times. A pilot is built on a hand-assembled dataset and works impressively. The pilot cannot be productionized because the hand assembly does not scale. The team spends three quarters building data infrastructure that was never scoped, while reporting against milestones that assumed a model problem. Sponsor confidence erodes, because from outside it looks like an AI project that cannot ship — and the engineering team, which is doing exactly the right work, gets the blame for a scoping failure that happened before they were consulted.

The diligence version

Six questions, asked of engineering rather than of the CEO. They take about an hour.

  1. For the AI capability in the plan, which systems hold the required data, and can it be reached without a human in the loop?
  2. What is the documented right to use that data for that purpose, and does it survive a change of control?
  3. Are outcomes recorded — not just actions, but whether the action was correct?
  4. Is there a single definition of your core entities, or does each system have its own?
  5. What is your current data engineering capacity, and what else is it committed to?
  6. Has anyone built a pilot on this data? What did it require to assemble, and could that be automated?

Question 6 is the most efficient. A pilot assembled by hand from three exports tells you the data layer is not ready, regardless of how good the pilot looked. Question 5 tells you whether there is anyone available to fix it — and the answer is frequently that the same two engineers are already committed to the platform migration also in the plan.

How to price it

Do not discount the AI thesis because the data is not ready. Reprice the timeline.

Where the five preconditions are met, an AI capability can move quickly, and the plan's assumptions are probably sound. Where two or three are missing, add two to four quarters and the associated engineering cost before the AI work starts producing anything, and check whether the value creation plan's return still clears the bar on that schedule. Where rights are the missing precondition, treat it as a legal finding as well as a technical one, because it may constrain what can be built at all rather than merely when.

And where the plan contains no data workstream whatsoever, that absence is the finding. It means the AI line was written by someone who has not done this before, and the number attached to it should be treated accordingly.


Data readiness is one of the preconditions assessed in a two-week AI assessment. Related reading: The AI Diligence Question Bank for the full question set, and Build, Buy, or Rent on why a data asset is the main reason to build anything. Book a free discovery call.

Underwriting an AI premium?

AI diligence and unit economics engagements start with a fixed-scope, fixed-fee two-week assessment.

AI Diligence & Unit Economics