AI Whitewashing: Telling Capability From a Wrapper

Greenwashing gave us a useful word for a familiar behavior: claiming an environmental virtue you have not earned because the market pays for it. We now need the same word for artificial intelligence, because the same thing is happening, and it is showing up in deal pricing.

AI whitewashing is the practice of presenting a product as AI-driven when the artificial intelligence involved is an API call to a foundation model that any competitor can make just as easily. It is not always deliberate. Founders genuinely believe they have built something differentiated, because from the inside it feels like they have. But when a PE firm pays a multiple premium for "AI capability," it matters enormously whether that capability is a moat or a line item on someone else's invoice.

I have watched this dynamic distort technical due diligence over the past two years. The pattern is consistent: management presentations lead with AI, the demo is genuinely impressive, and nobody in the room can tell whether what they just saw represents two years of engineering or two weeks. Here is how to find out.

Why the incentive is so strong right now

Companies positioned as AI-native command meaningful valuation premiums. In the deals I see, the spread between "software company" and "AI company" multiples runs anywhere from 2x to 5x EBITDA, and sometimes more in competitive processes.

That spread is the entire problem. When a label is worth several turns of EBITDA, sellers will find a way to qualify for it. The word "AI" appears in the CIM whether or not anything underneath it is defensible, and the burden falls on diligence to work out which.

Why it matters for PE: You are not trying to determine whether a company uses AI. Almost all of them now do. You are trying to determine whether their use of it is something a competitor could replicate in a quarter. That question has a price attached.

Test 1: Ask what happens if the API key is revoked

This is the fastest diagnostic I know. Ask the CTO directly: if your foundation model provider terminated your account tomorrow, what would your product do?

The answers sort companies quickly. A company with genuine capability will describe a migration path, fallback models, an abstraction layer they built precisely because they anticipated this, and a rough estimate of how long the switch would take. A company with a wrapper will describe an outage.

Neither answer is disqualifying on its own. Plenty of good businesses depend on a vendor. But the second answer tells you the AI is infrastructure the company rents, not an asset it owns, and it should be valued accordingly.

Test 2: Read the AI spend breakdown, not the AI narrative

Ask for the AI infrastructure costs split three ways: foundation model API charges, self-hosted inference compute, and training or fine-tuning.

That split is close to a fingerprint. A company whose AI spend is almost entirely third-party API charges has not built a model; it has built a prompt. A company with meaningful inference infrastructure and a real training line is doing something harder to copy. The ratio tells you more about technical depth than any architecture diagram in the data room.

Watch the absolute numbers too, because they set up the unit economics question. I have seen AI infrastructure run north of 70% of reported EBITDA at companies presenting themselves as high-margin software businesses.

Why it matters for PE: The spend breakdown is one of the few AI diligence artifacts that is difficult to dress up. Invoices are invoices.

Test 3: Ask to see the evaluation harness

Every team seriously building with language models has an evaluation suite: a set of test cases, a scoring method, and a record of how model changes affected quality over time. It is how you ship changes without breaking things silently.

Ask to see it. Ask how they knew their last prompt change was an improvement. Ask what their quality metric is and how it moved over the past two quarters.

Teams doing real work answer immediately and often enthusiastically, because this is the unglamorous part nobody asks about. Teams doing prompt engineering in production answer with anecdotes. The absence of an evaluation harness is one of the more reliable signals that AI is a feature bolted on rather than a discipline practiced.

Test 4: Find out whether the data is actually a moat

"We have proprietary data" is the most common defense against the wrapper accusation, and it is sometimes true. Test it with three questions.

Is the data genuinely exclusive, or is it available to anyone with the same customers? Is it in a form usable for training, or is it unstructured exhaust nobody has curated? And critically: is it being used? Plenty of companies hold genuinely valuable data and do nothing with it beyond the occasional dashboard.

A proprietary dataset that has never been used to fine-tune, retrieve against, or evaluate a model is an asset, but it is not evidence of AI capability. It is evidence of potential, which is worth something rather less.

Test 5: Look for model tiering

This one is subtle and it is my favorite, because it correlates with sophistication better than almost anything else.

Ask which model handles which request. A team that routes every request to the most capable and expensive model available has not thought carefully about the problem. A team that has segmented its traffic, established that most requests are simple and can be served by a small cheap model, and reserved the frontier model for the minority that genuinely needs it, has done real engineering and understands its own cost structure.

In the workloads I have analyzed, the distribution is remarkably consistent: roughly 70% of requests are simple, 25% moderate, and only about 5% genuinely require frontier capability. A company still sending all of it to the top-tier model is telling you two things at once — that its unit economics are worse than they need to be, and that nobody has looked.

What the premium should attach to

None of this argues against paying for AI capability. It argues for attaching the premium to the right thing.

Pay for proprietary data being actively used. Pay for an evaluation and deployment discipline that lets a team ship model changes safely. Pay for architecture that treats the model as a swappable component. Pay for a workflow so embedded in the customer's operation that replacing it means changing how they work.

Do not pay for the word. Do not pay for a demo. And be skeptical of a premium justified primarily by the sophistication of the model the company is calling, because that sophistication is available to their competitors at list price.

This is the shape of question that separates a real assessment from a demo — see representative engagements for how that plays out across a deal. The uncomfortable version, which I would ask in every process: if a competent team were given this company's customer list and six months, could they rebuild the AI capability? If the honest answer is yes, you are buying a customer list — which may still be a fine thing to buy, at customer-list prices.


Evaluating an AI-positioned target? Book a free discovery call to discuss what the diligence should cover.

Need help with cloud economics?

Book a free discovery call to discuss how I can help your fund or portfolio company.

Book a Discovery Call