When AI-Powered Quietly Breaks Gross Margin
Software has spent two decades teaching investors a comfortable assumption: marginal cost approaches zero. Once the product is built, another customer is close to free, gross margins settle somewhere north of 75%, and the business scales beautifully.
Adding a language model to that product breaks the assumption, and most of the financial models I see during diligence have not caught up.
AI inference is not a fixed cost that amortises. It is a variable cost that scales with usage, in a business whose pricing usually does not. That is a structural change to how the company makes money, and it can hollow out gross margin while revenue is growing nicely enough that nobody notices.
The shape of the problem
Consider the arithmetic in its simplest form. A customer service platform processes ten million support tickets a month and charges its customers fifty cents per ticket. Comfortable pricing, clear value, easy to explain.
The platform routes every ticket through a frontier language model. All-in AI infrastructure — model API charges, inference compute, fine-tuning — runs $3.6M annually. That is thirty-six cents per ticket.
Fourteen cents of gross margin per ticket, before sales, marketing, engineering, or general overhead. A twenty-eight percent gross margin on a business the market is pricing as software.
Now grow it. Volume doubles. Revenue doubles, which is the part that shows up in the board deck. AI costs also double, which is the part that does not, because it is buried in infrastructure rather than reported as cost of revenue. The margin percentage does not improve at all — and if competitive pressure forces a price reduction, it gets worse.
Why it matters for PE: This is the opposite of the operating leverage the model assumes. Scale does not fix it. Scale amplifies it.
AI cost as a share of EBITDA is the number to ask for
Absolute AI spend tells you very little. Three point six million dollars is unremarkable for a company with $50M of EBITDA and catastrophic for a company with $5M.
The ratio is what matters, and it is not usually presented. Ask for AI infrastructure cost as a percentage of EBITDA, and be prepared to compute it yourself from the invoices, because most targets have never framed it that way.
My rough calibration, from the AI-positioned companies I have looked at:
- Under 15% — normal. AI is a feature with a manageable cost.
- 15% to 40% — worth understanding in detail. Usually an optimization opportunity rather than a structural problem.
- Above 40% — the AI strategy and the business model are in tension. Something has to change during your hold period.
- Above 70% — reported EBITDA is not a meaningful number. Treat it as a starting point for a rebuild, not a multiple base.
The reported EBITDA problem
Here is where it becomes a pricing question rather than a technical one.
When AI costs are unsustainably high, reported EBITDA is artificially low — the company is spending more than it needs to on inference. That sounds like good news for a buyer: optimize the spend, EBITDA goes up, you bought cheaply.
The trap is that the seller has usually worked this out too, and will ask you to underwrite the optimized number. "True EBITDA is really $8M once we fix the model routing." Sometimes that is even correct. But you are being asked to pay today for savings that require six to twelve months of technical work you have not scoped, executed by a team that has not done it yet, with the outcome uncertain.
The discipline that works is the same one that works everywhere else in cloud economics: risk-adjust the projection and be explicit about the confidence level.
If $3M of AI optimization is identified, and the realistic capture is 70%, then $2.1M is the number that belongs in the model. Reported EBITDA of $5M becomes risk-adjusted EBITDA of $7.1M, not the $8M the seller is quoting. On a 10x multiple that is a $9M difference in what you should be willing to pay — and it is a difference you can defend in an investment committee, because the confidence level is stated rather than assumed.
What to ask for in the data room
Five requests, none of which should be controversial:
- Twelve months of AI infrastructure invoices, split between foundation model APIs, inference compute, and training.
- Cost per transaction — per ticket, per query, per document, whatever the unit is — trended monthly.
- The same figure alongside revenue per transaction, so the gross margin per unit is visible rather than inferred.
- The model routing policy, if one exists. Which model serves which class of request.
- The pricing model's sensitivity to volume. If a customer triples usage, does revenue triple, or is it capped by a subscription tier while costs keep climbing?
That last one catches a genuinely dangerous pattern: flat-rate pricing over a usage-scaled cost. It works until a large customer discovers the feature and uses it heavily, at which point your best logo becomes your worst margin.
The version of this that ends badly
The failure mode I have watched play out more than once: a company with strong revenue growth, a compelling AI story, and unit economics nobody stress-tested. The firm passes or gets outbid. Twelve to eighteen months later the company is raising a down round, having grown volume enthusiastically into a cost structure that punished every additional unit.
The optimizations that would have fixed it were identifiable at diligence. They usually are. What was missing was somebody asking what a transaction cost and what it earned, in the same sentence.
That is not an AI question. It is the oldest question in unit economics, arriving in unfamiliar clothes.
If you are sizing the optimization rather than the risk, the cloud savings calculator covers the non-AI half of the same estate.
Modeling an AI-positioned target? Book a free discovery call to talk through the diligence.