The AI Diligence Question Bank

Conventional technical diligence has a settled question set. Cloud diligence has one too — I published mine as a checklist. AI diligence does not, which is why most processes end up either accepting the management narrative or asking questions so general that any competent CTO can answer them without revealing anything.

What follows is the question bank I use. It is organized by what each group is testing rather than by system area, because the point is not to inventory the architecture. The point is to find out whether the AI premium in your model is supportable.

A note on method before the questions. Ask these of engineering, not of the CEO. Ask them in a working session rather than in writing. And pay as much attention to how quickly an answer arrives as to its content — the questions below are ones a team with genuine capability answers immediately, because they have already had the argument internally.

Group 1: Capability versus passthrough

The first thing to establish is what the company actually owns.

  1. Walk me through what happens between a user action and a model response. Which components are yours?
  2. Which models are you calling, from which providers, and at what version?
  3. Have you trained or fine-tuned a model? On what data, and what did it improve?
  4. If you fine-tuned, what is the measured delta against the base model on your own evaluation set?
  5. What proportion of your AI functionality would stop working if your primary model provider went down tomorrow?
  6. What is in your prompts that a competitor could not write?
  7. What does your retrieval layer index, and where did that corpus come from?
  8. Which parts of the system are engineering you would describe as hard?

A team with real capability answers question 4 with a number. A team without one either has not built an evaluation set or has not thought to measure the fine-tune against the base — and in the second case, the fine-tune may be providing nothing.

The answer to question 6 is the most diagnostic in the group. If the honest answer is domain vocabulary and formatting instructions, the prompt is not an asset.

Group 2: Data rights and provenance

This is where deals get repriced, and it is routinely under-examined because it sits between the technical and legal workstreams and neither team owns it.

  1. What data was used for training or fine-tuning, and what is the documented right to use it for that purpose?
  2. Does that right survive a change of control?
  3. For customer data used in training: what does the contract say, and what does the privacy policy say? Do they agree?
  4. Is any customer data used to improve a model that serves other customers? Do those customers know?
  5. What data flows to your model providers, under what terms, and with what retention?
  6. Are there provisions in your provider agreements about training on your inputs?
  7. Do you have data residency obligations, and does your inference path honor them?
  8. If a customer demanded deletion of their data from a trained model, what would you do?

Question 16 has no good answer for most companies, and that is the point. It establishes whether anyone has thought about the problem. Question 10 is the one that most often produces a material finding: rights granted to a named entity for internal purposes do not automatically transfer, and a company whose training corpus rests on such a grant has an asset that may not survive your acquisition of it.

Group 3: Unit economics

  1. What is your inference cost per active user, per month?
  2. What is your inference cost per transaction, or per unit of whatever you charge for?
  3. What proportion of cost of revenue is now inference?
  4. How has that proportion moved over the last four quarters?
  5. What happens to gross margin at three times current usage, holding pricing constant?
  6. Which customers are gross-margin negative on AI cost alone?
  7. Is your pricing per-seat, usage-based, or flat? Is usage bounded in the contract?
  8. What is the largest single-customer inference bill, and what do you charge that customer?

Most teams cannot answer 17 and 18 with confidence, which is itself the finding — a company that cannot attribute inference cost cannot manage it. Question 21 is the one that matters most for your model, and question 24 frequently surfaces a customer being served at a loss that nobody had isolated.

Group 4: Evaluation and quality

  1. How do you know the system is working? Show me the numbers.
  2. Do you have a held-out evaluation set? How large, and who built it?
  3. When you change a prompt or swap a model, what regression testing runs?
  4. What is your measured error rate, and how is an error defined?
  5. What happens when the model is confidently wrong? Who finds out?
  6. Have you had a quality incident? What was the cause and what changed after?
  7. Is there a human in the loop, and what fraction of volume do they touch?
  8. If a human is in the loop, is that cost in cost of revenue or in operating expense?

Question 32 catches a specific and common misstatement. Where AI output requires human review, that review is a cost of delivering the product. Companies frequently book it below the gross margin line, which flatters the margin on exactly the capability being sold at a premium.

Question 25 separates two populations cleanly. Teams with evaluation discipline reach for a dashboard. Teams without one describe their process.

Group 5: Dependency and switching cost

  1. If your primary provider raised prices 100 percent, what would you do, and how long would it take?
  2. Have you ever switched models in production? What happened?
  3. Is there an abstraction layer between your application and the model, or are provider calls scattered through the codebase?
  4. Which provider-specific features are you using that have no equivalent elsewhere?
  5. What is your committed spend with model providers, and what are the terms?
  6. Is any part of your capability dependent on a model that has already been announced as deprecated?

Question 34 is the best single question in this group. A team that has switched models in production has proven the abstraction works. A team that never has is asserting it.

Group 6: The organization

  1. Who owns AI cost? Not who reports it — who is accountable for it?
  2. If your most senior AI engineer left, what would stop?

How to read the answers

None of these questions are individually decisive. Read them in aggregate, and read for a pattern.

A company that answers Group 1 thinly, Group 3 vaguely, and Group 4 with demos is a distribution business with an API call in it. That can still be a good acquisition, and often is — but the premium should attach to the distribution, the workflow, and the customer relationships, all of which are real. It should not attach to the AI, because the AI is available to every competitor at list price.

A company that answers Group 2 uncertainly has a legal problem you are about to own.

A company that answers Group 3 precisely and Group 5 with evidence has probably already captured the easy optimization, which means you should not underwrite it a second time in your value creation plan.

And a company that answers all six groups well is unusual, defensible, and worth the premium. In that case the diligence has done its job in the other direction — it has given you a reason to be confident rather than a reason to discount.


This question bank is the working spine of an AI diligence engagement. Related reading: AI Whitewashing on the five tests behind Group 1, and When AI-Powered Quietly Breaks Gross Margin on modeling Group 3. Book a free discovery call to discuss a live process.

Underwriting an AI premium?

AI diligence and unit economics engagements start with a fixed-scope, fixed-fee two-week assessment.

AI Diligence & Unit Economics