Build, Buy, or Rent: Foundation Model Decisions

The model layer decision gets made badly in two directions. Some companies train when they should be calling an API, burning a year and a great deal of capital to arrive at something worse than the commercial option. Others call an API when they hold data that would make a tuned model materially better, and leave their only real advantage unused.

Both errors are expensive, and both are usually made without a written comparison. Here is the frame I use.

The four options

Rent. Call a commercial model over an API. Zero fixed cost, immediate availability, frontier capability, fully variable pricing. You depend on a vendor's roadmap, pricing, and availability, and you own nothing.

Rent and tune. Fine-tune a commercial model on your own data through the provider's tuning interface. Modest fixed cost, retains most of the operational simplicity of renting, and can produce meaningful quality gains on narrow tasks. You are still on the vendor's platform, and the tuned artifact is generally not portable.

Host open weights. Run an open-weight model on your own infrastructure, tuned or not. Substantial fixed cost in GPUs and in the engineering to operate them, but a variable cost that can be dramatically lower at high volume. Full control of the artifact, data path, and version. You now own an inference platform, which is a real operational commitment.

Train from scratch. Pretrain your own foundation model. For almost every company reading this, this is the wrong answer, and the cases where it is right are ones where the model is the entire company and the capital is raised specifically for it.

The default is rent, and the default is usually right

Start from renting and require a reason to move. This is not a counsel of timidity — it is an observation about where the frontier is. Commercial model capability improves faster than any individual company's ability to keep pace, and a company that trained a task-specific model eighteen months ago is frequently beaten today by a general model with a good prompt. Every dollar spent building at the model layer is a bet that the frontier will not come to you, and that bet has been losing consistently.

Renting also preserves optionality, which is worth more than it looks during a hold period. A company with an abstraction layer over three providers can respond to a price change or a capability jump in days.

Conditions that justify moving off rent

Four, and they are cumulative rather than alternative. One is usually not enough.

Proprietary data with labeled outcomes. Not just data — data paired with known-good results, in volume, on a task narrow enough for a model to learn. A company with ten years of expert decisions and the outcomes those decisions produced has something a general model does not. A company with a large document corpus and no labels usually wants retrieval, not tuning.

Volume high enough for the arithmetic to invert. Self-hosting has a fixed cost that only pays back above a threshold. Below it, renting is cheaper and simpler; above it, the variable cost saving can be very large. The threshold moves constantly, so compute it against current prices rather than remembered ones — and compute it on sustained load, since idle GPUs are the fastest way to make self-hosting more expensive than the API it replaced.

A hard constraint on the data path. Regulatory, contractual, or residency requirements that genuinely cannot be met by a commercial provider. Note the word genuinely: most providers now offer arrangements that satisfy most such requirements, and self-hosting justified by a constraint that a contract amendment would resolve is an expensive way to avoid a negotiation.

Latency or determinism requirements the API cannot meet. Real, occasionally. Frequently a proxy for an architectural problem elsewhere.

Absent at least two of these, renting is the answer, and a management team pushing to build should be asked which two they have.

What each option actually costs

The comparison that gets presented is usually API spend against GPU spend, which understates the build side by a wide margin. The costs that get left out:

  • Engineering to operate an inference platform — serving, batching, autoscaling, failover, upgrades. This is a standing team commitment, not a project.
  • The evaluation harness. Required in all cases, but self-hosting makes it load-bearing, because you no longer inherit a provider's quality guarantees.
  • Idle capacity. Reserved GPUs are paid for whether or not traffic arrives, which converts a variable cost into a fixed one. That is only an improvement if utilization is high and steady.
  • Falling behind. A self-hosted model is frozen at the capability of the weights you deployed. Keeping current means periodically redoing the work.
  • Recruiting and retention in a market where the people who can do this well are expensive and mobile.

Against which the build side has genuine benefits that are usually understated in the other direction: unit cost at scale, control over the data path, no vendor concentration risk, and an artifact that appears on your balance sheet rather than someone else's income statement.

The decision in diligence

When you find a company that has built at the model layer, the question is not whether building was fashionable. It is whether the four conditions were present and whether the arithmetic was done.

Ask to see the comparison that was made at the time. Ask what the tuned or self-hosted model measures against the current commercial frontier on their own evaluation set — not against the frontier as it stood when the decision was made. Ask what the fully loaded cost of the inference platform is, including the engineers, and compare it to the API bill it displaced.

Occasionally the answer is impressive: a company with a real data asset, a real volume threshold, and a measured quality advantage that a general model does not match. That is a defensible moat and it belongs in the premium.

More often the answer is that a capable engineering team built something interesting because they wanted to, the comparison was never written down, and the resulting system is now more expensive and slightly worse than the API it replaced. That is not a moat. It is technical debt with a research budget, and it should be priced as a cost to unwind rather than as an asset.

For portfolio companies

If you are operating rather than diligencing, the practical sequence is: rent first, instrument everything, and let the data tell you when a condition has been met. Build the evaluation harness before you need it, because it is the artifact that makes every subsequent decision reversible. Put an abstraction layer between your application and the provider on day one, when it costs nothing. And revisit the decision every couple of quarters, because the arithmetic genuinely does change that fast.


Model-layer decisions are one of the things a fractional AI CTO engagement owns. Related reading: AI Whitewashing on what counts as capability, and Model Tiering on getting most of the benefit without building anything. Book a free discovery call.

Underwriting an AI premium?

AI diligence and unit economics engagements start with a fixed-scope, fixed-fee two-week assessment.

AI Diligence & Unit Economics