Consider a composite team that signed a three-year GPU commitment in early 2024 at what everyone agreed was an excellent price. On-demand H100 capacity was clearing above $8 per GPU-hour at the time. They locked in around $2.90. Finance called it a win and moved on.

By mid-2026 that contract is the most expensive compute in their stack. On-demand rates for the same accelerator have fallen into the low single digits, spot-style tiers sit under $2, and a Blackwell-class node next door produces two to three times the tokens per dollar for the same inference workload. They are two years into a bet on a fixed rate for a fixed part, and the part moved.

The rent-versus-own decision trades duration risk (a fixed SKU at a fixed rate) against utilization risk (idle capacity or a shortage when you need it). Everything else is detail.

The Rate Card Is a Ladder, Not a Price

Every serious GPU provider now sells the same silicon at four or five different prices depending on how much certainty you are buying. Understanding the ladder is the first step, because vendors will quote you whichever rung makes their competitor look bad.

Tier Typical H100 SXM range What you are buying What you give up
Spot / preemptible $1.00 - $2.00 Cheapest per hour Eviction with minutes of notice, no capacity guarantee
On-demand $2.00 - $4.00 (neocloud), higher at hyperscalers Start and stop freely Highest steady-state cost, availability varies by region
Reserved (1-12 months) $1.80 - $2.80 Guaranteed allocation Pay for idle hours, limited SKU flexibility
Committed (1-3 years) $1.50 - $2.40 Best rate, priority access Duration risk, hard to exit
Owned See below Full control, resale option Capital, operations, obsolescence

Thunder Compute's market tracking shows the spread across providers is wide enough that two vendors quoting the same accelerator can differ by 2x, and the gap between hyperscaler list pricing and specialist neocloud pricing remains substantial. The list price at a major cloud is not the market price.

The interesting part is that these tiers are not really different products. They are different allocations of risk. Spot means the provider keeps the flexibility. A three-year commitment means you keep the risk. The price difference is the premium for who holds the bag when demand shifts.

What Owning Costs

Most rent-versus-buy decks compare a rental rate against a hardware purchase price. That comparison is wrong by roughly 60%, because the accelerator is not the cluster.

Here is the loaded model for an H100-class GPU in a real deployment:

Cost component Annual per GPU Notes
Depreciation $8,000 - $11,000 Loaded capex of $32k-$45k per GPU over 4 years. Includes InfiniBand or equivalent fabric, storage, racks, and integration, which typically add 40-60% on top of the bare accelerator
Colocation (space, power, cooling) $2,800 - $3,600 Roughly 1.4 kW average draw per GPU at realistic PUE, billed at $150-$250 per kW-month
Operations, spares, software $1,500 - $2,500 Site reliability staff, firmware management, failed-node replacement, scheduler and orchestration tooling
Financing $2,000 - $3,500 At 9-12% on the average outstanding balance
Total $16,000 - $19,000

Take the midpoint at $17,500 per year. Divide by 8,760 hours and you get about $2.00 per GPU-hour at 100% utilization. At 65% utilization that becomes $3.07. At 40% it becomes $5.00.

Compare that against a committed neocloud rate near $2.10, and the break-even utilization sits somewhere around 90-95%, a level most organizations and even many hyperscalers do not hit.

This is the reversal that catches finance teams off guard. In 2023 and 2024, when rentals cleared $8 per hour, ownership won at 30% utilization. The rental market has compressed faster than the cost of building, so the same spreadsheet now produces the opposite answer with the same methodology. If your buy case was built more than eighteen months ago, rerun it before signing anything. The same dynamic is reshaping who can supply capacity at all, which we covered in the AI compute buildout and its dependencies.

Utilization Is Not What Your Dashboard Says

The break-even math is only as good as your utilization number, and almost everyone overstates it. There are three different measurements and teams routinely conflate them:

  • Allocated utilization: what fraction of hours a GPU is assigned to a job. Usually high, usually meaningless.
  • Busy utilization: what fraction of assigned hours the GPU is actually executing kernels. Falls sharply with data loading stalls, checkpoint writes, and straggler synchronization.
  • Goodput: what fraction of wall-clock time produced work you kept. This is the only one that belongs in the model.

Meta published unusually honest numbers on this in the Llama 3 technical report. Over a 54-day pretraining window on a 16,384-GPU H100 cluster, they logged 466 job interruptions, 419 of them unexpected, with the large majority traced to hardware faults. Effective training time landed above 90%, but only because they had built substantial automated recovery tooling. A team without that tooling loses far more.

Production clusters running mixed workloads look worse. The Alibaba team's characterization of LLM development in their datacenter documented a wide gap between resource allocation and effective use across research and production jobs, driven by short-lived debugging runs, queuing delays, and failed jobs. Their data is a useful antidote to capacity plans built on the assumption that a purchased GPU is a working GPU.

Use your measured trailing-twelve-month goodput. If you do not measure it, that is the first project, not the contract negotiation.

Normalize on Tokens, Not Hours

The dollars-per-GPU-hour comparison has a structural flaw: it treats accelerator generations as interchangeable units, and they are not. A committed rate that looks cheap on an older part can be expensive on the metric your business actually pays for.

The right unit depends on the workload:

  • Inference: dollars per million output tokens at a fixed p95 latency and fixed context length. Latency has to be pinned or the number is meaningless, because you can always trade throughput for tail latency.
  • Training: dollars per effective training day, or dollars per unit of achieved model FLOPs.
  • Fine-tuning and batch: dollars per completed job at a target quality threshold.

MLCommons publishes MLPerf inference and training results with enough per-system detail to build a rough throughput ratio between generations for common model shapes. It will not match your workload exactly. It will tell you whether a generational jump is worth 20% more per hour or 200% more per hour, which is the decision you are making.

Run this calculation before you sign, then run it again with the assumption that a new generation ships halfway through your term. If a plausible next-generation part at plausible rental pricing beats your committed contract on cost per token, you have not locked in savings. You have locked in a ceiling on your own efficiency. This connects directly to how you should think about agent tokenomics and open model economics, where cost per token compounds across every agent step.

The Contract Terms That Decide the Outcome

The headline rate is the part vendors want you to negotiate, because it is the part they have modeled. The terms below move more money and get far less attention.

Ramp and acceptance testing. Capacity arrives in tranches. Define a ready-for-service test that the tranche must pass before billing starts: sustained NCCL all-reduce bandwidth above a stated threshold, a burn-in period at full load, and a maximum node failure rate during burn-in. Without this, you pay for weeks of a cluster you cannot use.

SLA definition. A 99.9% per-node availability SLA on a 512-node cluster means you should expect a node down more or less continuously. That is fine for inference and fatal for synchronous training. Negotiate a cluster-level availability or goodput commitment, and require a hot spare pool of 3-5% of node count with automated replacement.

SKU substitution rights. The single most valuable clause in a multi-year deal. Get the right to convert committed spend to a next-generation part at a defined conversion ratio when it becomes generally available. Providers will grant this more readily than you expect, because it keeps you from walking.

Dollar commitments over unit commitments. Commit to spend, not to GPU-hours of a named SKU. Dollar commitments flow across generations, regions, and instance types. Unit commitments strand you.

Burn-down and rollover. Unused committed hours should roll forward 30-90 days rather than expiring monthly. This alone can recover 5-10% of a commitment for teams with lumpy workloads.

Assignment and exit. Can you transfer the contract in an acquisition? Can you sublease unused capacity to another tenant? Is there termination for convenience with a defined fee schedule? Ask all three in writing.

Data gravity. Egress fees, storage tier pricing, and whether your fabric is dedicated or shared. A cheap compute rate attached to expensive egress is a retention mechanism, not a discount.

Counterparty Risk Is Now a Line Item

A three-year commitment is an unsecured credit exposure to your provider. Price it accordingly.

The capital intensity behind these rate cards is visible in public filings. Providers that report to the SEC disclose their debt structure, depreciation assumptions, and customer concentration, and you can read all of it through EDGAR full-text search. Three things to look for:

  1. Useful life assumptions. If a provider depreciates accelerators over six years while financing them over four, their pricing may be a function of accounting choice rather than sustainable cost. That gap closes eventually, and it closes on your renewal.
  2. Customer concentration. A provider where one or two customers represent the majority of revenue is a provider whose capacity commitments to you are subordinate to keeping those customers happy.
  3. Whether the datacenter is owned, leased long-term, or subleased. Subleased capacity means your provider has a counterparty risk of their own.

The consolidation pressure across specialist providers is already visible, and Vultr's analysis of neocloud consolidation lays out the shape of it: capital costs rising, rental prices falling, and margin compressing for anyone without either scale or a differentiated position. Some of your shortlist will not exist in their current form when your term ends.

Practical mitigations: split committed capacity across at least two providers even at a small price penalty, require migration assistance obligations on termination, keep your orchestration layer portable rather than built on provider-specific scheduling primitives, and hold a small on-demand relationship with a third provider so you have a warm path if you need one fast. Supply-side constraints upstream make this more than theoretical, as we discussed in the TSMC capacity planning piece.

Build a Portfolio, Not a Decision

Utilities stopped asking "should we build a plant or buy power" decades ago. They run a portfolio: baseload generation for the floor, mid-merit for the predictable swing, peakers and market purchases for the spikes. Compute planning should work the same way.

A workable allocation policy:

  • Commit to the P20 of trailing-twelve-month goodput. The capacity you have used at least 80% of days. This is genuinely baseload and deserves the best long-term rate.
  • Reserve to the P50 on shorter terms. One to six month reservations that you can reshape as the workload mix changes.
  • Burst above it on on-demand and spot. Accept the higher unit price for the top 20% of demand. Paying $3.50 per hour occasionally beats paying $2.10 per hour continuously for capacity you use a third of the time.
  • Own only what has a non-cost justification. Data residency requirements, regulated workloads, sovereignty constraints, or a genuine multi-year floor with staff to run it. Cost alone rarely gets you there in 2026.

Review the split quarterly against measured goodput, not against forecasts. Forecasts for agent workloads have been wrong in both directions and will keep being wrong. Epoch AI's tracking of compute price trends is a useful outside check on whether your assumed rate curve still resembles the market.

Measure Goodput Before You Commit

Most waste in client compute spend starts with an uninstrumented workload, so the utilization number in the business case remains an estimate that nobody revisits.

Our audit phase measures the thing the contract depends on: goodput by workload, token throughput at your real latency requirement, and where the idle hours are coming from. That produces a defensible P20 and P50 rather than a guess. From there the design and build phases focus on portability - orchestration and inference serving that can move between providers without a rewrite, so the commitment you sign is a pricing decision rather than a lock-in decision.

If you are staring at a multi-year commitment and the utilization number came from a slide rather than a metric, bring the workload data to a contract review before renewal.

Run the break-even calculation with your own measured goodput before the next renewal cycle. If it lands below 85%, the ownership case needs a reason that is not about cost, and the commitment needs substitution rights.