Ollama's Jeffrey Morgan has claimed that open-weight models now handle 80 to 90 percent of enterprise AI tokens. That number will show up in board decks this quarter. Very few of the people repeating it will have checked it against their own logs, and that matters, because a company where open models carry 85 percent of the tokens can still send most of its inference budget to a handful of frontier API calls.

That gap is what this post is about. Per-token prices keep falling, yet enterprise AI bills keep rising. 2026 inference-economics analyses report a roughly 280x drop in per-token costs over two years, while average enterprise AI spend rose from $1.2 million in 2024 to $7 million in 2026. Inference now makes up about 85 percent of the AI budget (Oplexa, inference economics coverage). Agentic workflows, longer contexts, and the end of subsidized frontier pricing all push volume up faster than prices come down. If you don't know your own token mix, you can't tell which of these forces is driving your bill.

Token Share and Spend Share Are Different Numbers

Volume and spend have come apart. Here is an illustrative month for a mid-sized deployment. The prices are blended placeholders, not quotes from any provider.

Bucket Tokens per month Blended price per 1M tokens (illustrative) Monthly cost Share of tokens Share of cost
Open-weight models 8.5B $0.20 $1,700 85% 15.9%
Frontier models 1.5B $6.00 $9,000 15% 84.1%
Total 10B - $10,700 100% 100%

Both statements are true for this company: "open models handle 85 percent of our tokens" and "frontier models account for 84 percent of our inference spend." Only the second one tells the CFO where the money goes.

The Larridin analysis of the split inference market describes this as two markets: commodity inference, where prices drop toward hardware cost, and frontier inference, where providers still have pricing power. Your budget sits in both, and the ratio between them is the number to manage. Public price data shows how wide the spread is. Compare the per-token rates on Anthropic's pricing page with open-model prices tracked on the Artificial Analysis model leaderboard, and the difference between the cheapest capable open model and a frontier reasoning model is often more than an order of magnitude.

A strategy built on the volume number tends to over-invest in migrating bulk work that was already cheap. Meanwhile the expensive frontier calls, which are often running in stages that don't need them, go untouched.

Why the Headline Number Needs a Methodology Check

Claims about open-model share are useful signals, but each one rests on choices you should understand before you quote it. Ask these questions of any token-share statistic, including your own:

  • Who is in the sample? A vendor that distributes local inference software sees companies that already run local models. Gateway rankings such as OpenRouter's public rankings show traffic through OpenRouter, which skews toward developers comparing models. Neither one is a census of enterprise usage.
  • Which tokens count? Input, cached input, output, and reasoning tokens have very different prices. A 90 percent share of input tokens can coexist with a much smaller share of output and reasoning tokens, which are usually the expensive ones.
  • Is local inference measured at all? Many companies don't meter self-hosted models per call. If open-model tokens are estimated while API tokens come from invoices, the comparison is skewed from the start.
  • Are embeddings and classifiers included? Embedding calls produce huge token counts at tiny prices. Including them pushes open-model share up without changing anything about spend.
  • What is the time window? A single batch backfill (re-embedding a document store, for example) can dominate a month of volume.

None of this means the 80-90 percent figure is wrong. It means the figure answers a question your board didn't ask. The board's question is where the money goes and whether it buys anything, and only your own ledger can answer that.

Build the Token Ledger Before the Strategy Debate

A token ledger is one table with one row per model call (or per aggregated call group for very high-volume stages). It joins usage to the business process that triggered it. Most teams already have the raw data scattered across provider invoices, gateway logs, and trace stores. The work is attaching workflow context to every call.

Where the data comes from

  • Gateways. If traffic goes through a proxy, use its native accounting. LiteLLM's spend tracking records cost per key, team, and custom tag. OpenRouter's usage accounting returns token counts and cost in the response, so you can log them inline.
  • Traces. If you instrument with OpenTelemetry, the GenAI semantic conventions define standard attributes for model name and input and output token counts. Add your own attributes for workflow and stage.
  • Self-hosted models. Your inference server reports tokens served. Cost has to be allocated: monthly GPU, serving, and on-call cost divided by tokens served. Store utilization next to the allocation, because a cluster running at 20 percent utilization has an effective per-token price five times higher than the same cluster running full. Our GPU-as-a-service economics post covers the rent-versus-own inputs.

A minimal schema

CREATE TABLE token_ledger (
  call_id            TEXT PRIMARY KEY,
  ts                 TIMESTAMPTZ NOT NULL,
  workflow           TEXT NOT NULL,      -- e.g. 'invoice_triage'
  stage              TEXT NOT NULL,      -- e.g. 'extract_line_items'
  bucket             TEXT,               -- 'bulk' | 'judgment' | 'not_llm' (set during review)
  model              TEXT NOT NULL,
  model_class        TEXT NOT NULL,      -- 'open_hosted' | 'open_api' | 'frontier'
  input_tokens       INTEGER NOT NULL,
  cached_input_tokens INTEGER DEFAULT 0,
  output_tokens      INTEGER NOT NULL,
  reasoning_tokens   INTEGER DEFAULT 0,
  cost_usd           NUMERIC(12,6) NOT NULL,
  cost_method        TEXT NOT NULL,      -- 'invoice' | 'list_price' | 'allocated'
  retry_of           TEXT,               -- call_id of the failed attempt, if any
  outcome            TEXT                -- 'accepted' | 'rejected' | 'escalated' | NULL
);

Three columns do most of the work. stage lets you see cost per step instead of per workflow, which is where waste shows up. cost_method records how confident you can be in each number, so nobody presents allocated self-hosting costs as invoice-grade data. retry_of exposes the spend that goes to failed attempts, which is often large in agent loops. The Zylos analysis of agent compute markets explains why agent workloads multiply call counts in ways chat workloads never did.

Add an outcome column if you can. Cost per call doesn't tell you much. Cost per accepted output is the figure that holds up in a budget meeting.

Three Buckets: Bulk Work, Judgment Work, and Work That Should Not Be an LLM Call

Once the ledger has a month of data, classify every workflow stage. Do this for stages, not whole workflows. A single invoice workflow can contain all three kinds of work.

Bucket What it looks like Default model choice Signal it is misplaced
Bulk Summarization, classification into known labels, extraction from messy text, first-draft generation, embedding Open-weight or small API models Frontier model in a stage where eval scores don't move when you swap in a smaller model
Judgment Multi-step planning, ambiguous policy decisions, code changes, synthesis across conflicting sources, anything where an error is costly Frontier models, sometimes with higher reasoning effort Open model in a stage with a high rejection or escalation rate
Not an LLM call Formatting, date parsing, lookups on known keys, routing on structured fields, schema validation, arithmetic Code, rules, or a database query Any LLM spend at all

The third bucket is usually the cheapest to fix and the easiest to miss, because those calls are each small and nobody notices them. A model that reformats JSON on every pass through an agent loop adds up across millions of runs. Our guide on keeping deterministic steps deterministic goes through how to spot these stages.

The judgment bucket deserves protection, not just cost pressure. If the ledger shows frontier spend concentrated in stages with high acceptance rates and expensive failure modes, that spend is doing its job. Cutting it to improve a share-of-tokens metric is how teams end up with cheaper agents that produce more escalations. MightyBot's breakdown of why enterprise AI projects blow budgets points to unplanned retries and rework, rather than list price, as the usual source of overruns. Downgrading a judgment stage to a weaker model produces exactly that kind of rework.

Moving work between bulk and judgment is a routing problem, and you should test it, not assume it. LLM routing in production covers routing patterns, and Agent Tokenomics shows how to use evals to check that a cheaper model holds quality before you switch.

Classification also tends to show that a large share of frontier spend sits in a few stages. In the illustrative month above, if two stages account for most of the $9,000 frontier line, the strategy discussion becomes two specific engineering questions instead of a debate about open versus closed models.

The Monthly Token Review

A ledger nobody reads goes stale. Set up a recurring 60-minute meeting with an engineering lead, the workflow owners, and someone from finance. Keep the agenda fixed so it doesn't turn into a debate about model news.

1. Two headline numbers, side by side (5 minutes). Open-model share of tokens and open-model share of cost, both compared with last month. If the two lines move in opposite directions, find out why.

2. Top ten stages by cost (20 minutes). This is where most of the value is. For each stage, review cost, call count, retry share, and outcome rate.

SELECT workflow, stage, model_class,
       COUNT(*)                                        AS calls,
       SUM(cost_usd)                                   AS cost,
       SUM(cost_usd) FILTER (WHERE retry_of IS NOT NULL) AS retry_cost,
       SUM(cost_usd) / NULLIF(COUNT(*) FILTER (WHERE outcome = 'accepted'), 0)
                                                       AS cost_per_accepted
FROM token_ledger
WHERE ts >= date_trunc('month', now()) - interval '1 month'
  AND ts <  date_trunc('month', now())
GROUP BY workflow, stage, model_class
ORDER BY cost DESC
LIMIT 10;

3. Bucket audit (15 minutes). Pick two or three stages and check whether their bucket assignment still holds. New model releases change the line between bulk and judgment, so a stage that needed a frontier model six months ago may not need one now. Our thinking budgets post covers the same question for reasoning effort.

4. Not-an-LLM candidates (10 minutes). List any stage where the output is fully determined by its input. Each one becomes a ticket to replace the call with code.

5. Decisions and owners (10 minutes). At most three changes per month, each with an owner, an eval plan, and an expected cost delta. Confirm the delta in next month's ledger.

The UD token cost framework and the Enterprise AI Tech FinOps framework both argue for attributing token cost to business units. The review is what turns that attribution into decisions. Without the stage-level agenda, these meetings drift into comparing per-token price charts, which is the one variable nobody in the room controls.

If your agents use a lot of tools, also look at input tokens per call. Tool output stuffed into context is a common hidden driver of input cost. Budgeting tool output and our cost engineering guide cover the fixes.

Getting the First Month of Ledger Data

Who should do what depends on where the data already lives:

  • You already run a gateway. Add workflow and stage tags to every request and export to a warehouse table nightly. This takes days, not weeks, and you don't need outside help.
  • Your traces are already in an observability platform. Use the platform's export if it carries token counts and custom attributes. Buying a separate cost tool usually adds little until you have tagged stages.
  • Calls are spread across teams, SDKs, and self-hosted clusters with no shared tagging. This is where the instrumentation work becomes a project, because someone has to define the stage taxonomy, set up allocation for self-hosted models, and get every team to adopt it.

OpenNash takes on the third case. We map workflows into stages, instrument the gateway and traces, build the ledger and the review queries in your warehouse, and hand over the whole system with documentation so your team owns it from then on. Any model move that comes out of the first review ships with an eval gate, so a cost cut can't quietly lower quality. If you want to scope this against your own stack, book a call with OpenNash and bring last month's invoices from every model provider.

Whether or not you bring in help, the first step is the same: pull last month's invoices from every provider and your self-hosted cluster costs, then list the five workflows you believe drive the most spend. Tag those five workflows' calls by stage this week, and hold the first review once 30 days of data are in.