Anthropic's prompt caching documentation prices a cache read at 0.1x the base input token rate. That one multiplier explains why teams in 2026 are deleting their vector databases. If re-reading a 150,000-token corpus costs a tenth of what it used to, why keep chunking, embedding, reranking, and debugging retrieval misses? Because a tenth is still more than zero, and the model doesn't read the middle of a long prompt as well as it reads the edges. Whether to stuff the context or retrieve can be calculated. Very few teams calculate it before they rip out the pipeline.
This post sets out that calculation: four variables, one inequality, and the measurements to collect before you change architectures.
Why "stuff the context or retrieve" became a real question
Two things changed. Context windows grew to 200K and then 1M tokens, so whole document sets fit in one prompt. Cached prefixes also brought down both the cost and the prefill latency of re-sending the same tokens. The Spheron guide to KV cache and prefix caching economics explains why. Once the attention state for a stable prefix is computed and stored, later requests skip most of the prefill work.
Neither change makes long context free. Comparisons from 2026, including Wire's roundup of the data, put the gap at high volume at under $10 per day to serve 100,000 queries with RAG, against roughly $10,000 per day to send full context every time. The same body of evidence finds that long context with caching can come out ahead for small, static knowledge bases under about 200,000 tokens. Both findings can be true together. Which one applies to you depends on corpus size, traffic shape, and how much a wrong answer costs.
Accuracy matters as much as price. Long context does better on analytical tasks that need several documents reasoned about together, which is what the Towards AI breakdown of when retrieval still wins describes. Retrieval does better on needle lookups in large, changing corpora. Most production workloads include both kinds of question, so a vendor benchmark won't settle it. Measuring your own traffic will.
The four variables that set the crossover
Here is the OpenNash position. The choice between long context and retrieval comes down to a threshold you can compute from four measurable inputs:
| Variable | What it measures | Where to get it |
|---|---|---|
| Corpus size (C) | Tokens you would stuff into the prompt | Tokenize the actual document set |
| Cache hit rate (h) | Share of queries that read a warm cached prefix | Provider usage logs, or simulate from request timestamps against the cache TTL |
| Retrieved context (k) | Tokens RAG sends per query, plus retrieval infra cost | Current pipeline traces |
| Accuracy by position | Answer correctness when the relevant fact sits at different depths | A position sweep on your corpus (described below) |
Reuse rate shows up in the formula through h. With a five-minute cache TTL, a support assistant answering a query every few seconds keeps the prefix warm all day. An internal research tool used in bursts has to rewrite the cache at 1.25x the base rate each time the cache goes cold, according to the same Anthropic pricing page. The one-hour TTL option costs 2x on the write.
The cost per query for each approach, with p as the base input price per token:
def long_context_cost(C, h, p, read_mult=0.10, write_mult=1.25):
# Every query pays for the full corpus, cached or not
return p * C * (h * read_mult + (1 - h) * write_mult)
def rag_cost(k, p, retrieval_infra_per_query):
return p * k + retrieval_infra_per_query
def expected_cost(token_cost, miss_rate, cost_per_wrong_answer):
return token_cost + miss_rate * cost_per_wrong_answer
If you set the token costs equal and ignore accuracy, you get the break-even corpus size:
C* = (k + r/p) / (0.1h + 1.25(1 - h))
Even with a perfect hit rate (h = 1), the denominator is 0.1. So on token cost alone, long context breaks even only when the corpus is under about 10x the size of what RAG retrieves per query. With 8,000 retrieved tokens and a 100% hit rate, that's 80,000 tokens. At a 90% hit rate it falls to about 37,000. Caching gives you a 10x discount and nothing more.
On token cost alone, then, retrieval wins for nearly any corpus worth the name. Long context starts to win once you include the cost of wrong answers.
Worked example: where the line moves
The numbers below are illustrative. They show how the inputs interact and are not benchmarks from any real deployment. Assumptions:
- Base input price: $3 per million tokens. Cached read: $0.30/M. Cache write: $3.75/M.
- RAG sends 8,000 tokens per query plus $0.002 of retrieval infrastructure, which comes to $0.026 per query.
- A wrong answer costs $5 in human rework, escalation, or a customer redo.
- Miss rates are assumed for illustration. In practice, you get them from your own position sweep and retrieval evals.
| Scenario | Corpus | Cache hit rate | Long-context token cost | RAG token cost | LC miss rate | RAG miss rate | LC expected cost | RAG expected cost |
|---|---|---|---|---|---|---|---|---|
| A: small, steady traffic | 50K | 95% | $0.024 | $0.026 | 2% | 5% | $0.124 | $0.276 |
| B: mid-size, steady traffic | 150K | 95% | $0.071 | $0.026 | 2% | 6% | $0.171 | $0.326 |
| C: mid-size, bursty traffic | 150K | 50% | $0.304 | $0.026 | 2% | 6% | $0.404 | $0.326 |
| D: large corpus | 1M | 95% | $0.473 | $0.026 | 8% | 6% | $0.873 | $0.326 |
The math for scenario B: a cached read of 150,000 tokens at $0.30/M is $0.045. A cache write is $0.5625. At a 95% hit rate, 0.95 × $0.045 + 0.05 × $0.5625 = $0.071. Add 2% × $5 = $0.10 and you get $0.171.
What the table shows:
- Scenario B is the main counter-intuitive result. Long context costs about 2.7x more per query in tokens and still wins on expected cost, because it misses fewer answers and misses are expensive. A team comparing token bills alone would pick RAG and lose money.
- Scenario C uses the same corpus with different traffic. When half the queries hit a cold cache, the token cost rises more than 4x, from $0.071 to $0.304, and RAG wins again. Your traffic pattern matters as much as your corpus.
- Scenario D is where long context stops winning. A million-token prompt costs more than RAG at every hit rate, and accuracy at that length often gets worse (more on that below). Error cost can't rescue it.
- When the cost of a wrong answer approaches zero, as with a casual internal FAQ, RAG wins every row on cost. Accuracy only moves the line when mistakes cost something.
Watch the pricing tier too. Some providers charge a higher per-token rate once input passes a length threshold. Check your model's price sheet before you put a 400K-token prompt into the formula at the standard rate.
Accuracy at position: the variable everyone guesses at
Liu et al.'s "Lost in the Middle" showed that models use information placed at the start or end of a long context more reliably than information in the middle. Accuracy across positions forms a U-shaped curve. Later work on RULER went further and found that many models' effective context length, meaning the length at which they still perform well on harder tasks than single-needle lookup, falls well short of their advertised window.
For the crossover math, this means the long-context miss rate isn't a constant. It changes with corpus size and with where the relevant fact sits. A 150K corpus in a 1M window may perform well. The same model at 800K may not. Any decision that skips this measurement is a guess, however detailed the spreadsheet.
Run a position sweep before you commit:
- Pick 50-100 facts from your own corpus whose answers you can check automatically, such as a contract clause, a policy threshold, or a SKU attribute.
- Place each fact's source document at ten depths in the stuffed prompt, from 0% to 100%.
- Ask the same questions at each depth and score correctness. Include multi-hop questions that need two documents, because those are where long context should beat retrieval.
- Run the same question set through your current RAG pipeline and score retrieval and generation separately. Our RAG evaluation guide covers how to split those failures apart.
- Repeat at 1x, 2x, and 4x your current corpus size so you can see where accuracy starts to drop off.
The result is a curve rather than a single number. Use the worst-case depth, not the average, as the long-context miss rate in the formula, because your important facts won't all sit at the start of the prompt.
Prompt order also helps on the cost side. Put the most stable, most frequently referenced documents at the start of the prefix so they stay cached. Put volatile material and the user's question at the end. That keeps the cache valid when content changes and puts the question where models attend best.
When retrieval stays mandatory regardless of the math
Some constraints remove long context from consideration before any cost calculation:
- Per-user permissions. If two users may see different subsets of the corpus, they can't share one cached prefix. You end up caching a separate prefix per permission set, which drives the hit rate down, or you filter at retrieval time. Retrieval with access filtering is the safer default. Permission-aware agents covers the design.
- High churn. Any edit to the cached prefix invalidates the cache from that point on. A corpus that changes hourly will keep paying the 1.25x write rate.
- Corpus growth. A knowledge base at 120K tokens today may reach 600K within a year. Run the formula on next year's corpus size as well as today's.
- Agent loops. In a multi-step agent, the stuffed corpus shares the window with tool outputs and conversation history. A 150K corpus plus unbudgeted tool output can push the critical facts into the weak middle of the context. Long sessions also need compaction, and compaction can quietly drop parts of a corpus you thought stayed in context.
The Zylos research notes on context window management recommend treating the window as a budget rather than a container. That matches what the numbers above show.
The hybrid most teams should build
For most production workloads, the answer is retrieve-then-stuff. Retrieval narrows the corpus to a coarse working set: one customer's account history, one product's documentation, or one deal's contract bundle. That working set becomes a cached prefix, and the model reasons over all of it.
This setup keeps the useful property of each approach. Retrieval handles scale, permissions, and churn. Long context handles cross-document reasoning inside a working set small enough to keep accuracy high. It also improves the cache hit rate, because questions about the same account within a TTL window reuse the same prefix.
A practical way to build it:
- Shard by entity rather than by chunk. Retrieve whole documents or document bundles keyed to a customer, case, or product. Avoid 500-token fragments that cut a clause in half.
- Cap the working set at the depth where your position sweep shows accuracy holding up. Go below that cap, not above it.
- Let the agent fetch more when it needs to. An agentic retrieval tool can pull in a document outside the working set when the first pass falls short, so you don't have to stuff for the worst case.
- Log cache reads and writes per query. Hit rate is the input most likely to drift as traffic changes, and it moves cost more than any other variable.
Running the crossover on your own workload
OpenNash scopes this decision in an audit: tokenize the corpus, pull request timestamps to simulate cache hit rates under your TTL, run the position sweep on your documents, and put a dollar figure on a wrong answer with the team that handles the rework. The output is a per-workload recommendation. Some workloads stay on RAG, some move to cached long context, and most get the hybrid with a stated working-set cap. The build includes the eval harness, so the decision gets re-checked as the corpus grows instead of being made once and forgotten.
Some teams don't need outside help. If your corpus is under 50K tokens, traffic is steady, and wrong answers are cheap, stuff the context, turn on caching, and move on. If you have a large RAG pipeline and are under pressure to "just use the 1M window," book a call and bring three numbers: your corpus token count, your peak and off-peak query rate, and what a wrong answer costs you today. We'll work through the crossover formula with you before anyone deletes a vector index.