LangChain moved its coding agent from the bottom of the Terminal Bench 2.0 rankings into the top five without changing the model. It only changed the harness, according to Beam's write-up on harness engineering. The same piece cites Princeton's CORE-Bench, where one base model scored 42% under one scaffold and 78% under another. That's a 36-point swing from code that has nothing to do with model weights.
Those results explain why every agent team now ships harness changes weekly. They also point to a measurement problem. Most teams have no reliable way to tell whether a given harness change caused the improvement they're celebrating. The tool set changed on Tuesday, the provider pushed a model update on Thursday, someone added twelve tasks to the eval suite on Friday, and on Monday the dashboard is up four points. Nobody can say which of those changes produced the gain.
Our position is that harness changes need the same experimental discipline as model changes. Pin the model, fix sampling, run every task several times, compare both versions on the same tasks, and report a confidence interval before anyone claims a lift. The rest of this post covers the method and works through a full example.
Why Harness Wins Get Credited to the Wrong Change
The harness is everything around the model: the system prompt, the tool definitions, retry policy, context management, memory, verification steps, and the loop that decides when the agent is done. A recent preprint on harness engineering for language agents treats this layer as the place where control, agency, and runtime behavior get decided. That matches what anyone who has built one sees. In a production agent, most of the code is harness code, and most of the bugs live there too.
We made the strategic version of this argument in Open Source Just Caught Up to Opus 4.5: Why the Moat Is the Harness. The gap between frontier and open-weight models has narrowed on many coding and reasoning benchmarks, so the scaffold is where the remaining differentiation sits. If the harness is your moat, you need to know which changes widened it and which ones only looked good for a week.
Attribution gets muddled for three structural reasons:
- Model updates arrive without asking. Unless you call a dated snapshot, the model behind your API alias can change under you. A harness change that coincides with a silent model refresh will get credit for the model's gain.
- Agent runs are noisy. The same task with the same harness can pass on one run and fail on the next. We covered this in Agent Run Variance: Why pass@k Hides the Failures That Matter. A single run per task gives you a coin flip on many tasks.
- Eval suites grow. Teams add tasks as they find failures, which is good practice. It also means last month's score and this month's score were computed on different denominators.
Any one of these can produce a four-point swing with no real change in quality. When all three hit in the same week, the dashboard is close to meaningless.
The Experimental Contract: What to Pin Before You Run Anything
Before running either harness, write down what stays fixed. Most of these items are cheap to control and expensive to discover later.
| Variable | What to fix | Why it matters |
|---|---|---|
| Model | Dated snapshot ID, never a floating alias | Silent model updates are the most common confounder |
| Sampling | Temperature, top_p, max tokens, reasoning effort | Different sampling changes the pass-rate distribution |
| Seed | Fixed where the provider supports it | Cuts some variance, though APIs rarely guarantee determinism |
| Task set | Frozen list of task IDs, hashed | Prevents the changing-denominator problem |
| Environment | Container image, tool versions, fixture data | A changed fixture can fail tasks for reasons unrelated to the harness |
| Grader | Grader version and, for LLM judges, the judge model snapshot | Grader drift looks exactly like harness improvement |
| k (runs per task) | Same k for both arms, set before the run | Choosing k after seeing results is a quiet form of p-hacking |
| Decision rule | Threshold for shipping, written in advance | Stops people from rationalizing a noisy result |
Even with temperature at zero and a fixed seed, most hosted models aren't fully deterministic, and tool calls add their own variation through network timing, search results, and file system state. That's why k matters. You're measuring each task's pass rate under the harness, and one pass/fail outcome is a poor estimate of that rate.
Anthropic's post on raising the bar on SWE-bench Verified is a useful reference because it treats the scaffold as a first-class, documented artifact. It describes the prompt, the tools, and the loop alongside the score. If you can't describe your harness that precisely, you can't make a clean comparison between two versions of it. Version the harness the way you'd version a service: a commit hash, a config file, and a changelog entry for each experiment.
Worked Example: Same Model, Two Tool Sets, 50 Tasks, 5 Runs Each
Here's a scenario with illustrative numbers chosen so the math is easy to check. A team runs an internal data-analysis agent. Harness A uses a single general run_sql tool. Harness B splits it into list_tables, describe_table, and run_query, with tighter output truncation. The model, sampling, and environment are pinned. The suite has 50 frozen tasks, and each harness runs every task 5 times, for 250 runs per arm.
Raw results (illustrative):
| Harness A | Harness B | |
|---|---|---|
| Passing runs | 155 / 250 | 171 / 250 |
| Pass rate | 62.0% | 68.4% |
| Difference | +6.4 points |
The headline says B wins by 6.4 points. The per-task breakdown tells you more:
| Task behavior | Tasks | Net change in passing runs |
|---|---|---|
| No change in pass count | 31 | 0 |
| Improved under B (+2 runs each) | 12 | +24 |
| Regressed under B (-1 run each) | 6 | -6 |
| Regressed under B (-2 runs) | 1 | -2 |
| Total | 50 | +16 |
A gain of 16 runs out of 250 matches the 6.4-point difference. Nineteen tasks moved. Twelve got better and seven got worse. Read that regression list before shipping, because those seven tasks may share a cause, such as a query pattern the split tools handle badly.
To get the interval, compute each task's pass-rate difference (B minus A, in steps of 0.2 because k = 5), then bootstrap over tasks:
import numpy as np
def paired_bootstrap(a, b, n_boot=10_000, seed=0):
"""a, b: arrays of shape (n_tasks, k) with 0/1 outcomes."""
rng = np.random.default_rng(seed)
diffs = b.mean(axis=1) - a.mean(axis=1) # per-task difference
n = len(diffs)
boots = np.array([
diffs[rng.integers(0, n, n)].mean() for _ in range(n_boot)
])
lo, hi = np.percentile(boots, [2.5, 97.5])
return diffs.mean(), lo, hi
In this example the per-task differences have a standard deviation of about 0.21, so the standard error of the mean difference is about 0.21 / √50 ≈ 0.029. The paired bootstrap returns roughly +6.4 points, 95% CI [+0.8, +12.0]. The lower bound clears zero, but only just. It's a real but modest result, and the regressions are worth investigating before rollout.
Now run the same data unpaired, treating the two arms as independent samples of tasks. A suite that mixes easy and hard tasks will typically have per-task pass rates with a standard deviation around 0.4 (illustrative). The standard error of the difference becomes √(0.16/50 + 0.16/50) = 0.08, and the 95% interval widens to roughly [-9.3, +22.1]. The data hasn't changed, but the unpaired analysis can't tell whether B helped at all.
Pairing works because most of the variance in an agent eval comes from tasks differing in difficulty. When each task is compared against itself, that variance cancels. Evan Miller's Adding Error Bars to Evals argues for the same approach when comparing models: when two systems answer the same questions, analyze the paired differences, and resample answers to reduce within-question noise. The same reasoning applies when the thing you're changing is the scaffold.
METR's work on measuring how long a task AI agents can complete is also worth reading here. It runs multiple attempts per task and derives its uncertainty by bootstrapping over the task structure instead of treating every run as independent. If your tasks come in families (ten variants of the same SQL pattern, for example), bootstrap over families, not individual tasks. Otherwise the interval will be narrower than the evidence supports.
Traps That Fake a Lift
Once the paired method is in place, most bad claims come from a few recurring mistakes.
New harness, new tasks. This is the most common one. The team builds Harness B, adds 15 tasks that motivated the change, runs B on the expanded suite, and compares the score against A's number from last month on the old suite. The new tasks were often written with B's design in mind, so they flatter it. Re-run A on the full current suite, or restrict the comparison to task IDs both versions saw. If you can't re-run A because the old harness no longer builds, that's a reason to keep old harness versions runnable.
Simultaneous model change. If the model alias moved during the experiment, the result is confounded. Check the response metadata for the model version on every run and fail the experiment if the two arms used different snapshots.
Grader drift. If an LLM judge scores outputs and the judge model or rubric changed between runs, you measured the grader. Pin the judge snapshot and validate it against human labels on a sample, as Hamel Husain's evals FAQ recommends for any model-based evaluator.
Picking k or the metric after the fact. Running k = 3, looking at the result, then running two more because it was close is optional stopping, and it inflates false positives. The same goes for switching from pass@1 to pass@3 because the second looks better. Set k and the metric in advance.
Ignoring cost and latency. A harness that adds a verification loop might gain 5 points while doubling tokens per task. Report cost per task and p95 latency for both arms with the same paired method. A small lift that doubles spend can be a bad trade.
Leaderboard thinking. Small gaps on public leaderboards often fall inside the error bars, a point we made in Leaderboard Gaps Smaller Than Error Bars. Your internal suite works the same way. A one-point gain on 50 tasks with k = 1 is noise until proven otherwise.
When the Model and Harness Change Together: Run a 2x2
Sometimes you can't separate the changes. A new model ships, and the harness has to change to work with it, perhaps because the new model handles parallel tool calls differently. Teams usually report one number for "new model plus new harness" and credit whichever change they're more excited about.
A 2x2 design costs four arms of runs and answers the question directly:
| Old harness | New harness | |
|---|---|---|
| Old model | Baseline | Harness effect alone |
| New model | Model effect alone | Combined |
From these four cells, you can read:
- Harness effect: new harness minus old harness, averaged across both models.
- Model effect: new model minus old model, averaged across both harnesses.
- Interaction: whether the harness helps more on one model than the other.
The interaction term is where surprises show up. A context-compaction policy might help a smaller model a lot and do nothing for a larger one with a bigger context window. A harness built around a specific model's quirks may hurt when you swap models. That matters for anyone routing across providers. If your harness gains depend on one model, they disappear when you switch.
Run the same paired bootstrap within each comparison, using the same tasks, the same k, and per-task differences. With 50 tasks and k = 5, the four arms total 1,000 runs. On a typical agent eval, that costs less than one engineer-day, and far less than shipping a harness rewrite that only looked good because the model improved in the same release.
Turning This Into a Weekly Harness Release Gate
The method only pays off if it runs on every harness change automatically. Teams that already treat their harness as the product, as described in AI Coding Harnesses: Stop Writing Code, Start Designing the System Around It, can add it to CI with a few components:
- A frozen eval manifest. Task IDs, a content hash per task, the environment image, and the grader version. Changing the manifest bumps its version, and comparisons across manifest versions are blocked.
- A pinned model config. Snapshot ID and sampling parameters in a file checked into the repo. The runner rejects floating aliases.
- A paired runner. Runs the baseline harness (the current main branch) and the candidate harness on the same manifest with the same k. Store 0/1 outcomes per run, plus cost and latency.
- A pre-written decision rule. For example, ship if the 95% paired-bootstrap lower bound is above zero, no task family regresses by more than 10 points, and median cost per task rises less than 20%. Pick thresholds that fit your business and write them down before running anything.
- A regression report. The list of tasks that got worse, with links to traces. This is often more useful than the headline number, because it shows what the change broke.
Business stakeholders get a defensible answer to "did the agent get better, and why?" Engineers get a gate that stops a lucky week from turning into a shipped regression.
Building the Harness Experiment Pipeline With OpenNash
Most teams already have the parts of this setup: an eval suite, some traces, and a CI pipeline. What they usually lack is the pairing, the pinning, and a decision rule that people can't argue around after the fact. OpenNash builds that layer as part of production agent work. We start with an audit of how your team currently attributes agent improvements, then design the manifest, model pinning, and release gate around your workflows. The build ships into your CI, and your team owns the runner, the dashboards, and the thresholds after handoff.
Not every team needs outside help. If you have one agent and a small suite, the bootstrap function above and a frozen task list will get you most of the value this week. The pipeline becomes worth formalizing once several agents share a harness, or once harness and model changes start landing in the same release.
To start, take your last harness change that "improved the score." Re-run the old and new versions on the identical task set with a pinned model and k = 5, then compute the paired interval. If the lower bound doesn't clear zero, you've found a lift you credited without evidence. If you'd like help turning that check into a standing release gate, book a call and bring the eval suite you have now.