A typical agent run in 2026 touches three or four external providers before it returns anything: a model API, a retrieval index, a couple of business tools behind MCP, and sometimes a second model for routing or judging. Most eval suites assume all of them return clean, fast, correctly shaped data every time. Production doesn't. According to Zylos Research's survey of chaos engineering for agent systems, ReliabilityBench mapped agent reliability across three dimensions and found that rate limiting was the single most damaging fault type. A boring HTTP 429 did more harm than anything clever.
That result should change how teams test agents. Your eval suite measures whether the agent is smart. A fault injection suite measures whether the system survives contact with its dependencies. These are different questions, and a team that answers only the first one is shipping a system whose failure behavior nobody has observed.
Evals Measure Intelligence, Fault Injection Measures Survivability
The standard eval loop, the one Hamel Husain's evals FAQ describes well, starts with error analysis on real traces, builds custom evaluators for your failure modes, and runs them on a cadence. That discipline is correct, and we have written about it in production evals for agentic systems. But eval datasets are built from traces where the dependencies mostly worked. The failure modes they capture are reasoning failures: wrong tool picked, wrong field extracted, wrong policy applied.
Dependency failures produce a different class of incident:
- The model provider returns 429 on step 4 of a 7-step plan, and the agent restarts from step 1, doubling cost.
- A CRM tool returns the first 30% of a record because a pagination limit kicked in, and the agent treats the partial record as complete.
- A tool's upstream API adds a wrapper object, so
customer.tieris nowdata.customer.tier. The agent readsundefinedand confidently reports the customer has no tier. - Retrieval returns zero documents because an index rebuild is in progress, and the agent answers from parametric memory without saying so.
- An OAuth token expires 40 minutes into a long-running task, and every later write fails silently.
None of these show up in a happy-path eval. All of them show up in postmortems. Our breakdown of why AI agents fail in production covers the broader taxonomy. Fault injection is how you find these failures on your own schedule instead of your customer's.
The Principles of Chaos Engineering still give the right frame. Define a steady state, hypothesize that it holds under realistic disruption, vary real-world events, and keep the blast radius small. For agents, "steady state" needs a behavioral definition (more on that below), and "real-world events" means the specific ways model APIs, tools, and retrieval break.
The Agent Fault Injection Matrix
Start with six fault classes. They cover most dependency incidents we see discussed in public agent postmortems, and each one maps to a concrete injection technique.
| Fault | Layer | How to inject | Acceptable behavior | Unacceptable behavior |
|---|---|---|---|---|
| Model 429 / 503 | Model API | Raise provider error with retry-after |
Backoff honoring retry-after, resume from checkpoint |
Retry storm, plan restart, abandoned task with no signal |
| Truncated tool output | Tool | Cut response to a fraction of its length | Detect incompleteness, re-fetch or flag partial | Treat partial data as complete |
| Wrong-schema tool response | Tool | Rename, nest, or drop fields | Validation error, retry or escalate | Read missing field as null and proceed |
| Empty or poisoned retrieval | Retrieval | Return [], or insert stale/contradictory docs |
Say "no source found," cite conflict, escalate | Answer from model memory as if grounded |
| Latency past budget | Any | Sleep beyond the step's timeout | Timeout, degrade, or hand off with state | Hang, or blow the run's overall time budget |
| Credential expiry mid-task | Tool / Model | Return 401 after N successful calls | Stop writes, surface auth failure, preserve state | Keep "succeeding" with no-op writes |
A few notes on getting these realistic.
Rate limits should match real provider behavior. Providers document their limit structure and headers; Anthropic's rate limits documentation, for example, describes per-minute request and token limits and the retry-after header returned with a 429. Your injected 429s should carry the same headers, because the point of the test is whether your client honors them. A bare exception with no header tests a code path that never runs in production.
Schema drift is the cheapest fault to inject and the one most often missed. If you already enforce output contracts with schema validation on tool responses, this test confirms the validator fires and that the agent does something sensible with the error. If you don't, this test will show you why you should.
Poisoned retrieval differs from empty retrieval. Empty retrieval tests whether the agent admits ignorance. Poisoned retrieval (a stale pricing doc, two policies that contradict each other) tests whether it notices conflict. Our post on RAG poisoning covers the adversarial version; for chaos testing, plain staleness is enough to start.
The MCP tool-call plane deserves its own experiments. This cloudandsre.com walkthrough on chaos engineering for MCP argues for breaking the tool-call plane directly, including server disconnects and slow tool listings. If your agent discovers tools at runtime, a tool that disappears mid-session is a realistic fault.
Where to Inject: A Wrapper at Every Dependency Boundary
Infrastructure chaos tools like Gremlin, Chaos Mesh, and AWS Fault Injection Simulator work well at the process, network, and infrastructure layers. They can add latency to a pod or drop packets to an endpoint. They can't easily return a CRM record with a renamed field or slip a stale document into the top-k results. Those faults live at the application layer, and as Flasqo's guide to API chaos testing points out, API-level injection is where you control the exact shape of the failure.
The simplest approach is a seeded wrapper around every dependency call. Here's a minimal version:
import random
import time
from dataclasses import dataclass
class ProviderError(Exception):
def __init__(self, status, retry_after=None):
self.status = status
self.retry_after = retry_after
@dataclass
class Fault:
target: str # "model", "tool:crm_lookup", "retrieval"
kind: str # "429", "503", "truncate", "schema_drift", "empty", "latency", "auth_expiry"
rate: float # probability per call
class FaultInjector:
def __init__(self, faults, seed=None, auth_expiry_after=20):
self.faults = faults
self.rng = random.Random(seed)
self.calls = 0
self.auth_expiry_after = auth_expiry_after
self.injected = [] # record every fault for grading
def wrap(self, target, call):
def wrapped(*args, **kwargs):
self.calls += 1
for f in self.faults:
if f.target == target and self.rng.random() < f.rate:
self.injected.append((self.calls, target, f.kind))
return self._apply(f, call, args, kwargs)
return call(*args, **kwargs)
return wrapped
def _apply(self, f, call, args, kwargs):
if f.kind in ("429", "503"):
raise ProviderError(status=int(f.kind), retry_after=2)
if f.kind == "auth_expiry" and self.calls > self.auth_expiry_after:
raise ProviderError(status=401)
if f.kind == "latency":
time.sleep(30)
result = call(*args, **kwargs)
if f.kind == "truncate" and isinstance(result, str):
return result[: len(result) // 3]
if f.kind == "schema_drift" and isinstance(result, dict):
return {"data": result} # extra nesting the agent did not expect
if f.kind == "empty":
return []
return result
Three design choices matter more than the code itself.
- Seed everything. A fault you can't reproduce is an anecdote. With a fixed seed, a failing run becomes a regression test, which also makes it a candidate for your replay tooling.
- Record what you injected. The
injectedlist is the ground truth for grading. Without it, you can't tell whether the agent handled a fault well or simply never hit one. - Wrap at the boundary, not inside the agent. The agent loop should not know it's under test. If your harness has a single place where model, tool, and retrieval calls pass through, injection is a configuration change. If it doesn't, building that choke point is worth doing for observability alone.
Start by running your existing eval set through the wrapper at a 5-10% per-call fault rate (an illustrative starting point; tune it to your run length). You don't need new test cases on day one. The same tasks under degraded dependencies will produce a very different trace distribution.
Grade Behavior, Not Output
This is where most teams' first chaos experiment goes wrong. They inject a fault, run their usual output evaluator, see a high score, and conclude the agent is resilient. But a high-quality answer produced after a retrieval failure is often the worst outcome, because it means the agent made something up and sounded confident doing it.
Grade each faulted run on what the agent did in response to the fault, using the trace:
| Grade | What the trace shows | Verdict |
|---|---|---|
| Recovered | Retried with backoff, honored retry-after, resumed without repeating completed side effects |
Pass |
| Degraded | Completed a reduced task and flagged what was missing to the user or caller | Pass |
| Escalated | Stopped, preserved state, routed to a human or a fallback queue with context | Pass |
| Failed loud | Returned a clear error with no partial side effects | Pass, with a note |
| Thrashed | Retry storm, plan restart, or token spend far above baseline | Fail |
| Fabricated | Produced a confident answer that ignored or hid the fault | Hard fail |
The steady-state hypothesis for an agent then reads something like: "Under a 10% dependency fault rate, zero runs fabricate, under 2% thrash, and median cost stays within 1.5x of baseline." Those thresholds are illustrative; set yours from your own baseline traces. What matters is that the hypothesis concerns behavior and cost, not answer quality.
Detecting fabrication automatically is easier than it sounds when you know what you injected. If the injector emptied retrieval on a run, any answer that cites a source or states a fact from the knowledge base is suspect. If a tool returned a truncated record, any claim about fields past the cut is invented. An LLM judge can do this comparison, but a deterministic check against the injection log catches the bulk of cases at no model cost.
Instrumentation makes this work at scale. The OpenTelemetry GenAI semantic conventions standardize span attributes for model calls, such as the operation name, requested model, and token usage, alongside the general error.type attribute. Add your own attribute for the injected fault (we use something like chaos.fault.kind) to the same spans, and your grading becomes a query over traces rather than a manual review. Our guide to agent traces and tool calls covers which fields to capture so this query is possible.
From Staging to Production Without Hurting Anyone
Staging is where you should start, and staging is also where you'll get false confidence. Staging rarely has real provider rate limits, since your test key sees a fraction of production traffic. It rarely has realistic data staleness, because someone reseeded the index last week. And it never has the latency profile of a provider during a regional incident. The research behind this post is consistent on one point: chaos experiments that only run in staging tend to validate the staging environment rather than the system.
Move to production in stages:
- Shadow traffic first. Replay a sample of real requests against a faulted copy of the agent with all write tools stubbed. You get production-shaped inputs with no customer exposure.
- Internal users next. Enable low-rate injection for employee-facing runs only. Read-only faults (latency, empty retrieval, truncation) come before anything touching writes.
- Small production slice with a kill switch. A fixed small percentage of runs, read-only fault classes only, with automatic abort if the fabrication or thrash rate crosses your threshold. The agent kill switch patterns apply directly here; a chaos experiment is a planned scope escape, and it needs the same containment.
Two production rules are non-negotiable. Never inject faults into runs that perform irreversible writes (payments, outbound customer email, record deletion) until you've proven idempotency and checkpointing in lower environments. And never run an experiment you can't stop in under a minute.
Long-running agents need one more test that short ones don't. Credential expiry and provider outages mid-task only matter if the run is long enough to hit them. If your agents run for tens of minutes, test whether they resume from a checkpoint after a fault or start over. Durable execution with checkpointing turns a 503 at step 9 from a full rerun into a single retried step, and fault injection is how you verify that it does.
This suite sits alongside two others. Agent simulation testing stresses the agent from the user side with difficult, ambiguous, or adversarial conversations. Fault injection stresses it from the dependency side. Combining them is where the expensive bugs live: a frustrated simulated user plus a 429 storm is a realistic Monday morning.
A Two-Week Plan for a First Fault Suite
For teams that have none of this yet, here is the order that pays off fastest:
- Days 1-3: Find or build the single choke point where model, tool, and retrieval calls pass through. Add the seeded wrapper with all fault rates at zero.
- Days 4-5: Turn on model 429 and 503 injection alone. Fix backoff and
retry-afterhandling first, since this is the fault class ReliabilityBench found most damaging. - Week 2: Add schema drift and empty retrieval. Wire the injection log into your trace store and write the deterministic fabrication check.
- End of week 2: Write the steady-state hypothesis, run the full eval set at a low fault rate in CI, and fail the build on any fabricated run.
Teams with an SRE practice for agents already have most of the pieces: error budgets, runbooks, on-call. The fault suite gives those budgets a way to be tested before the pager goes off.
OpenNash builds this kind of harness as part of production agent work: mapping each dependency boundary during the audit, defining behavioral pass criteria and escalation paths in design, and handing the fault suite over with the codebase so your team owns it after deployment. If your agent depends on three or more external providers and you have never watched it run while one of them fails, book a call with OpenNash and bring one workflow. We'll write its fault matrix together and pick the first three experiments to run.