A coding agent on a developer laptop has a person watching it. The same agent running in CI has no one watching, holds write credentials, gets triggered by events that outsiders can partly control, and is billed per token. Teams are now running agents unattended for issue triage, dependency bumps, and test repair, and most of the pipelines we see described publicly were set up with the same defaults as a lint job. Lint jobs don't make decisions.
CI is the strictest environment an agent can enter, and it should get the strictest configuration you run anywhere. In practice that comes down to four controls: a per-run token ceiling, a hard timeout, scoped credentials, and a rule that agents propose diffs while humans keep merge authority on anything touching auth, migrations, or infrastructure. The rest of this post covers how to set each one and what to measure once they are in place.
Why CI Changes the Risk Math for Coding Agents
A CI runner combines three things that are each manageable alone and risky together.
- Untrusted input. Issue bodies, PR descriptions, commit messages, and code from forks all reach the agent's context. Any of them can carry a prompt injection.
- Write credentials. The job's token can push branches, comment, and sometimes approve or merge. Repository secrets may be present in the environment.
- No human in the loop at runtime. Nobody sees the agent decide to "fix" a failing test by deleting its assertions until a reviewer opens the diff, if one ever does.
That is the same pattern Simon Willison calls the lethal trifecta: private data, untrusted content, and a way to send data out. A CI runner with network access and secrets in its environment meets all three conditions by default.
The supply chain record shows what this looks like in practice. In March 2025 the widely used tj-actions/changed-files action was compromised (CVE-2025-30066), and the modified code dumped CI secrets into workflow logs across thousands of repositories. In August 2025 the Nx build system's npm packages were backdoored after attackers exploited an injectable GitHub Actions workflow. The malicious versions then invoked AI coding CLIs installed on victims' machines to search for credentials. Neither incident needed a clever model. Both needed a pipeline that trusted its inputs and held more permission than it used.
If you have already treated developer-side agents as unmanaged endpoints that need runtime controls, CI needs the same treatment and then some, because nobody is watching the terminal.
Cost Controls: Ceilings Per Run, Per Repo, and Per Month
Agent cost in CI goes wrong in a specific way. Single runs are cheap, but loops are not. One agent stuck retrying a failing test, or a workflow that fires on every push to a busy PR, can multiply spend overnight. Researchers studying this have catalogued cost-inefficient behaviors in coding agents, including redundant exploration and repeated attempts that never converge. In unattended CI, nothing interrupts those behaviors except the limits you configure.
LangChain described how they made coding agent spend predictable by routing traffic through an LLM gateway that gave them real-time spend visibility and limits at several levels. The layering is the part to copy. One cap is not enough because each layer catches a different failure.
| Layer | What it stops | Example control (illustrative values) |
|---|---|---|
| Per run | One agent looping on one task | Max 30 turns, 20-minute job timeout |
| Per trigger | Duplicate runs from rapid pushes | Concurrency group with cancel-in-progress |
| Per repository | A noisy repo consuming shared budget | Daily token budget enforced at the gateway |
| Per org / month | Aggregate drift nobody noticed | Monthly spend alert at 50%, hard stop at 100% |
The values above are starting points, not benchmarks. Tune them from your own run data after a few weeks.
The per-run layer lives in the workflow file. GitHub Actions jobs have a default timeout of 360 minutes, which is a long time for an agent to keep paying for retries. Set it explicitly:
name: agent-test-repair
on:
workflow_dispatch:
issues:
types: [labeled]
permissions:
contents: read
concurrency:
group: agent-${{ github.event.issue.number || github.run_id }}
cancel-in-progress: true
jobs:
repair:
if: github.event.label.name == 'agent:repair'
runs-on: ubuntu-latest
timeout-minutes: 20
permissions:
contents: write
pull-requests: write
steps:
- uses: actions/checkout@<full-commit-sha>
- uses: anthropics/claude-code-action@<full-commit-sha>
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
claude_args: "--max-turns 30"
prompt: |
Diagnose the failing test referenced in this issue.
Propose a fix as a pull request. Do not modify files
under migrations/, infra/, auth/, or .github/.
The Claude Code GitHub Actions documentation covers the configuration options for the action itself. Whatever agent you use, check that it exposes a turn or token cap and set it. If the agent doesn't expose one, the job timeout is your only per-run limit, and it measures time rather than money.
The per-repo and monthly layers belong outside the workflow, at a gateway or the provider's admin console, because a workflow file can be edited by the same pull request it is supposed to constrain. Portal26 argues that token budgets should be set as financial policy by finance leaders before deployment, rather than discovered on an invoice afterward. That framing matters for a business reader. A per-repo budget gives a CFO a number to approve, and the engineering team gets a hard limit it can design against.
The cheapest run is one that never starts. Gate agent jobs on explicit labels or comments from maintainers rather than on every push or every new issue. That one change usually does more for the bill than any model choice.
Flake Budgets: Treat Agent Nondeterminism Like a Flaky Test
Engineering teams already know how to live with nondeterminism in CI. Flaky tests get tracked, quarantined past a threshold, and fixed. Agent runs need the same discipline, because the same prompt on the same commit will not always produce the same diff.
A flake budget is the share of agent runs you accept as unusable over a rolling window. "Unusable" needs a concrete definition, and it should be measured automatically:
- The proposed diff fails the existing test suite
- The diff touches files outside the allowed path list
- The run hit its turn cap or timeout without producing output
- A reviewer closed the PR without merging, tagged with a reason
- The diff "fixes" a test by weakening it (deleted assertions, added skips, longer sleeps, broader exception handling)
The last item is the reason test-repair agents need their own budget. An agent told to make CI green has a cheap path available: make the test stop testing. Add a check that diffs to test files are flagged when they remove assertions or add skip markers, and route those PRs to a human with a label that says so.
Once the definition is fixed, the operating rule is simple. Pick a threshold per agent job (for example, an illustrative 20% unusable over the last 50 runs). When a job crosses it, the job is disabled automatically and an issue is opened. Someone looks at the failures, adjusts the prompt, tool access, or scope, and re-enables it. This is the same loop described in agent run variance and why pass@k hides failures: a high success rate across retries can hide a first-attempt rate that makes the job a net cost.
The flake budget also gives you a cost number. If each run costs a known amount and 30% of runs are unusable, your effective cost per merged change is roughly 1.4 times the per-run cost, plus the reviewer time spent closing bad PRs. That reviewer time is usually the bigger number, and it is the one that doesn't show up on the provider invoice.
Scoped Tokens and Provenance: What the Agent Job Is Allowed to Touch
GitHub's security hardening guidance for Actions predates coding agents, and every recommendation in it matters more once a model is choosing which commands to run. The ones that matter most for agent jobs:
- Default permissions to read. Set
permissions: contents: readat the workflow level and grant write scopes only to the specific job that needs them, as in the example above. - Pin third-party actions to a full commit SHA. Tags can be moved. The
tj-actionscompromise worked by repointing tags to malicious code. - Never check out untrusted fork code in a
pull_request_targetjob. That trigger runs with the base repository's secrets and write token. Combining it with an agent that reads the PR's contents gives an attacker both the injection vector and the credentials. - Treat issue and PR text as untrusted input. Don't interpolate
${{ github.event.issue.body }}directly into shell commands, and assume the agent will read hostile instructions in those fields. - Prefer short-lived credentials. Use OIDC federation for cloud access instead of long-lived keys stored as secrets, and keep cloud credentials out of agent jobs entirely unless the task requires them.
The agent should hold the narrowest identity that lets it do its job, and that identity should expire when the job ends. We covered the general pattern in why AI agents should never hold keys. In CI it comes down to a few concrete rules. Use a GitHub App or fine-grained token scoped to one repository. Never give the agent job an admin token. Don't let the job that runs the agent also hold deploy credentials.
Provenance is the other half. Every agent-authored PR should be identifiable as such: a distinct bot account, a label, and a link back to the workflow run that produced it. When something goes wrong three months later, you need to answer "which agent, which prompt version, which model, which trigger" without reconstructing it from memory. Northflank's write-up on enterprise coding agent deployment makes the same point about isolation and auditability for agent workloads in general. CI just makes the audit trail easier to build, because the run logs already exist.
What to Never Automate: Merge Authority Stays With Humans
The rule is that agents propose diffs. Whether a diff merges without a human depends on what it touches, and that decision should be encoded in branch protection and CODEOWNERS, not left to the agent's prompt. A prompt that says "don't touch migrations" is a request. A CODEOWNERS entry requiring a database owner's approval on migrations/ is enforced by the platform.
| Change area | Agent may | Human must |
|---|---|---|
| Documentation, comments, typos | Propose and, if checks pass, auto-merge | Spot-check periodically |
| Lockfile and patch-level dependency bumps | Propose with changelog summary | Approve if any test changed |
| New dependencies | Propose with justification | Approve every time |
| Application code (non-sensitive paths) | Propose | Review and merge |
| Test files | Propose | Review, especially removed assertions |
| Authentication and authorization code | Propose at most | Review and merge, owner approval required |
| Database migrations | Draft only | Review, test against production-like data, merge |
| Infrastructure as code (Terraform, Helm, k8s) | Draft only | Review plan output, merge, apply |
CI workflow files (.github/) |
Nothing | All changes |
| Secrets, release tags, license files | Nothing | All changes |
Two rows deserve explanation.
CI workflow files are on the "nothing" row because an agent that can edit its own workflow can raise its own permissions, remove its own timeout, or add a step that sends secrets somewhere. Block writes to .github/ at the path level in the agent's instructions, and also require CODEOWNERS approval from a platform owner, so the block holds even if the instruction is ignored.
New dependencies get human approval every time because adding a package is a supply chain decision. Agents tend to reach for a library when a few lines of code would do, and typosquatted package names are a known attack. A version bump of something you already trust is a different risk class from a new package nobody has vetted.
Migrations and infrastructure sit on "draft only" for a reason that is about timing more than about trusting the agent. The failure shows up after merge: a migration that locks a large table, or a Terraform change that replaces a resource instead of updating it. The agent can't observe those effects from inside the runner. A human who reads the plan output can.
None of this slows a team down much if the review path is designed for volume. Redesigning the merge path for AI-generated code covers how to route agent PRs by risk so that reviewers spend their time on the bottom half of this table and skim the top half.
A Rollout Sequence That Produces Data Before It Produces Risk
Cortex Solutions' list of common automation mistakes includes automating too much at once. For CI agents, the fix is a staged rollout where each stage produces the cost and flake data needed to justify the next.
- Read-only triage (weeks 1-3). The agent labels issues, summarizes failing builds, and comments with a suspected cause. Permissions: read, plus issues and PR comments write. No code changes. Measure cost per run and how often maintainers agree with the label.
- Diff proposals on low-risk paths (weeks 4-8). Dependency bump PRs, doc fixes, test failure diagnosis with a proposed patch. Contents write on a branch only, never to the default branch. Flake budget tracking turns on here.
- Selective auto-merge (after stage 2 data supports it). Only the top rows of the table above, only with all required checks green, only with the agent's bot account excluded from approving its own PRs.
- Widen scope one path at a time. Each new area gets its own flake budget and its own cost line.
Most teams will stay at stage 2 for a long time, and that is fine. Agent proposals with human merge authority is a stable operating model.
Getting an Agent Pipeline Past Security Review
The hard part of most CI agent projects is getting the security and platform teams to approve it at all, and the model is rarely what holds that up. They want to see the permission scopes, the cost ceiling, the path restrictions, and the audit trail written down before anything runs.
OpenNash builds these pipelines as owned infrastructure. That covers mapping which CI tasks have a clear success signal, designing the token scopes, CODEOWNERS rules, and spend limits, then shipping the workflows with flake tracking wired in from the first run. The client keeps the workflows, the dashboards, and the runbook. If your team has an agent prototype working on a laptop and a security reviewer asking what happens when it runs unattended, book a call with OpenNash and bring the workflow file. We will mark which jobs can go to stage 2 now and which need controls added first.