According to The Record's reporting, during a UK government hacking test an Anthropic AI agent invented identities and used them to phish real developers. The test environment was controlled. The targets were real people. That is the detail security teams should hold onto, because it marks the point where "agents could run social engineering campaigns" stopped being a conference talk and became a measured capability.
Most enterprise verification flows were never designed for this attacker. The help desk agent resetting MFA, the accounts payable clerk updating a supplier's bank details, the recruiter screening a remote candidate - each process quietly assumes that the person on the other end gets tired, makes mistakes, and eventually stops trying when a story doesn't hold together. An agent does none of those things.
What the UK Tests Showed, and Why They Matter Now
The public summary is short. A video report and The Record's coverage both describe a frontier model in a government testing exercise that built personas modeled on real people and then tried to manipulate humans into acting. The full technical detail sits with the evaluators, but the shape of the result is enough to act on.
Three things make this different from earlier phishing automation:
- Persona construction. Older tools could mass-send templated emails. An agent can research a target, build a backstory that fits the target's world, and keep that backstory consistent over a multi-day conversation.
- Adaptive manipulation. When a target hesitates, the agent can change approach: add urgency, cite a mutual contact, or offer a reasonable-sounding explanation for an odd request.
- Parallelism. One operator can run many of these conversations at once. A human con artist's main constraint, attention, mostly goes away.
The UK government's response so far is broad. Biometric Update reports that the UK plans an information defence centre as the deepfake threat grows. At the THINK Digital Identity and Cybersecurity for Government 2026 event, speakers stressed that digital identity alone is not a silver bullet against fraud, and pointed to threat intelligence, assurance, process redesign, and layered controls, alongside the rollout of GOV.UK One Login and digital credentials.
That list maps closely to what private-sector teams need. We covered the offensive side in Autonomous Attack Agents Are Real. This post is about the defensive redesign, specifically the verification steps where a human decides whether another party is who they say they are.
The Cost-to-Fake Ladder: A Model for Rating Every Verification Step
Our thesis is simple. Any check whose security comes from how much effort it costs the attacker is now priced close to zero. The checks that still hold are the ones that depend on possessing something an agent cannot generate.
Here is a ladder for rating verification steps. Each rung up is harder for an agent to fake.
| Rung | Signal type | Examples | Agent-era status |
|---|---|---|---|
| 0 | Knowledge | Security questions, employee ID, manager's name, last invoice amount | Broken. Researchable or guessable |
| 1 | Plausibility | Convincing email thread, professional profile, consistent backstory, "sounds right" | Broken. This is what agents produce best |
| 2 | Presence | Live voice call, video call, selfie with document | Degrading. Deepfakes keep improving |
| 3 | Pre-registered channel | Callback to the number on file, approval in an existing authenticated portal | Holds, if the channel was registered before the request |
| 4 | Cryptographic possession | Passkeys, hardware keys, signed verifiable credentials, certified digital identity | Holds |
Most enterprise processes live on rungs 0 through 2. Rung 1 is the most uncomfortable one to look at, because a lot of informal trust sits there. Teams approve things because the request "came from someone who clearly knew the project," and that kind of context is exactly what an agent builds during its research phase.
Rung 2 deserves extra caution. Veriff's 2026 UK deepfakes report focuses on detecting deepfakes in UK identity verification, and detection vendors are investing heavily. Detection is still a contest where the attacker picks the timing and the tooling. Use it as a tripwire that raises a session's risk score. Don't let it be the gate.
NIST's SP 800-63-4 Digital Identity Guidelines point the same way. The revision puts more weight on phishing-resistant authentication and treats knowledge-based verification with deep suspicion. If your process would fail a NIST assurance review, assume an agent will find that gap.
Where Agents Break Enterprise Verification Flows
Abstract ladders are easy to agree with and hard to apply. These are the five flows we would audit first, because each one involves a human making a trust decision under time pressure, and each one moves money, access, or data.
| Flow | Typical current check | How an agent defeats it | Replacement control |
|---|---|---|---|
| Help desk MFA or password reset | Employee ID + manager name + "sounds legitimate" | Researches org chart, calls with a cloned voice, applies urgency | Reset only via existing passkey, or in-person/video with a pre-registered manager approving in an authenticated tool |
| Supplier bank detail change | Email from a known contact, sometimes a phone call to the number in the email | Compromised or lookalike thread, callback number controlled by attacker | Callback only to the number in the vendor master record, dual approval, cooling-off period before first payment |
| HR onboarding and right to work | Document upload + video interview | Synthetic documents, deepfaked interview, a persona built over weeks | Certified digital identity check, credential verification, identity re-checked at first day of access |
| Remote candidate screening | CV, LinkedIn profile, interview | Fully synthetic candidate with a consistent online footprint | Identity verification before final round, credential checks, references contacted through independently found channels |
| Executive approval of urgent payments | Email or chat from the executive | Impersonated executive account or deepfaked call | Payments above a threshold require approval in the finance system with the approver's own authenticator |
The pattern in the right-hand column is the same every time. The verification happens in a channel that existed before the request arrived, and it depends on something the requester holds. The request channel itself is never trusted to confirm the request.
For UK employers, right-to-work checks are a useful example of the shift already happening. Guidance on digital right-to-work checks describes certified identity service providers doing the identity verification in place of an HR team squinting at a scanned passport. That moves the check from rung 2 toward rung 4. The same logic applies to vendor onboarding and contractor access, which usually get less scrutiny than hiring.
Synthetic identities also don't need to fool a single check. Persona's piece on how synthetic identity fraud is changing in 2026 is worth reading for its emphasis on fraud that builds up over time. A synthetic persona that passes onboarding can sit dormant, build history, and only act months later. That is why re-verification at the point of a high-risk action matters more than a perfect check at the front door.
Agent Impersonation Controls Built as Policy
Telling help desk staff to "be more careful" fails against an attacker who never gets tired. The controls need to live in systems, where a person under pressure cannot waive them.
A workable approach is to express verification requirements as policy attached to actions, with no reliance on how believable the requester seems. A simplified example:
# verification-policy.yml
actions:
mfa_reset:
min_rung: 4 # existing passkey or hardware key
fallback:
min_rung: 3
requires: [manager_approval_in_idp, 24h_delay]
forbid_same_channel_confirmation: true
vendor_bank_change:
min_rung: 3
callback_source: vendor_master_record # never the request
approvals: 2
hold_before_first_payment: 72h
payment_over_threshold:
threshold_gbp: 25000 # illustrative
min_rung: 4
approver_auth: own_authenticator_in_erp
signals:
deepfake_detector:
effect: raise_risk_and_require_review # never auto-approve
The threshold value above is illustrative. The structure is what matters. Each sensitive action has a minimum rung, the confirmation channel cannot be the request channel, and detection signals can only add friction. They can never remove it.
Three design rules make this hold up:
- Same-channel confirmation is forbidden. If the request arrived by email, confirming by replying to that email proves nothing. If it arrived by phone, the callback goes to the number on file.
- Delays on irreversible actions. A 24 to 72 hour hold on bank changes or MFA resets costs little in normal operations and takes away the urgency attackers depend on.
- Humans approve inside authenticated systems. A manager "approving" a reset by saying yes on a call is rung 2. The same manager approving in the identity provider with their own passkey is rung 4.
Your own agents need the same discipline. If you deploy customer-facing or internal agents, they should identify themselves as agents and act through scoped, short-lived credentials. That way a malicious agent impersonating your agent has nothing useful to steal. We covered the credential side in Agent Identity Management: Why AI Agents Should Never Hold Keys. Agents that read inbound email or tickets are also targets for the manipulation itself, which is the problem described in our prompt injection playbook. An attacker's agent talking to your support agent is social engineering between two machines.
A Composite Example: The Supplier Bank Change
This scenario is a composite, assembled from common invoice fraud patterns. It is not a specific incident.
A mid-sized manufacturer's accounts payable team gets an email from a long-standing supplier's finance contact. The message picks up an existing thread about a delayed shipment, mentions the right purchase order numbers, and explains that the supplier is moving banks after an acquisition. A PDF on letterhead is attached. When the clerk hesitates, a follow-up arrives within minutes. It is polite, gives the name of the supplier's new CFO, and includes a phone number "for any questions." The clerk calls, a calm voice confirms everything, and the change goes through.
Each check in that story sits on rung 0 to 2. The thread context is researchable or comes from a compromised mailbox. The letterhead is trivial to produce. The quick, patient follow-up is something an agent does easily. The phone number came from the attacker.
Under the policy above, the same request fails without anyone needing to spot a fake:
- The callback goes to the number in the vendor master record, which reaches the real finance contact, who knows nothing about a bank move.
- Even if that call were somehow compromised, a second approver has to sign off in the ERP with their own authenticator.
- A 72-hour hold before the first payment to new details gives the real supplier time to notice the missing payment.
None of these steps needs AI. They are old controls that many companies relaxed because the old attacker rarely had the patience to beat them. That reasoning no longer holds.
How to Run the Verification Audit This Month
You don't need a large program to start. A focused audit over one to two weeks gives you a ranked list of what to fix.
- Pick five flows. Use the table above as a starting list and swap in whatever moves the most money or access in your business.
- Map every verification step. For each flow, write down every point where someone decides the counterparty is legitimate, including informal ones like "I recognized their name."
- Rate each step on the ladder. Label it rung 0 to 4. Be strict. A video call is rung 2 no matter how confident the interviewer felt.
- Find the weakest gate on each irreversible action. The flow is only as strong as the lowest rung that can authorize the action on its own.
- Write the replacement control as policy. Name the system that enforces it, the owner, and the exception path. If the exception path is "a senior person can override verbally," you have recreated the hole.
- Test with a simulated attacker. Red-team the redesigned flow with a scripted persona. If you already run agent simulation tests, the same harness can play the attacker.
For regulated teams, line each rung up with the assurance levels in NIST SP 800-63-4 or the UK digital identity trust framework your sector follows. That gives auditors a familiar reference point and keeps the redesign from looking like an internal opinion.
Where OpenNash Fits in a Verification Redesign
Plenty of organizations can run this audit with internal security and operations staff, and if your identity provider and ERP already support passkeys, approval workflows, and holds, the work is mostly configuration and training. Buy or configure first in that case.
Custom work makes sense when the verification logic spans systems that don't talk to each other, such as a help desk, an identity provider, an ERP, and an HR platform, or when you are deploying your own agents into those flows and need them to apply the policy consistently. OpenNash works in that gap. We audit the flows, design the policy and human approval points, build the enforcement into your existing systems, and hand over full ownership with audit trails your team can read.
If you have one flow you suspect is sitting on rung 1, such as supplier bank changes or help desk resets, bring it to a 30-minute working session. We will rate each step on the ladder with you and leave you with a written replacement control for the weakest gate, whether or not you build it with us.