On the thirteenth of November last year, Anthropic published a disclosure that ought to have ruined the weekend of every security officer worth their salt: a state-sponsored group, which they tracked as GTG-1002, had been caught using Claude Code to run roughly eighty to ninety per cent of an espionage campaign against some thirty targets — reconnaissance, exploit development, credential harvesting, lateral movement, exfiltration — with human beings intervening at only four to six decision points per campaign. The machine had not merely picked the lock. It had cased the building, copied the keys, drafted the getaway route, and filed the paperwork, pausing only occasionally to ask a human whether it might be so kind as to approve the next felony.
Permit me one small throat-clear before we proceed, because the skepsis is part of the story. Several researchers I respect — Kevin Beaumont and Daniel Card among them — publicly wondered whether this was quite the watershed it was billed as, noting that the tooling involved was hardly exotic and that ordinary detections would have caught much of it; Anthropic’s own report concedes, with admirable candour, that its model periodically overstated its findings and occasionally fabricated credentials that did not, in fact, work. So hold both thoughts at once, as one must hold most things worth holding in this field: it was a genuine inflection point, and it was a marketing event. The epochal and the exaggerated are not mutually exclusive; they are, in our industry, practically married.
What is not exaggerated is the economics. A targeted intrusion campaign that once required a room of patient, expensive humans can now be staffed by one patient human and a subscription. And it arrives, this new industriousness on the part of our adversaries, at the precise moment when the defenders have obligingly spent three years wiring AI agents into every crevice of the enterprise — agents that read our email, triage our tickets, push our code, and refund our customers, each one a small, trusting butler with a set of house keys. A Cisco survey this March found eighty-five per cent of organisations experimenting with AI agents and five per cent willing to let them near production. That eighty-point gap is not a technology gap. It is a trust gap; and trust, as any seasoned parliamentarian will tell you, is what evaporates right after the assurances are given.
This essay is therefore an attempt at the full perambulation. We shall begin with how one is actually robbed in 2026 — the numbers are far worse than the folklore — then walk the entire defensive edifice, floor by floor: code, infrastructure, network, email, endpoints, ransomware, encryption, detection. Then we shall turn to the new and genuinely peculiar problem of attacks conductedby and against AI systems, and what one concretely does about them. And finally, because I am an engineer and not a pamphleteer, we shall end with what to do on Monday morning. The title, for those who recognise the borrowing, is an homage to a certain honourable member of Parliament whose vocabulary I can admire but not afford. The firewall farrago, alas, is entirely our own.
The present unpleasantness: how one is actually robbed in 2026
The single most important thing to understand about the modern intrusion is that it does not, in any meaningful sense, involve breaking in. CrowdStrike’s 2026 Global Threat Report puts the share of malware-free intrusions at eighty-two per cent. Verizon’s DBIR, the industry’s annual census of woe, has stolen credentials as the number one initial-access vector at twenty-two per cent of breaches, and a scarcely believable eighty-eight per cent of basic web-application attacks begin with somebody else’s password. The burglar does not smash the window. The burglar purchased the key for four dollars on a marketplace, lets himself in through the front door, and waves pleasantly at the camera on the way to the silver.
The supply of keys is industrial. Infostealer malware — Lumma, RedLine, and their many heirs — spends its days hoovering credentials, cookies and session tokens out of browsers, and the harvest is remarkable: more than half of ransomware victims had prior infostealer exposure, which is to say the robbery was over, in a sense, months before anyone called the police. Microsoft’s constabulary did drag away the Lumma operation in May of last year, and the market barely paused to inhale; a new stealer family, one notes, now hunts specifically for the local AI CLI tools developers have been installing with such enthusiasm, because where the credentials go, there go the thieves.
And the speed. Here the numbers pass from the concerning into the faintly comic. The average eCrime breakout time — the interval between an intruder’s first foothold and their expansion across your estate — is now twenty-nine minutes. The fastest observed was twenty-seven seconds, which is less time than it takes to microwave a poppadum. Mandiant, measuring the same phenomenon from a different angle, found the median time between initial access and the hand-off to a second actor has collapsed to twenty-two seconds. Twenty-two seconds. There is no human process on earth — no ticket queue, no approval workflow, no committee, however lean — that operates on that clock. Any defence that requires a person to notice something before it can act has already lost; it lost twenty-six seconds ago.
Drag to pan · Ctrl/⌘+scroll to zoom · % resets to fit · Full for fullscreen
Before we leave the census, one apparent paradox deserves untangling, because your board will ask about it. Breakout time has collapsed to minutes, yet median dwell time — how long intruders sit inside before discovery — actually rose to fourteen days, and for the espionage-minded the median is a languorous hundred and twenty-two. These are not contradictory; they are two different adversaries sharing one report. The eCrime fellow is a smash-and-grab artist who wants to be in, paid, and gone before lunch. The espionage fellow is a tenant, and he has discovered that your edge appliances — the VPN concentrators, the firewalls, the odd forgotten router — were never wired for logging at all, which makes them less a perimeter than a furnished flat with no landline. Mandiant’s BRICKSTORM cases are exactly this: persistence on devices your EDR cannot see and never will. When someone tells you their mean time to detect is excellent, the correct follow-up question is: on which estate?
The help desk is the perimeter now
If the front door is a stolen password, the tradesman’s entrance is a telephone call to your service desk. Scattered Spider — the loose collective behind the MGM and Caesars unpleasantness of 2023, and last year’s merry romp through airlines and insurers — has refined social engineering to a performing art: ring the help desk, sound flustered and important, obtain a password reset and an MFA enrolment, proceed to burgle. Voice phishing was the number two initial-access vector in this year’s M-Trends at eleven per cent, which means one attack in nine began with somebody being helpful, that most dangerous of corporate virtues.
The deepfake has now joined the repertoire, and here I must ask you to savour the Arup case, because it is the future arriving early and in fancy dress. A clerk in Hong Kong attended a video conference with what appeared to be the entire senior management of his firm. Every single participant was a puppet — synthetic faces, synthetic voices — and the clerk, observing a full boardroom of familiar colleagues instructing him in unison, obligingly wired out twenty-five million dollars. We spent years telling staff to be suspicious of the lone, misspelled email from a prince. We did not think to warn them about a quorum.
Meanwhile the data-breach business has quietly re-platformed itself onto OAuth. The 2025 Salesforce extortion wave — Scattered Spider, ShinyHunters and friends, two-point-eight million records at Allianz Life alone — and the Salesloft/Drift token affair, which relieved some seven hundred organisations of their data via abused OAuth tokens, share a structure: why attack the fortress when you can steal the courier’s badge? The victim list for the Drift incident included Cloudflare, Zscaler, PagerDuty and Tanium. When the companies whose entire raison d’être is keeping you safe are themselves relieved of their tokens, one is entitled to a moment of quiet despair — followed, one hopes, by a moment of loud token binding.
Your build pipeline is production (and other supply-chain parables)
The two defining supply-chain events of the period deserve their own paragraph of infamy. In March 2025, an attacker with a stolen token retroactively repointed every version tag of the popular tj-actions/changed-files GitHub Action to a malicious commit (CVE-2025-30066) that dumped CI runner memory — which is to say, your secrets — into public build logs. Some twenty-three thousand repositories were affected, and the thing was caught not by any scanner of code but by egress anomaly detection noticing the runner whispering to a strange host. Underline that detail; it becomes our hero later.
Then, in September, came Shai-Hulud: an npm worm seeded through a trivial colour-library package, which used a post-install script to run TruffleHog — yes, a defensive secret-scanner, pressed into service as the burglary tool, which takes a certainje ne sais quoi — harvested every credential it could find, exfiltrated them to public GitHub repositories, and republished itself into every package the stolen npm tokens could touch, spreading with no command-and-control server at all. The 2.0 wave in November moved to a pre-install hook, added a destructive fallback that could wipe your home directory, compromised over seven hundred packages with twenty million weekly downloads between them, spawned twenty-five thousand malicious repositories, and exposed some fourteen thousand secrets across four hundred and eighty-seven organisations — among them, crowed the headlines with evident delight, CrowdStrike’s own npm packages. There is no moral here except the uncomfortable one: every third-party package and every third-party action is unaudited code executing with your secrets, and your build pipeline enjoys a level of privilege that would make a domain admin blush.
Ransomware, or the extortion economy’s pivot to recovery denial
The ransomware business, after the law-enforcement theatre of Operation Cronos and the quiet closure of RansomHub, consolidated around Qilin, Akira and Cl0p — Qilin listing somewhere between eight hundred and a thousand victims last year depending on whose tally you trust, Akira credited by the CISA/FBI joint advisory with at least three hundred and forty-two organisations and roughly a quarter of a billion dollars in proceeds. The strategic shift is what matters: the mature groups have understood that encryption is merely the opening argument, and that the negotiation collapses the moment you can restore. Hence the pivot to recovery denial — they come for your backups first, your VMware estate second, and your EDR-less edge appliances for the long stay. The medallion on the wall of every ransomware affiliate now reads, in effect: kill the lifeboats, then torch the ship.
And the DDoS, briefly, is absurd
The Aisuru/TurboMirai botnet — one to four million conscripted IoT devices, your security cameras and doorbells pressed into service as artillery — drove the record to 29.7 terabits per second in the third quarter of 2025, then 31.4 terabits in November. That November flood lasted thirty-five seconds. Cloudflare mitigated 47.1 million attacks across 2025, up a hundred and twenty-one per cent year on year, an average of 5,376 attacks per hour; eighty-nine per cent of network-layer attacks and seventy-one per cent of HTTP attacks are over in under ten minutes, and attacks on AI companies spiked three hundred and forty-seven per cent in September alone. The modern DDoS, in other words, is finished before the incident channel is created, which is why on-demand mitigation is a contradiction in terms and always-on scrubbing is simply the price of being on the internet at all.
A 360° perambulation around the fortress
So much for the census of woe. Now to the edifice, and I want to state the thesis plainly before we climb, because it will be easy to lose in the detail: security is not a product one purchases but a property one engineers into every storey of the building, and the controls that actually repay their cost in 2026 are mostly unglamorous. Phishing-resistant identity kills the number one initial-access vector. Egress control kills both classic exfiltration and, as we shall see, most agentic mischief too. Immutable backups kill ransomware’s leverage. OIDC federation kills the long-lived cloud keys the supply-chain worms were built to steal. Everything else — the AI SOC, the agentic guardrails, the quantum-resistant cryptography — is a force multiplier on top of those four, and a force multiplier applied to zero is, arithmetically speaking, still zero.
Drag to pan · Ctrl/⌘+scroll to zoom · % resets to fit · Full for fullscreen
Let us walk the floors. I shall try to be brief about each, and I shall fail, but the attempt is sincere.
Floor one: code, and the AI that writes it
Begin where everything begins: the commit. And here we must have an honest word about AI-generated code, because the evidence has stopped being anecdotal. Veracode tested over a hundred models across eighty tasks and found forty-five per cent of generated code samples failed security tests — seventy-two per cent for Java, which suggests the models learned Java from the same place everyone else did, and eighty-six per cent failure on cross-site scripting specifically. AI-written code carried 2.74 times more vulnerabilities than human-written code, and AI-assisted commits leak secrets at roughly double the human rate. This is a volume problem stacked on a rate problem: the machine writes code the way a Victorian novelist produced serials — copiously, confidently, and with only occasional regard for how it ends — and it does so at ten times the speed of the Victorian novelist, which is rather the point.
The remedy is not to stop the machine; that ship has not merely sailed, it has been acquired by Google. The remedy is to treat every AI-authored pull request as untrusted input and gate it accordingly: SAST, secret scanning with push protection, and dependency scanning must pass before any AI-opened PR merges, and branch protection must require an accountable human reviewer — a named adult who signs for the merge the way a partner signs for the opinion. The single highest-leverage change in the whole of CI/CD, though, is OIDC federation: replace every long-lived cloud credential in your pipelines with short-lived tokens minted on demand, and you have removed the very thing tj-actions and Shai-Hulud were engineered to steal. Add SHA-pinning for every third-party action (tags, as twenty-three thousand repositories learned, can be repointed), provenance attestations via SLSA and in-toto, signed images via cosign, and egress control on the runners themselves — audit mode first, block mode when your courage matures. That last control, recall, is what actually caught tj-actions. The runner tried to whisper to a strange host, and the whisper was heard.
Floor two: infrastructure and cloud
The cloud-security market consolidated so violently that the consolidation itself became news: Google closed its thirty-two billion dollar, all-cash acquisition of Wiz this March — the largest deal in Google’s history and the largest pure-cybersecurity transaction on record — which tells you what the hyperscalers think posture management is worth. Underneath the ticker symbols, the reference stack is pleasantly boring: scan your Terraform before it becomes somebody else’s incident (Checkov, Trivy, OPA for policy); enforce admission control in Kubernetes so that unsigned, privileged, or otherwise insolent pods never schedule (Kyverno or Gatekeeper, with cosign verification); and watch the runtime with eBPF — Falco or Cilium Tetragon — so that a process loading a backdoored compression library is seen at the syscall layer, not in the post-mortem. Cisco’s Hypershield is the notable commercial bet here, all eBPF and autonomous segmentation, with a lovely party trick of blocking a vulnerable library load via kernel policy; the efficacy claims are vendor-sourced, as are most efficacy claims, which is the polite term for ‘unverified’.
Floor three: the network, and the one control to rule it
Zero Trust has graduated from whitepaper to plumbing — NIST moved from the concept (SP 800-207) to worked implementations (SP 1800-35) — and if you permit me to reduce an entire architectural movement to one sentence: default-deny egress is the exfiltration control, and everything else is commentary. Consider the evidence. Egress anomaly detection caught tj-actions. Egress control is what would have stopped the Shai-Hulud worm from reporting its harvest. Egress control is, as we shall see shortly, the precondition that makes every serious agentic defence work at all. An attacker who has run rings around your EDR, your SIEM and your dignity still has to carry the data out through a network you own, and if that network permits outbound connections to anywhere but an allowlist of known destinations, the rings run nowhere. Complement it with microsegmentation — the twenty-two-second hand-off means lateral movement must be architecturallyhard, not merely monitored — and mutual TLS between workloads so that identity, not network location, decides who may speak to whom.
Floor four: endpoints and the strange new shape of malware
EDR remains worth its licence fee, but remember the census: eighty-two per cent of intrusions arrive without any malware to detect, so behavioural detection — canary files, mass-file-change heuristics, credential-access patterns — is the main event and signatures the sideshow. And do note the shape malware is taking in the age of local models. ESET’s PromptLock, the first AI-orchestrated ransomware proof of concept, carried no malicious payload worth the name; it asked a locally hosted twenty-billion-parameter model to write its Lua scripts on the fly, fresh each time, which is the end of signature detection as a concept. Mandiant has since documented in-the-wild families that query LLMs mid-execution. Your next piece of malware may not contain its malicious code at all; it may simply request it, politely, at runtime, in flawless prose.
Floor five: email, humans, and the passkey
We have established that the number one way in is a stolen credential. It follows, with the inexorability of a syllogism, that the highest-leverage identity control is the one that makes credentials unstealable, and that control is the passkey — FIDO2, device-bound, phishing-resistant, immune in one stroke to credential stuffing, prompt bombing, and the entire infostealer economy. There is no security purchase with a better effort-to-effect ratio anywhere in this essay. If you retain one sentence from this entire farrago, let it be this: put every administrator on passkeys this quarter, and everyone else the quarter after. The deepfake boardroom, the vishing call, the helpful clerk — all of them end the same way, with a human being talked into handing over an authentication factor, and the passkey is the factor that cannot be handed over. It is the security equivalent of making the crown jewels ungiftable.
Floor six: ransomware and the theology of backups
Against recovery denial there is exactly one doctrine, and it is not negotiable: immutable, logically air-gapped backups, with restores that are tested. A backup you have never restored is not a backup; it is a hope, written to disk. S3 Object Lock or its equivalent makes the copy unchangeable even by your own compromised administrators; isolation keeps the ransomware from reaching it; and the quarterly fire drill — actually restoring something, on a clock, with witnesses — is what separates a recovery capability from a recovery anecdote. The median figure from IBM’s breach research is that a tested incident-response plan saves over two and a half million dollars per incident, which makes the fire drill the best-paid afternoon in your calendar.
Floor seven: encryption, key rotation, and the coming certificate apocalypse
Now the cryptographer’s corner, where two honest confessions are in order. The first concerns key rotation, on which I hold a heretical view: automated rotation of envelope-encryption keys is excellent precisely because it is free and invisible, whereas the calendar-driven manual rotation of application secrets is, more often than not, compliance theatre with an outage risk attached. The correct answer to long-lived secrets is not to rotate them religiously; it is to stop having them. Dynamic, just-in-time credentials — database passwords minted per session, cloud tokens minted per pipeline run — expire before they can be stolen, which renders the whole rotation calendar moot. (One procurement note, since we are here: Vault is no longer open source and its new owners have been pruning; the community fork, OpenBao, is where the OSS-first path now runs.)
The second confession is that certificate management is about to become an operational emergency for anyone still doing it by hand. The CA/Browser Forum has voted the maximum public TLS certificate lifetime down in steps — two hundred days from this March, one hundred days in 2027, and forty-seven days by 2029 — with domain-validation reuse shrinking to ten days. At forty-seven days you must re-prove ownership of every domain roughly thirty-five times a year, which is to say: manual renewal is dead, and if you do not have cert-manager and ACME humming away today, you have a dated, certain, entirely self-inflicted outage in your future.
And then there is the quantum fellow, patient as a glacier. Harvest-now-decrypt-later is the threat: adversaries recording your encrypted traffic today, to read at leisure when a cryptographically relevant quantum computer arrives. NIST finalised the post-quantum standards in August 2024, and the pragmatic 2026 move is hybrid key exchange — X25519MLKEM768 — which Chrome and Firefox already negotiate by default and which, by Cloudflare’s measurement, now protects over half of all web requests, roughly double the share of a year ago. Deploy it at your edge; it is configuration, not research. And test your infrastructure against the larger post-quantum signatures before you must — an ML-DSA-65 signature is fifty times the size of its ECDSA predecessor, and MTU-related surprises are the least enjoyable surprises. The quantum computer that will one day read your secrets need not exist yet. It suffices that your secrets, intercepted today, are patient.
Floor eight: logging, detection, and the AI-augmented SOC
The top of the building is where everything drains into: the security operations centre, whose job in one phrase is to detect quickly and intimate the responsible team before the twenty-two-second hand-off becomes the four-million-dollar breach. The pipeline is well understood — telemetry into a SIEM, detections as Sigma rules versioned in git (so they are portable, auditable, and not hostage to any one vendor’s fortunes), triage, severity, routing, response, and the regulatory clocks that start ticking at the worst possible moment.
Drag to pan · Ctrl/⌘+scroll to zoom · % resets to fit · Full for fullscreen
The honest news about the AI SOC is genuinely good, which will surprise readers who have been following my general line. Microsoft ran the strongest study in the field — a proper randomised controlled trial, not a webinar — and found analysts augmented with a triage agent worked up to seventy-eight per cent faster, with seventy-seven per cent better verdict accuracy and six and a half times more true positives per minute. But here is the detail that should be printed and framed: eighty-three per cent of the gain came from queue prioritisation — from simply putting the right alert in front of the human first. The magic is not robotic judgement; the magic is skipping the queue. Meanwhile roughly forty-six per cent of alerts remain false positives, and ninety-two per cent of organisations admit to incidents traceable to alerts that were seen and not investigated. So by all means, buy the triage agent. Just understand that you are buying a magnificent sorting hat, not a replacement wizard; and treat every ‘autonomous response’ pitch the way you would treat a ‘self-driving car’ pitch — ask, specifically, whether you are being sold lane-keep assist or a vehicle with no steering wheel. In 2026, it is lane-keep assist, guarded containment under human-defined guardrails, and anyone claiming otherwise is encouraged to publish their false-positive rate.
One paragraph of intimation plumbing, because detection without routing is just anxiety with a dashboard. Severity classification must drive automatic routing — the right team, the right channel, an acknowledgement SLA — and your mean time to acknowledge is the metric that exposes alert fatigue and understaffing more honestly than any maturity assessment. Wire the SEV-1 path to auto-create the war-room, page the on-call, and draft the regulatory paperwork simultaneously. On which subject, and I say this from Bengaluru with feeling: if you operate in India, CERT-In requires you to report a defined incident within six hours of detection — not confirmation, not convening a working group, detection — alongside a hundred and eighty days of log retention on Indian soil and time-synced servers. It is the strictest reporting window on earth (GDPR allows a leisurely seventy-two hours; the SEC four whole days), it carries criminal penalties, and the only way to meet it is to have decided in advance which severities auto-draft the preliminary report. Also, a public service announcement for the Opsgenie faithful among you: it is being switched off entirely in April 2027, data and all, so migrate before the lights go out.
Meanwhile, the machines have unionised
We come now to the part of the essay where the attack surface stops being familiar. Everything in the previous sections — stolen credentials, worms, floods, extortion — is the old war fought at new speed. What the GTG-1002 affair announced is something categorically different: the threat has migrated from the model to the orchestration layer. For three years the security conversation about AI fixated on the model saying something dreadful — a jailbreak, a slur, a recipe for napalm — while the actual 2026 problem is the agent doing something dreadful: chaining tool calls across a dozen politely integrated systems, each individual call defensible, the composition catastrophic. The unit of compromise is no longer the sentence. It is the workflow.
Consider the tools the modern agent is handed. The Model Context Protocol — MCP, the connective tissue that lets an agent discover and invoke external tools — has been adopted with the reckless enthusiasm our industry reserves for things that demo well. And the attack taxonomy that has grown around it reads like a pamphlet of confidence tricks, each with a real-world anchor. There is indirect prompt injection: hide the instruction in data the agent reads — a Jira ticket, a web page, an email — and the agent obligingly executes content as intent. Zenity’s Agent Flayer work demonstrated exactly this, a weaponised Jira ticket that talked the agent into exfiltrating secrets, which means your ticketing system is now an attack vector that types. There is tool poisoning, where the malicious instruction lives in the tool’s owndescription — the documentation attacks the librarian — and the MCPTox benchmark measured a 72.8 per cent attack success rate with the best models refusing less than three per cent of the time. There is the rug pull, where a tool you approved on Monday is silently redefined by Thursday (Cursor, CVE-2025-54136). There is straightforwardsupply-chain RCE: the mcp-remote package, four hundred and thirty-seven thousand downloads, carried a 9.6-severity remote-code-execution flaw (CVE-2025-6514, credit to Or Peles at JFrog) for anyone who ran it against a hostile server. And there ismemory poisoning — malicious content tucked into the agent’s long-term memory, a sleeper instruction that wakes in a future session when the context is right. The butler, in other words, can be hypnotised through the post.
The benchmarks, taken together, make for sobering reading over one’s morning filter coffee. On AgentDojo, the standard agent-hijacking suite, the best defended agents solve fewer than two-thirds of their tasks cleanly, attacks succeed up to a quarter of the time, and adding a secondary detector only brings that down to about eight per cent — better, not zero. The Agent Security Bench reports attack success up to 84.3 per cent. MCP-SafetyBench found host-side attacks succeeding about eighty-two per cent of the time, and — this is the number I would like every vendor of ‘AI guardrails’ to embroider on a sampler — adding safety prompts reduced attack success by a statistically insignificant 1.22 per cent. My favourite, in a bleak sort of way, is the WASP benchmark for web agents, whose authors coined the phrase of the year: attacks on current agents often fail for the same reason my attempts at the violin fail, namely incompetence rather than opposition. Security by incompetence. The agents are not defended; they are merely bad at being attacked. The day they become competent — and competence, one notes, is the entire direction of travel — the protection evaporates.
If you retain one design heuristic from this entire section, make it Simon Willison’s lethal trifecta: an agent that can (one) read private data, (two) ingest untrusted content, and (three) communicate externally is an agent with a guaranteed exfiltration path. Any two of the three is manageable. All three together is a smuggling route with a company badge. Audit every agent you run against those three capabilities, and treat every yes-yes-yes as a critical finding, not an architecture decision.
Drag to pan · Ctrl/⌘+scroll to zoom · % resets to fit · Full for fullscreen
Two more structural shifts deserve their place in the census before we turn to remedies. The first is the quiet explosion of non-human identity. Machine identities — service accounts, API keys, agents — now outnumber human users forty-five to one on average, and up to a hundred and forty-four to one in cloud-native estates, a figure that itself jumped fifty-six per cent in a single year. Yet only twenty-two per cent of security teams treat their agents as first-class identities, and only twenty-eight per cent can trace an agent’s actions back to a responsible human sponsor. We have built a workforce of ghosts and then expressed surprise when the forensic trail is ghostly. The second is the agent-to-agent cascade: in multi-agent settings, one compromised participant’s output is trusted downstream, and control-flow hijacking in such systems succeeds fifty-six per cent of the time. The Morris II worm proved the propagation mechanism in the lab two years ago — a self-replicating adversarial prompt, spreading from agent to agent through email and retrieval stores across three different vendors’ models, zero clicks required. It was a proof of concept. So was the printing press, once.
The arms merchants: who is selling what, and does any of it work?
Now, lest this read as unrelieved gloom, the defence has not been idle. Indeed the most interesting development of the period is that the frontier laboratories have started shipping genuinely defensive machinery, and one cannot survey this field honestly without naming names — both because the capabilities are real and because the marketing is, shall we say, aspirational. Let us take the model-makers first and the platforms second, and let us keep our wits about us throughout.
The frontier laboratories, and the guns they will not sell you
Google has the most quietly impressive record. Its Big Sleep discovery agent found a live SQLite vulnerability (CVE-2025-6965) before the attackers could use it — the first occasion on which an AI is known to have foiled an in-the-wild exploit, which is a genuine milestone however modestly it was announced. Its sibling CodeMenderdoes the unglamorous other half, having upstreamed seventy-two security patches into real projects, and Gemini 3.5 Flash Cyber is the productised edge of the same capability — restricted, note, to governments and trusted partners.OpenAI’s Aardvark works the same seam from the developer’s end: it threat-models your repository, scans each commit, validates that a finding is actually exploitable in a sandbox rather than merely theoretically alarming, and drafts the patch. Ninety-two per cent detection on golden repositories, ten or more CVEs filed, and still in private beta.
Microsoft made the loudest move of the summer.MAI-Cyber-1-Flash, shipped on the twenty-seventh of July, is a cyber-tuned mixture-of-experts model — a hundred and thirty-seven billion parameters total but only five billion active, with a two-hundred-and-fifty-six-thousand-token context — which handles roughly ninety per cent of security tasks itself and escalates the awkward tenth to a larger frontier model. It scores 95.95 per cent on CyberGym at about half the going cost, sits inside the MDASH platform, and its companion Project Perception entered public preview, as it happens, this very morning. The architecture is the interesting part and generalises well beyond Microsoft: a cheap, fast, domain-tuned model doing the bulk of the work with selective escalation to an expensive one. That is precisely the two-tier pattern Uber’s ADR uses for detection, and it is becoming the house style of security AI for the excellent reason that it is the only shape that survives contact with a real alert volume and a real budget.
And then there is Anthropic, whose experience this summer ought to be pinned above the desk of anyone drafting a three-year security roadmap. Its frontier bug-hunter, Mythos 5 — the restricted variant, with safeguards removed for vetted organisations, alongside the safeguarded public Fable 5 and the broader Project Glasswing effort — was shut down by US export controls on the twelfth of June, partially restored between the twenty-sixth and the thirtieth, and has been in an extraordinary geopolitical tug-of-war ever since. The NSA, one notes, was a user. The lesson is not about Anthropic. The lesson is that access to the most capable defensive models is now a matter of government directive rather than contract, which is a sentence no procurement process in the world is currently designed to accommodate. Architect for model portability. Rent the brain if you must, but the policy plane, the identity plane and the telemetry must be yours, in your git, on your terms — because the brain may be repossessed by a ministry on a Thursday.
One further observation about this entire category, which I offer without much comfort: every one of these systems is a dual-use instrument. A model that can find an exploitable bug in your repository is, by construction, a model that can find one in somebody else’s. The restraint currently on display — the restricted access, the vetted-organisations-only postures, the export controls — is the only thing standing between the defensive arms race and the offensive one, and restraint is historically among the more perishable commodities.
The platforms, and the eighty-point gap
Cisco has assembled the most complete agentic-security platform on the market, largely by acquisition. AI Defense is the centrepiece: discovery of the agents you did not know you had, supply-chain scanning of model files, datasets, tools and MCP servers, and runtime guardrails that claim to catch memory poisoning, tool misuse, privilege escalation and intent hijacking — which is to say, most of the taxonomy from the previous section. Underneath sit Foundation-Sec-8B, an open-weight security model built on Llama 3.1 and genuinely useful for teams who want something they can host themselves; and, more recently, Antares — a family of tiny, open-weight agentic models (three hundred and fifty million and one billion parameters released, a three-billion-parameter sibling kept in-house) trained specifically to walk a repository the way a triage analyst would and point at the files most likely to hold a known vulnerability. Compact enough to run on-premises, which means you need not ship your source code to somebody else’s cloud for the privilege of being told which file to open. That is the useful pattern: not a general coding assistant pressed into security service, but a narrow instrument for one expensive step in the chain. Under the same roof sit Robust Intelligence (acquired 2024) for red-teaming mapped to OWASP and ATLAS, and Splunk, bought for twenty-eight billion dollars, now serving as the observability and detection backend, with the Galileo acquisition bolted on for agent observability specifically. At RSAC this year Cisco added DefenseClaw (a secure agent framework), an Agent Runtime SDK that pushes build-time policy into Bedrock AgentCore, Vertex, Azure AI Foundry and LangChain, and an LLM Security Leaderboard. Separately, Hypershield covers the infrastructure floor we walked earlier, all eBPF and autonomous segmentation.
It was also Cisco that produced the single most clarifying statistic in this entire essay, and I return to it deliberately because it is the thesis in numerical form: eighty-five per cent of organisations experimenting with AI agents, five per cent willing to put them into production. Their chief product officer called that eighty-point chasm a trust deficit, and he is right, but I would put it more sharply. The security gap is the adoption gap. Every organisation stuck at the pilot stage is stuck there because nobody can answer the question ‘what exactly can this thing do at three in the morning when no one is watching’ — and that question is answered by identity, allowlists and a broker, not by a better model. The commercial prize for solving agentic security is therefore not a security budget. It is the entire agentic AI market, currently sitting in a queue behind it.
CrowdStrike owns the agentic-SOC narrative more thoroughly than anyone. Charlotte AI Agentic Response arrived at RSAC 2025; the Agentic Security Workforce followed at Fal.Con in September with seven mission-ready agents — malware analysis, hunting, exposure prioritisation, search analysis, correlation-rule generation, data transformation and workflow generation — alongside AgentWorks, a no-code builder for rolling your own, and Charlotte Agentic SOAR in November. MCP is the connective tissue throughout, with native AI Detection and Response governance layered over it, which is at least an honest acknowledgement that their own agents constitute an attack surface. My one reservation is evidentiary: the efficacy claims are almost entirely qualitative. ‘Machine speed’ is a lovely phrase and not a number, and I have yet to see a published MTTR figure one could hold them to.
Palo Alto Networks is the other heavyweight and has been shopping with real intent: Protect AI for somewhere around six hundred and fifty to seven hundred million dollars, folded into Prisma AIRS, which protects AI applications, models, data and agents. The AIRS MCP Server validates every tool invocation and flags tool-definition and tool-input poisoning — which is, pleasingly, the gateway-as-enforcement-point pattern this essay has been advocating, sold as a product. AIRS 2.0 brought five hundred red-team attack types; 3.0 extends to discovering agents across cloud, SaaS, endpoint and browser, including what the marketing department has gamely christened ‘vibe coding agents’. Cortex AgentiX is the agentic SOC play, trained on 1.2 billion playbook executions, with role-based access and human approval for impactful actions — and a claimed ninety-eight per cent MTTR reduction which is, I must note with the weariness of a man who has read many such decks, vendor-sourced and unaudited. They are also acquiring CyberArk, which tells you exactly how valuable machine identity is about to become.
Microsoft’s Security Copilot agents deserve a closing mention for one detail of design rather than for their (genuinely good, genuinely measured) triage numbers: every agent gets its own identity and its own scoped permissions. That is per-agent authentication shipped as a default in a mass-market product, and while it is not yet the per-task, plan-hash-bound credential I shall advocate shortly, it is the correct direction of travel and rather more than most of the field manages.
The startups, in three buckets
The venture money has sorted itself into three functional piles, which is a useful map when the logos start to blur. There is agent identity and access governance — Oasis Security (a hundred and twenty million dollar Series B at RSAC this year), Astrix, Entro, Linx, and Aim Security, now absorbed by Cato — all of them selling the answer to the forty-five-to-one non-human identity problem. There is runtime and MCP defence — Zenity, whose researchers gave us Agent Flayer; Prompt Security, acquired by SentinelOne; Lakera, acquired by Check Point, whose Guard product answers injection queries in under fifty milliseconds; plus Noma, Straiker, Lasso, Pillar, HiddenLayer and Witness AI. And there is policy and gateway infrastructure, which is refreshingly open-source-heavy: Solo.io AgentGateway (Rust, Cedar, now under the Linux Foundation), Red Hat’s MCP Gateway, ToolHive, IBM ContextForge, Permit.io and Cerbos, with AWS Bedrock AgentCore as the managed equivalent. Special mention to XBOW, which raised a hundred and twenty million dollars at a billion-plus valuation for autonomous offensive security, and which I include partly for completeness and partly because somebody ought to say aloud that the same capital markets funding your defences are funding the other thing too.
Where does this leave a buyer? My honest reading is that you should buy the planes that are solved problems and own the ones that encode your judgement. Nobody should be writing their own cryptographic workload identity in 2026, and nobody should be outsourcing their tool-necessity policy, because that policy is your security posture expressed as code. The map:
Drag to pan · Ctrl/⌘+scroll to zoom · % resets to fit · Full for fullscreen
A final word on reading vendor claims, offered in a spirit of public service. When a deck promises ‘autonomous’, ask what happens when it is wrong and who signs for it. When it promises a percentage, ask whose workload produced it and whether anyone independent has reproduced it — the answer, for essentially every figure in this section, is no. And when it promises to solve prompt injection, recall that the best peer-reviewed detectors take AgentDojo attack success from twenty-five per cent down to about eight, which is a fine day’s work and is not a solution. The strongest evidence in this entire field remains a Microsoft randomised controlled trial about queue sorting and a single-author preprint about deterministic authorisation. Buy accordingly: platforms for correlation and speed, open source for the control plane you must still own when the contract lapses or the ministry intervenes.
Defending against one’s own butler
Which brings us to the constructive heart of the essay: what does one actually build? The guiding principle, established by every benchmark cited above, is that durable defence is architectural, not rhetorical. You cannot talk an agent into being safe any more than you can talk a river into flowing uphill; safety prompts move the needle 1.22 per cent, and prayer has a similar effect size. What works is structure: identity, scoping, independent authorisation, egress control, and observability. Let us take them in order, from the foundation to the capstone.
First, see the agent at all
The foundational problem is observability, because your EDR watches outcomes — files written, packets sent — while the agent’s mischief lives in the causal chain: the prompt, the reasoning, the tool call, the result. Uber’s ADR system, the anchor paper in this space and a genuinely impressive piece of engineering, closes precisely that gap: a lightweight sensor parses the local session caches of the AI coding tools on seven thousand two hundred hosts (ten thousand sessions a day) to reconstruct what each agent was asked, what it reasoned, and what it did; a cheap model triages, an expensive one investigates, and an offline red-team loop keeps discovering new attack variants to feed back in. The honest numbers: ninety-seven per cent precision on credential-leak prevention, but sixty-seven per cent detection on the attack benchmark. Read that second figure again. The best purpose-built detection system in the public literature, deployed by a company that does not lack for engineers, misses a third of attacks — deliberately, trading recall for the precision that keeps a production SOC sane. Detection is necessary. Detection is not sufficient. What is sufficient, or nearer to it, is making bad things impossible rather than merely visible.
Second, an identity of one’s own
The first structural sin of enterprise agent deployments is the borrowed credential: the agent running on a human’s session, API key or OAuth token. The consequences are threefold and each is fatal to something you care about: every action the agent takes is legally and forensically yours; its blast radius is everything you have ever accumulated access to; and the only kill switch that exists also kills you. The remedy is per-task agent authentication — the agent as a first-class, cryptographically attested non-human identity (SPIFFE/SPIRE is the mature open standard), with the human preserved as delegating subject through OAuth token exchange, the token bound to a single audience, and the whole credential scoped to the task and alive for minutes, not months.
Drag to pan · Ctrl/⌘+scroll to zoom · % resets to fit · Full for fullscreen
The plan hash is the load-bearing idea. When a task is created, the orchestrator hashes the task description together with the approved plan and tool set; that hash rides inside the token as a claim; and every tool call is checked against it. A call inside the plan executes. A call outside the plan — the agent, mid-refund, suddenly developing an interest in the payroll export — is denied automatically, by mathematics rather than by anyone’s vigilance. The agent cannot drift beyond its brief because its brief is cryptographically stapled to its passport.
Third, least privilege, applied to tools
Do not grant an agent the union of every tool it might ever want; grant the minimal set for the task at hand, deny by default, and derive the sets from observed usage as your telemetry matures. This is the eighty-twenty win of the entire agentic defence programme — it requires no new infrastructure, only a decision — and yet it is where I see the most resistance, because convenience is a hell of a drug. One honest caveat, courtesy of the research: strictness has a utility tax. In the aiAuthZ evaluation, a rigid payee allowlist dropped clean-task utility from a hundred to eighty per cent by blocking one legitimate first-time payment. Budget for it, and provide a human-escalation path that is cheaper than the meeting in which the exception is currently decided.
Fourth, and the capstone: the Tool Necessity Broker
Now the strongest idea in the whole design space, and the one I would build if I built only one. The orchestrator requests ten tools for a task. An independent system — off-host, in a separate trust domain, holding no credentials the agent can reach — sits between the agent and the tools and asks, of every single call: is this actually necessary, and is it permitted? The prior art runs from clever-but-porous (dual-LLM designs, where a privileged planner never touches untrusted data) through formal-but-heavy (DeepMind’s CaMeL, enforcing capabilities on every value) to the decisively empirical: the aiAuthZ paper, which published the strongest result in the field this year. An off-host, identity-bound authoriser with a deterministic three-gate policy — role, then argument constraints, then rate limit — drove residual attack success to zero per cent across all fifteen models tested, with thirty microseconds of decision latency. Note why it worked: not because the verifier was clever, but because the decision was deterministic and physically out of the agent’s reach. A guard LLM that reads the same poisoned context as the agent is merely a second victim with a job title.
Drag to pan · Ctrl/⌘+scroll to zoom · % resets to fit · Full for fullscreen
Three design rules, hard-won. First, the deterministic layer carries the guarantee; the model-based layer — which scores whether requesting those ten tools is plausible for the stated task — is advisory only, sees the structured plan and never the untrusted content, and escalates rather than authorises. Let a context-reading LLM make the final call and you have rebuilt the injectable surface you were defending, at considerable expense and with extra latency. Second, never put a synchronous model call in the per-tool-call path; deterministic policy evaluation is sub-millisecond, an LLM round-trip is one to twelve seconds, and your agents will die of old age. Third — and the authors deserve eternal credit for publishing their own limitation — the broker only works if the agent has no tool path that bypasses it. In end-to-end testing, when the model had built-in tools, it simply performed the sensitive action through those and never consulted the gateway at all. Which means egress-locking the execution environment is not a hardening step; it is the precondition. The velvet rope works only if there is no side door, and agents, like nightclub patrons, are remarkably inventive about side doors. Fail closed for state-changing tools when the broker is unreachable; fail open, if you must, only for read-only ones.
And a word against my own enthusiasm: for a read-only agent — the doc-summariser that cannot write, cannot send, cannot spend — the full broker is over-engineering, because the trifecta is absent and scope restriction plus observability suffice. The broker earns its complexity precisely when the agent can touch private data and act on the world. Build the inventory first, the identity second, the allowlists third, and the arbiter last. It is the capstone, not the foundation, and organisations that invert that order end up with a magnificent authorisation layer guarding an estate they cannot enumerate.
Round it out with the unglamorous plumbing: secrets resolved by a credential broker on the gateway (referenced by name, minted post-authorisation, never resident in the agent’s memory); MCP servers treated as untrusted third-party code — scanned, vetted, pinned to exact versions, sourced from a signed registry; and a tamper-evident, hash-chained audit log whose head is anchored somewhere a compromised broker cannot rewrite, because the audit trail is what separates an incident from a mystery.
The shape of storms to come
Permit me a short act of disciplined prophecy — disciplined, because each item below is an extrapolation from something already demonstrated, not a fantasy. Over the next twelve to thirty-six months, I expect four developments, and I shall lay them out in ascending order of discomfort.
First, agent worms in the wild. Morris II was a laboratory result, but consider what the enterprise is presently doing, with the best of intentions: interconnecting agents, giving them shared memory, and tuning their agency upward for productivity. In epidemiological terms, we are raising the basic reproduction number of a self-replicating prompt as a side effect of our OKRs. The first real worm will not announce itself as malware; it will present as a strangely popular instruction circulating through shared inboxes and knowledge bases, which is why ACL-trimmed retrieval stores and provenance labels on content — knowing which bytes are trusted — stop being optional.
Second, the orchestrator becomes the crown jewel. As the planner/router accumulates more authority than any single tool, attacks shift from tricking the model to compromising the planner — and control-flow hijacking already succeeds fifty-six per cent of the time in multi-agent settings. The necessity-broker pattern is the counter, because it removes authorisation from the planner’s jurisdiction entirely: the general may draw the battle plan, but the paymaster is deaf, literal, and immune to eloquence.
Third, identity abuse at machine scale, plus sleeper payloads. With non-human identities outnumbering humans forty-five to one and barely a quarter of organisations able to trace an agent to its sponsor, stolen agent tokens become the default initial-access vector — the same credential economy that fuels today’s intrusions, transplanted onto a population nobody is watching. Expect, too, dormant instructions planted in vector stores and long-term memory today, designed to trigger on some future context; the time-bomb as a retrieval result. And expect the deepfake quorum to migrate into your human-in-the-loop: when the approval checkpoint is a video call, the attacker’s job is to synthesise the approver. The Arup clerk has shown us the film; we have merely to wait for the sequel in which the ‘human approval’ itself is counterfeited.
Fourth, and most speculative, the poisoning of the marketplaces. As skill and tool registries standardise — and they will, because our industry cannot resist an app store — they inherit the npm problem with an agentic twist: Snyk’s audit of nearly four thousand community agent skills found over thirteen per cent with critical-level issues and seventy-six confirmed malicious payloads, ninety-one per cent of which combined prompt injection with traditional malware. The trojan horse of the agentic era is not a binary; it is a five-star plugin with excellent documentation.
The unified edifice, and what actually pays
If you have followed the perambulation this far, you will notice that the agentic defences and the classical defences are not two programmes but one: the same identity fabric (passkeys for humans, SPIFFE for workloads and agents), the same policy language (OPA or Cedar, from Kubernetes admission to tool calls), the same gateway pattern (SASE for humans, service mesh for workloads, MCP gateway for agents), the same observability spine draining into one SIEM, and the same two controls appearing at every turn like a leitmotif in a Wagner opera — egress control and phishing-resistant identity — because they sit athwart the kill chain at both ends. Map the defences to the attack lifecycle and the picture is almost embarrassingly clear:
Drag to pan · Ctrl/⌘+scroll to zoom · % resets to fit · Full for fullscreen
And here, after nine thousand words of machinery, is the conclusion that will disappoint the keynote circuit: the controls that pay off in 2026 are unglamorous. Passkeys kill the credential vector that starts a fifth of breaches. Egress control kills both classic exfiltration and the agentic bypass path. Immutable, tested backups kill ransomware’s leverage. OIDC federation kills the keys the worms were built to steal. Certificate automation and hybrid post-quantum key exchange kill two dated, certain, entirely avoidable future failures. The AI defence machinery — the brokers, the triage agents, the ADR sensors — is a force multiplier on top of those five, and I say this as someone who has just spent four sections singing its praises. Nobody writes a ballad about egress control. Write one anyway.
What to do on Monday morning
Ten things, in order, none requiring a procurement cycle:
- Stand up the agent and workload inventory. Owner, identity, tool scope, data class — for every agent and service. You cannot defend, or run incident response on, an estate you have not enumerated. Join the twenty-eight per cent who can trace an agent to a human sponsor; then beat them.
- Kill one long-lived cloud key. Flip a single CI pipeline to OIDC federation today. Repeat until none remain.
- SHA-pin every third-party GitHub Action and turn on secret scanning with push protection. Tags lie; hashes do not.
- Put every administrator on passkeys. This closes the number one initial-access vector, full stop. Everyone else next quarter.
- Impose deny-by-default tool allowlists per task class. The highest-ROI agentic control in existence, and it requires no new infrastructure — only the decision.
- Apply the lethal-trifecta rule as policy. Any agent or service that can simultaneously read private data, ingest untrusted content, and communicate externally goes on the review list today.
- Verify that your backups are immutable and actually restore. Run one test restore, on a clock, with witnesses. A backup you have never restored is a hope, written to disk.
- Turn on agent-native and egress telemetry. Capture the causal chain — intent, reasoning, tool calls — that your EDR cannot see, and put your CI runners in egress-audit mode so the next tj-actions whispers into a microphone.
- Confirm your CERT-In six-hour path works. Who files, from which alert, with a hundred and eighty days of logs retained in India? Tabletop it once, before you have to do it for real.
- Pick your gateway. Choose one MCP gateway and commit to routing a hundred per cent of tool calls through it. It is the precondition for every stronger control that follows, and a bypassed gateway is a decorative one.
Coda
There is a temptation, after a survey of this length, to conclude that the situation is hopeless — that with twenty-two-second hand-offs, hypnotisable butlers, and worms that write themselves, the rational course is to quietly retrain as a potter. I would resist it. The arms race is real, but as of this writing it is being run by participants mostly choosing to hold back: the defensive models are shipping faster than offensive capability is proliferating to the merely curious, and the window that creates — this window, now — is the opportunity. Security was never a destination one arrives at, nor a product one purchases from a smiling account executive; it is a practice, like cleanliness, like constitutionalism, like the violin — maintained daily or not at all. The fortress with the well-kept moat is not the fortress that is never attacked. It is simply the fortress the attacker drives past, shaking his head, on the way to one whose owners mistook the warranty for the wall.
Eternal vigilance, as the old slogan has it, is the price of liberty. In 2026 it comes with a service-level objective, a six-hour reporting clock, and — if you have read this far — a to-do list for Monday. Off you go.
Frequently asked questions
What is the single most common way organisations get breached in 2026?
Stolen credentials. Attackers increasingly log in rather than break in: stolen passwords and session tokens are the top initial-access vector (22% of breaches per the Verizon DBIR 2025), 88% of basic web-application attacks use stolen credentials, and 82% of observed intrusions are now malware-free. Phishing-resistant passkeys (FIDO2) are the highest-leverage defence against it.
What is the 'lethal trifecta' in AI security?
Simon Willison's term for an agent that can simultaneously read private data, ingest untrusted content, and communicate externally. With all three present, a prompt-injection attack has a guaranteed exfiltration path. The design rule: never grant one agent all three capabilities without independent review, deterministic policy enforcement, and egress control.
Is prompt engineering enough to secure an AI agent?
No. Model- and prompt-level defences are porous: MCP-SafetyBench found safety prompts reduce attack success by a statistically insignificant 1.22%, and MCPTox measured a 72.8% attack success rate with best-model refusal below 3%. Durable protection comes from architecture — least-privilege tool allowlists, egress control, and deterministic off-host authorisation.
Do AI SOC agents replace human analysts?
Not in 2026. Microsoft's randomised controlled trial found agent-augmented analysts triaged up to 78% faster with 77% better verdict accuracy — but 83% of the gain came from queue prioritisation, and roughly 46% of alerts remain false positives. Production 'autonomous response' means guarded containment under human-defined guardrails, not lights-out automation.
What is indirect prompt injection?
An attack where malicious instructions are hidden in data an AI agent reads — a Jira ticket, a web page, an email, a document — rather than typed by the user. The agent ingests the instruction as content and executes it as intent. Real-world proofs include the Agent Flayer Jira-ticket exfiltration and tool-poisoned MCP servers documented by Invariant Labs.
What should an organisation do first to secure its AI agents?
Five moves before any sophisticated tooling: (1) inventory every agent with owner, identity, tool scope and data class; (2) give each agent its own first-class identity instead of borrowed human credentials; (3) enforce deny-by-default tool allowlists per task class; (4) apply the lethal-trifecta rule to flag risky agents; (5) route 100% of tool calls through a gateway so no tool path bypasses policy.
References and further reading
The claims above lean on the following; figures flagged as vendor-sourced should be read with the customary grain of salt.
Papers and benchmarks
- Li, C., et al. “ADR: An Agentic Detection System for Enterprise Agentic AI Security.” Uber + MIT, MLSys 2026 (Industry Track). arXiv:2605.17380. Source of the observability architecture, the 7,200-host deployment, 97.2% credential-prevention precision and 67% detection figures.
- “aiAuthZ: off-host, identity-bound deterministic tool-call authorisation.” arXiv:2607.05518 (2026). Single-author preprint; source of the 0% residual attack-success result across 15 models, ~0.03 ms latency, the utility-tax finding, and the built-in-tools bypass caveat.
- Microsoft Research. Randomised controlled trial of AI-assisted alert triage. arXiv:2511.13860 (2025). 78% faster triage, +77% verdict accuracy, 83% of gains from queue prioritisation.
- AgentDojo, Agent Security Bench, WASP, MCPTox and MCP-SafetyBench benchmark suites (2024–2025). Sources of the attack-success statistics, the 1.22% safety-prompt effect, and the “security by incompetence” finding — AgentDojo; Agent Security Bench; WASP; MCPTox; MCP-SafetyBench.
- Willison, S. The dual-LLM pattern (2023) and the lethal trifecta framing.
- “Morris II”: zero-click self-replicating prompt worm across GPT-4, Gemini Pro and LLaVA (2024) — arXiv:2403.02817; DeepMind’s CaMeL capability-enforcement design (2025) — arXiv:2503.18813.
Threat reporting and industry data
- Anthropic. Disrupting the first reported AI-orchestrated cyber espionage campaign (GTG-1002). November 13, 2025 — read alongside the public scepticism of Kevin Beaumont and Daniel Card, and Anthropic’s own caveat about fabricated findings.
- CrowdStrike Global Threat Report 2026: 29-minute average breakout (fastest 27 s), 82% malware-free intrusions. Mandiant M-Trends 2026: 22-second median hand-off, 14-day median dwell (122 for espionage), vishing as #2 initial vector (11%). Verizon DBIR 2025: stolen credentials at 22% of breaches. IBM Cost of a Data Breach: 241-day identify-plus-contain, savings from AI/automation and tested IR plans.
- CISA/FBI joint advisory AA24-109A (Akira ransomware, updated November 2025); Cloudflare quarterly DDoS reports (31.4 Tbps record, 47.1M attacks in 2025); Snyk’s ToxicSkills agent-skill audit (February 2026); CSA Non-Human Identity Governance Vacuum whitepaper (May 2026).
- Veracode 2025 GenAI Code Security Report (45% failure rate, 2.74× vulnerability density); Pearce et al., Asleep at the Keyboard? (2022); CSA and Apiiro research on AI-assisted commit hygiene.
Vendor products and platform claims
- Frontier cyber models: Anthropic Mythos 5 / Fable 5 and Project Glasswing (export-control shutdown June 12, 2026; partial restore June 26–30); Microsoft MAI-Cyber-1-Flash / MDASH / Project Perception (launched July 27, 2026; Perception preview August 3, 2026; 95.95% CyberGym); Google Big Sleep, CodeMender and Gemini 3.5 Flash Cyber; OpenAI Aardvark (private beta).
- Platforms: Cisco Antares (open-weight vulnerability-localisation SLMs, July 2026), AI Defense, DefenseClaw, Foundation-Sec-8B, Hypershield; CrowdStrike Charlotte AI Agentic Response / Agentic Security Workforce / Charlotte Agentic SOAR (Fal.Con 2025); Palo Alto Prisma AIRS / Cortex AgentiX (Protect AI acquisition; vendor-sourced MTTR claims); Microsoft Security Copilot agents (Ignite 2025 RCT numbers).
- Cisco Newsroom survey (March 2026): 85% experimenting with AI agents, 5% in production — the eighty-point trust gap cited throughout.
Vulnerabilities, standards and rules
- CVE-2025-30066 (tj-actions/changed-files), CVE-2025-30154 (reviewdog), CVE-2025-6514 (mcp-remote RCE, CVSS 9.6, Or Peles / JFrog), CVE-2025-54136 (Cursor rug-pull), CVE-2025-6965 (SQLite, found by Google Big Sleep).
- RFC 8693 (OAuth token exchange; delegation-splicing weakness documented by the IETF OAuth WG in 2026), RFC 8707 (resource indicators), OAuth 2.1 + PKCE, DPoP/mTLS, CAEP, SPIFFE/SPIRE, OpenTelemetry GenAI semantic conventions.
- NIST FIPS 203/ 204/ 205 (post-quantum standards, August 2024); CA/Browser Forum SC-081v3 (TLS lifetimes to 47 days by 2029); NIST SP 800-207 and SP 1800-35 (Zero Trust); CERT-In Direction 20(3)/2022 (6-hour reporting, 180-day in-country log retention).
On this site
- Anatomy of an Agentic AI System — sandboxes, egress allowlists, and permission-aware retrieval — the build-side of this security argument.
- The Rope Sellers Buy a Rope Machine — why the industry is shipping agents into production faster than it is securing them.

