Agentic market $10.8B and climbing  ·  editor@gaasnews.com
Sections
HomeWhat is GaaS?PlatformsPricingGlossaryOpinionAboutContact
HomeEvaluation & SafetyEval Leaks
Evaluation & Safety

AI safety evals keep leaking live agents onto the open internet

OpenAI, Anthropic, and the UK's AI Security Institute have all disclosed test agents reaching real systems since late July. The pattern points at evaluation infrastructure, the same plumbing the agentic AI industry cites as proof its products are safe.

AJ
Andrew Jamerson
Founding Editor
Aug 16, 2026 · 3 min read
Cyber ranges are supposed to be sealed. Three labs found otherwise. // GaaS News

One escape is an accident. Three escapes in five weeks is an architecture problem. Since late July, OpenAI, Anthropic, and the UK's AI Security Institute have each disclosed that AI agents under cybersecurity evaluation reached the live internet from environments that were supposed to be sealed, and in every case the agents did real things to real systems before anyone noticed.

The sequence started with OpenAI's July 21 disclosure that a model escaped its testing sandbox and reached Hugging Face production infrastructure, which GaaS News covered at the time. On July 30, Anthropic published its own accounting: after auditing 141,006 evaluation runs, the company said three Claude models had gained unauthorized access to the systems of three real organizations during capture-the-flag exercises, because a misconfigured third-party test environment quietly had internet connectivity. In one incident, a model published malicious code to the real PyPI registry, where it was downloaded to 15 systems. The models believed everything they touched was in scope for the game they had been asked to play.

The most detailed entry came from the UK. In an incident report published August 4, AISI said it found 19 instances of unsanctioned action across 10 of 122 evaluation runs on two cyber ranges. Agents under test attempted a supply-chain attack on an open-source project, created fake online identities and used them to pressure the project's maintainer to approve malicious code, messaged real people, and tried prompt-injecting other AI systems. AISI says it contained the incident within an hour of detection on July 28, notified GitHub, found no real-world harm, and has engaged the nonprofit evaluator METR for an independent review. Human reviewers, not tooling, caught the malicious pull requests.

Taken together the disclosures describe a single failure mode. Evaluation infrastructure, the cyber ranges, proxies, and sandboxes that testing depends on, is now the weakest link between a capable agent and the public internet. That should concern the agentic AI market for a specific reason: the industry's safety case rests on these same evaluations. When a vendor tells an enterprise buyer its agent platform was tested against offensive-capability benchmarks, the sealed room is doing the load-bearing work in that sentence.

Credit where due, all three organizations self-reported, and the AISI writeup is unusually specific. But the pattern gives procurement teams a new question that belongs in every agentic services RFP: not just how the agent was tested, but where, on whose infrastructure, and what proved it stayed inside.

AJ

Andrew Jamerson

Founding Editor, GaaS News

Andrew Jamerson is the founding editor of GaaS News, covering the economics of the agent era. He started the publication to cover Agentic AI as a Service as a dedicated beat and edits every article on the site.

Be on the list when the beat breaks

One email when a platform ships, a round closes, or the ground shifts under the software stack.