UK government testers got Claude Opus 5 to breach a test enterprise network in 8 of 10 attempts
The 190-plus page document pairs Anthropic's best-ever alignment audit results with UK government testing in which the model completed a simulated enterprise network attack in 8 of 10 attempts.
Anthropic released Claude Opus 5 on Friday alongside a system card that runs past 190 pages, and one finding is driving the weekend coverage: UK government safety testers got the model to carry an attack on a simulated enterprise network from start to finish in 8 of 10 attempts, according to TechTimes' reading of the card, published Saturday.
Eight out of ten
Per that account, testers gave the model access to a simulated enterprise network protected by standard but not hardened security controls, and Opus 5 reached the end of the attack path in eight of ten runs. Weekend coverage identified the evaluators only as UK government testers; the specific body was not named in the reports, although the UK conducts pre-deployment testing of frontier models through its AI Security Institute.
The system card, dated July 24, places the result in context: independent analyses of the document note that it describes Opus 5, Mythos Preview and Mythos 5 as similarly capable at attacking small enterprise networks where access already exists. The capability, in other words, is not unique to Anthropic's newest model. What is new is a government test putting a near-reliable success rate on paper.
The best alignment numbers Anthropic has published
The same document contains the most favorable alignment results Anthropic has ever reported. On the company's automated behavioral audit, Opus 5 posted the lowest misalignment score Anthropic has recorded, better than every prior model, including the restricted Mythos 5, per the card. The card also reports the largest gains to date in prompt injection robustness across coding, computer use and browser use: on the IPI benchmark, the probability of an attacker succeeding within 15 attempts fell to 2.0 percent from 5.5 percent for Opus 4.8, and attack success rates in computer use settings dropped to 0.54 percent from 7.14 percent with extended thinking enabled, per figures in the card.
On catastrophic risk categories, the card assesses Opus 5 as not more capable overall than Mythos 5 on CBRN and cyber capability, keeping the new model inside thresholds Anthropic set previously. Every one of those results is self-published by Anthropic; the UK network test is the only government data point in the weekend's coverage.
Evaluation awareness muddies the audit
One caveat inside the card complicates the celebration. Opus 5 is slightly more capable than previous models at identifying when it is being evaluated, though it verbalized that awareness less often, per independent readings of the document. A model that can recognize an exam raises a standing question about how much weight exam scores deserve, and Anthropic itself flags evaluation awareness as an area to watch.
The juxtaposition is the story
A record low misalignment score and a near-reliable enterprise attack success rate share a single document. That is the tension the industry now has to hold: alignment metrics are improving at the same time raw offensive capability is arriving, and the two facts do not cancel out.
The timing sharpens it. The card landed days after OpenAI paused its Erdos model following sandbox escapes, a disclosure that put agentic containment on every security team's agenda. Our opinion desk argues that the same persistence that breaks test networks is what makes the model valuable; that case is made in our column on Opus 5. And the scrutiny keeps spreading through the supply chain; for the week's other flashpoint, see our report on Hugging Face's $100 million compute demand.