Open models are four to seven months behind the cyber frontier, and nearly free
The UK AI Security Institute measured GLM-5.2 and DeepSeek V4-Pro against frontier closed models on 70 cyber tasks. The gap is narrowing, the cost collapse is dramatic, and refusals barely hold.
- The UK AI Security Institute finds the open-weight cyber gap has narrowed to 4 to 7 months behind the closed frontier, down from 6 to 10 months through most of 2025.
- On AISI's 32-step cyber range, a 100 million token attack run cost about $85 with Opus 4.5 or 4.6, about $46 with GLM-5.2, and $1.19 with DeepSeek V4-Pro.
- AISI found the open models' safety measures largely ineffective, with refusals bypassed by simply retrying, and says it will test Kimi K3 by the end of July.
The distance between open weights and the closed frontier keeps collapsing, and the UK's evaluators have now put numbers on it. The AI Security Institute published an evaluation finding that leading open-weight models GLM-5.2 and DeepSeek V4-Pro perform on cyber tasks comparably to frontier closed models released only 4 to 7 months before them, a sharp narrowing from the 6 to 10 month gap AISI measured through most of 2025, per the institute's report, with The Decoder covering the findings.
The head-to-heads are specific. GLM-5.2, released in June, performed comparably to Anthropic's Opus 4.6 and OpenAI's GPT-5.3-Codex, both released four months earlier, across all four difficulty levels of AISI's 70-task cyber suite. DeepSeek V4-Pro matched Opus 4.5, released five months before it.
The cost table is the story
On AISI's 32-step cyber range, The Last Ones, which simulates an end to end corporate network attack, a 100 million token run cost roughly $85 with Opus 4.5 or 4.6, about $46 with GLM-5.2, and $1.19 with DeepSeek V4-Pro. Per task, that is $0.28 for V4-Pro against $12.50 for Opus 4.5. Near-frontier offensive capability is not just catching up, it is approaching free, the same economics that left 99.9 percent of scanned agent deployments unpatched in a market that cannot afford defense at attacker prices.
Safeguards that fold on the second try
AISI also tested the models' guardrails and found them "largely ineffective." DeepSeek V4-Pro occasionally refused reverse engineering tasks, but retrying the request was enough to get through. That finding lands in a week when Alibaba promised open weights for a 2.4 trillion parameter flagship and Moonshot scheduled Kimi K3's weights for July 27; AISI says it will test K3 by the end of the month. The institute's running argument with the frontier labs has been about closed models shipping past their evals. The open-weight version of that argument has no ship gate at all.