Agentic market $10.8B and climbing  ·  editor@gaasnews.com
Sections
HomeWhat is GaaS?PlatformsPricingGlossaryOpinionAboutContact
HomeEvaluation & SafetyAISI Cyber Gap
Evaluation & Safety

Open models are four to seven months behind the cyber frontier, and nearly free

The UK AI Security Institute measured GLM-5.2 and DeepSeek V4-Pro against frontier closed models on 70 cyber tasks. The gap is narrowing, the cost collapse is dramatic, and refusals barely hold.

AJ
Andrew Jamerson
Founding Editor
Jul 19, 2026 · 4 min read
Illustration: the gap on the graph keeps shrinking. // GaaS News
TL;DR
  • The UK AI Security Institute finds the open-weight cyber gap has narrowed to 4 to 7 months behind the closed frontier, down from 6 to 10 months through most of 2025.
  • On AISI's 32-step cyber range, a 100 million token attack run cost about $85 with Opus 4.5 or 4.6, about $46 with GLM-5.2, and $1.19 with DeepSeek V4-Pro.
  • AISI found the open models' safety measures largely ineffective, with refusals bypassed by simply retrying, and says it will test Kimi K3 by the end of July.

The distance between open weights and the closed frontier keeps collapsing, and the UK's evaluators have now put numbers on it. The AI Security Institute published an evaluation finding that leading open-weight models GLM-5.2 and DeepSeek V4-Pro perform on cyber tasks comparably to frontier closed models released only 4 to 7 months before them, a sharp narrowing from the 6 to 10 month gap AISI measured through most of 2025, per the institute's report, with The Decoder covering the findings.

The head-to-heads are specific. GLM-5.2, released in June, performed comparably to Anthropic's Opus 4.6 and OpenAI's GPT-5.3-Codex, both released four months earlier, across all four difficulty levels of AISI's 70-task cyber suite. DeepSeek V4-Pro matched Opus 4.5, released five months before it.

The cost table is the story

On AISI's 32-step cyber range, The Last Ones, which simulates an end to end corporate network attack, a 100 million token run cost roughly $85 with Opus 4.5 or 4.6, about $46 with GLM-5.2, and $1.19 with DeepSeek V4-Pro. Per task, that is $0.28 for V4-Pro against $12.50 for Opus 4.5. Near-frontier offensive capability is not just catching up, it is approaching free, the same economics that left 99.9 percent of scanned agent deployments unpatched in a market that cannot afford defense at attacker prices.

Safeguards that fold on the second try

AISI also tested the models' guardrails and found them "largely ineffective." DeepSeek V4-Pro occasionally refused reverse engineering tasks, but retrying the request was enough to get through. That finding lands in a week when Alibaba promised open weights for a 2.4 trillion parameter flagship and Moonshot scheduled Kimi K3's weights for July 27; AISI says it will test K3 by the end of the month. The institute's running argument with the frontier labs has been about closed models shipping past their evals. The open-weight version of that argument has no ship gate at all.

AJ

Andrew Jamerson

Founding Editor, GaaS News

Andrew Jamerson is the founding editor of GaaS News, covering the economics of the agent era. He started the publication to cover Agentic AI as a Service as a dedicated beat and edits every article on the site.

Be on the list when the beat breaks

One email when a platform ships, a round closes, or the ground shifts under the software stack.