Agentic market $10.8B and climbing  ·  editor@gaasnews.com
Sections
HomeWhat is GaaS?PlatformsPricingGlossaryOpinionAboutContact
HomeOpinionTwo Agent Economies
Opinion

The $50 token and the $4.40 token are building different futures

Fable 5 and GLM 5.2 sit one point apart on a coding benchmark and eleven times apart on price. Both numbers are telling you the truth.

AJ
Andrew Jamerson
Founding Editor
Jul 9, 2026 · 4 min read
Illustration: the price of intelligence, forked. // GaaS News

The cost of intelligence forked last week, and I have not stopped thinking about the shape of the fork. At the top, Anthropic priced Claude Fable 5 at 10 dollars per million input tokens and 50 per million output, exactly double its own Opus 4.8. At the bottom, Z.ai's GLM 5.2 landed at 1.40 and 4.40, MIT licensed, within a single point of Opus on the FrontierSWE agentic coding benchmark. Vercel says GLM's daily token volume grew roughly 27 times in its first full week.

You can read that as one market with a wide spread. I think that undersells what happened. Those are two different bets about what an agent is for, and they will build two different economies.

The margin math decides, not the benchmark

Start with the businesses on our beats. A GaaS vendor selling outcomes, like HubSpot at 50 cents per resolved conversation, lives or dies on the token cost underneath each outcome. At GLM prices, the tokens under a 50 cent resolution cost pennies. At Fable prices they can eat most of the coin. Outcome pricing, the business model this whole industry is converging on, quietly assumes the cheap end of the fork exists.

In practice nobody serious will run everything on either model. The winning architecture shows up in every engineering blog I read: route the bulk of the steps to the cheap model, escalate the hard steps to the frontier one, cache aggressively. Anthropic clearly knows it, which is why prompt caching carries a 90 percent discount. The 50 dollar price is not for your workload. It is for the five percent of your workload that decides whether the other 95 percent was worth doing.

The part that connects to the report card

Now the uncomfortable observation. I wrote about the FLI safety index handing the agent economy's suppliers a C+ at best. Go look up where the cheap end of this fork scored. Z.ai took a D-. DeepSeek took an F. The models winning the volume war on price are, on the only public report card we have, the least careful. Some of the discount is genuine efficiency and open weights. I would not bet that all of it is.

None of that makes open models untouchable. It makes routing a safety decision as well as a cost decision. The step where an agent touches customer money probably should not run on the F student, however good the per-token price looks that quarter.

What most people overlook is that the router is quietly becoming the product. Whoever decides which step runs on which model now controls cost, quality and risk in one place. The vendors can see it too, which is why both ends of the fork are fighting for default status inside orchestration stacks. NemoClaw pitching itself at a tenth of the cost is exactly that fight.

For two years this industry priced intelligence as if it were one thing. As of last week it is officially two things: a commodity you meter and a specialist you consult. Most working agents will be built from both. The builders who thrive will be the ones who know, step by step, which one they are paying for and why.

Opinion columns reflect the personal views of the author. Our reporting on the stories referenced here lives on the linked pages.

AJ

Andrew Jamerson

Founding Editor, GaaS News

Andrew Jamerson is the founding editor of GaaS News, covering the economics of the agent era. He started the publication to cover Agentic AI as a Service as a dedicated beat and edits every article on the site.

Be on the list when the beat breaks

One email when a platform ships, a round closes, or the ground shifts under the software stack.