Agentic market $10.8B and climbing  ·  editor@gaasnews.com
Sections
HomeWhat is GaaS?PlatformsPricingGlossaryOpinionAboutContact
HomePricing & ModelsThe Price of Intelligence Splits
Pricing & Models

The price of agent intelligence split in two this week

Anthropic meters its flagship at $10 in and $50 out per million tokens. A Chinese open model does near-frontier agent work for a tenth of that. Both things are the story.

AJ
Andrew Jamerson
Founding Editor
Jul 8, 2026 · 4 min read
Illustration: the ceiling rises and the floor drops in the same news cycle. // GaaS News
TL;DR
  • Anthropic extended Claude Fable 5's subscription-included access to July 12, announced July 7. After that, Fable 5 usage runs on prepaid credits at $10 per million input tokens and $50 per million output tokens, double Opus 4.8's $5 and $25.
  • Included access covers up to 50% of weekly usage limits on Pro, Max, Team and select Enterprise plans. Prompt caching carries a 90% input discount.
  • Meanwhile Z.ai's GLM 5.2, MIT licensed, prices at $1.40 in and $4.40 out and lands within a point of Opus 4.8 on agentic coding benchmarks.
  • Vercel reports GLM 5.2's daily token volume grew roughly 27x in its first full week.

The cost curve under every agent business bent twice this week, in opposite directions. At the top, Anthropic extended included access to Claude Fable 5, its most capable model, through July 12. When the window closes, Fable 5 usage moves to prepaid credits at $10 per million input tokens and $50 per million output, the highest pricing Anthropic has published for a generally available model and exactly double Opus 4.8's $5 and $25. Subscribers on Pro, Max, Team and select Enterprise plans keep Fable 5 access up to 50% of their weekly limits, and prompt caching still cuts input costs by 90%. Anthropic's stated position, per Forbes, is that it wants Fable 5 back inside subscriptions as soon as capacity allows.

The floor fell at the same time

The same week, CNBC reported US companies routing agent workloads toward Chinese open-weight models as frontier costs climb. The poster child is Z.ai's GLM 5.2: MIT licensed, priced around $1.40 per million input and $4.40 per million output, and within a single point of Opus 4.8 on the FrontierSWE agentic coding benchmark, 74.4 against 75.1. Vercel says GLM 5.2's daily token volume grew roughly 27x in its first full week on the platform, a figure The Information also ties to Zhipu exploring its own custom chip.

Harpreet Arora, Vercel's head of agentic infrastructure, described the buying logic to CNBC: "When a task doesn't need the best model, teams are beginning to route it to the cheapest one that's good enough, and the recent wave of models coming out of China is winning that trade."

Routing is the new architecture

Put the two moves together and the shape of 2026 agent economics is visible. Frontier intelligence is scarce, capacity-rationed and metered like a utility. Good-enough intelligence is abundant and getting cheaper by the month. The winning agent vendors run both: reasoning-heavy steps on the expensive model, volume steps on the cheap one, with the router deciding in between.

The margin math under outcome pricing

For anyone selling outcomes at a fixed price, the resolution fees we covered in HubSpot's per-resolution pricing, this split is the whole profit and loss statement. Your revenue per task is fixed. Your cost per task now ranges across an order of magnitude depending on routing skill. The seat died first. Now even the model behind the agent is priced by the metered unit, at both ends of the market at once.

Sources: Forbes, TechTimes, CNBC, Investing.com.

Last fact-checked: Jul 8, 2026 by Andrew Jamerson
AJ

Andrew Jamerson

Founding Editor, GaaS News

Andrew Jamerson is the founding editor of GaaS News, covering the economics of the agent era. He started the publication to cover Agentic AI as a Service as a dedicated beat and edits every article on the site.

Be on the list when the beat breaks

One email when a platform ships, a round closes, or the ground shifts under the software stack.