Agentic market $10.8B and climbing  ·  editor@gaasnews.com
Sections
HomeWhat is GaaS?PlatformsPricingGlossaryOpinionAboutContact
HomePricing & ModelsCheap Tokens, Costly Agents
Pricing & Models

Cheap tokens do not mean cheap agents, and the math says why

Last week three vendors cut prices in a day. This week the analysis landed: agent spend is volume times attempts times tokens times price, and only one of those four is falling.

AJ
Andrew Jamerson
Founding Editor
Jul 14, 2026 · 4 min read
Illustration: the rate card is not the bill. // GaaS News
TL;DR
  • Falling per-token prices do not translate into cheaper enterprise agents, a July 13 Forbes analysis argues, because agentic systems multiply token consumption through planning, retrieval, tool calls, validation, and retries.
  • The cost formula has four independent variables: task volume, attempts per task, tokens per attempt, and effective token price. Vendors only compete on the last one.
  • A Goldman Sachs forecast cited in the analysis projects token consumption multiplying 24 times between 2026 and 2030, to 120 quadrillion tokens a month.

The week after three vendors repriced agent intelligence in a single day, the more useful document arrived: a Forbes analysis by Janakiram MSV arguing that none of those cuts guarantee a smaller agent bill. The reason is structural. Agents do not consume tokens the way chatbots do; they plan, retrieve, call tools, validate, and retry, and every one of those steps multiplies consumption.

Four variables, one discount

The analysis reduces agent economics to a formula worth pinning above any procurement desk: "Total spend equals task volume, multiplied by attempts per task, multiplied by tokens per attempt, multiplied by effective token price, plus tool and infrastructure costs." The price war touches only the last multiplier. Meta's Llama API sits at $1.25 per million input tokens and $4.25 per million output; OpenAI's GPT-5.6 ladder runs from Luna at $1 and $6 up to Sol at $5 and $30. Meanwhile a Goldman Sachs forecast cited by Forbes projects token consumption growing 24 times by 2030, to 120 quadrillion tokens a month. Falling unit prices against exploding volume is how utility bills work, not how costs shrink.

Reliability is the real rate card

The sharpest line in the analysis is about failure: "A more expensive model that resolves 90% of tickets on the first attempt can cost less per resolved ticket than a cheap model that resolves 40%." Attempts per task is a pricing variable that never appears on a pricing page, and it is the one that decides whether per-resolution pricing or per-token pricing wins the enterprise. It also explains why buyers now watch capacity dramas like Anthropic's rolling Fable 5 extensions so closely: when the most reliable model's availability is uncertain, the cheap-token fallback quietly doubles the attempts column, and the discount pays for itself in retries.

AJ

Andrew Jamerson

Founding Editor, GaaS News

Andrew Jamerson is the founding editor of GaaS News, covering the economics of the agent era. He started the publication to cover Agentic AI as a Service as a dedicated beat and edits every article on the site.

Be on the list when the beat breaks

One email when a platform ships, a round closes, or the ground shifts under the software stack.