Google cuts Gemini Flash output pricing with 3.6 Flash launch
A cheaper, more token-efficient Flash tier lowers the per-task cost of running production agents. The benchmark gains are Google's own.
- Google launched Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output, down from 3.5 Flash's $9 output.
- Per Google's own benchmarks, the model uses about 17 percent fewer output tokens and needs fewer reasoning steps and tool calls on multi-step workflows.
- It shipped alongside a cheaper Flash-Lite tier, and Google confirmed it has begun its "most ambitious pre-training run yet" for Gemini 4.
Google just made the middle of its model lineup cheaper to run agents on. On July 21 it launched Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens, below the prior 3.5 Flash, which ran $9 per million output. For teams paying by the token to keep production agents looping, the output cut is the line item that moves.
Cheaper twice over
The sticker price is only half of it. Per Google's own benchmarks, 3.6 Flash uses about 17 percent fewer output tokens than 3.5 Flash and needs fewer reasoning steps and tool calls on multi-step workflows. That compounds with the lower rate: you pay less per token and emit fewer tokens per task. The figures are self-reported, so treat the 17 percent as a vendor claim until independent runs confirm it, but the direction is the point. Token efficiency, not sticker price alone, is where the real agent savings hide.
The knowledge cutoff also advances from January 2025 to March 2026, which matters more for agents than for chat because agents act on the state of the world rather than describe it.
The benchmark line
Google's cited gains over 3.5 Flash: DeepSWE coding at 49 percent versus 37 percent, MLE Bench at 63.9 percent versus 49.7 percent, OSWorld-Verified computer use at 83 percent versus 78.4 percent, and GDPval-AA at 1421 versus 1349. Those are Google's numbers on Google's harness, so the usual caveat applies. Third-party trackers such as Artificial Analysis will settle how much of it survives contact with independent testing. The model is available immediately across the Gemini app, Search, Google AI Studio, Vertex AI, and Android Studio, per the launch writeup.
A two-model day
This did not ship alone. Google released it alongside a cheaper Gemini 3.5 Flash-Lite tier aimed squarely at high-throughput agentic work, which we cover separately in our Flash-Lite breakdown. The pairing is the story: Google is segmenting the mid tier by price and latency rather than shipping one Flash for everyone. It also confirmed it has begun its "most ambitious pre-training run yet" for Gemini 4, which reads as a signal that the current cuts are a holding pattern, not the ceiling.
The move fits the pattern we tracked when the 3.5 slip kicked off a price war. Each Flash release now arrives cheaper and more efficient than the last, and the competitors respond in kind.
What this means for the agent economy
The per-task cost of a production agent is set by the mid tier, not the flagship, and that floor keeps dropping. A cheaper, more token-frugal Flash means the arithmetic on running agents at scale gets easier every quarter, which is precisely the unit economics that make outcome-priced GaaS defensible. The vendors racing to the bottom on Flash pricing are, whether they mean to or not, subsidizing everyone building agents on top of them.