Gemini 3.5 Flash-Lite pushes agent pricing below a dollar
Sub-dollar input pricing on a model built for agentic search and document processing pushes the floor for high-volume per-task agent economics.
- Google released Gemini 3.5 Flash-Lite at $0.30 per million input tokens and $2.50 output, positioned for high-throughput, low-latency agentic work.
- Per Google's benchmarks, it posts Terminal-Bench 2.1 at 54 percent versus 31 percent and SWE-Bench Pro at 54.2 percent, beating even the larger Gemini 3 Flash.
- It shipped in the same July 21 announcement as Gemini 3.6 Flash, but targets a different price and latency tier.
The cheapest model in Google's refresh is the one built for agents to hammer. In the same July 21 rollout that brought Gemini 3.6 Flash, Google released Gemini 3.5 Flash-Lite at $0.30 per million input tokens and $2.50 per million output, and positioned it explicitly for high-throughput, low-latency work including agentic search and document processing. When an agent is reading thousands of pages or firing a search loop hundreds of times, input price is the number that decides whether the workload is viable.
Same day, different tier
To be clear about the accounting: this is a distinct product from the Gemini 3.6 Flash launch we cover separately, not a rebadge of it. The two shipped together, along with a Gemini 3.5 Flash Cyber variant, but they aim at different points on the price and latency curve. Flash-Lite is the floor. Read the two pieces together and you have the shape of Google's mid-market strategy, but do not double-count them as one release.
Google framed the whole rollout as a "surprise" arriving amid delays to its larger Pro model. Shipping a cheap, fast agentic tier while the flagship slips is a telling choice about where the near-term revenue is.
The benchmark jump
Per Google's own benchmarks, Flash-Lite posts large gains over the March 3.1 Flash-Lite: Terminal-Bench 2.1 at 54 percent versus 31 percent, SWE-Bench Pro at 54.2 percent, and OSWorld-Verified at 74.0 percent versus 65.1 percent. The SWE-Bench Pro number is the eyebrow-raiser, because at 54.2 percent it beats even the larger Gemini 3 Flash's 49.6 percent, which is not how model tiers are supposed to order. That is a self-reported figure on a smaller, cheaper model, so it is exactly the claim independent testing should probe first. Trackers like Artificial Analysis will show whether it holds. The model is live in the Gemini app and Search, per 9to5Google.
The agentic floor
What makes Flash-Lite matter is the workload it was built for. Agentic search and document processing are the high-volume, low-margin tasks where token price is destiny, and $0.30 per million input tokens is the kind of number that turns a workflow from too expensive to run into a default. This is the concrete version of a trend we have tracked across the mid tier, from the Gemini 3.5 slip and the price war it triggered onward. Every quarter, the model class purpose-built for agent grunt work gets cheaper and faster at the same time.
What this means for the agent economy
Per-task economics for high-volume agents live and die on input pricing, and Flash-Lite just reset the floor. A model tuned for search and document loops at $0.30 per million input tokens makes a whole category of previously marginal agent workloads pencil out, which is where GaaS margins actually get made. The flashy flagship benchmarks get the headlines, but this is the tier that quietly decides which agent businesses are solvent.