Google's new Gemini Flash models cut the price of frontier AI again
Google DeepMind announced three new models on July 21, 2026: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a cyber-specialized Gemini 3.5 Flash Cyber. The headline for the flagship of the batch, 3.6 Flash, is efficiency: the company says it uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65% fewer on the DeepSWE coding benchmark, while scoring higher across the board — 49% versus 37% on DeepSWE, 63.9% versus 49.7% on MLE Bench, and 83.0% versus 78.4% on OSWorld-Verified.
Pricing is where the abundance story lands. Gemini 3.6 Flash runs at $1.50 per million input tokens and $7.50 per million output. The smaller Flash-Lite is far cheaper — $0.30 input and $2.50 output — and DeepMind reports it now beats larger prior-generation models on several tasks while running at 350 output tokens per second.
Because fewer tokens do more work at lower per-token prices, the effective cost of a given task keeps falling. That is the mechanism steadily pushing intelligence toward a commodity: automation, drafting, coding, and analysis get cheaper for anyone with an API key.
The caveats are real. All benchmark figures come from Google and cite the vendor's own or third-party indices, not independent evaluation, and named customers are company-cited. DeepMind also says it has begun pre-training Gemini 4, so the cadence isn't slowing.
Source: Google DeepMind
MANY MINDED