the feed MANY MINDED · THE BRIEF
KNOWLEDGE · forward · impact 3/5 · 2026-07-24 · Google DeepMind

Google's new Gemini Flash models cut the price of frontier AI again

DeepMind's latest Flash models post higher benchmark scores while trimming token use and cost, per the company's own figures.

Google DeepMind announced three new models on July 21, 2026: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a cyber-specialized Gemini 3.5 Flash Cyber. The headline for the flagship of the batch, 3.6 Flash, is efficiency: the company says it uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65% fewer on the DeepSWE coding benchmark, while scoring higher across the board — 49% versus 37% on DeepSWE, 63.9% versus 49.7% on MLE Bench, and 83.0% versus 78.4% on OSWorld-Verified.

Pricing is where the abundance story lands. Gemini 3.6 Flash runs at $1.50 per million input tokens and $7.50 per million output. The smaller Flash-Lite is far cheaper — $0.30 input and $2.50 output — and DeepMind reports it now beats larger prior-generation models on several tasks while running at 350 output tokens per second.

Because fewer tokens do more work at lower per-token prices, the effective cost of a given task keeps falling. That is the mechanism steadily pushing intelligence toward a commodity: automation, drafting, coding, and analysis get cheaper for anyone with an API key.

The caveats are real. All benchmark figures come from Google and cite the vendor's own or third-party indices, not independent evaluation, and named customers are company-cited. DeepMind also says it has begun pre-training Gemini 4, so the cadence isn't slowing.

Source: Google DeepMind