the feed MANY MINDED · THE BRIEF
KNOWLEDGE · forward · impact 3/5 · 2026-07-28 · Google DeepMind

Google's new Flash models cut token costs while raising benchmark scores

DeepMind's Gemini 3.6 Flash and lighter siblings post higher benchmark scores at lower token usage, pushing the price of machine cognition down.

Google DeepMind announced three new models on July 21, 2026: Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber-focused 3.5 Flash Cyber. The headline claim is efficiency. According to the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash, and on the DeepSWE coding benchmark it showed up to 65% token reduction. It is priced at $1.50 per million input tokens and $7.50 per million output tokens.

The scores climbed as well. DeepMind reports 3.6 Flash improving on 3.5 Flash across several benchmarks — 49% versus 37% on DeepSWE, 63.9% versus 49.7% on MLE Bench, and 83.0% versus 78.4% on OSWorld-Verified. The cheaper 3.5 Flash-Lite, at $0.30 input and $2.50 output per million tokens, runs at 350 output tokens per second and reportedly beats a larger prior model on some coding tasks. The Cyber variant pairs with a code security agent called CodeMender.

Cheaper tokens and fewer tokens per task compound: the effective cost of a unit of useful machine reasoning keeps falling. That matters for knowledge work broadly, since lower inference costs let more applications run affordably at scale, and the security-tuned model aims at defensive code review.

The caveats are standard. Performance figures come from DeepMind's own testing and the benchmarks it cites, and customer endorsements are company-attributed. Gemini 3.5 Pro is still in partner testing and Gemini 4 has only entered pre-training.

Source: Google DeepMind