Google's new Flash models cut token costs while raising benchmark scores
Google DeepMind announced three new models on July 21, 2026: Gemini 3.6 Flash, 3.5 Flash-Lite, and a cyber-focused 3.5 Flash Cyber. The headline claim is efficiency. According to the Artificial Analysis Index, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash, and on the DeepSWE coding benchmark it showed up to 65% token reduction. It is priced at $1.50 per million input tokens and $7.50 per million output tokens.
The scores climbed as well. DeepMind reports 3.6 Flash improving on 3.5 Flash across several benchmarks — 49% versus 37% on DeepSWE, 63.9% versus 49.7% on MLE Bench, and 83.0% versus 78.4% on OSWorld-Verified. The cheaper 3.5 Flash-Lite, at $0.30 input and $2.50 output per million tokens, runs at 350 output tokens per second and reportedly beats a larger prior model on some coding tasks. The Cyber variant pairs with a code security agent called CodeMender.
Cheaper tokens and fewer tokens per task compound: the effective cost of a unit of useful machine reasoning keeps falling. That matters for knowledge work broadly, since lower inference costs let more applications run affordably at scale, and the security-tuned model aims at defensive code review.
The caveats are standard. Performance figures come from DeepMind's own testing and the benchmarks it cites, and customer endorsements are company-attributed. Gemini 3.5 Pro is still in partner testing and Gemini 4 has only entered pre-training.
Source: Google DeepMind
MANY MINDED