Google introduces Gemini 3.6 Flash and two new 3.5 Flash variants

Google just cut the bill for agent loops and kept the sharpest knife behind a velvet rope. On July 21, the company launched Gemini 3.6 Flash, plus 3.5 Flash-Lite and a gated 3.5 Flash Cyber stack inside CodeMender. The pitch is blunt: production agents need fewer tokens, lower latency, and reliability that survives multi-step tool use.
What 3.6 Flash changes vs 3.5 on tokens and price
3.6 Flash is the new workhorse. Google says it improves coding, knowledge work, and multimodal performance while burning fewer output tokens than 3.5 Flash. On the Artificial Analysis Index, it cites a 17% cut in output tokens; on Datacurve’s DeepSWE, drops as high as 65% in some runs. Price: $1.50 per million input tokens and $7.50 per million output, billed as cheaper than 3.5 Flash per agentic task.
Google also cites DeepSWE (49% vs 37%), MLE Bench (63.9% vs 49.7%), OSWorld-Verified (83.0% vs 78.4%), and GDPval-AA v2 (1421 vs 1349). Computer use is now a built-in client-side tool in the Gemini API and Gemini Enterprise. Safety notes stress harder CBRN and cyber-offense refusals. Vendor numbers, not a third-party audit โ but the intent is clear: same agent jobs, shorter traces, lower unit cost.
Flash-Lite’s 350 tok/s pitch for cheap agents
3.5 Flash-Lite is the throughput play. Artificial Analysis clocks it at 350 output tokens per second, priced at $0.30 / $2.50 per million input/output. Google aims it at high-volume agentic search, document processing, and subagent swarms where latency and cost dominate.
Thinking levels let builders trade depth for speed. Versus 3.1 Flash-Lite, Google cites Terminal-Bench 2.1 (54% vs 31%), GDM-MRCR v2 (72.2% vs 60.1%), and GDPval-AA v2 (1140 vs 642). On some evals Flash-Lite even beats older 3 Flash, including SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%). The intended stack: 3.6 Flash as planner, Flash-Lite as cheap swarm.
Why Flash Cyber stays in a limited pilot
3.5 Flash Cyber is fine-tuned on 3.5 Flash for finding and fixing vulnerabilities at a lower price per token than larger models. Inside CodeMender, multiple Cyber agents collaborate on a single report; Google says the combo is competitive on CyberGym.
Dual-use is the constraint. The model will be available only to governments and trusted partners through CodeMender in a limited-access pilot. Defenders get a head start; broad public access stays off the table. Google will sell cheaper general agents today, and sell cyber capacity only to people it already trusts. Competitors who ship open cyber tooling will move faster โ and take more blame when something escapes.
Gemini 3.5 Pro testing and Gemini 4 pre-training notes
Beyond Flash, Google says Gemini 3.5 Pro is testing with partners and will go broad when ready. In parallel, the team has started its “most ambitious” pre-training run yet for Gemini 4. No ship date, no eval table โ just a signal that the next training wave is already burning compute while Flash absorbs the agent economics fight.
3.6 Flash and 3.5 Flash-Lite are live in the Gemini API, AI Studio, Android Studio, Gemini Enterprise, and the Gemini app; Flash-Lite is also rolling into Search. Analysis: the release is less about one IQ jump than about making agent fleets cheap enough to run all day. Cyber stays gated because automated patching and exploit hunting are the same capability with different customers. Builders get cheaper loops. Defenders get a pilot.



