Accessibility Adjustments

Use these optional tools to adjust reading and display preferences. These tools cannot resolve every accessibility barrier. Please contact the website owner if you need assistance.

  • Text adjustments
  • Content scaling 100%
  • Font size 100%
  • Line height 100%
  • Letter spacing 100%
  • Colour adjustments
  • Orientation adjustments

Google introduces Gemini 3.6 Flash and two new 3.5 Flash variants

Listen to this article

Google just cut the bill for agent loops and kept the sharpest knife behind a velvet rope. On July 21, the company launched Gemini 3.6 Flash, plus 3.5 Flash-Lite and a gated 3.5 Flash Cyber stack inside CodeMender. The pitch is blunt: production agents need fewer tokens, lower latency, and reliability that survives multi-step tool use.

What 3.6 Flash changes vs 3.5 on tokens and price

3.6 Flash is the new workhorse. Google says it improves coding, knowledge work, and multimodal performance while burning fewer output tokens than 3.5 Flash. On the Artificial Analysis Index, it cites a 17% cut in output tokens; on Datacurve’s DeepSWE, drops as high as 65% in some runs. Price: $1.50 per million input tokens and $7.50 per million output, billed as cheaper than 3.5 Flash per agentic task.

Google also cites DeepSWE (49% vs 37%), MLE Bench (63.9% vs 49.7%), OSWorld-Verified (83.0% vs 78.4%), and GDPval-AA v2 (1421 vs 1349). Computer use is now a built-in client-side tool in the Gemini API and Gemini Enterprise. Safety notes stress harder CBRN and cyber-offense refusals. Vendor numbers, not a third-party audit โ€” but the intent is clear: same agent jobs, shorter traces, lower unit cost.

Flash-Lite’s 350 tok/s pitch for cheap agents

3.5 Flash-Lite is the throughput play. Artificial Analysis clocks it at 350 output tokens per second, priced at $0.30 / $2.50 per million input/output. Google aims it at high-volume agentic search, document processing, and subagent swarms where latency and cost dominate.

Thinking levels let builders trade depth for speed. Versus 3.1 Flash-Lite, Google cites Terminal-Bench 2.1 (54% vs 31%), GDM-MRCR v2 (72.2% vs 60.1%), and GDPval-AA v2 (1140 vs 642). On some evals Flash-Lite even beats older 3 Flash, including SWE-Bench Pro (54.2% vs 49.6%) and OSWorld-Verified (74.0% vs 65.1%). The intended stack: 3.6 Flash as planner, Flash-Lite as cheap swarm.

Why Flash Cyber stays in a limited pilot

3.5 Flash Cyber is fine-tuned on 3.5 Flash for finding and fixing vulnerabilities at a lower price per token than larger models. Inside CodeMender, multiple Cyber agents collaborate on a single report; Google says the combo is competitive on CyberGym.

Dual-use is the constraint. The model will be available only to governments and trusted partners through CodeMender in a limited-access pilot. Defenders get a head start; broad public access stays off the table. Google will sell cheaper general agents today, and sell cyber capacity only to people it already trusts. Competitors who ship open cyber tooling will move faster โ€” and take more blame when something escapes.

Gemini 3.5 Pro testing and Gemini 4 pre-training notes

Beyond Flash, Google says Gemini 3.5 Pro is testing with partners and will go broad when ready. In parallel, the team has started its “most ambitious” pre-training run yet for Gemini 4. No ship date, no eval table โ€” just a signal that the next training wave is already burning compute while Flash absorbs the agent economics fight.

3.6 Flash and 3.5 Flash-Lite are live in the Gemini API, AI Studio, Android Studio, Gemini Enterprise, and the Gemini app; Flash-Lite is also rolling into Search. Analysis: the release is less about one IQ jump than about making agent fleets cheap enough to run all day. Cyber stays gated because automated patching and exploit hunting are the same capability with different customers. Builders get cheaper loops. Defenders get a pilot.

Maya Chen
Maya Chen

Maya Chen is focused on covering AI models, research, and the evidence behind new capabilities. Maya follows model launches, benchmarks, open weights, and scientific uses of AI with one question in mind. What changed, and how would we know? The voice is curious and exacting, with a soft spot for elegant technical ideas and little patience for a leaderboard without context.

Leave a Reply

Your email address will not be published. Required fields are marked *

Gravatar profile