Image credits: Google
![]()
Google just shipped three new AI models in a single announcement, and buried inside the blog post is a line that matters more than any benchmark chart: the company has started pretraining Gemini 4, calling it its most ambitious pretraining run yet. That’s a bigger deal than it sounds, and I think most coverage buried the lede.
The headline release is Gemini 3.6 Flash, Google’s new workhorse model built on top of Gemini 3.5 Flash from May. It’s been barely two months since that release, and Google’s already moving on. According to Google’s own announcement, 3.6 Flash reduces output token usage by 17% compared to 3.5 Flash based on the Artificial Analysis Index, with some benchmarks like DeepSWE showing reductions up to 65%. That’s not a small tweak. It’s the kind of efficiency jump that actually changes how much a company pays to run agents at scale.
Pricing backs that up too. Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens, down from $9 per million output tokens for 3.5 Flash. On the coding side, the model delivers higher precision with fewer unwanted code edits and reduced execution loops than its predecessor. I’ve been following Google’s Flash lineup for a while, and honestly, this is the fastest cadence I’ve seen them push updates at.
Gemini 3.5 Flash-Lite steals the show on speed
The second release is where things get genuinely interesting for anyone building high-volume apps. Gemini 3.5 Flash Lite generates around 350 output tokens per second, according to artificial analysis, making it the fastest model in the Gemini 3.5 lineup, and it’s priced at just $0.30 per million input tokens and $2.50 per million output tokens.
What I find interesting here is how aggressively Google is positioning this model against its own older releases, not just competitors. Compared to 3 Flash, it shows real gains on Terminal-Bench 2.1 at 54% versus 31%, on long context tasks measured by GDM-MRCR v2 at 72.2% versus 60.1%, and on real-world task execution via GDPval-AA v2 at 1140 versus 642. Those aren’t cherry-picked marginal wins. That’s a genuine generational jump for a model meant to be cheap and disposable at scale.
After digging into this more closely, I can tell you the rollout is broad too. Gemini 3.6 Flash and 3.5 Flash-Lite are rolling out to the Gemini app, Google Search, Google AI Studio, Android Studio, and the Gemini Enterprise app, so this isn’t a developer-only preview. Regular users are getting touched by this update whether they notice it or not.
The third model nobody outside of enterprise circles will use
Tucked alongside the two consumer-facing releases is Gemini 3.5 Flash Cyber, a security-focused model built specifically to hunt vulnerabilities. Sources suggest this is Google’s answer to the cybersecurity-focused models rivals have been shipping lately. Access to Gemini 3.5 Flash Cyber is currently restricted to governments and trusted partners through a limited-access pilot program, and the reasoning is straightforward: a model this good at finding security holes is also a model that’s good at finding ways to exploit them.
The numbers on this one are wild, honestly. Flash Cyber found 55 unique V8 issues in testing, compared to 47 for 3.5 Flash and just 36 for a leading rival’s flagship model. What most articles missed is that this quietly signals Google isn’t just racing on chatbot quality anymore. It’s a race on who can automate vulnerability discovery fastest, and that’s a much scarier competition to watch from the outside.
Where’s Gemini 3.5 Pro?
Here’s what’s interesting and probably the real story behind today’s announcement. Google promised Gemini 3.5 Pro back at its I/O conference in May, and that release still hasn’t happened. Reports point to internal delays tied to the Pro model underperforming on internal coding benchmarks, forcing engineering teams back to the drawing board for retraining. Industry insiders hint that today’s Flash trio exists partly to paper over that gap and keep Google looking competitive while the flagship model stays stuck in partner testing.
I actually think this is the right call from a business standpoint, even if it’s a little awkward optically. Shipping cheaper, faster models while the big one cooks keeps developers engaged instead of drifting toward competitors. And there’s real pressure to stay visible right now. In the months since Gemini Pro was last updated in February, OpenAI released GPT-5.5 and began rolling out GPT-5.6, while Anthropic launched Claude Opus 4.8 and Claude Sonnet 5 and expanded access to its Fable 5 model. That’s an intense release pace across the whole industry, and Google can’t afford to look quiet for six straight months.
Gemini 4 is already cooking
Now for the part I think deserves more attention than it got. If the current trajectory holds, Gemini 4 could end up being Google’s most significant model launch since the original Gemini debuted. Google says it has started training Gemini 4, describing it as its most ambitious pretraining run yet, a note that came both in the official blog post and in a confirming post from a senior Google AI staffer on X.
Community reaction has been split. Some see it as a meaningful milestone for Google’s next flagship, while skeptics argue that starting pretraining now only underscores how far away Gemini 4 actually is, especially with 3.5 Pro still unfinished. This is one of those things I genuinely got excited about the moment I saw it, but I’ll admit the skeptics have a point. Pretraining runs for frontier models take months, sometimes longer, before anything close to a public release shows up.
What does this mean for the next six to twelve months? Expect Gemini 3.5 Pro to finally land at some point, likely positioned as a stopgap flagship while Gemini 4 works through its training pipeline in the background. Developers building on Flash-tier models right now are getting real, tangible savings today. The bigger flagship war is still very much unwritten, and Google just told everyone it’s starting from scratch to fight it.