---
title: "Google Ships Three New Gemini Models. The One Everyone's Waiting For Still Isn't One of Them."
date: 2026-07-22
tags: ["gemini","google","deepmind","benchmarks","industry","claude-code"]
categories: ["Industry"]
summary: "Google shipped Gemini 3.6 Flash, 3.5 Flash-Lite, and a government-only 3.5 Flash Cyber model on July 21 — but Gemini 3.5 Pro, promised for June and still nowhere, wasn't among them. Independent benchmarking firm Artificial Analysis found the flagship-adjacent 3.6 Flash didn't actually improve on its own Intelligence Index, even as Google's own coding-specific numbers (DeepSWE, SWE-bench Pro) looked genuinely better."
---


![Google Ships Three New Gemini Models. The One Everyone's Waiting For Still Isn't One of Them.](/images/gemini-flash-trio-no-pro-still-missing.png)

Google released three new Gemini models on July 21. None of them is Gemini 3.5 Pro. That's the fourth missed window for the model Google itself said in May was "already being used internally" and would ship "next month" — and the gap between what Google shipped this week and what it didn't is a useful lens on how the coding-model race actually works right now.

## What Actually Shipped

**Gemini 3.6 Flash** is positioned as Google's new "workhorse" — the model most developers will actually touch day to day. It cuts output token usage by roughly 17% versus 3.5 Flash, dropping the effective cost of a task even though the sticker price ($1.50/$7.50 per million input/output tokens) sits close to its predecessor's. On Google's own coding benchmarks the gains look real: DeepSWE jumped to 49% from 37%, MLE-Bench (ML research tasks) hit 63.9% versus 49.7%, and OSWorld-Verified (computer-use) rose to 83% from 78.4%.

**Gemini 3.5 Flash-Lite** is the budget tier, priced at $0.30/$2.50 per million tokens, and posted the more dramatic jump: Terminal-Bench 2.1 went from 31% to 54%, and it edged out the full Gemini 3 Flash on SWE-bench Pro (54.2% versus 49.6%) — a small model beating a bigger one from the same family on a hard coding benchmark, which is a genuinely interesting result if it holds up under independent scrutiny.

**Gemini 3.5 Flash Cyber** is the odd one out: a version fine-tuned specifically for finding and patching security vulnerabilities, available only to governments and "trusted partners" through a limited pilot. It's Google's answer to the same defensive-security use case Anthropic has been building out with Claude Security and [Project Glasswing](/posts/claude-mythos-preview-project-glasswing-zero-days/) — except gated behind a much narrower access list from day one, with no public benchmark numbers attached.

## The Independent Check Tells a More Complicated Story

Here's where it gets interesting for anyone deciding what to actually trust. Independent benchmarking firm Artificial Analysis ran its own evaluation and reported something Google's press materials didn't lead with: Gemini 3.5 Flash-Lite gained 11 points on Artificial Analysis's Intelligence Index, but **Gemini 3.6 Flash did not improve in intelligence over 3.5 Flash at all** — despite the real coding-specific gains Google published on DeepSWE and MLE-Bench.

That's not a contradiction so much as a reminder of something this blog keeps having to point out about vendor benchmark selection: a model can post genuine, verifiable improvement on the specific evals a company chooses to publish while showing no improvement — or even a slight regression — on a broader, independently-run measure of general capability. It's the same pattern that undercut Grok 4.5's launch claims in July and Meta's Muse Spark 1.1 numbers the same month — self-reported or narrowly-scoped benchmarks that don't fully survive contact with a third party running its own harness. Google's coding numbers here look real. The claim that 3.6 Flash is a broad step up doesn't, at least not yet.

## Still No Pro, Still No Reason That Holds Up

The actual news, though, is what's missing. Gemini 3.5 Pro — the model that handles the complex reasoning and coding work the Flash tier isn't built for — remains stuck in "partner testing." Logan Kilpatrick, DeepMind's product lead, told press the team is "currently testing Gemini 3.5 Pro with partners and hopes to land soon." That's the fourth version of essentially the same sentence since May: "next month" in May, a slip through June, a July 17 date that also came and went, and now no date attached to "soon" at all.

Bloomberg's July 16 reporting is still the closest thing to a real explanation on record: Google retrained Gemini specifically to close its coding gap in late June, and the internal evaluation of that retrain was disappointing enough that Google held the model back rather than ship it — [covered here in detail](/posts/gemini-3-5-pro-delayed-coding-retrain-disappointing/). Nothing in this week's announcement contradicts or updates that account. If anything, shipping three Flash-tier models around the hole where Pro should be reads as confirmation: Google has plenty to say about the tier it can ship, and nothing new to say about the tier it can't.

Kilpatrick's other line worth noting: DeepMind has "already started its most ambitious pre-training run yet" for Gemini 4. Read generously, that's a lab looking past a stuck release toward the next real leap. Read less generously, it's a lab hedging by talking about the model after the one that's currently failing to ship.

## The Cadence Gap Is the Real Story

None of this happens in a vacuum. The same week Google worked through its fourth Pro delay, Moonshot AI's Kimi K3 — a 2.8-trillion-parameter open-weight model launched July 16 — got popular enough that Moonshot had to pause new subscription signups within 48 hours because demand outstripped its GPU capacity. Whatever else is true about the frontier coding race in mid-2026, it isn't short on labs that can actually ship a flagship-class coding model on schedule.

Anthropic's cadence over the same stretch is the comparison that matters for anyone reading this as a Claude Code user rather than a spectator. Claude Sonnet 5 shipped June 30 as a clean default-model upgrade — no retrain drama, no partner-testing purgatory, just a new model that got better on SWE-bench Pro (63.2%, up from Sonnet 4.6's 58.1%) and went straight into every plan the same day. Fable 5 came back online July 1 after its export-control suspension. Opus 4.8 has held a public, independently-checkable SWE-bench Pro lead (69.2%) since May 28. None of that required a "hopes to land soon" quote from a product lead.

Alphabet reports Q2 2026 earnings later today, July 22, and analysts will be listening for anything more concrete on Gemini 3.5 Pro than what Kilpatrick gave press this week. Until there's an actual GA date from Google's own blog — not a leak, not a "next month," not a partner-testing update — the honest read is that the company best resourced to challenge Anthropic on coding still can't tell you when its answer is arriving.

---

**Sources:**
- [TechCrunch: Google releases three new Gemini models — but no 3.5 Pro](https://techcrunch.com/2026/07/21/google-releases-three-new-gemini-models-but-no-3-5-pro/)
- [9to5Google: Gemini 3.6 Flash launch](https://9to5google.com/2026/07/21/gemini-3-6-flash-launch/)
- [Artificial Analysis on X: Gemini 3.6 Flash and 3.5 Flash-Lite benchmark results](https://x.com/ArtificialAnlys/status/2079596244339707956)
- [Bloomberg: Google Gemini Launch Delayed as Tech Falls Short of Internal Goals (July 16)](https://www.bloomberg.com/news/articles/2026-07-16/google-gemini-launch-delayed-as-tech-falls-short-of-internal-goals)

