↓ Skip to main content
  1. Articles/

Gemini 4 Argon Claims the Benchmark Crown. Almost Nobody Can Use It

·884 words·5 mins·
Florent Clairambault
Author
Florent Clairambault
CTO & software engineer — writing daily about spec-driven development and agentic coding

Gemini 4 Argon Claims the Benchmark Crown. Almost Nobody Can Use It

Google DeepMind announced Gemini 4 Argon on September 30, and the framing was a familiar one: a new frontier model that retakes the benchmark lead. The headline number is 77.9% on DeepSWE v1.1, against 74.2% for Claude Opus 5.5 and 74.1% for GPT-6 Astra. Read the footnotes and the story changes. This is a restricted preview, the scores are vendor-reported, and on the benchmark closest to what agents actually do all day, Argon loses to Opus.

What Google actually shipped
#

Argon is not generally available. Per the launch coverage, it is rolling out first to “a set of trusted cyber defenders” through Google’s Fairwind Program, with US government pre-release participation. Paid API customers and Google AI Ultra subscribers come next, with no published timeline.

As of the announcement, reports say there is no public model ID and no cloud listing. It was not in Vertex AI, Gemini CLI, OpenRouter, Cursor or the GitHub Copilot catalog. You cannot point your agent at it today.

Pricing is $2 per million input tokens and $10 per million output tokens during an introductory period of unpublished length, rising to $4/$20 afterwards. Cached input gets a 95% discount. Note that the $2/$10 figure is the same as GPT-6.1 Sol and Claude Sonnet 5.5, but Argon’s price is a promotion, not a list price.

The numbers, side by side
#

The figures below are as reported by VentureBeat and DataCamp from Google’s launch materials:

BenchmarkGemini 4 ArgonClaude Opus 5.5GPT-6 Astra
DeepSWE v1.177.9%74.2%74.1%
Terminal-Bench 4.057.4%66.4%58.2%
Vals Index68.9%67.0%63.1%
GraphWalks (1M tokens)84.2%66.8%71.8%
FrontierSWE v255.0%n/a65.5%

Google says Argon wins 13 of 19 benchmarks. That is a respectable claim. It is also a claim about which 19 benchmarks were chosen.

Two losses matter for developers. On FrontierSWE v2, Argon trails GPT-6 Astra by more than ten points. On Terminal-Bench 4.0, it trails Opus 5.5 by nine points. Terminal-Bench is the one that exercises what a coding agent does: navigating a shell, recovering from failed commands, and finishing long tasks without supervision. We covered why that leaderboard is the one to watch; Argon’s result reinforces it.

Why “nobody has reproduced it” matters
#

DataCamp’s write-up puts the caveat bluntly: nobody outside the Fairwind cohort has reproduced a single score. That is not a scandal. It is how restricted previews work. But it means every number above is a marketing number until independent evaluators get API access.

We have seen this pattern. Launch-day benchmark leads shrink, sometimes by a lot, once third parties run the same suites with their own harnesses. DeepSWE v1.1 is a narrow slice of software work, and a 3.7-point lead on one benchmark, from the vendor, with no independent replication, is not a reason to re-platform anything.

The cost claim deserves the same scepticism. Coverage cites $1.99 per task for Argon against $3.26 for Astra, about 60% of the cost. That is at introductory pricing, and it is a per-task figure on a benchmark. Double the list price after the promo and the saving mostly disappears.

The 1M output token detail
#

One genuinely interesting spec: Argon reportedly supports 1 million output tokens, up from a 64K limit on previous Gemini models. DataCamp notes the input context window is not specified.

If it holds up, that is a real capability for long-horizon generation: whole-repo migrations, large generated test suites. But output length is only useful if the model stays coherent across it, and no agent harness will want a single million-token generation anyway. Claude Code’s loop of plan, edit, test and compact is built to avoid exactly that failure mode. A long output window is a feature for a chat product. An agent wants short, verifiable steps.

What this means for your workflow
#

Three practical takeaways:

  1. Do not switch on a press release. If you run a spec-driven workflow, you already have the tool for evaluating a new model: your own specs and your own test suite. Run Argon against them when it ships. That tells you more than any leaderboard.
  2. Model leadership is rotating monthly. Opus 5.5, GPT-6 Astra, GPT-6.1 Sol, now Argon. Whichever lab leads this week will not lead in six weeks. That argues for a harness that is model-agnostic at the spec layer and strict at the permission layer.
  3. Terminal-native evals are the honest ones. Where a model has to operate a real shell for hours, Opus 5.5 still leads this table. Cursor and other IDE-centric tools will add Argon to their model pickers the day it ships, but a model picker is not an autonomy story. The harness is.

Credit where it is due: Google is moving fast, the 1M output window and the cyber-defender rollout are notable, and a vendor-reported win on DeepSWE is still a signal. If independent numbers confirm it, Argon will be a serious option for agentic work. Until then, it is a press release with a price tag.

Sources
#

Related