
Benchmark saturation was the story of the spring: three frontier models parked at ~80% on SWE-bench Verified and we all asked what came next. Google’s Android Bench 2.0, published September 17, is one answer, and it is humbling. The original Android Bench topped out around 91%. The new long-horizon set has a best pass rate of about 28%.
What changed#
Version 1.0 measured “incremental changes to existing repositories.” Version 2.0 targets, in Google’s words, tasks that “take an engineer multiple days or even a week to complete.” There are 30 of them across four streams, per tBreak’s summary:
- App creation: multi-screen apps built from design specifications.
- Migrations: library swaps and architecture restructuring.
- Feature additions: widgets, Picture-in-Picture and similar platform features.
- Cross-platform conversions: Flutter or React Native apps rewritten as native Android.
Scoring is no longer binary. Google reports a strict pass rate (fully completed and validated) alongside a completion rate for partial progress, built from functionality, regression checks, requirements coverage and visual fidelity. Validation goes beyond unit tests: runtime checks, database inspection, accessibility-tree analysis and multimodal visual review.
The numbers#
Per Android Central’s reporting, OpenAI’s GPT-6 Astra leads at 28.0%, with Claude Fable 5.1 reported at 22.7% and Gemini 3.8 Flash at 8%. 9to5Google confirms Astra on top and lists Kimi K3, Qwen 3.8 Max and GPT-6 among the other entrants. I could only verify the full ordering through secondary coverage, so treat the mid-table positions as provisional until you read Google’s leaderboard yourself.
Two details matter more than the ranking:
- Cross-platform conversions never hit 100%. No model fully passed them, and the leaders top out around 80% completion. That gap between “80% done” and “passes” is where every real engineering team lives.
- Harness design changed outcomes. Google evaluated agent implementations as well as bare models (for example Gemini 3.8 Flash inside its Antigravity agent) and found that “harness design positively impacts developer outcomes.” The scaffold around the model is part of the score.
Where agents fail, and why it matters#
The failure profile is consistent and instructive. Models are strong on deterministic transformations: Java-to-Kotlin, Retrofit-to-Ktor, dependency-injection framework swaps. They handle these even at scale. They are weak on:
- Runtime validation. The code compiles and looks right; the app crashes on a device path nobody exercised.
- Breaking framework changes and libraries newer than the training data.
- Modifying existing systems versus greenfield writing. New code beats refactoring.
tBreak’s phrasing of the residual problem is the one I’d frame on a wall: models frequently deliver a convincing first implementation and leave “architectural cleanup, edge cases and visual polish to a human.”
The spec-driven reading#
This is the argument for Spec-Driven Development in benchmark form. Look at what separates the 91% tasks from the 28% tasks: it isn’t raw coding ability. It’s verifiability and ambiguity. A Java-to-Kotlin conversion has an unambiguous definition of done. “Port this Flutter app so it feels native” does not, and the agent fills in the blanks with plausible guesses.
A practical response, for anyone running week-sized work through Claude Code or any terminal agent:
## Migration spec: checkout flow, Flutter -> native
### Done means
- All 14 screens in /docs/screens.md render at parity (screenshot diff < 2%)
- `./gradlew connectedCheck` passes on API 29 and API 35 emulators
- No new lint baseline entries
### Out of scope
- Payment SDK replacement (tracked separately)
### Validation the agent must run itself
- Launch on emulator, drive the 3 golden paths, attach logsExecutable acceptance criteria turn “runtime validation,” the exact category Android Bench says agents fail, into something the agent can loop against instead of guessing. That is the whole thesis: the spec is the verification harness, and long-horizon autonomy scales with how well you can write it.
The harness point cuts our way#
Google’s finding that scaffolding moves scores lines up with what Claude Code users see daily: subagents, hooks, plan mode and persistent context are not garnish. An IDE-bound assistant that asks for approval on every step cannot even attempt a week-long task; a terminal-native agent that can run emulators, read logs and iterate for hours at least gets to play. Note that Android Bench’s headline model is not Anthropic’s, and that is fine. The lesson isn’t “my model won.” It’s that 28% is the floor for a single unassisted attempt, and the gap to production is closed by specs, verification loops and harness, which is exactly where Claude Code has been investing.
Caveats#
- 30 tasks is a small sample; a couple of tasks moving flips rankings.
- The benchmark is Android-specific, and Google both authors it and fields a competing model, so read the leaderboard with that in mind.
- Pass rate is strict by design. Completion rates paint a rosier, arguably more realistic picture of how much human time an agent saves.
Takeaways#
- Stop quoting 80%+ SWE-bench numbers as evidence agents can own a week of work. On long-horizon tasks the best result is 28%.
- Invest in machine-checkable acceptance criteria before you invest in a better model.
- Evaluate harness and model together; Google’s own data says the harness matters.
- Use agents aggressively on deterministic migrations today, and keep humans on runtime validation and polish until the numbers move.
