---
title: "Android Bench 2.0: The Best Model Passes 28% of Multi-Day Tasks. Your Spec Is the Missing 72%."
date: 2026-09-29
tags: ["benchmarks","android","long-horizon","spec-driven-development","harness","gpt-6","claude-fable"]
categories: ["Industry","Spec-Driven Development"]
summary: "Google's Android Bench 2.0 (Sept 17) swaps small patches for 30 week-sized engineering tasks, and the top pass rate falls from roughly 91% to 28%. The gap is a useful map of where autonomous agents still need a human-written spec and a good harness."
---


![Android Bench 2.0: The Best Model Passes 28% of Multi-Day Tasks. Your Spec Is the Missing 72%.](/images/android-bench-2-long-horizon-28-percent-spec-driven-development.png)

Benchmark saturation was the story of the spring: three frontier models parked at ~80% on SWE-bench Verified and we all asked what came next. [Google's Android Bench 2.0](https://android-developers.googleblog.com/2026/09/android-bench-2-long-horizon-tasks.html), published September 17, is one answer, and it is humbling. The original Android Bench topped out around 91%. The new long-horizon set has a best pass rate of about **28%**.

## What changed

Version 1.0 measured "incremental changes to existing repositories." Version 2.0 targets, in Google's words, tasks that "take an engineer multiple days or even a week to complete." There are 30 of them across four streams, [per tBreak's summary](https://tbreak.com/android-bench-2-0-ai-agents-week-long-app-builds/):

1. **App creation**: multi-screen apps built from design specifications.
2. **Migrations**: library swaps and architecture restructuring.
3. **Feature additions**: widgets, Picture-in-Picture and similar platform features.
4. **Cross-platform conversions**: Flutter or React Native apps rewritten as native Android.

Scoring is no longer binary. Google reports a strict **pass rate** (fully completed and validated) alongside a **completion rate** for partial progress, built from functionality, regression checks, requirements coverage and visual fidelity. Validation goes beyond unit tests: runtime checks, database inspection, accessibility-tree analysis and multimodal visual review.

## The numbers

Per [Android Central's reporting](https://www.androidcentral.com/apps-software/android-os/android-bench-2-0), OpenAI's GPT-6 Astra leads at 28.0%, with Claude Fable 5.1 reported at 22.7% and Gemini 3.8 Flash at 8%. [9to5Google](https://9to5google.com/2026/09/17/android-bench-2-0/) confirms Astra on top and lists Kimi K3, Qwen 3.8 Max and GPT-6 among the other entrants. I could only verify the full ordering through secondary coverage, so treat the mid-table positions as provisional until you read Google's leaderboard yourself.

Two details matter more than the ranking:

- **Cross-platform conversions never hit 100%.** No model fully passed them, and the leaders top out around 80% completion. That gap between "80% done" and "passes" is where every real engineering team lives.
- **Harness design changed outcomes.** Google evaluated agent implementations as well as bare models (for example Gemini 3.8 Flash inside its Antigravity agent) and found that "harness design positively impacts developer outcomes." The scaffold around the model is part of the score.

## Where agents fail, and why it matters

The failure profile is consistent and instructive. Models are strong on **deterministic transformations**: Java-to-Kotlin, Retrofit-to-Ktor, dependency-injection framework swaps. They handle these even at scale. They are weak on:

- **Runtime validation.** The code compiles and looks right; the app crashes on a device path nobody exercised.
- **Breaking framework changes** and libraries newer than the training data.
- **Modifying existing systems** versus greenfield writing. New code beats refactoring.

tBreak's phrasing of the residual problem is the one I'd frame on a wall: models frequently deliver a convincing first implementation and leave "architectural cleanup, edge cases and visual polish to a human."

## The spec-driven reading

This is the argument for Spec-Driven Development in benchmark form. Look at what separates the 91% tasks from the 28% tasks: it isn't raw coding ability. It's **verifiability and ambiguity**. A Java-to-Kotlin conversion has an unambiguous definition of done. "Port this Flutter app so it feels native" does not, and the agent fills in the blanks with plausible guesses.

A practical response, for anyone running week-sized work through Claude Code or any terminal agent:

```markdown
## Migration spec: checkout flow, Flutter -> native
### Done means
- All 14 screens in /docs/screens.md render at parity (screenshot diff < 2%)
- `./gradlew connectedCheck` passes on API 29 and API 35 emulators
- No new lint baseline entries
### Out of scope
- Payment SDK replacement (tracked separately)
### Validation the agent must run itself
- Launch on emulator, drive the 3 golden paths, attach logs
```

Executable acceptance criteria turn "runtime validation," the exact category Android Bench says agents fail, into something the agent can loop against instead of guessing. That is the whole thesis: the spec is the verification harness, and long-horizon autonomy scales with how well you can write it.

## The harness point cuts our way

Google's finding that scaffolding moves scores lines up with what Claude Code users see daily: subagents, hooks, plan mode and persistent context are not garnish. An IDE-bound assistant that asks for approval on every step cannot even attempt a week-long task; a terminal-native agent that can run emulators, read logs and iterate for hours at least gets to play. Note that Android Bench's headline model is not Anthropic's, and that is fine. The lesson isn't "my model won." It's that **28% is the floor for a single unassisted attempt**, and the gap to production is closed by specs, verification loops and harness, which is exactly where Claude Code has been investing.

## Caveats

- 30 tasks is a small sample; a couple of tasks moving flips rankings.
- The benchmark is Android-specific, and Google both authors it and fields a competing model, so read the leaderboard with that in mind.
- Pass rate is strict by design. Completion rates paint a rosier, arguably more realistic picture of how much human time an agent saves.

## Takeaways

1. Stop quoting 80%+ SWE-bench numbers as evidence agents can own a week of work. On long-horizon tasks the best result is 28%.
2. Invest in machine-checkable acceptance criteria before you invest in a better model.
3. Evaluate harness and model together; Google's own data says the harness matters.
4. Use agents aggressively on deterministic migrations today, and keep humans on runtime validation and polish until the numbers move.

## Sources

- [Android Bench 2.0: Pushing the frontier with challenging long-horizon tasks (Android Developers Blog)](https://android-developers.googleblog.com/2026/09/android-bench-2-long-horizon-tasks.html)
- [Android Bench methodology](https://developer.android.com/bench/methodology/2)
- [9to5Google: Android Bench 2.0](https://9to5google.com/2026/09/17/android-bench-2-0/)
- [tBreak: Android Bench 2.0 tests AI agents on week-long app builds](https://tbreak.com/android-bench-2-0-ai-agents-week-long-app-builds/)
- [Android Central: Android Bench 2.0](https://www.androidcentral.com/apps-software/android-os/android-bench-2-0)

