
Two watch items this blog has tracked separately for the past week just resolved into the same story. On September 5, this blog noted that GPT-6 Astra’s self-reported Terminal-Bench 4.0 score (57.7%) and Claude Fable 5.1’s self-reported score (55.8%) were both still absent from the independent leaderboard — vendor claims, not yet confirmed. As of this week, snorkel.ai’s Terminal-Bench 4.0 leaderboard lists both models with real, independently-run scores. The headline isn’t who won. It’s that the two scores are close enough that “won” is the wrong word.
The numbers#
| Rank | Model | Resolution Rate | Agent |
|---|---|---|---|
| 1 | GPT-6 Astra | 58.2% ± 2.8 | Codex |
| 2 | Fable 5.1 | 57.9% ± 3.8 | Claude Code |
| 3 | Opus 5 | 51.8% ± 3.4 | Claude Code |
| 4 | Fable 5 | 44.5% ± 3.8 | Claude Code |
| 5 | GLM-5.3 | 41.8% ± 3.2 | Claude Code |
| 6 | GPT-5.6 Sol | 37.3% ± 3.8 | Codex |
| 9 | Grok 4.6 | 20.3% ± 3.1 | Grok Build |
| 12 | Grok 4.5 | 12.4% ± 2.6 | Grok Build |
| 13 | Sonnet 5 | 12.4% ± 3.1 | Claude Code |
A 0.3-point gap between first and second place, against error margins of ±2.8 and ±3.8 respectively, means the confidence intervals overlap almost entirely. Both models land within a point of their own self-reported figures — Astra actually scored half a point higher independently than OpenAI claimed (58.2% vs. 57.7%), and Fable 5.1 came in two points higher than Anthropic’s own number (57.9% vs. 55.8%). That’s worth stating plainly, because this blog has spent months declining to repeat self-reported benchmark claims from Anthropic, OpenAI, xAI, and Meta alike without independent confirmation. This time, both labs’ numbers held up. The lesson isn’t “trust self-reported scores now” — it’s that these two happened to be accurate, which is different from being verified in advance.
What the table actually shows is a genuine coin flip between two different labs’ flagship models, run through two different agent harnesses (OpenAI’s Codex harness for Astra, Claude Code for Fable 5.1). That’s an important caveat on its own: Terminal-Bench 4.0 measures a model-plus-harness combination, not a model in isolation. Anthropic’s own Sonnet 5 ties for last place on this same board at 12.4%, run through the same Claude Code harness that put Fable 5.1 in second — a reminder, as this blog noted when the leaderboard first populated on September 2, that “Claude” isn’t one performance profile. Tier, effort setting, and harness all move the number as much as the underlying model does.
A leaderboard that won’t sit still#
One more thing worth flagging for accuracy: the leaderboard’s own metadata lists its “last updated” date as August 28, 2026 — the day Terminal-Bench 4.0 launched. GPT-6 Astra didn’t exist until September 3. Either that date field doesn’t reflect the actual row-level update history, or Snorkel’s page is serving a stale timestamp alongside live data. This is consistent with what this blog documented in detail on August 30 about the predecessor Terminal-Bench 3.0: tbench.ai’s own “Continuous Benchmarks” philosophy means the table is a living scoreboard with no fixed, citable snapshot, not a dated report you can screenshot and trust six weeks later. Anyone citing a Terminal-Bench number should note the exact date they pulled it, because the same URL can plausibly show something different next week.
The other standing disclosure applies here too: Terminal-Bench 4.0’s Open Benchmarks Grants program is funded in part by OpenAI, Anthropic, Z.ai, and SpaceX AI — three of the labs whose models occupy the top four spots on this table. That doesn’t make the numbers wrong. The Codex and Claude Code harnesses are third-party-verifiable, and the task-level methodology (66 tasks, 8-hour timeouts, saturated tasks removed) is genuinely more rigorous than the opacity that plagued Terminal-Bench 3.0 all last month. But a leaderboard whose infrastructure is funded by its own top entrants deserves the same skepticism this blog has applied to every self-reported number that preceded it.
What this actually settles, and what it doesn’t#
This resolves the specific watch item: both GPT-6 Astra’s and Claude Fable 5.1’s Terminal-Bench 4.0 claims are now independently corroborated, within a small margin of the vendors’ own numbers. It does not settle which model is “better” for real agentic coding work, because Terminal-Bench measures completion rate on a fixed task suite under a specific harness and timeout — it says nothing about Astra’s Critical-risk cybersecurity classification (covered on this blog September 5), nothing about token efficiency or real-world cost per task, and nothing about either model’s behavior outside this one benchmark’s task distribution. Grok 4.5 and 4.6 remain absent from Scale AI’s SWE-bench Pro leaderboard entirely — a separate, still-open watch item — even as they now have confirmed (low) Terminal-Bench 4.0 scores using the Grok Build agent. Different benchmark, different question, still no independent SWE-bench Pro number for xAI’s flagship coding claims.
The practical takeaway for anyone picking a model for agentic coding work this week: Astra and Fable 5.1 are close enough on this measure that the choice should come down to pricing, safety posture, and harness maturity — not a half-point leaderboard gap that sits inside the margin of error.
Sources: Terminal-Bench 4.0 leaderboard (primary, fetched directly Sept 6, 2026); Terminal-Bench 4.0 methodology (primary); this blog’s prior coverage — GPT-6 Astra Is OpenAI’s First ‘Critical’-Risk Model (Sept 5, 2026), Terminal-Bench 4.0 Got Real Numbers (Sept 2, 2026), Terminal-Bench 4.0 Replaces Unverified 3.0 (Aug 30, 2026).
