---
title: "Terminal-Bench 3.0 Never Got Verified. Terminal-Bench 4.0 Just Replaced It."
date: 2026-08-30
tags: ["terminal-bench","benchmarks","claude-code","evals","ai-coding-tools"]
categories: ["Industry"]
summary: "Terminal-Bench 4.0 shipped August 28, quietly superseding version 3.0 — a benchmark this blog spent four weeks declining to cite because its live leaderboard kept returning contradictory numbers. The primary source finally explains why: 'continuous benchmarks' don't have a fixed scoreboard, they have a timestamp."
---


![Terminal-Bench 3.0 Never Got Verified. Terminal-Bench 4.0 Just Replaced It.](/images/terminal-bench-4-0-replaces-unverified-3-0.png)

For most of August, this blog has been tracking a small, nagging discrepancy: secondary sites kept citing Claude Opus 5 at "42.7%" on Terminal-Bench 3.0, sometimes "43.5%," while the primary leaderboard itself was either unreachable, redirecting in circles, or rendering client-side in a way that resisted direct verification. Every watch-item recheck landed on the same conclusion — hold the line, don't cite a number nobody can confirm firsthand.

On August 28, the answer arrived, just not the one anyone was expecting. Terminal-Bench 3.0 didn't get fixed. It got replaced. Terminal-Bench 4.0 shipped less than a month after 3.0 launched, and the primary source now explains exactly why this blog's numbers never lined up: there was never a single, fixed "Terminal-Bench 3.0 score" to confirm. There was only a live table that kept changing underneath the people trying to screenshot it.

## What actually changed in 4.0

Per [the team's own announcement](https://www.tbench.ai/news/terminal-bench-4-0), authored by Ryan Marten on behalf of the Terminal-Bench project, version 4.0 is a maintenance release dressed up as a major version bump — which turns out to be the whole point. Three concrete changes:

- **Uniform 8-hour agent timeout.** The team calibrated task resource limits using methodology borrowed from Anthropic's own infrastructure-noise research, rather than per-task guesswork that let some agents time out on trivial variance.
- **Eight tasks removed.** Two for saturation (defined precisely: "all classes within all families of the latest generation of models solve it 5/5 times"), two over model refusals, two because solutions had leaked publicly, two for quality or compatibility bugs.
- **Nineteen tasks patched** for flakiness or misspecification, flagged either by users or by the team's own leaderboard runs.

None of that is dramatic on its own. What's notable is the versioning logic behind it. Terminal-Bench's [companion philosophy post](https://www.tbench.ai/news/continuous-benchmarks), published the same day as version 3.0 back on July 30, lays out the actual policy: patch versions don't touch scores, minor versions let old runs be re-graded without new rollouts, and major versions — like this one — mean the environment changed enough that every model has to be re-run from scratch. "Benchmarks are software and should be maintained like software," the post states. Terminal-Bench isn't a fixed exam anymore. It's a rolling release.

## Why this blog's numbers never matched

Here's the part that actually resolves the month-long watch item. Digging into the archived Terminal-Bench 3.0 launch post directly — the actual primary source, not an aggregator's re-scrape — turns up a leaderboard snapshot from the day of that launch:

| Model | Harness | Score |
|---|---|---|
| GPT-5.6 Sol | Codex | 34.4% |
| Fable 5 | Claude Code | 33.8% |
| Opus 4.8 | Claude Code | 21.1% |
| GPT-5.6 Terra | Codex | 20.8% |
| Grok 4.5 | Cursor CLI | 17.8% |
| Sonnet 5 | Claude Code | 14.6% |
| GPT-5.6 Luna | Codex | 14.3% |
| GLM 5.2 | Claude Code | 5.1% |

Notice what's missing: Opus 5. It isn't there, because Terminal-Bench 3.0 launched July 30 — before Opus 5's own run had been added to the live table. The "42.7%" figure this blog kept declining to cite almost certainly reflects a later scrape of the same URL, after Anthropic's newer flagship got tested against the same task set. Both numbers were, technically, "Terminal-Bench 3.0 scores." They just described two different moments of a table that never stopped moving.

That's not a gotcha against any single aggregator — it's a structural property of what Terminal-Bench now is. A continuously-versioned, live-updating leaderboard doesn't have a single citable state the way a dated PDF or a locked-in academic paper does. Every "current standing" claim implicitly needs a timestamp, and most outlets quoting these numbers don't provide one. This blog's own caution — treat secondary-aggregator agreement as insufficient without a direct, dated primary fetch — held up better than it might have looked over the past four weeks. It wasn't stalling. It was correctly identifying that the ground kept shifting.

## The funding consortium is the more interesting story

Buried in the fine print of both announcements is a detail worth more attention than the leaderboard churn itself. Terminal-Bench 3.0 credits **Modal, Anthropic, OpenAI, Google, Scale AI, Snorkel, Turing, gNucleus AI, Boolean AI, and Handshake AI** as sponsors. Terminal-Bench 4.0's credits shift to **OpenAI, Anthropic, Z.ai, SpaceX AI, and the Laude Institute**.

Read that second list again in light of this week's other story: OpenAI just told SpaceX it's cutting off Cursor's model access over trust concerns tied to Elon Musk's companies — and SpaceX AI is now co-funding, alongside OpenAI and Anthropic, the shared infrastructure used to grade all their agents against each other. The labs that won't sell each other API access are still perfectly happy to jointly bankroll the referee. That's not hypocrisy so much as basic self-interest: nobody wants to be graded by a benchmark only their rivals paid for, so everyone chips in, feuds notwithstanding.

## The actual lesson for anyone reading a leaderboard number

Terminal-Bench's "continuous benchmark" model is a real improvement over the old static-benchmark playbook — the one where a leak, a saturation ceiling, or a subtly broken task grader could quietly poison a number for a year before anyone noticed. Reward hacking gets caught faster when production runs keep probing the task set. Saturated tasks get retired instead of turning into a permanent free 5%. That's good methodology, and it's consistent with the same "evals as living infrastructure" thinking Anthropic has applied to its own model cards.

But it comes at a cost that benchmark-quoting coverage — including, at times, this blog's own caution-first approach — hasn't fully adapted to: there is no longer a single number to look up. There's a live table, a version, and a fetch time, and any of those three missing from a citation should be treated as a red flag. The next time a headline says "Model X scores Y% on Terminal-Bench," the honest response isn't "confirmed" or "unconfirmed." It's "as of when?"

Terminal-Bench 4.0 is barely two days old as of this writing, and no independent aggregator has yet published a stable re-run across the full model set — which means the same caution applies immediately. This blog will hold the same line it held for version 3.0: no specific 4.0 number gets cited here until it's pulled directly from a dated, live fetch of the primary leaderboard, not a secondary site's snapshot of one.

---

**Sources:**
- [Terminal-Bench 4.0 — tbench.ai](https://www.tbench.ai/news/terminal-bench-4-0)
- [Continuous Benchmarks — tbench.ai](https://www.tbench.ai/news/continuous-benchmarks)
- [Terminal-Bench 3.0 — tbench.ai](https://www.tbench.ai/news/terminal-bench-3-0)
- [Ryan Marten on X: Terminal-Bench 4.0 announcement](https://x.com/ryan_marten/status/2093523335972036657)
- [Terminal-Bench news archive](https://www.tbench.ai/news)

