
What Opus 5 actually is#
Opus 5 slots into the Opus tier at unchanged pricing: $5 per million input tokens, $25 per million output tokens — identical to Opus 4.8. A Fast mode runs at 2x the base price for roughly 2.5x the speed, aimed at high-volume agentic workloads where latency, not cost, is the binding constraint.
Anthropic’s own framing is unusually direct about where Opus 5 sits in the lineup: “a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price.” Fable 5 costs $10/$50 per million tokens. Opus 5 is being sold explicitly as the efficient alternative to Anthropic’s own flagship, not as a replacement for it.
The benchmark story backs that positioning up. On CursorBench 3.2 at max effort, Opus 5 lands within 0.5% of Fable 5 — at half the cost per task. On OSWorld 2.0, it outperforms Fable 5 outright, at a third of Fable 5’s best-case cost. On Frontier-Bench v0.1, Anthropic’s newer coding-and-knowledge-work eval, Opus 5 more than doubles Opus 4.8’s score and claims state-of-the-art. It also tops GDPval-AA v2 and posts a 3x lead over the next-best model on ARC-AGI 3, plus a 1.5x lead on Zapier’s AutomationBench at matched cost.
One number is conspicuously absent from Anthropic’s own announcement: SWE-bench Verified. Third-party trackers (BenchLM’s July leaderboard, cross-checked against several outlet reports) put Opus 5 at roughly 96% Verified and around 79% on SWE-bench Pro — competitive, but Anthropic didn’t lead with it. That’s consistent with something this blog flagged back in April: the SWE-bench-plateau problem. When five frontier labs are all clustered in the mid-to-high 90s on Verified, the benchmark stops discriminating between models, and vendors know it. Frontier-Bench, CursorBench, and OSWorld 2.0 are what a model release looks like once the old scoreboard has saturated.
Against Mythos 5 — Anthropic’s restricted, ~100-org cybersecurity-focused tier — Opus 5 stays deliberately behind: weaker on offensive cybersecurity exploitation and biology research, though close on vulnerability identification. That gap is a feature, not an oversight. Anthropic is drawing the line between “generally available frontier coding model” and “restricted dual-use capability” more sharply than any prior release, and pricing/positioning both models to keep that line intact.
The safety framing: fewer refusals, not fewer guardrails#
Opus 5 ships with cybersecurity classifiers roughly 85% less restrictive than Fable 5’s — a real loosening, aimed at the false-positive problem that’s dogged Fable 5 since launch (this blog covered Fable 5’s Cyber Jailbreak Severity framework routing legitimate debugging queries to Opus 4.8 back in July). The model can now do vulnerability finding without tripping the same guardrails, while binary scanning, penetration testing, and exploit generation stay blocked. A new Cyber Verification Program lets enterprises and researchers get elevated access with attestation.
Anthropic also claims Opus 5 is its “most aligned model to date” per internal automated behavioral audits, adhering to Claude’s Constitution more consistently than Opus 4.8, Sonnet 5, or Fable 5. Take vendor self-grading for what it’s worth — but the direction (loosen refusals selectively while tightening the underlying alignment target) is the harder, more interesting engineering problem than simply cranking capability up or down uniformly, and it’s the same pattern Anthropic used to walk back Fable 5’s over-restrictive cyber classifier in July.
Claude Code: default within 24 hours#
Claude Code v2.1.219 shipped the same day as Opus 5 and made it the default Opus model immediately — no staged rollout, no opt-in flag. The release also bundles:
- 1M-token context window on
claude-opus-5, Fast mode priced at $10/$50 per million tokens sandbox.network.strictAllowlist— a new setting that silently denies non-allowlisted hosts for sandboxed commands instead of prompting, tightening the network-egress control surface this blog has tracked sincesandbox.credentialsshipped in June- Nested subagent depth increased to 3 (from 1) by default — subagents spawned at depth-2+ now forward properly through the stream-json output
DirectoryAddedhook, firing after/add-diror the SDK’sregister_repo_rootmid-session- Opus 4.7 dropped entirely from Fast mode; only Opus 5 and Opus 4.8 remain eligible
A same-day maintenance release, v2.1.220, followed on July 25 with bug fixes — meaning the “sprint pause” this blog flagged as a watch item on July 24 lasted exactly zero extra days.
The tell in the launch quotes#
The part of this release worth dwelling on isn’t the benchmark table — it’s who Anthropic put in the press materials. Devin’s Scott Wu praised Opus 5’s debugging and root-cause-analysis strength. JetBrains’ Denis Shiryaev called out its planning judgment. And Cursor’s CTO, Sualeh Asif, is quoted directly: “Claude Opus 5 delivers near Fable 5 intelligence at Opus speed and cost. We are excited to see how developers use it.”
That’s Cursor — the IDE this blog has spent a year arguing is a wrapper around someone else’s intelligence, not an autonomous engine of its own — publicly celebrating an Anthropic model upgrade on Anthropic’s own announcement page. It’s a small data point, but it’s the cleanest evidence yet for the model-layer-versus-harness-layer thesis: Cursor, Devin, and JetBrains all ship products whose ceiling moves when Anthropic ships a model, and all three said so, on the record, the same day. Claude Code doesn’t need someone else’s harness to benefit from Opus 5 — it had the new default running in production before the ink on the newsroom post was dry.
Sources: Anthropic — Introducing Claude Opus 5, Claude Code changelog, Fortune — Anthropic debuts Claude Opus 5, CNBC — Claude Opus 5 rivals Fable 5, cheaper, BenchLM SWE-bench Verified Leaderboard, July 2026, MarkTechPost — Meet the New Claude Opus 5
