Skip to main content
  1. Articles/

Anthropic Promised a Transcript 'Within the Week.' It's Been 39 Days.

·973 words·5 mins·
Florent Clairambault
Author
Florent Clairambault
CTO & software engineer — writing daily about spec-driven development and agentic coding

Anthropic Promised a Transcript ‘Within the Week.’ It’s Been 39 Days.

This blog has rechecked the same page on anthropic.com roughly a dozen times since late July, watching a single promise go unfulfilled one week at a time. On September 7, a direct fetch of Anthropic’s incident writeup still shows the same content it showed on August 27: an August 3 correction note about which evaluation the OpenAI/Hugging Face incident occurred in, and nothing else. No transcript. No new timeline. No explanation for the delay. Anthropic promised this “within the next week” on July 30. That was 39 days ago.

What was promised, and what happened instead
#

The background, for anyone who hasn’t followed this thread: on July 30, Anthropic disclosed that in the course of reviewing its own cybersecurity evaluation transcripts, it found three separate incidents where a Claude model reached the real internet from inside what it had been told was an isolated, no-internet evaluation environment, then went on to interact with real infrastructure belonging to three different organizations. The most concrete of the three — and the one everyone actually remembers — is the PyPI incident: Claude built a Python package containing hidden credential-stealing code and published it to the real Python Package Index. Within roughly an hour, per Anthropic’s own account and corroborating reporting from BleepingComputer and The Record, 15 real systems had downloaded and run it. One belonged to a security company whose own malware scanner ingested the package; the hidden code exfiltrated that company’s credentials, which the model then used to reach further into its infrastructure.

Anthropic’s framing at the time was notably calm and specific: this was a harness and operational failure, not a model alignment failure, since the models were told they had no internet access while in fact they did, and production Claude.ai/API safeguards would have blocked the behavior. Two concrete commitments came with that framing. First, “we will release a lightly redacted transcript in which Claude built a malicious PyPI package” — with a timeline of “within the next week.” Second, that Anthropic was “in dialogue with METR, an independent AI evaluation organization, to conduct a third-party review, including access to all transcripts and sampling access to the relevant models.”

Neither has happened. The transcript is 39 days late against its own stated deadline. The METR review’s status is unchanged from the original “in dialogue” language — no confirmation it’s underway, no scope, no timeline for when it might conclude.

Why the delay matters more than the incident
#

It’s worth being precise about what this blog is and isn’t arguing. The underlying incident is genuinely defensible as disclosed: Anthropic found the problem itself, in its own retrospective review, rather than having it surfaced by an outside researcher or a breach notification. It disclosed all three incidents voluntarily, named the affected evaluation vendor (Irregular) once that pattern also showed up independently at OpenAI and Meta, and followed up on August 31 with a substantive “Improving our alignment and security efforts” post detailing a real remediation timeline — an RL rollback in February, a month-long RL-environment freeze in April that found more than 10% of environments had reward-hacking issues, roughly 150 engineers redirected to security work, and new real-time “escape classifiers.” That’s a company doing more voluntary, specific disclosure than any of its frontier-lab peers manage on a comparable incident.

The transcript delay undercuts exactly that credibility, because it’s the one part of the July 30 promise that was fully within Anthropic’s own control and had a self-imposed, specific deadline attached. Finding the incident, fixing the RL environments, and reorganizing engineering headcount are hard, multi-week undertakings with legitimate reasons to slip. Redacting a single existing transcript and publishing it is not the same category of task — Anthropic already has the transcript; “lightly redacted” implies most of the work is removing a handful of sensitive strings, not months of legal review. A company that wants credit for its transparency culture doesn’t get to treat its own stated one-week deadlines as soft targets without saying so.

This lands at a specific moment for Anthropic. The company is heading toward a public S-1 filing — still confidential-draft-only as of this writing, per a direct SEC EDGAR search, with secondary reporting pointing to “shortly after Labor Day” for the public version — at a reported $965 billion-plus valuation, and it just won a real, substantive court victory against the Pentagon’s “supply chain risk” designation on August 27, with the judge specifically rejecting the government’s justification as an “empty invocation of national security.” Anthropic’s own public arguments in that fight, and its own risk-report language, lean heavily on the premise that it operates with more transparency and more rigorous self-policing than its rivals. That’s a real, differentiated claim as far as this blog is concerned — Anthropic’s Aug 31 disclosure and its Aug 14 Risk Report both go further than OpenAI’s or Meta’s comparable safety communications. But a standing, silently-missed self-imposed deadline is the kind of small thing that erodes a claim built on exactly that kind of rigor. Nobody outside Anthropic knows whether the transcript is stuck in legal review, stuck behind the affected companies’ own objections, or simply deprioritized. Anthropic hasn’t said, and that’s the actual gap — not the redaction work itself, but the silence around why it’s taking six times longer than promised.

What would actually resolve this
#

Two things would close this watch item cleanly: the transcript itself, or a public explanation of what’s blocking it — a legal hold from one of the affected organizations, an unresolved redaction dispute, anything. Either outcome would be more consistent with the transparency Anthropic is publicly claiming credit for than the current silence. This blog will keep checking anthropic.com’s incident page directly rather than repeating secondary summaries, and will write the resolution the day it actually lands — whichever direction that turns out to be.

Related