
Meta shipped Muse Spark 1.3 on September 2, its third coding-focused update to the Muse line since April’s closed-source pivot. The official announcement leads with efficiency gains — roughly 20% fewer tool calls and 25% fewer tokens per task than Muse Spark 1.2 — and Mark Zuckerberg’s own framing, per VentureBeat, calls it “frontier performance almost too cheap to meter.” Artificial Analysis’s independent numbers tell a more complicated story: the model that’s actually available to developers ties for its performance tier rather than leading it, and it got measurably more expensive to run per task than its predecessor, even though the per-token price didn’t move.
Two models, one name#
Muse Spark 1.3 ships in two configurations. The one anyone can actually call through the Meta Model API or Muse Code is “xhigh.” A second, higher-reasoning configuration called “max reasoning” exists and has published benchmark numbers, but per Artificial Analysis’s own independent review, it remains in “limited preview for Meta’s partners” — not available through any public API, with no announced access path or date for general release.
The Artificial Analysis Intelligence Index numbers, run independently rather than taken from Meta’s own materials, show why the distinction matters:
| Model | Intelligence Index | Availability |
|---|---|---|
| Claude Fable 5.1 (max) | 66 | Public |
| Claude Opus 5 (max) | 63 | Public |
| Muse Spark 1.3 (max reasoning) | 62 | Limited partner preview only |
| Muse Spark 1.3 (xhigh) | 61 | Public |
| GPT-5.6 Sol (max) | 61 | Public |
| Grok 4.6 (high) | 61 | Public |
The publicly available xhigh configuration ties GPT-5.6 Sol and Grok 4.6 at 61 — a respectable result, but a tie, not the “frontier performance” the launch framing implies. The number that would actually put Muse Spark 1.3 in clear second place behind Fable 5.1 — 62, ahead of GPT-5.6 Sol and Grok 4.6 by a full point — belongs to the max reasoning configuration nobody outside Meta’s partner program can call. This is the same pattern this blog flagged when Meta launched Muse Spark 1.1 in July, where an 11-point gap opened up between Meta’s self-reported Terminal-Bench 2.1 score and Vals AI’s independent rerun: the number in the launch post and the number a third party can actually reproduce keep landing in different places.
The efficiency claim doesn’t survive the cost math either#
Meta’s own headline — 20% fewer tool calls, 25% fewer tokens per task — describes a relative comparison inside Meta’s own benchmark suite. Artificial Analysis’s independent cost-per-task figures, run against its own Intelligence Index task set, show the opposite direction: the cost of completing an average task on the publicly available xhigh configuration rose from $0.40 on Muse Spark 1.2 to $0.55 on Muse Spark 1.3 — a 37.5% increase — driven by roughly 57% more input tokens per task on agentic evaluations, with output tokens up only modestly. Per-token pricing didn’t change at all: $1.25/$4.25 per million input/output tokens, the same rate Muse Spark 1.2 shipped at in August.
Both things can be true at once, and probably are: Meta’s own task suite may genuinely show fewer tool calls and tokens on the specific workloads it measured, while Artificial Analysis’s different, independently-run task suite shows the model reasoning more (and therefore consuming more input tokens) to hit a higher score. That’s not necessarily a red flag on its own — more input-token consumption for better accuracy is a defensible tradeoff. But it directly undercuts the “almost too cheap to meter” framing Zuckerberg attached to the launch, and it means the actual dollar cost of running Muse Spark 1.3 against a real, third-party-measured workload went up, not down, generation over generation. Muse Spark 1.3 is still the cheapest model in its performance tier by a wide margin — $0.55 per task against GPT-5.6 Sol’s $0.95 and Grok 4.6’s $0.94, per Artificial Analysis’s own comparison — so the value case survives. The “cheaper than before” claim does not.
The Contributor tier is still the more interesting story#
Underneath both configurations sits the same tradeoff this blog covered when Muse Spark 1.2 launched the Contributor tier in August: developers can run the identical publicly available model at $0.10 input/$0.20 output per million tokens — roughly 12.5x cheaper on input and over 20x cheaper on output than the standard $1.25/$4.25 rate — in exchange for granting Meta rights to train on their prompts and completions. Meta told reporters the Contributor tier is seeing “meaningful double digit” adoption among developers using the API, which means a real, non-trivial share of Muse Spark 1.3 usage is now feeding Meta’s own training pipeline in exchange for a steep discount. That tradeoff structure carried over unchanged from 1.2 to 1.3, and it remains the more consequential design decision in this release than either the Intelligence Index score or the efficiency claims — a growing share of a Meta Muse Spark developer’s code and prompts, not just their subscription fee, is the actual price of the cheap tier.
The pattern holding across the Muse Spark line#
Three releases in, a consistent shape has emerged. Muse Spark 1.1 launched with a headline Terminal-Bench number that an independent lab couldn’t reproduce within 11 points. Muse Spark 1.2 tightened that gap to roughly 3 points and shipped a real, useful terminal agent (Muse Code) alongside it. Muse Spark 1.3 doesn’t have a reproducibility gap in the same sense — Artificial Analysis’s numbers aren’t contradicting Meta’s benchmark claims, they’re measuring a different, gated configuration than the one Meta’s marketing leads with. That’s a subtler problem than an outright benchmark miss, but it’s the same underlying habit: leading with the best number Meta has, whether or not it’s the number a developer can actually get. Claude Fable 5.1 tops the same Artificial Analysis Intelligence Index at a publicly available 66, with no gated “max reasoning” tier held back for partners — the number Anthropic advertises is the number anyone with an API key can call today. That’s the more durable kind of credibility, and it’s the standard this blog will keep holding every lab’s launch claims to, Meta’s included.
