
Three days after Kimi K3 claimed the largest open-weight coding model ever released, Alibaba answered with a bigger parameter count and almost nothing else. Qwen3.8-Max-Preview, announced July 19, is a 2.4-trillion-parameter multimodal model — the Qwen team’s first to cross the 1-trillion mark with vision, video, and document understanding built in. Alibaba’s own framing: “second only to Fable 5.” There is, as of this writing, no benchmark table backing that claim.
What Alibaba Actually Disclosed#
The preview is live now through Alibaba Cloud’s Token Plan, Qoder, and QoderWork, priced at 10% of standard rates during the preview window — an unusually aggressive discount even by the open-weight sector’s race-to-the-bottom pricing norms. Alibaba says the model should outperform its predecessor, Qwen3.7-Max, “especially in coding and complex productivity tasks such as full-stack development, data analysis, and office workflows.” Open weights are promised “soon,” with no date, no named license, and no Hugging Face repository yet — a notably vaguer commitment than Moonshot gave for Kimi K3’s July 27 weight drop under a named Modified MIT license.
What’s missing is the more interesting story. Alibaba has not disclosed how many of the 2.4 trillion parameters are active per inference pass — a load-bearing detail for any mixture-of-experts model, since it’s the active-parameter count, not the total, that determines inference cost and realistic deployment hardware. Every recent Qwen coding release this blog has tracked — 3-Coder-Next’s 80B-A3B, 3.6-27B’s dense architecture, 3.6-Max-Preview’s undisclosed-but-implied sparse design — has led with an architecture story. Qwen3.8-Max led with a headline parameter count and a self-graded ranking instead. No SWE-bench Verified, no SWE-bench Pro, no Terminal-Bench, no independent leaderboard placement of any kind.
“Second Only to Fable 5” Is an Unfalsifiable Claim Today#
That phrase is doing a lot of work in Alibaba’s own materials, and it can’t currently be checked against anything. Fable 5 has a public track record on this blog: 46.3% on Cognition’s FrontierCode benchmark against Opus 4.8’s 34.3% and GPT-5.5’s 25.5%, plus whatever ground it’s ceded to Kimi K3 on Terminal-Bench 2.1 and the Arena.ai Frontend Code arena, as covered yesterday. Qwen3.8-Max has published none of those same benchmarks, so “second only to Fable 5” is a claim about relative rank on an unpublished internal test suite, made by the company selling the model. That’s not automatically false — Alibaba’s Qwen team has shipped genuinely competitive open-weight models on a near-monthly cadence for two years, and Qwen3.6-Max-Preview did back up a similar top-tier claim with real SWE-bench Pro numbers (58.4%) at launch. But the pattern this blog has flagged repeatedly this year — GPT-5.6 Sol’s METR eval-gaming, Grok 4.5’s self-tested comparison chart, Meta’s Muse Spark 1.1 Terminal-Bench gap versus Vals AI’s independent rerun — is specifically vendor-reported numbers arriving ahead of, or instead of, independent verification. Qwen3.8-Max skips even the vendor-reported-numbers step and asks for trust on a bare ranking claim alone.
The Real Story Is the Cadence, Not the Model#
Put the two releases side by side and the more useful signal isn’t either model individually — it’s how fast China’s open-weight labs are now iterating against each other. Moonshot shipped Kimi K3 on July 16 as a direct, benchmarked challenge to Fable 5. Alibaba answered three days later with a bigger total parameter count and an explicit positioning claim against the same model, timed closely enough that it reads as a direct response to Kimi K3’s momentum rather than a coincidental release-calendar overlap. That’s a meaningfully faster competitive-response cycle than the roughly monthly cadence this blog tracked through GLM-5.1 → GLM-5.2, or Kimi K2.6 → K2.7-Code earlier this year — those were each labs iterating on their own roadmap, not visibly reacting to a rival’s release inside the same week.
For teams evaluating open-weight options for cost-sensitive or on-prem agentic coding, the practical takeaway today is: do nothing yet. A model with no published benchmark, no active-parameter disclosure, and no license name attached is not something to plan a migration around, however large the total parameter count sounds. The genuinely useful comparison point remains Kimi K3, which at least shipped Terminal-Bench, DeepSWE, and Arena.ai numbers alongside its announcement — flawed and partly self-reported as some of those are — and whose actual open weights land July 27. Qwen3.8-Max is worth revisiting once Alibaba publishes a benchmark table, names a license, or the full (non-preview) model actually ships.
What to Actually Do#
Don’t factor Qwen3.8-Max into any evaluation matrix yet — there’s nothing to benchmark against. If you’re tracking the open-weight coding landscape for a near-term decision, Kimi K3’s July 27 weight release remains the nearer, better-documented milestone; this blog will cover the independent SWE-bench Pro numbers that follow it. Revisit Qwen3.8-Max once Alibaba backs “second only to Fable 5” with an actual benchmark table — until then, treat the claim as marketing, not a data point.
Sources:
- MarkTechPost — Alibaba Previews Qwen3.8-Max, a 2.4 Trillion-Parameter Multimodal Model, Days After Moonshot’s Kimi K3 Open-Weight Launch
- The Decoder — Alibaba’s Qwen takes on Kimi K3 with open-weight Qwen 3.8, says model is “second only to Fable 5”
- Officechai — Alibaba Announces 2.4 Trillion-Parameter Open-Weight Qwen 3.8, Says It’s Second Only To Fable 5
- Kimi K3: A 2.8-Trillion-Parameter Open-Weight Model Just Beat Fable 5 on Terminal-Bench — sdd.sh, 2026-07-19
