---
title: "The Lab Anthropic Accused of Laundering Claude's Answers Just Shipped Its Best Model Yet"
date: 2026-09-23
tags: ["xiaomi","mimo","open-weight","distillation","benchmarks","claude","china"]
categories: ["AI Tools","Industry"]
summary: "Twelve days after Anthropic named Xiaomi in a threat report accusing it of routing requests through Claude to harvest its output, Xiaomi shipped MiMo-V2.6 — an MIT-licensed, trillion-parameter model that's now #1 on Artificial Analysis' open-weight Intelligence Index at roughly a tenth of Opus 5.5's price. Xiaomi still hasn't responded to the allegation, and its own benchmark table shows a real gap behind frontier models on the hardest current agentic tests."
---


![The Lab Anthropic Accused of Laundering Claude's Answers Just Shipped Its Best Model Yet](/images/xiaomi-mimo-v2-6-open-weight-launch-post-distillation-report.png)

[Twelve days ago, this blog covered Anthropic's September threat intelligence report](/2026/09/anthropic-moonshot-deepseek-xiaomi-distillation-report/), which accused Moonshot of routing nearly 300,000 requests through 5,380 fraudulent accounts to harvest Claude's answers and reasoning traces under the Kimi brand — and named DeepSeek and Xiaomi as engaged in similar behavior. None of the three companies has directly rebutted the specific numbers. On September 22, Xiaomi answered the news cycle in the only language that actually moves the conversation: it shipped a model.

MiMo-V2.6 — Pro and Flash variants, MIT licensed, [announced on Xiaomi's own MiMo blog](https://mimo.xiaomi.com/mimo-v2-6) — is a 1.02 trillion parameter (42B active) sparse mixture-of-experts model with a 1M-token context window and native text, image, video, and audio support. Within the same news cycle, [Artificial Analysis](https://artificialanalysis.ai/models/mimo-v2-6-pro) put MiMo-V2.6 Pro at the top of its open-weight Intelligence Index — a score of 46, ranked #1 of 114 comparably sized open-weight models, well clear of the category median of 18.

## The price gap is the real headline

Pro runs $0.435 input / $0.87 output per million tokens — roughly a tenth of [Claude Opus 5.5's $4/$20](/2026/09/claude-opus-5-5-launch-cheaper-faster-terminal-bench/), which shipped the same day. Flash is cheaper still, at $0.14/$0.28. A "Pro-UltraSpeed" tier runs $4.35/$8.70 for lower latency. Xiaomi says the Pro checkpoint's training run cost roughly $2.6 million and finished in under six days; Flash reportedly cost under $900,000 on the same timeline. Whether or not those figures capture the full cost of the underlying infrastructure and research effort that made a six-day run possible, the sticker price alone is the kind of number that keeps showing up in this blog's coverage of open-weight Chinese labs — Kimi K3, Qwen3.8-Max, and Z.ai's GLM-5.3 have all made similar plays this year, undercutting Western frontier pricing by an order of magnitude while claiming benchmark parity.

## What Xiaomi's own numbers actually say

Model cards from labs racing to top an aggregate leaderboard are worth reading past the headline score, and MiMo-V2.6's own published table is more honest than most about where the gap actually is. Xiaomi's evaluation compares MiMo-V2.6 directly against Claude Opus 5, GPT-5.6 Sol, and Claude Fable 5:

| Benchmark | MiMo-V2.6 Pro | Claude Opus 5 | GPT-5.6 Sol | Claude Fable 5 |
|---|---|---|---|---|
| DeepSWE v1.1 | 71.9 | 74.0 | 73.0 | 70.0 |
| Terminal Bench 4.0 | 34.9 | 49.0 | 39.9 | 42.4 |
| Terminal Bench 2.1 | 89.9 | 89.1 | 88.8 | 84.3 |
| ExploitGym | 17.8 | 22.1 | 30.3 | 28.4 |
| ExploitBench | 47.9 | 70.0 | 78.5 | 78.0 |

The pattern is consistent: on the older, easier Terminal-Bench 2.1, MiMo-V2.6 is essentially tied with — even slightly ahead of — Opus 5. On Terminal-Bench 4.0, the current and considerably harder version of the same benchmark, it trails Opus 5 by 14 points and every other listed model by several. Same story on the offensive-security benchmarks: MiMo-V2.6 lags noticeably behind Opus 5, GPT-5.6 Sol, and Fable 5 on both ExploitGym and ExploitBench. An aggregate "Intelligence Index" score built across dozens of tasks can genuinely lead the open-weight pack while the model still meaningfully underperforms frontier closed models on the specific benchmarks that matter most for agentic coding work — both things are true here at once, and Xiaomi's own table is the source for both.

One more line in that table deserves a flag rather than a pass: MiMo-V2.6 Pro scores 94.0 on "CyberGym" and 80.2 on "MiMo Cyber Bench" — its own house benchmark, with no Claude or GPT-5.6 comparison figures listed for either row. A self-authored cybersecurity benchmark with no independent comparison point is exactly the kind of number [this blog has learned to treat as a placeholder, not evidence](/2026/08/opus-5-swe-bench-pro-phantom-score/) — doubly so for a capability category where the honest, disclosed comparison rows (ExploitGym, ExploitBench) show the model well behind the frontier, not ahead of it.

## The allegation nobody's addressing

What makes this release more than a routine open-weight drop is the timing against Anthropic's own report. Xiaomi is one of the labs Anthropic says routed user requests through Claude to harvest its output and reasoning traces — the same mechanism [this blog has now tracked across Moonshot, DeepSeek, and, in June, Alibaba's Qwen team](/2026/06/anthropic-alibaba-distillation-attack-25000-accounts/). China's Commerce Ministry called the broader accusations baseless; Xiaomi itself has offered no specific rebuttal of the claim. A model card doesn't settle a distillation allegation one way or the other — training provenance isn't something a benchmark table can prove or disprove — but it's a reminder of the standing advice from the last time this blog covered this beat: a self-reported benchmark table tells you nothing about how a model was actually built, and 2026 has produced repeated instances of that provenance mattering more than the score.

## Why this belongs in the same conversation as Claude Code

None of this is an argument against using MiMo-V2.6 for the workloads it's actually good at — an MIT-licensed, trillion-parameter model at a tenth of frontier pricing is a legitimate option for high-volume, cost-sensitive agentic tasks, and its own numbers show it's genuinely competitive on several of them. The argument is narrower: "tops an aggregate leaderboard" and "matches Claude Code's default model on the hardest current agentic benchmark" are different claims, and this release is a clean example of a lab's own disclosure making that distinction for you, if you read past the top-line rank.

**Sources**: [Xiaomi MiMo blog: MiMo-V2.6](https://mimo.xiaomi.com/mimo-v2-6) · [Hugging Face model card: MiMo-V2.6-Pro-RL](https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL) · [Artificial Analysis: MiMo-V2.6-Pro](https://artificialanalysis.ai/models/mimo-v2-6-pro) · [The Next Web, Sept 22, 2026](https://thenextweb.com/news/xiaomi-mimo-v2-6-open-weight-model-anthropic-distillation) · This blog's prior coverage of [Anthropic's September distillation report](/2026/09/anthropic-moonshot-deepseek-xiaomi-distillation-report/)

