Skip to content
← All news
4 min read

Xiaomi put a trillion parameter model on Hugging Face under MIT

MiMo-V2.6-Pro is a sparse mixture of experts with 1.02 trillion total parameters, 42 billion active, a one million token context, and the most permissive license in the category.

The license is the surprise, not the size. MIT on a trillion parameter model is close to unheard of.

Xiaomi's MiMo team published MiMo-V2.6 on Hugging Face on September 20, 2026. The headline model, MiMo-V2.6-Pro-RL, is a sparse mixture of experts with 1.02 trillion total parameters and 42 billion active per token, a context length of one million tokens, and native FP8 quantization. The license on the model card is MIT.

The release is a family, not a single file. MiMo-V2.6-Flash-RL runs 309 billion total parameters with 15 billion active, and there is a 9 billion parameter distill built on Qwen3.5 for people who want something that fits on one card. Xiaomi also published a technical report as a PDF alongside the weights.

The card's own benchmark table puts Pro at 71.9 on DeepSWE v1.1, 89.9 on Terminal Bench 2.1, 76.9 on Toolathlon-Verified and 94.0 on CyberGym, with comparison columns against Claude Opus 5, GPT-5.6 Sol and Claude Fable 5. Secondary coverage, including Latent Space's AINews, reports that it tops open-weight rankings and puts the reinforcement learning post-training cost at roughly 2.28 million euros for Pro and 740,000 for Flash. Those cost and ranking figures are from the writeups, not the model card, and we have not verified them against the report.

MIT is the unusual part

Size gets the headline and the license is the actual news. Most frontier-scale open-weight releases arrive under a bespoke agreement with a revenue threshold, an acceptable-use annex, or a research-only clause. Alibaba's Qwen-Image-2.1 shipped on the same day under the Qwen Research License, which The Decoder reports bars commercial use without a separate agreement. MIT has none of that: use it, modify it, sell what you build, keep the copyright notice.

Here is the catch. A permissive license on a 1.02 trillion parameter model does not make it self-hostable in any ordinary sense. The 42 billion active parameters set the compute per token, but the full weight set still has to sit in memory somewhere, and that is a multi-node GPU bill, not a card in a workstation. The license removes the legal blocker and leaves the hardware one exactly where it was. For most teams the practical artefact in this release is the 9 billion parameter distill.

Why a build studio cares

When a client asks whether they can run a model themselves, the honest answer is usually a license question disguised as an infrastructure question, and it is normally the license that kills it. MIT flips which half of that sentence is hard. What we would actually reach for here is the distill: 9 billion parameters, MIT, trained on output from a model that scores in the nineties on Terminal Bench, is a very different proposition from a 9 billion parameter model trained the usual way. That is the pattern worth watching across this whole category, because a permissively licensed big model is mostly a factory for the small models the rest of us can afford to run.

Next step: read the Pro model card and the technical report linked from it, then try the 9 billion distill against whatever classifier you are currently paying an API for. If you want help sizing that comparison, write to us at hello@gattyworks.com.

XiaomiOpen WeightsLLMSelf-HostingXiaomiMiMoOpenWeightsOpenSourceAIHuggingFaceMITLicenseLLMMixtureOfExpertsSelfHostingAICoding

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.
Book a call