Skip to content
← All news
4 min read

Alibaba shipped Qwen3.8-Max with the benchmarks its preview did not have

A 2.4 trillion parameter model at $2 per million input tokens, with a first-ever promise to open-weight a Max class Qwen. The weights are not out yet.

Three weeks ago it launched with no numbers at all. Now there is a full benchmark table and a price.

Alibaba released Qwen3.8-Max on August 3, 2026: a 2.4 trillion parameter mixture of experts model with roughly 95 billion active parameters per token, a 1 million token context window, and native text, image, and video input. API pricing is $2.00 per million input tokens and $6.00 per million output, with cached input at $0.25. When we wrote about this model in July, none of those numbers existed.

What the preview was missing

The July preview was notable for what it did not include. Alibaba claimed the model trailed only Claude Fable 5 and published not a single benchmark to support it, which is why we said at the time it was worth tracking specifically to see whether the claim survived contact with independent testing. The launch answers that. Alibaba now publishes OSWorld-Verified at 86.1, ahead of the numbers it cites for GPT-5.6 Sol Max at 83.2 and Gemini 3.1 Pro at 76.2, GPQA Diamond at 92.6, PaperBench at 93.0, and DeepSWE 1.1 at 56.6 against 21.6 for the previous generation. Independent evaluator Artificial Analysis scores it 58 on its Intelligence Index.

The open weights are a promise, not a download

The genuinely new commitment is that Alibaba says it will open-weight a Max class Qwen for the first time, alongside a smaller Qwen3.8-27B, and it put the release roughly a week out from launch. As of today the weights are not published, the license text is not public, and Artificial Analysis still lists the model as proprietary. One observer has reported apparent geographic restrictions in the license covering the USA, EU, UK, and Korea, which we could not confirm against any published text. Until the license lands, treat the open weights as announced rather than available.

The benchmark numbers disagree with each other

Two things are worth holding at arm's length. Terminal-Bench 2.1 is reported at 86.6 in vendor-sourced coverage and 67.4 elsewhere, a gap too large to split, so treat the high number as a claim rather than a result. And the widely shared line that Qwen3.8-Max became the best overall model on Artificial Analysis's agentic index was true for a window and then was not: it led Claude Opus Max 55.4 to 55.3, then a previously planned methodology update landed and the order flipped to Opus 59.2 against Qwen 58.4. If you are checking that ranking, check it live, because it has already moved once.

Why a build studio cares

We pick models for client builds on cost per useful result, not on launch-day positioning. A frontier-class model at $2 per million input tokens with a million token context changes that arithmetic if the numbers hold, and it changes it a great deal more if the weights actually ship under a license we can deploy. Both of those are conditional right now. What we will do is exactly what we said in July: wait for the license and for independent evals, then re-run our own comparison rather than adopting a benchmark table written by the vendor selling the model.

Next step: read Alibaba's launch post and the Artificial Analysis model page for the independent scores. If you want help choosing a model for a build on evidence rather than launch-day claims, write to us at hello@gattyworks.com.

AlibabaOpen WeightsModel SelectionQwenAlibabaOpenWeightsAgenticAILLMAIBenchmarksModelSelectionAICodingMachineLearningAI

Ready to know?

Send what you want checked or built. Fixed scope, price, and date in writing inside 24 hours, or the website or audit fee on your first project is refunded in full.

24 clock hours. Weekends included.