Alibaba has released what it’s calling the most powerful model in its Qwen series to date — and the numbers behind it are genuinely massive. Qwen3.8-Max packs 2.4 trillion parameters, arriving just months after a separate Chinese lab, Moonshot, released its own 2.8 trillion-parameter model, Kimi K3.
What Makes Qwen3.8-Max Different
Parameters are the internal values a model learns during training that shape how it processes information and generates responses — more parameters generally (though not always) means more capability. For context, leading US labs like Anthropic don’t publicly disclose exact parameter counts, though industry experts have estimated Claude Opus 4.8 sits somewhere between 1.5 and 2 trillion parameters.
Alibaba claims Qwen3.8-Max shows particularly strong performance in autonomous coding and long-horizon tasks — meaning it can work independently on complex projects over extended periods without constant human check-ins. In internal testing, the company says the model autonomously executed a real-world software engineering project spanning 16 continuous days.
The model is also built as a multimodal system, meaning it can process more than just text — Alibaba states it can ingest hundred-page documents, entire TV series, or even 100-hour livestreams, converting them into searchable, interactive knowledge you can query afterward.
How It Stacks Up Against Competitors
Using the independent Arena.ai benchmarking platform, Alibaba reports Qwen3.8-Max ranked fifth in the Text Arena leaderboard and second in the Vision Arena leaderboard — with Anthropic’s Claude models ranking ahead of it in both categories.
A Fresh Controversy Attached to the Launch
The release lands amid an ongoing dispute between Anthropic and Chinese AI labs. In June, Anthropic’s Head of Policy formally accused Alibaba of what she called the largest known “distillation attack” on Anthropic’s models to date — distillation being a technique where one company’s AI outputs are used to help train a competing model. Anthropic argues this lets Chinese labs benefit from the enormous cost and research investment that goes into training frontier models without bearing those same costs themselves, and has previously made similar accusations against other Chinese labs including DeepSeek and Moonshot.
Alibaba has not issued a public response specifically addressing the distillation accusations tied to this release.
The Bigger Picture
This launch fits into a broader pattern industry analysts have been tracking: US AI labs are generally seen as maintaining a narrow lead on the most advanced frontier models, while Chinese labs are closing the gap faster than expected, particularly on cost-efficient models and real-world deployment speed. This dynamic sits within a wider period of economic tension between the two countries, with trade and technology disputes affecting the broader relationship.
Read More :- Meta, Microsoft, Nvidia & IBM Sign Open Letter to Protect Open-Source AI | Affitronix
Frequently Asked Questions
How many parameters does Qwen3.8-Max have?
2.4 trillion parameters, according to Alibaba, making it one of the largest publicly known AI models released to date.
Is Qwen3.8-Max better than Claude or GPT models?
On independent Arena.ai rankings, Qwen3.8-Max ranked below Anthropic’s Claude models in both text and vision benchmarks, though direct comparisons vary by specific task and benchmark used.
What is AI model distillation?
Distillation is a technique where one AI model’s outputs are used to help train or improve a second model — it’s a common practice in AI research, though its use across company or national boundaries without authorization has become a point of dispute.




