📊 Full opportunity report: Where Does Qwen3.8-Max Stand In AI Rankings After Latest Results? on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Alibaba has confirmed that its Qwen3.8-Max model features 2.4 trillion parameters and outperforms several competitors in key benchmarks. The model’s open weights will be available next week, marking a significant milestone in large-scale AI models.
Alibaba has officially published the benchmark results for its Qwen3.8-Max model, confirming it as a leading large language model with 2.4 trillion parameters. The company also announced that the open weights will be released next week, making it the largest open-weight model to date, and positioning it high in AI performance rankings.
On August 3, Alibaba revealed the full benchmark table for Qwen3.8-Max, which had previously been previewed in July without detailed data. The model, built on the Qwen3.5 architecture, features approximately 95 billion active parameters per query and employs sparse mixture-of-experts technology. It demonstrates top-tier performance in several benchmarks, including a score of 86.6 on Terminal-Bench 2.1, surpassing models like Claude Opus 4.8 and Fable 5, but slightly below GPT-5.6 Sol at 88.8. In paper-based benchmarks, it leads with a score of 93.0 on PaperBench and excels in multimodal and agentic tasks, such as OSWorld-Verified at 86.1 and Parametric CAD Bench at 91.5.
Alibaba’s results also show significant improvements over its previous models in agentic tasks, with the model achieving a jump from 21.6 to 56.6 in DeepSWE, and from 40.7 to 73.5 in FrontierSWE, indicating substantial progress in long-horizon agent work. The company confirmed that the 2.4 trillion parameters are sparse, with the active set around 95 billion, and that the open weights will be available next week, although the licensing terms remain unpublished.
For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.
▲ All performance figures: Alibaba’s own harnessThe claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.
“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.
“Qwen3.8 is going open-weight” describes three things with very different deployment realities.
OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.
A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.
The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.
Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.
- The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
- More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
- If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
- The 27B sibling could become the best local agent model on hardware people already own.
- Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
- The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
- “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
- Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
and it says “second only” depends entirely on which row you read.
Implications of Alibaba’s Benchmark Performance
The confirmation of Qwen3.8-Max’s 2.4 trillion parameters and its high benchmark scores mark a major milestone in large language model development. The release of open weights will allow researchers and developers to experiment with the largest openly available model, potentially accelerating innovation and competition in the AI field. The performance improvements, especially in agentic tasks, suggest advancements in AI's ability to perform complex, long-horizon reasoning, which could impact applications across industries.
However, the fact that the full licensing details remain unpublished raises questions about accessibility and commercial use. The distinction between the flagship model and the smaller, more deployable 27B version also highlights ongoing debates about scalability and practical deployment of such large models.
large language model AI training hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on Alibaba’s AI Model Development
Alibaba’s Qwen series has been under development for several years, with the company gradually revealing its capabilities through stealth previews and selective benchmarking. The initial preview of Qwen3.8-Max in July generated significant attention due to its size and performance, but lacked detailed data until the recent full benchmark release. The model’s architecture builds on the earlier Qwen3.5, featuring sparse mixture-of-experts technology, which allows for scaling to larger parameter counts while maintaining efficiency. Previous models like Kimi K3 and the anonymous “kaleb” model have contributed to Alibaba’s reputation for pushing large-scale AI models into the public eye.
The recent benchmarks, including the high scores on Terminal-Bench and PaperBench, position Alibaba’s model among the top performers globally, although it remains slightly behind GPT-5.6 Sol in some areas. The upcoming release of open weights signals an important shift toward more open AI development, contrasting with the more closed approaches of some competitors.
"We are pleased to publish the full benchmark results and will release open weights next week, enabling broader research and deployment."
— Alibaba spokesperson

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Aspects of Qwen3.8-Max’s Deployment
Details about the licensing terms for the open weights remain unpublished, raising questions about how accessible and usable the model will be for different users. It is also unclear whether the 2.4 trillion parameters will be fully activated in all deployment scenarios, given the substantial hardware requirements. Additionally, the performance of the smaller 27B version in practical, real-world applications has not yet been benchmarked or publicly detailed, leaving its utility and competitiveness uncertain.
As an affiliate, we earn on qualifying purchases.
Next Steps for Alibaba’s AI Model Ecosystem
Alibaba plans to release the open weights of Qwen3.8-Max next week, enabling researchers and developers to experiment with the model directly. The company is expected to clarify licensing terms at that time. Meanwhile, further benchmarking, especially of the 27B variant, is anticipated to evaluate its practical deployment potential. Industry analysts will closely monitor how the model performs in real-world applications and whether Alibaba’s open approach influences broader AI development trends.

AI Systems Performance Engineering: Optimizing Model Training and Inference Workloads with GPUs, CUDA, and PyTorch
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the significance of Alibaba’s Qwen3.8-Max benchmark results?
The results position Qwen3.8-Max among the top large language models, especially in multimodal and agentic tasks, and mark a milestone as the largest open-weight model to date. The upcoming open release could accelerate AI innovation.
When will the open weights of Qwen3.8-Max be available?
Alibaba announced that the open weights will be released next week, though licensing details are still to be clarified.
How does Qwen3.8-Max compare to other models like GPT-5.6 or Fable 5?
In benchmark scores, Qwen3.8-Max outperforms models like Claude Opus 4.8 and Fable 5 on several tests but remains slightly behind GPT-5.6 Sol at 88.8 on Terminal-Bench.
What are the main limitations of Qwen3.8-Max based on current data?
Its performance on software engineering benchmarks like SWE-bench Pro is still below some competitors, and details about licensing and deployment scalability are still unclear.
Source: ThorstenMeyerAI.com