Qwen3.8-Max’s AI Metrics: Breaking Down The Numbers Behind The Claims

📊 Full opportunity report: Qwen3.8-Max’s AI Metrics: Breaking Down The Numbers Behind The Claims on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba announced the full benchmark results for its Qwen3.8-Max model, confirming it has 2.4 trillion parameters and outperforms some competitors on key tests. The open weights will be available next week, marking a significant step in large-scale open AI models. However, claims about its overall superiority are based on selective benchmarks.

Alibaba has publicly released the full benchmark table and confirmed the specifications of its Qwen3.8-Max model, including its 2.4 trillion parameters and multimodal capabilities, after weeks of speculation and stealth previews. This marks the first time the model’s detailed performance metrics are openly available, with open weights scheduled to ship next week.

On August 3, Alibaba confirmed that Qwen3.8-Max features approximately 2.4 trillion total parameters with roughly 95 billion active parameters per query. The model employs sparse mixture-of-experts architecture based on Qwen3.5 and supports multimodal inputs — text, images, and videos — with text output. The benchmark table, published on Alibaba’s own platform, shows the model outperforming several competitors on key tests such as Terminal-Bench 2.1 (86.6) and PaperBench (93.0), but trailing behind GPT-5.6 Sol (88.8) on some measures.

Alibaba demonstrated the model’s capabilities by reproducing research results and outperforming its own previous generation (Qwen3.7-Max) on long-horizon agentic tasks, notably improving scores on DeepSWE (from 21.6 to 56.6) and FrontierSWE (from 40.7 to 73.5). The open weights will be available next week, though the licensing terms remain undisclosed, and the full model requires multi-node deployment due to its size.

At a glance
reportWhen: announced August 3, 2023
The developmentAlibaba officially published detailed benchmark data for Qwen3.8-Max after two weeks of speculation, confirming its size and performance metrics.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Alibaba's Benchmark Disclosure

The detailed benchmark release confirms Alibaba’s position as a major player in large-scale AI development, especially with the largest open-weight model to date. The performance on multimodal and agentic tasks indicates progress in AI capabilities, potentially influencing industry standards and future model development. However, the selective nature of the benchmarks and the lack of a full licensing framework mean claims of overall superiority should be viewed cautiously.

Amazon

large language model AI open weights

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba’s AI Model Launch Strategy

Alibaba’s AI model journey has involved stealth previews, anonymous leaks, and strategic disclosures. The initial teaser, called 'kaleb,' was revealed during the July World AI Conference, with the company claiming it was 'second only to Fable 5.' The model’s specifications remained unconfirmed until the recent benchmark publication, which followed two weeks of media coverage driven by Alibaba’s selective reveal tactics. Previous models like Kimi K3 and Qwen3.7-Max laid the groundwork for this release, with Alibaba gradually increasing transparency.

"We are committed to transparency and will release open weights next week, enabling broader access and experimentation."

— Alibaba spokesperson

ESP32 Basic Starter Ai Chatbot Kit Development Board USB-C Dual Core Microcontroller Support AP/STA/AP+STA Compatible with Arduino IDE IoT for Beginners Engineers

ESP32 Basic Starter Ai Chatbot Kit Development Board USB-C Dual Core Microcontroller Support AP/STA/AP+STA Compatible with Arduino IDE IoT for Beginners Engineers

  • High Performance Dual-Core CPU: Equipped with dual-core processor and Type-C USB
  • Rich Peripheral Support: Includes SPI, LCD, Camera, UART, I2C, and more
  • Wireless Connectivity: Supports Wi-Fi 2.4 GHz and Bluetooth 5 (LE)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Open-Weight Licensing and Deployment Challenges

It remains unclear what the licensing terms for the 2.4 trillion-parameter weights will be, and whether they will be fully open-source under permissive licenses. The size of the model also means it requires multi-node deployment, limiting immediate accessibility for individual researchers or smaller organizations. The impact of the model’s agentic improvements on real-world applications is still under evaluation, as the long-term robustness of these gains is unconfirmed.

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

Local LLM Inference Optimization: A Comprehensive Guide to Quantization, Hardware Acceleration, and Efficient Private AI Deployment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Model Access and Evaluation

Alibaba will release the open weights next week, allowing the community to test and benchmark the model independently. Further detailed evaluations and comparisons are expected to follow, especially on tasks where the model trails behind competitors like Fable 5. Additionally, the industry will watch for licensing details and deployment frameworks that determine how broadly this model can be adopted outside Alibaba’s ecosystem.

MINISFORUM MS-S1 MAX Mini AI Workstation PC, AMD Ryzen AI Max+ 395 (16C/32T),RDNA3.5 GPU,64GB LPDDR5 2TB SSD Mini PC,Dual M.2 PCIe 4.0, PCIe x16 Slot, USB4 V2(80Gbps)& Dual 10GbE, 320W PSU,Wi-Fi 7

MINISFORUM MS-S1 MAX Mini AI Workstation PC, AMD Ryzen AI Max+ 395 (16C/32T),RDNA3.5 GPU,64GB LPDDR5 2TB SSD Mini PC,Dual M.2 PCIe 4.0, PCIe x16 Slot, USB4 V2(80Gbps)& Dual 10GbE, 320W PSU,Wi-Fi 7

  • High-Performance AMD Ryzen AI Max+ 395: 16-core, 32-thread CPU with RDNA 3.5 GPU
  • Powerful AI Computing Capabilities: 126 TOPS system output for demanding AI tasks
  • 64GB LPDDR5x-8000MT/s Memory: Shared high-bandwidth memory for smooth performance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

When will Alibaba release the open weights for Qwen3.8-Max?

The open weights are scheduled to be released next week, on August 10, 2023.

How does Qwen3.8-Max compare to other large models?

According to Alibaba’s benchmark table, Qwen3.8-Max outperforms models like Claude Opus 4.8 and Fable 5 on several tests but trails behind GPT-5.6 Sol on some key metrics. Its agentic capabilities have shown significant improvement over previous versions.

What are the licensing terms for the open weights?

The licensing details remain undisclosed, and it is unclear whether the weights will be open-source or have restrictions. This will be clarified upon release next week.

What are the main limitations of Qwen3.8-Max?

The model’s size requires multi-node deployment, limiting accessibility. Its performance on certain benchmarks, like SWE-bench Pro, indicates room for improvement. The long-term robustness of agentic and multimodal capabilities is still under assessment.

Source: ThorstenMeyerAI.com

You May Also Like

Jobs Being Taken Over by AI: Your Essential How-To

Uncover the strategies to thrive in a job market reshaped by AI, ensuring your career remains relevant and secure.

OpenAI's Software Revolutionizes Device Tasks

Delve into how OpenAI's software is reshaping device tasks with unprecedented precision, revolutionizing efficiency and productivity in ways you never imagined.

Exploring Scriptures: What Does the Bible Say About AI?

In our rapidly changing society, the topic of artificial intelligence (AI) is…

Stability AI Unveils Groundbreaking AI Tool for 3D Model Generation

We are thrilled to introduce the groundbreaking AI solution, Stable 3D, from…