📊 Full opportunity report: AI Comparison: How Qwen3.8-Max Matches Up Against Fable 5 on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Alibaba has released detailed benchmark data for its Qwen3.8-Max model, claiming it is ‘second only to Fable 5.’ The model shows strong performance in multimodal and agentic tasks but lags behind in certain software engineering benchmarks. The open weights will be available next week, with a smaller, more deployable 27B version also announced.

Alibaba’s Qwen3.8-Max has been officially launched with comprehensive benchmark data, confirming it as ‘second only to Fable 5’ in several key performance areas. The model’s open weights will be released next week, marking a significant milestone in open AI model deployment.

Two weeks after its stealth preview, Alibaba publicly shared detailed benchmark results for Qwen3.8-Max, a 2.4 trillion-parameter model built on the Qwen3.5 architecture. The model features a roughly 95 billion active parameter count per query, utilizing sparse mixture-of-experts technology, and is multimodal, handling text, images, and videos.

The benchmark table, tested on Alibaba’s own infrastructure, ranks the model at 86.6 on Terminal-Bench 2.1, surpassing Claude Opus 4.8 and Claude Fable 5, but behind GPT-5.6 Sol at 88.8. It scores highest on PaperBench (93.0) and excels in multimodal and agentic tasks, such as OSWorld-Verified (86.1) and Parametric CAD Bench (91.5). However, it trails significantly in deep software engineering benchmarks like SWE-bench Pro (67.7) and FrontierSWE (73.5), compared to Fable 5’s 80.0 and 88.8 respectively.

Alibaba also demonstrated impressive long-horizon reasoning improvements, doubling agentic scores from previous versions, but the claim that it is ‘second only to Fable 5’ applies selectively, based on the specific benchmarks chosen for comparison.

At a glance
reportWhen: announced August 3, 2023; benchmarks pu…
The developmentAlibaba announced the full benchmark results for its Qwen3.8-Max model, confirming it as a top contender against Fable 5 and other leading models.
AI DISPATCH · REALITY CHECK Released 3 Aug 2026
Alibaba’s Qwen3.8-Max leaves preview
Second Only to Fable 5?

For fifteen days the claim ran without a benchmark table. Today Alibaba published the table, the active-parameter count, and a weights timeline. The numbers are genuinely strong on the rows Alibaba chose — and twelve to fifteen points behind on the rows it didn’t.

▲ All performance figures: Alibaba’s own harness
2.4T / 95B
Total / active parameters (MoE)
~1M
Context window · 131K max output
Text+Img+Video
Multimodal in · text out
“Next week”
Open weights · licence unpublished
01
Fifteen days from slogan to spec sheet

The claim shipped on a Sunday. The evidence shipped two weeks later. In between, the claim did its work.

17 Jul
Moonshot releases Kimi K3
2.8T parameters; rattles US tech stocks, later suspends new subscriptions under demand.
18 Jul
“kaleb” appears on Code Arena
Anonymous model introduces itself as “Claude” — a distillation artifact — and is identified within a day by a Qwen tokenizer quirk.
19 Jul
WAIC preview: “second only to Fable 5”
No benchmark table, no model card, no licence, no active-parameter count. Paid preview at 10% of standard pricing.
20 Jul
Shares rise as much as 5.4%
The market prices the claim, not the table.
3 Aug
General availability + full benchmark table
95B active confirmed; 2.4T weights and a Qwen3.8-27B checkpoint promised for next week. Licence still unwritten.
02
The table, both halves

“Second only to Fable 5” is true on the rows Alibaba chose and false on the rows it didn’t. Both halves below are from the same release.

Where it leads
Terminal-Bench 2.1 · agentic terminal work
Qwen3.8-Max
86.6
GPT-5.6 Sol
88.8
Fable 5
84.6
OSWorld-Verified · computer use — plus PaperBench 93.0, CAD Bench 91.5
Qwen3.8-Max
86.1
Where it trails — the rows the slogan skips
SWE-bench Pro · deep software engineering
Qwen3.8-Max
67.7
Fable 5
80.0
FrontierSWE · frontier coding agents
Qwen3.8-Max
73.5
Fable 5
88.8
The real jump: one generation of agentic gains vs Qwen3.7-Max
DeepSWE 1.1
21.6 → 56.6
FrontierSWE
40.7 → 73.5
JobBench
31.3 → 53.4
03
Three artifacts, three different facts

“Qwen3.8 is going open-weight” describes three things with very different deployment realities.

Hosted API
Live today

OpenAI- and DashScope-compatible — a base-URL change to A/B against your current backend.

2.4T weights
“Next week” · no licence yet

A multi-node datacenter artifact. At 95B active, no single machine serves it. A flag planted, not a deployment option.

Qwen3.8-27B
Announced · no benchmarks yet

The checkpoint that fits real hardware. Whether the agentic gains survive distillation is the question that decides whether next week matters.

04
Bull and bear

Three Chinese frontier releases in seventeen days, each measured against the same export-controlled model. The contest is real; it is not the same thing as your workload.

Bull
  • The generation jump is real and consistent across a dozen agentic rows, with a stated mechanism: RL-environment scaling.
  • More disclosure than Kimi K3 shipped — full table, active-parameter count, weights timeline.
  • If 2.4T lands under a permissive licence, the ceiling of “open weight” moves permanently.
  • The 27B sibling could become the best local agent model on hardware people already own.
Bear
  • Every number is Alibaba’s harness. Independent testing already tempered Kimi K3’s launch claims substantially.
  • The paying use case still belongs to Fable 5 — twelve to fifteen points on deep software engineering.
  • “Next week” comes from a company that sat on a finished benchmark table for fifteen days.
  • Until the licence text exists, “going open-weight” is a press strategy, not a property of the model.
The claim ran for fifteen days without evidence. Now the evidence exists —
and it says “second only” depends entirely on which row you read.

Implications of Benchmark Results for AI Leadership

The detailed performance data confirms Alibaba's ambition to position Qwen3.8-Max as a top-tier AI model, especially in multimodal and agentic tasks, which are critical for practical AI applications. The release of open weights next week will enable wider testing and deployment, potentially reshaping competitive dynamics among leading AI labs. However, the model's weaknesses in software engineering benchmarks highlight ongoing challenges in achieving balanced AI capabilities.

For industry stakeholders and developers, the availability of a 2.4 trillion-parameter open-weight model signals a new level of accessibility for large-scale AI, though the hardware requirements remain high. The smaller 27B checkpoint, suitable for local deployment, offers immediate utility for practical applications, pending validation of its agentic performance.

Amazon

AI model benchmark comparison

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on Alibaba's AI Model Development

Alibaba's AI model development has been marked by strategic previews and stealth releases, culminating in the recent full benchmark disclosure. The company initially teased the model with a slogan claiming it was 'second only to Fable 5,' without releasing detailed data, which generated significant industry speculation. The model, Qwen3.8-Max, was identified through a community-led reverse engineering effort and confirmed during the World AI Conference in Shanghai on July 19.

Since then, Alibaba has gradually revealed benchmark scores, emphasizing multimodal and agentic capabilities, and positioning Qwen3.8-Max as a flagship in its AI lineup. The upcoming release of open weights aims to challenge existing models and expand access to large-scale AI models outside of proprietary ecosystems.

Amazon

multimodal AI development tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unconfirmed Aspects of Model Licensing and Deployment

It remains unclear what the exact licensing terms for the 2.4 trillion-parameter weights will be, as Alibaba has not yet published the license details. The hardware requirements for running the full model are also not specified, raising questions about practical deployment. Additionally, the agentic performance of the smaller 27B checkpoint has yet to be publicly benchmarked or validated.

Amazon

AI software engineering benchmarks

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Testing and Industry Adoption

Alibaba will release the open weights next week, prompting independent testing and comparison. The AI community will evaluate the 27B checkpoint's performance in real-world scenarios, especially its agentic capabilities. Further benchmark updates and potential licensing clarifications are expected in the coming weeks, shaping the competitive landscape.

Amazon

AI model deployment hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What are the main strengths of Qwen3.8-Max according to Alibaba?

Qwen3.8-Max excels in multimodal tasks, agentic reasoning, and long-horizon reasoning, outperforming many models on specific benchmarks like PaperBench and OSWorld-Verified.

When will the open weights for Qwen3.8-Max be available?

The open weights are scheduled for release next week, enabling broader access and testing by the AI community.

How does Qwen3.8-Max compare to Fable 5 overall?

It ranks just below Fable 5 in several deep software engineering benchmarks but surpasses it in multimodal and agentic tasks, according to Alibaba's benchmark data.

What are the hardware requirements for running Qwen3.8-Max?

The full 2.4 trillion-parameter model requires multi-node data center infrastructure; the smaller 27B checkpoint is designed for high-memory single machines.

What remains uncertain about Alibaba's model release?

The licensing terms, practical deployment details, and agentic performance of the smaller checkpoint are still unconfirmed.

Source: ThorstenMeyerAI.com

You May Also Like

ByteDance’s Plan To Dominate AI – Financial Times

The Financial Times reports ByteDance is pursuing a strategic plan to become a dominant force in artificial intelligence, though details remain unconfirmed.

The Memento Constraint: Why Continual Learning Is the Trillion-Dollar Bottleneck Nobody Is Pricing

A16z highlights the ‘Memento’ constraint in AI, revealing that solving continual learning could reshape the trillion-dollar enterprise AI sector by 2028.

When-to-replace planner for data center equipment

A new workflow for data center capacity planning tests a tool that recommends optimal replacement timing for servers, UPS, and cooling units based on asset data.

Netcity Telecom Surges In Global Coverage

Netcity Telecom reports a surge in its international network coverage, with 23 mentions indicating rapid expansion across multiple regions.