Qwen

Qwen 3.8 Max

QwenFlagship
ThinkingTool UseVisionStructured Output

About this model

Qwen3.8 Max is the flagship model in Alibaba's Qwen3.8 series, the general-availability successor to the Qwen3.8 Max Preview. It is a multimodal reasoning model intended for complex reasoning, visual understanding,...

Performance Tier

Flagship

Qwen 3.8 Max is a flagship model from Qwen : the most capable in their lineup.

Best-in-class model from this provider. Highest performance across benchmarks, ideal for demanding tasks.

Pricing

This model is included in Elosia plans
Moderate

Moderate cost. A balanced choice for regular use without constant cap watching.

Typeper 1M tokens
Input (prompt)$2.00
Output (completion)$6.00
Cache read$0.250
Cache write$2.50

Capabilities

Context Length1.0M
Max Output Tokens131K
TokenizerQwen
Inputtext, image, video
Outputtext
Release DateAugust 3, 2026

Benchmarks

General Intelligence
MMLU
Not reported
GPQA Diamond
92.6%
Mathematics
MATH-500
Not reported
Programming
HumanEval
Not reported
SWE-bench Verified
Not reported
Reasoning
IFEval
Not reported
Humanity's Last Exam
43.6%
Multimodal
MMMU-Pro
82.3%
Agentic
SWE-bench Pro
67.7%
Terminal-Bench 2.1
86.6%

Where does Qwen 3.8 Max stand?

Compare its performance index against every other model.

View the performance leaderboard

Recommended Use Cases

CodingAnalysisResearchGeneral ChatData Extraction

Strengths

  • Agentic behaviour carried by a 2.4T-parameter mixture of experts, only a fraction of which is engaged on each token
  • Tool calling with structured outputs, and up to 128K of output in a single response
  • Native video input alongside text and images, combined with a 1M-token context window
  • Vision goes past captioning: object detection and counting are part of the reported evaluation, on top of image and video inputs

Limitations

  • Weak factual reliability: on a question whose answer it does not know, it tends to produce a confident answer rather than abstain
  • Software engineering on real repositories trails its reasoning results, and the weights are closed, with no checkpoint published for self-hosting or auditing

Frequently asked questions

Resources

This model may use your data for training

Similar Models