Alibaba
Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model’s overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026.
Qwen model releasesRank #202 across 600
No rank
Rank #253 across 379
Percentile score by analysis domain.
* Cost is inverted: lower input, output, and blended prices rank higher.
Higher bars mean stronger relative placement.
Streaming speed is not measured for this model yet.
| Metric | Domain | Value | Rank |
|---|---|---|---|
| Artificial Analysis Intelligence Index | overall | 25.0 | #202 |
| Artificial Analysis Math Index | math | 82.3 | #58 |
| MMLU-Pro | reasoning | 82.4% | #63 |
| GPQA | reasoning | 77.6% | #185 |
| Humanity's Last Exam |
| reasoning |
| 12.0% |
| #184 |
| LiveCodeBench | coding | 53.5% | #138 |
| SciCode | coding, reasoning | 38.7% | #184 |
| Blended Price | cost | $2.40/M | #253 |
| Input Price | cost | $1.20/M | #246 |
| Output Price | cost | $6.00/M | #262 |
| Value Index | cost, overall | 10.4 | #243 |