Alibaba
Qwen3-Max-Preview shows substantial gains over the 2.5 series in overall capability, with significant enhancements in Chinese-English text understanding, complex instruction following, handling of subjective open-ended tasks, multilingual ability, and tool invocation; model knowledge hallucinations are reduced.
Qwen model releasesRank #258 across 600
No rank
Rank #252 across 379
Percentile score by analysis domain.
* Cost is inverted: lower input, output, and blended prices rank higher.
Higher bars mean stronger relative placement.
Streaming speed is not measured for this model yet.
| Metric | Domain | Value | Rank |
|---|---|---|---|
| Artificial Analysis Intelligence Index | overall | 19.2 | #258 |
| Artificial Analysis Math Index | math | 75.0 | #80 |
| MMLU-Pro | reasoning | 83.8% | #44 |
| GPQA | reasoning | 76.4% | #202 |
| Humanity's Last Exam |
| reasoning |
| 9.3% |
| #231 |
| LiveCodeBench | coding | 65.1% | #98 |
| SciCode | coding, reasoning | 37.0% | #214 |
| Blended Price | cost | $2.40/M | #252 |
| Input Price | cost | $1.20/M | #245 |
| Output Price | cost | $6.00/M | #261 |
| Value Index | cost, overall | 8.0 | #274 |