Alibaba
Compared with the snapshot as of September 23, 2025, the Qwen-3 series Max model in this release achieves an effective integration of thinking and non-thinking modes, resulting in a comprehensive and substantial improvement in the model’s overall performance. In thinking mode, the model simultaneously supports web search, web information extraction, and a code interpreter tool, enabling it to tackle more complex and challenging problems with greater accuracy by leveraging external tools while engaging in slow, deliberative reasoning. This version is based on a snapshot taken on January 23, 2026.
Qwen model releasesRank #141 across 600
No rank
No rank
Percentile score by analysis domain.
Higher bars mean stronger relative placement.
Streaming speed is not measured for this model yet.
| Metric | Domain | Value | Rank |
|---|---|---|---|
| Artificial Analysis Intelligence Index | overall | 31.7 | #141 |
| GPQA | reasoning | 86.1% | #83 |
| Humanity's Last Exam | reasoning | 26.2% | #84 |
| SciCode | coding, reasoning | 43.1% | #105 |