Best Models
Gemini 3.1 Pro Preview is currently the highest-ranked model for MMLU-Pro in the available data.
Use the top-ranked model as a starting point, not an automatic decision.
Compare the top choices on price, speed, availability, and your own constraints.
Open the linked result page and methodology before quoting the ranking.
MMLU-Pro measures broad knowledge and reasoning with harder multiple-choice questions than classic MMLU.
| Rank | Model | Provider | Value |
|---|---|---|---|
| #1 | Gemini 3.1 Pro Preview | 91.2% | |
| #2 | Gemini 3 Pro Preview (high) | 89.8% | |
| #3 | Gemini 3 Pro Preview (low) | 89.5% | |
| #4 | Claude Opus 4.5 (Reasoning) | Anthropic | 89.5% |
| #5 | Gemini 3 Flash Preview (Reasoning) | 89% |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Source: Artificial Analysis MMLU-Pro. Updated: Jul 16, 2026.
| #6 | Claude Opus 4.5 (Non-reasoning) | Anthropic | 88.9% |
| #7 | Gemini 3 Flash Preview (Non-reasoning) | 88.2% |
| #8 | Claude 4.1 Opus (Reasoning) | Anthropic | 88% |
| #9 | Claude 4.5 Sonnet (Reasoning) | Anthropic | 87.5% |
| #10 | MiniMax-M2.1 | MiniMax | 87.5% |
| #11 | GPT-5.2 (xhigh) | OpenAI | 87.4% |
| #12 | Claude 4 Opus (Reasoning) | Anthropic | 87.3% |
| #13 | GPT-5 (high) | OpenAI | 87.1% |
| #14 | GPT-5.1 (high) | OpenAI | 87% |
| #15 | GPT-5 (medium) | OpenAI | 86.7% |
| #16 | Grok 4 | SpaceXAI | 86.6% |
| #17 | GPT-5 Codex (high) | OpenAI | 86.5% |
| #18 | DeepSeek V3.2 Speciale | DeepSeek | 86.3% |
| #19 | Gemini 3.1 Flash-Lite | 86.2% |
| #20 | Gemini 2.5 Pro | 86.2% |