Beste Modelle
Gemini 3.1 Pro Preview ist derzeit das bestplatzierte Modell für MMLU-Pro in den verfügbaren Daten.
Nutze den Erstplatzierten als Ausgangspunkt, nicht als automatische Entscheidung.
Vergleiche die ersten Modelle bei Preis, Geschwindigkeit und Verfügbarkeit.
Öffne die verlinkte Ergebnisseite und Methodik, bevor du das Ergebnis zitierst.
MMLU-Pro measures broad knowledge and reasoning with harder multiple-choice questions than classic MMLU.
| Rang | Modell | Anbieter | Wert |
|---|---|---|---|
| #1 | Gemini 3.1 Pro Preview | 91,2% | |
| #2 | Gemini 3 Pro Preview (high) | 89,8% | |
| #3 | Gemini 3 Pro Preview (low) | 89,5% | |
| #4 | Claude Opus 4.5 (Reasoning) | Anthropic | 89,5% |
| #5 | Gemini 3 Flash Preview (Reasoning) | 89% |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Quelle: Artificial Analysis MMLU-Pro. Aktualisiert: 16.07.2026.
| #6 | Claude Opus 4.5 (Non-reasoning) | Anthropic | 88,9% |
| #7 | Gemini 3 Flash Preview (Non-reasoning) | 88,2% |
| #8 | Claude 4.1 Opus (Reasoning) | Anthropic | 88% |
| #9 | Claude 4.5 Sonnet (Reasoning) | Anthropic | 87,5% |
| #10 | MiniMax-M2.1 | MiniMax | 87,5% |
| #11 | GPT-5.2 (xhigh) | OpenAI | 87,4% |
| #12 | Claude 4 Opus (Reasoning) | Anthropic | 87,3% |
| #13 | GPT-5 (high) | OpenAI | 87,1% |
| #14 | GPT-5.1 (high) | OpenAI | 87% |
| #15 | GPT-5 (medium) | OpenAI | 86,7% |
| #16 | Grok 4 | SpaceXAI | 86,6% |
| #17 | GPT-5 Codex (high) | OpenAI | 86,5% |
| #18 | DeepSeek V3.2 Speciale | DeepSeek | 86,3% |
| #19 | Gemini 3.1 Flash-Lite | 86,2% |
| #20 | Gemini 2.5 Pro | 86,2% |