Best Models
GPT-5 (high) is currently the highest-ranked model for AIME in the available data.
Use the top-ranked model as a starting point, not an automatic decision.
Compare the top choices on price, speed, availability, and your own constraints.
Open the linked result page and methodology before quoting the ranking.
AIME measures advanced competition math performance. Higher percentages indicate more solved problems.
| Rank | Model | Provider | Value |
|---|---|---|---|
| #1 | GPT-5 (high) | OpenAI | 95.7% |
| #2 | Grok 4 | SpaceXAI | 94.3% |
| #3 | o4-mini (high) | OpenAI | 94% |
| #4 | Qwen3 235B A22B 2507 (Reasoning) | Alibaba | 94% |
| #5 | Grok 3 mini Reasoning (high) | SpaceXAI | 93.3% |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Source: Artificial Analysis AIME. Updated: Jul 16, 2026.
| #6 |
| GPT-5 (medium) |
| OpenAI |
| 91.7% |
| #7 | Qwen3 30B A3B 2507 (Reasoning) | Alibaba | 90.7% |
| #8 | o3 | OpenAI | 90.3% |
| #9 | DeepSeek R1 0528 (May '25) | DeepSeek | 89.3% |
| #10 | Gemini 2.5 Pro | 88.7% |
| #11 | GLM-4.5 (Reasoning) | Z AI | 87.3% |
| #12 | Gemini 2.5 Pro Preview (Mar' 25) | 87% |
| #13 | Llama Nemotron Super 49B v1.5 (Reasoning) | NVIDIA | 86% |
| #14 | o3-mini (high) | OpenAI | 86% |
| #15 | MiniMax M1 80k | MiniMax | 84.7% |
| #16 | EXAONE 4.0 32B (Reasoning) | LG AI Research | 84.3% |
| #17 | Gemini 2.5 Flash Preview (Reasoning) | 84.3% |
| #18 | Gemini 2.5 Pro Preview (May' 25) | 84.3% |
| #19 | Qwen3 235B A22B (Reasoning) | Alibaba | 84% |
| #20 | GPT-5 (low) | OpenAI | 83% |