Best Models
GPT-5 (high) is currently the highest-ranked model for Math in the available data.
Use the top-ranked model as a starting point, not an automatic decision.
Compare the top choices on price, speed, availability, and your own constraints.
Open the linked result page and methodology before quoting the ranking.
MATH-500 measures mathematical problem solving on a curated set of competition-style questions.
| Rank | Model | Provider | Value |
|---|---|---|---|
| #1 | GPT-5 (high) | OpenAI | 99.4% |
| #2 | o3 | OpenAI | 99.2% |
| #3 | Grok 3 mini Reasoning (high) | SpaceXAI | 99.2% |
| #4 | GPT-5 (medium) | OpenAI | 99.1% |
| #5 | Claude 4 Sonnet (Reasoning) | Anthropic | 99.1% |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Source: Artificial Analysis MATH-500. Updated: Jul 16, 2026.
| #6 |
| Grok 4 |
| SpaceXAI |
| 99% |
| #7 | o4-mini (high) | OpenAI | 98.9% |
| #8 | GPT-5 (low) | OpenAI | 98.7% |
| #9 | Gemini 2.5 Pro Preview (May' 25) | 98.6% |
| #10 | o3-mini (high) | OpenAI | 98.5% |
| #11 | Qwen3 235B A22B 2507 (Reasoning) | Alibaba | 98.4% |
| #12 | Llama Nemotron Super 49B v1.5 (Reasoning) | NVIDIA | 98.3% |
| #13 | DeepSeek R1 0528 (May '25) | DeepSeek | 98.3% |
| #14 | Claude 4 Opus (Reasoning) | Anthropic | 98.2% |
| #15 | Gemini 2.5 Flash (Reasoning) | 98.1% |
| #16 | Gemini 2.5 Flash Preview (Reasoning) | 98.1% |
| #17 | Gemini 2.5 Pro Preview (Mar' 25) | 98% |
| #18 | MiniMax M1 80k | MiniMax | 98% |
| #19 | Qwen3 235B A22B 2507 Instruct | Alibaba | 98% |
| #20 | GLM-4.5 (Reasoning) | Z AI | 97.9% |