Best Models
GPT-5.6 Sol (max) is currently the highest-ranked model for LiveBench Mathematics in the available data.
Use the top-ranked model as a starting point, not an automatic decision.
Compare the top choices on price, speed, availability, and your own constraints.
Open the linked result page and methodology before quoting the ranking.
Average score across the Mathematics tasks in the latest versioned LiveBench public release. It stays separate from other math benchmarks because the task mix and evaluation protocol differ.
| Rank | Model | Provider | Value |
|---|---|---|---|
| #1 | GPT-5.6 Sol (max) | OpenAI | 96.2% |
| #2 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 96% |
| #3 | GPT-5.5 (xhigh) | OpenAI | 95.9% |
| #4 | GPT-5.6 Sol (xhigh) | OpenAI | 95.5% |
| #5 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Source: Artificial Analysis LiveBench Mathematics. Updated: Jul 16, 2026.
| Anthropic |
| 95.3% |
| #6 | GPT-5.5 (high) | OpenAI | 95.2% |
| #7 | GPT-5.6 Terra (max) | OpenAI | 94.9% |
| #8 | GPT-5.4 (xhigh) | OpenAI | 94.1% |
| #9 | Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | 92.9% |
| #10 | Gemini 3.1 Pro Preview | 91% |
| #11 | GPT-5.4 nano (xhigh) | OpenAI | 91% |
| #12 | Grok 4.5 (high) | SpaceXAI | 90.8% |
| #13 | GLM-5.2 (max) | Z AI | 89.8% |
| #14 | GPT-5.6 Terra (xhigh) | OpenAI | 89.5% |
| #15 | GPT-5.2 Codex (xhigh) | OpenAI | 88.8% |
| #16 | Gemini 3.5 Flash (high) | 88.2% |
| #17 | GPT-5.6 Luna (max) | OpenAI | 87.2% |
| #18 | Muse Spark 1.1 (xhigh) | Meta | 87.1% |
| #19 | GPT-5.6 Luna (xhigh) | OpenAI | 86.3% |
| #20 | Qwen3.7 Max | Alibaba | 85.2% |