Current LiveBench Mathematics category score.
LiveBench is designed to reduce contamination by publishing versioned releases with newer tasks. Easy Benchmarks calculates the Mathematics category exactly from the task list in the matching public categories file and keeps the release date with every imported score.
Test type: Average score across the Mathematics tasks declared by the selected LiveBench release.
26 models have this metric.
Current leader: GPT-5.6 Sol (max)
Project links
This app automatically imports the newest public LiveBench release that has both a versioned score table and category definition file.
Top models ranked by LiveBench Math.
| Rank | Model | Creator | Value | Speed | Blended Price |
|---|---|---|---|---|---|
| #1 | GPT-5.6 Sol (max) | OpenAI | 96.2% | 67.4 tok/s | $11.25/M |
| #2 |
| Anthropic |
| 96.0% |
| 69.4 tok/s |
| $20.00/M |
| #3 | GPT-5.5 (xhigh) | OpenAI | 95.9% | 73.9 tok/s | $11.25/M |
| #4 | GPT-5.6 Sol (xhigh) | OpenAI | 95.5% | 57 tok/s | $11.25/M |
| #5 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | Anthropic | 95.3% | 54.3 tok/s | $10.00/M |
| #6 | GPT-5.5 (high) | OpenAI | 95.2% | 58.4 tok/s | $11.25/M |
| #7 | GPT-5.6 Terra (max) | OpenAI | 94.9% | 170.1 tok/s | $5.63/M |
| #8 | GPT-5.4 (xhigh) | OpenAI | 94.1% | 154.5 tok/s | $5.63/M |
| #9 | Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | 92.9% | 69.1 tok/s | $4.00/M |
| #10 | Gemini 3.1 Pro Preview | 91.0% | 123.1 tok/s | $4.50/M |
| #11 | GPT-5.4 nano (xhigh) | OpenAI | 91.0% | 161.4 tok/s | $0.463/M |
| #12 | Grok 4.5 (high) | SpaceXAI | 90.8% | 122.8 tok/s | $3.00/M |
| #13 | GLM-5.2 (max) | Z AI | 89.8% | 149.2 tok/s | $2.15/M |
| #14 | GPT-5.6 Terra (xhigh) | OpenAI | 89.5% | 117.6 tok/s | $5.63/M |
| #15 | GPT-5.2 Codex (xhigh) | OpenAI | 88.8% | 125.5 tok/s | $4.81/M |
| #16 | Gemini 3.5 Flash (high) | 88.2% | 241.2 tok/s | $3.38/M |
| #17 | GPT-5.6 Luna (max) | OpenAI | 87.2% | 243.6 tok/s | $2.25/M |
| #18 | Muse Spark 1.1 (xhigh) | Meta | 87.1% | 124.3 tok/s | $2.00/M |
| #19 | GPT-5.6 Luna (xhigh) | OpenAI | 86.3% | 185.7 tok/s | $2.25/M |
| #20 | Qwen3.7 Max | Alibaba | 85.2% | 198.9 tok/s | $3.75/M |
| #21 | Grok 4.3 (high) | SpaceXAI | 84.3% | 112.5 tok/s | $1.56/M |
| #22 | Kimi K2.6 | Kimi | 84.3% | 39.2 tok/s | $1.71/M |
| #23 | Qwen3.6 Plus | Alibaba | 83.7% | 53.3 tok/s | $1.13/M |
| #24 | Kimi K2.7 Code | Kimi | 79.6% | 43 tok/s | $1.71/M |
| #25 | GPT-5.4 mini (xhigh) | OpenAI | 78.5% | 167.5 tok/s | $1.69/M |
| #26 | MiniMax-M3 | MiniMax | 76.9% | 94.8 tok/s | $0.525/M |