Melhores modelos
GPT-5.6 Sol (max) é atualmente o modelo mais bem classificado para LiveBench Mathematics nos dados disponíveis.
Use o modelo líder como ponto de partida, não como decisão automática.
Compare as melhores opções por preço, velocidade, disponibilidade e limites.
Abra a página de resultados e a metodologia antes de citar o ranking.
Average score across the Mathematics tasks in the latest versioned LiveBench public release. It stays separate from other math benchmarks because the task mix and evaluation protocol differ.
| Posição | Modelo | Fornecedor | Valor |
|---|---|---|---|
| #1 | GPT-5.6 Sol (max) | OpenAI | 96,2% |
| #2 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 96% |
| #3 | GPT-5.5 (xhigh) | OpenAI | 95,9% |
| #4 | GPT-5.6 Sol (xhigh) | OpenAI | 95,5% |
| #5 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Fonte: Artificial Analysis LiveBench Mathematics. Atualizado: 16 de jul. de 2026.
| Anthropic |
| 95,3% |
| #6 | GPT-5.5 (high) | OpenAI | 95,2% |
| #7 | GPT-5.6 Terra (max) | OpenAI | 94,9% |
| #8 | GPT-5.4 (xhigh) | OpenAI | 94,1% |
| #9 | Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | 92,9% |
| #10 | Gemini 3.1 Pro Preview | 91% |
| #11 | GPT-5.4 nano (xhigh) | OpenAI | 91% |
| #12 | Grok 4.5 (high) | SpaceXAI | 90,8% |
| #13 | GLM-5.2 (max) | Z AI | 89,8% |
| #14 | GPT-5.6 Terra (xhigh) | OpenAI | 89,5% |
| #15 | GPT-5.2 Codex (xhigh) | OpenAI | 88,8% |
| #16 | Gemini 3.5 Flash (high) | 88,2% |
| #17 | GPT-5.6 Luna (max) | OpenAI | 87,2% |
| #18 | Muse Spark 1.1 (xhigh) | Meta | 87,1% |
| #19 | GPT-5.6 Luna (xhigh) | OpenAI | 86,3% |
| #20 | Qwen3.7 Max | Alibaba | 85,2% |