Mejores modelos
GPT-5.6 Sol (max) es actualmente el modelo mejor clasificado para LiveBench Mathematics en los datos disponibles.
Usa el modelo líder como punto de partida, no como una decisión automática.
Compara las mejores opciones por precio, velocidad, disponibilidad y tus límites.
Abre la página de resultados y la metodología antes de citar el ranking.
Average score across the Mathematics tasks in the latest versioned LiveBench public release. It stays separate from other math benchmarks because the task mix and evaluation protocol differ.
| Posición | Modelo | Proveedor | Valor |
|---|---|---|---|
| #1 | GPT-5.6 Sol (max) | OpenAI | 96,2% |
| #2 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 96% |
| #3 | GPT-5.5 (xhigh) | OpenAI | 95,9% |
| #4 | GPT-5.6 Sol (xhigh) | OpenAI | 95,5% |
| #5 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Fuente: Artificial Analysis LiveBench Mathematics. Actualizado: 16 jul 2026.
| Anthropic |
| 95,3% |
| #6 | GPT-5.5 (high) | OpenAI | 95,2% |
| #7 | GPT-5.6 Terra (max) | OpenAI | 94,9% |
| #8 | GPT-5.4 (xhigh) | OpenAI | 94,1% |
| #9 | Claude Sonnet 5 (Adaptive Reasoning, Xhigh Effort) | Anthropic | 92,9% |
| #10 | Gemini 3.1 Pro Preview | 91% |
| #11 | GPT-5.4 nano (xhigh) | OpenAI | 91% |
| #12 | Grok 4.5 (high) | SpaceXAI | 90,8% |
| #13 | GLM-5.2 (max) | Z AI | 89,8% |
| #14 | GPT-5.6 Terra (xhigh) | OpenAI | 89,5% |
| #15 | GPT-5.2 Codex (xhigh) | OpenAI | 88,8% |
| #16 | Gemini 3.5 Flash (high) | 88,2% |
| #17 | GPT-5.6 Luna (max) | OpenAI | 87,2% |
| #18 | Muse Spark 1.1 (xhigh) | Meta | 87,1% |
| #19 | GPT-5.6 Luna (xhigh) | OpenAI | 86,3% |
| #20 | Qwen3.7 Max | Alibaba | 85,2% |