Best Models
GPT-5.6 Sol (max) is currently the highest-ranked model for Reasoning & Knowledge in the available data.
Use the top-ranked model as a starting point, not an automatic decision.
Compare the top choices on price, speed, availability, and your own constraints.
Open the linked result page and methodology before quoting the ranking.
GPQA is a difficult graduate-level science benchmark. Higher percentages indicate stronger expert reasoning performance.
| Rank | Model | Provider | Value |
|---|---|---|---|
| #1 | GPT-5.6 Sol (max) | OpenAI | 94.1% |
| #2 | Gemini 3.1 Pro Preview | 94.1% | |
| #3 | GPT-5.5 (xhigh) | OpenAI | 93.5% |
| #4 | GPT-5.5 (high) | OpenAI | 93.2% |
| #5 | GPT-5.6 Sol (xhigh) | OpenAI | 93.1% |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Source: Artificial Analysis GPQA. Updated: Jul 16, 2026.
| #6 |
| Grok 4.5 (high) |
| SpaceXAI |
| 93.1% |
| #7 | MiniMax-M3 | MiniMax | 92.9% |
| #8 | GPT-5.6 Sol (high) | OpenAI | 92.8% |
| #9 | GPT-5.5 (medium) | OpenAI | 92.6% |
| #10 | GPT-5.6 Sol (medium) | OpenAI | 92.6% |
| #11 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) | Anthropic | 92.6% |
| #12 | GPT-5.6 Terra (max) | OpenAI | 92.5% |
| #13 | Qwen3.7 Max | Alibaba | 92.3% |
| #14 | Gemini 3.5 Flash (high) | 92.2% |
| #15 | Gemini 3.5 Flash (medium) | 92.1% |
| #16 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) | Anthropic | 92% |
| #17 | GPT-5.4 (xhigh) | OpenAI | 92% |
| #18 | GPT-5.3 Codex (xhigh) | OpenAI | 91.5% |
| #19 | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) | Anthropic | 91.4% |
| #20 | GPT-5.6 Luna (max) | OpenAI | 91.1% |