DeepSeek
DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) ManifoldConstrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability
DeepSeek release notesRank #144 across 600
Rank #63 across 266
Rank #133 across 379
Percentile score by analysis domain.
* Cost is inverted: lower input, output, and blended prices rank higher.
Higher bars mean stronger relative placement.
| Metric | Domain | Value | Rank |
|---|---|---|---|
| Artificial Analysis Intelligence Index | overall | 31.2 | #144 |
| Artificial Analysis Coding Index | coding | 59.4 | #63 |
| GPQA | reasoning | 71.7% | #251 |
| Humanity's Last Exam | reasoning | 7.7% | #257 |
| SciCode |
| coding, reasoning |
| 42.4% |
| #113 |
| Output Speed | speed | 60.1 tok/s | #125 |
| Time to First Token | speed | 1.12s | #77 |
| Blended Price | cost | $0.544/M | #133 |
| Input Price | cost | $0.435/M | #169 |
| Output Price | cost | $0.870/M | #108 |
| Value Index | cost, overall | 57.4 | #81 |