DeepSeek
DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) ManifoldConstrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability
DeepSeek release notesRank #54 across 600
Rank #69 across 266
Rank #134 across 379
Percentile score by analysis domain.
* Cost is inverted: lower input, output, and blended prices rank higher.
Higher bars mean stronger relative placement.
| Metric | Domain | Value | Rank |
|---|---|---|---|
| Artificial Analysis Intelligence Index | overall | 43.1 | #54 |
| Artificial Analysis Coding Index | coding | 58.7 | #69 |
| GPQA | reasoning | 90.5% | #34 |
| Humanity's Last Exam | reasoning | 33.5% | #47 |
| SciCode |
| coding, reasoning |
| 46.4% |
| #75 |
| Output Speed | speed | 56.1 tok/s | #133 |
| Time to First Token | speed | 1.04s | #70 |
| Blended Price | cost | $0.544/M | #134 |
| Input Price | cost | $0.435/M | #170 |
| Output Price | cost | $0.870/M | #109 |
| Value Index | cost, overall | 79.2 | #52 |