DeepSeek
DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) ManifoldConstrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability
DeepSeek release notesRank #48 across 600
Rank #64 across 266
Rank #135 across 379
Percentile score by analysis domain.
* Cost is inverted: lower input, output, and blended prices rank higher.
Higher bars mean stronger relative placement.
| Metric | Domain | Value | Rank |
|---|---|---|---|
| Artificial Analysis Intelligence Index | overall | 44.3 | #48 |
| Artificial Analysis Coding Index | coding | 59.4 | #64 |
| GPQA | reasoning | 88.8% | #56 |
| Humanity's Last Exam | reasoning | 35.9% | #40 |
| SciCode |
| coding, reasoning |
| 50.0% |
| #46 |
| Output Speed | speed | 59.1 tok/s | #128 |
| Time to First Token | speed | 1.10s | #74 |
| Blended Price | cost | $0.544/M | #135 |
| Input Price | cost | $0.435/M | #171 |
| Output Price | cost | $0.870/M | #110 |
| Value Index | cost, overall | 81.4 | #50 |