DeepSeek
DeepSeek-V4 series incorporate several key upgrades in architecture and optimization: (1) a hybrid attention architecture that combines Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) to improve long-context efficiency; (2) ManifoldConstrained Hyper-Connections (mHC) that enhance conventional residual connections; (3) and the Muon optimizer for faster convergence and greater training stability
DeepSeek release notesRank #172 across 600
No rank
Rank #47 across 379
Percentile score by analysis domain.
* Cost is inverted: lower input, output, and blended prices rank higher.
Higher bars mean stronger relative placement.
| Metric | Domain | Value | Rank |
|---|---|---|---|
| Artificial Analysis Intelligence Index | overall | 28.7 | #172 |
| GPQA | reasoning | 71.6% | #252 |
| Humanity's Last Exam | reasoning | 7.0% | #275 |
| SciCode | coding, reasoning | 37.3% | #210 |
| Output Speed |
| speed |
| 101.9 tok/s |
| #87 |
| Time to First Token | speed | 0.94s | #59 |
| Blended Price | cost | $0.175/M | #47 |
| Input Price | cost | $0.140/M | #57 |
| Output Price | cost | $0.280/M | #34 |
| Value Index | cost, overall | 164.0 | #16 |