DeepSeek
DeepSeek-V3.1 is post-trained on the top of DeepSeek-V3.1-Base, which is built upon the original V3 base checkpoint through a two-phase long context extension approach, following the methodology outlined in the original DeepSeek-V3 report. We have expanded our dataset by collecting additional long documents and substantially extending both training phases. The 32K extension phase has been increased 10-fold to 630B tokens, while the 128K extension phase has been extended by 3.3x to 209B tokens. Additionally, DeepSeek-V3.1 is trained using the UE8M0 FP8 scale data format to ensure compatibility with microscaling data formats.
DeepSeek release notesRank #236 across 600
No rank
Rank #169 across 379
Percentile score by analysis domain.
* Cost is inverted: lower input, output, and blended prices rank higher.
Higher bars mean stronger relative placement.
Streaming speed is not measured for this model yet.
| Metric | Domain | Value | Rank |
|---|---|---|---|
| Artificial Analysis Intelligence Index | overall | 21.0 | #236 |
| Artificial Analysis Math Index | math | 49.7 | #141 |
| GPQA | reasoning | 73.5% | #233 |
| Humanity's Last Exam | reasoning | 6.3% | #298 |
| LiveCodeBench |
| coding |
| 57.7% |
| #124 |
| SciCode | coding, reasoning | 36.7% | #219 |
| Blended Price | cost | $0.840/M | #169 |
| Input Price | cost | $0.560/M | #190 |
| Output Price | cost | $1.68/M | #156 |
| Value Index | cost, overall | 25.0 | #146 |