NVIDIA
The model is an auto-regressive vision language model that uses an optimized transformer architecture. The model enables multi-image reasoning and video understanding, along with strong document intelligence, visual Q&A and summarization capabilities.
NVIDIA model catalogRank #418 across 600
No rank
Rank #83 across 379
Percentile score by analysis domain.
* Cost is inverted: lower input, output, and blended prices rank higher.
Higher bars mean stronger relative placement.
| Metric | Domain | Value | Rank |
|---|---|---|---|
| Artificial Analysis Intelligence Index | overall | 9.0 | #418 |
| Artificial Analysis Math Index | math | 75.0 | #79 |
| MMLU-Pro | reasoning | 75.9% | #149 |
| GPQA | reasoning | 57.2% | #368 |
| Humanity's Last Exam |
| reasoning |
| 5.3% |
| #342 |
| LiveCodeBench | coding | 69.4% | #76 |
| SciCode | coding, reasoning | 26.2% | #376 |
| Output Speed | speed | 79.4 tok/s | #100 |
| Time to First Token | speed | 4.92s | #132 |
| Blended Price | cost | $0.300/M | #83 |
| Input Price | cost | $0.200/M | #97 |
| Output Price | cost | $0.600/M | #91 |
| Value Index | cost, overall | 30.0 | #129 |