NVIDIA
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 is a large language model (LLM) trained by NVIDIA. The model employs a hybrid Mixture-of-Experts architecture, utilizing interleaved Mamba-2 and MoE layers, along with select Attention layers. The Lightning 3.5 model is released alongside a number of speculative decoding methods for faster text generation. The model has 3B active parameters and 30B parameters in total.
NVIDIA model catalogRank #331 across 687
Rank #228 across 330
Rank #28 across 432
Percentile score by analysis domain.
* Cost is inverted: lower input, output, and blended prices rank higher.
Higher bars mean stronger relative placement.
| Metric | Domain | Value | Rank |
|---|---|---|---|
| Artificial Analysis Intelligence Index | overall | 12.9 | #331 |
| Artificial Analysis Coding Index | coding | 26.8 | #228 |
| GPQA | reasoning | 74.3% | #277 |
| Humanity's Last Exam | reasoning | 10.6% | #298 |
| SciCode |
| coding, reasoning |
| 32.1% |
| #178 |
| Output Speed | speed | 267.8 tok/s | #9 |
| Time to First Token | speed | 0.46s | #17 |
| Blended Price | cost | $0.108/M | #28 |
| Input Price | cost | $0.070/M | #32 |
| Output Price | cost | $0.220/M | #33 |
| Value Index | cost, overall | 119.4 | #26 |