Melhores modelos
Mercury 2 é atualmente o modelo mais bem classificado para velocidade nos dados disponíveis.
Use o modelo líder como ponto de partida, não como decisão automática.
Compare as melhores opções por preço, velocidade, disponibilidade e limites.
Abra a página de resultados e a metodologia antes de citar o ranking.
Median generated output tokens per second. Higher means the model streams completions faster after it starts responding.
| Posição | Modelo | Fornecedor | Valor |
|---|---|---|---|
| #1 | Mercury 2 | Inception | 880,1 tok/s |
| #2 | Granite 4.0 H Small | IBM | 439,5 tok/s |
| #3 | Granite 3.3 8B (Non-reasoning) | IBM | 414,9 tok/s |
| #4 | LFM2.5-VL-1.6B | Liquid AI | 392,7 tok/s |
| #5 | Step 3.7 Flash | StepFun | 385,2 tok/s |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Fonte: Artificial Analysis Output Speed. Atualizado: 16 de jul. de 2026.
| #6 | HyperNova 60B 2605 | Multiverse Computing | 350,9 tok/s |
| #7 | LFM2.5-8B-A1B | Liquid AI | 343,4 tok/s |
| #8 | Nemotron 3 Nano Omni 30B A3B Reasoning | NVIDIA | 320 tok/s |
| #9 | gpt-oss-120b (low) | OpenAI | 296,3 tok/s |
| #10 | Gemini 3.1 Flash-Lite | 291,6 tok/s |
| #11 | Llama 3.1 Nemotron Instruct 70B | NVIDIA | 286,3 tok/s |
| #12 | NVIDIA Nemotron Nano 12B v2 VL (Reasoning) | NVIDIA | 282,5 tok/s |
| #13 | gpt-oss-20b (low) | OpenAI | 260,9 tok/s |
| #14 | Step 3.5 Flash 2603 | StepFun | 256,9 tok/s |
| #15 | Nova Micro | Amazon | 249,8 tok/s |
| #16 | Step 3.5 Flash | StepFun | 247,4 tok/s |
| #17 | GPT-5.6 Luna (max) | OpenAI | 243,6 tok/s |
| #18 | Gemini 3.5 Flash (high) | 241,2 tok/s |
| #19 | Gemini 2.5 Flash-Lite (Reasoning) | 237,8 tok/s |
| #20 | Gemini 3.5 Flash (medium) | 234,7 tok/s |