Melhores modelos
Command A+ é atualmente o modelo mais bem classificado para Time to First Token nos dados disponíveis.
Use o modelo líder como ponto de partida, não como decisão automática.
Compare as melhores opções por preço, velocidade, disponibilidade e limites.
Abra a página de resultados e a metodologia antes de citar o ranking.
Median time to first token. Lower means the model starts responding sooner, even if total generation speed differs.
| Posição | Modelo | Fornecedor | Valor |
|---|---|---|---|
| #1 | Command A+ | Cohere | 0,16 s |
| #2 | North Mini Code | Cohere | 0,18 s |
| #3 | NVIDIA Nemotron Nano 12B v2 VL (Reasoning) | NVIDIA | 0,22 s |
| #4 | Tiny Aya Global | Cohere | 0,22 s |
| #5 | Llama Nemotron Super 49B v1.5 (Reasoning) | NVIDIA |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Fonte: Artificial Analysis Time to First Token. Atualizado: 16 de jul. de 2026.
| 0,26 s |
| #6 | Llama Nemotron Super 49B v1.5 (Non-reasoning) | NVIDIA | 0,26 s |
| #7 | Gemini 2.5 Flash-Lite (Non-reasoning) | 0,32 s |
| #8 | Phi-4 Mini Instruct | Microsoft | 0,32 s |
| #9 | Hermes 3 - Llama-3.1 70B | Nous Research | 0,34 s |
| #10 | Command A | Cohere | 0,34 s |
| #11 | Cogito v2.1 (Reasoning) | Deep Cogito | 0,35 s |
| #12 | Phi-4 Multimodal Instruct | Microsoft | 0,37 s |
| #13 | Ministral 3 3B | Mistral | 0,37 s |
| #14 | Gemma 3n E4B Instruct | 0,38 s |
| #15 | Grok Build 0.1 0616 | SpaceXAI | 0,4 s |
| #16 | Qwen3.5 4B (Reasoning) | Alibaba | 0,4 s |
| #17 | Mistral Small 3.2 | Mistral | 0,4 s |
| #18 | NVIDIA Nemotron 3 Nano 30B A3B (Non-reasoning) | NVIDIA | 0,41 s |
| #19 | gpt-oss-20b (high) | OpenAI | 0,41 s |
| #20 | Qwen3.5 4B (Non-reasoning) | Alibaba | 0,43 s |