Skip to content
Easy Benchmarks
Best ModelsToolsProvidersUpdatesOpen Full Benchmarks
EnglishFrançaisEspañolDeutschPortuguês

Easy Benchmarks

Tell us what you want to achieve. Easy Benchmarks turns current quality, price, and speed data into a simple shortlist.

Benchmark data is attributed to Artificial Analysis. Model metadata and catalog pricing may also use Vercel AI Gateway data when available.

Latest ReportBenchmark GlossaryMethodologyllms.txt

Best Models

Best AI Model for Speed

Mercury 2 is currently the highest-ranked model for Speed in the available data.

Open Full BenchmarksAI Model Finder
Current Leader
Mercury 2
Value
880.1 tok/s
Dataset Coverage
314
Updated
Jul 16, 2026

How to use this result

  1. 1

    Start with the leader

    Use the top-ranked model as a starting point, not an automatic decision.

  2. 2

    Check the best alternatives

    Compare the top choices on price, speed, availability, and your own constraints.

  3. 3

    Verify the source

    Open the linked result page and methodology before quoting the ranking.

Top Models

Median generated output tokens per second. Higher means the model streams completions faster after it starts responding.

RankModelProviderValue
#1Mercury 2Inception880.1 tok/s
#2Granite 4.0 H SmallIBM439.5 tok/s
#3Granite 3.3 8B (Non-reasoning)IBM414.9 tok/s
#4LFM2.5-VL-1.6BLiquid AI392.7 tok/s
#5Step 3.7 FlashStepFun385.2 tok/s

Data & Evidence

This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.

Source: Artificial Analysis Output Speed. Updated: Jul 16, 2026.

View the live benchmark resultsView the original result pageRead the benchmark methodology

Related Pages

Best AI Model for OverallBest AI Model for CodingBest AI Model for MathBest AI Model for Reasoning & KnowledgeBest AI Model for Price & ValueBest AI Model for Artificial Analysis Math Index
#6HyperNova 60B 2605Multiverse Computing350.9 tok/s
#7LFM2.5-8B-A1BLiquid AI343.4 tok/s
#8Nemotron 3 Nano Omni 30B A3B ReasoningNVIDIA320 tok/s
#9gpt-oss-120b (low)OpenAI296.3 tok/s
#10Gemini 3.1 Flash-LiteGoogle291.6 tok/s
#11Llama 3.1 Nemotron Instruct 70BNVIDIA286.3 tok/s
#12NVIDIA Nemotron Nano 12B v2 VL (Reasoning)NVIDIA282.5 tok/s
#13gpt-oss-20b (low)OpenAI260.9 tok/s
#14Step 3.5 Flash 2603StepFun256.9 tok/s
#15Nova MicroAmazon249.8 tok/s
#16Step 3.5 FlashStepFun247.4 tok/s
#17GPT-5.6 Luna (max)OpenAI243.6 tok/s
#18Gemini 3.5 Flash (high)Google241.2 tok/s
#19Gemini 2.5 Flash-Lite (Reasoning)Google237.8 tok/s
#20Gemini 3.5 Flash (medium)Google234.7 tok/s