Best Models
Gemma 3n E4B Instruct is currently the highest-ranked model for Output Price in the available data.
Use the top-ranked model as a starting point, not an automatic decision.
Compare the top choices on price, speed, availability, and your own constraints.
Open the linked result page and methodology before quoting the ranking.
Provider price for generating 1M output tokens. Output tokens are often priced higher than input tokens, so this can dominate long responses.
| Rank | Model | Provider | Value |
|---|---|---|---|
| #1 | Gemma 3n E4B Instruct | $0.04 / 1M | |
| #2 | Ministral 3 3B | Mistral | $0.10 / 1M |
| #3 | Granite 4.1 8B | IBM | $0.10 / 1M |
| #4 | Llama 3.1 Instruct 8B | Meta | $0.10 / 1M |
| #5 | Llama 3.2 Instruct 1B | Meta | $0.10 / 1M |
This answer uses the exact same metric and live database ranking as the linked Easy Benchmarks leaderboard.
Source: Artificial Analysis Output Price. Updated: Jul 16, 2026.
| #6 | Sarvam 30B (high) | Sarvam | $0.11 / 1M |
| #7 | Nova Micro | Amazon | $0.14 / 1M |
| #8 | HyperNova 60B 2605 | Multiverse Computing | $0.14 / 1M |
| #9 | Llama 3 Instruct 8B | Meta | $0.145 / 1M |
| #10 | Ministral 3 8B | Mistral | $0.15 / 1M |
| #11 | Qwen3.5 4B (Non-reasoning) | Alibaba | $0.15 / 1M |
| #12 | Qwen3.5 9B (Reasoning) | Alibaba | $0.15 / 1M |
| #13 | Qwen3.5 4B (Reasoning) | Alibaba | $0.15 / 1M |
| #14 | Llama 3.2 Instruct 3B | Meta | $0.15 / 1M |
| #15 | Solar Mini | Upstage | $0.15 / 1M |
| #16 | NVIDIA Nemotron Nano 9B V2 (Reasoning) | NVIDIA | $0.16 / 1M |
| #17 | Sarvam 105B (high) | Sarvam | $0.17 / 1M |
| #18 | NVIDIA Nemotron Nano 9B V2 (Non-reasoning) | NVIDIA | $0.195 / 1M |
| #19 | gpt-oss-20b (low) | OpenAI | $0.20 / 1M |
| #20 | gpt-oss-20b (high) | OpenAI | $0.20 / 1M |