Text-to-Video
State-of-the-art video generation across quality, cost, and latency. Grok Imagine is x.AI's most powerful video-audio generative model yet. Bring an image to life, start from a simple text prompt, or even refine a complex cinematic sequence.
1,232
Jan 2026
video
$0.05/s
Catalog
Category rows come directly from Artificial Analysis when the endpoint exposes category-level Elo scores.
| Category | Elo | 95% CI | Appearances |
|---|---|---|---|
| Cartoon and anime | 1,334 | -24/24 | 797 |
| Fantasy | 1,324 | -30/30 | 533 |
| Action | 1,322 | -22/22 | 909 |
| 3D animation | 1,309 | -40/40 | 257 |
| Sci Fi | 1,305 | -20/20 | 1,178 |
| Sports | 1,301 | -24/24 | 736 |
| Fashion | 1,284 | -26/26 | 676 |
| Multi-scene | 1,281 | -29/29 | 593 |
| Long prompt |
| 1,279 |
| -22/22 |
| 1,022 |
| People | 1,275 | -11/11 | 3,695 |
| Specific location or era | 1,266 | -20/20 | 1,111 |
| Buildings | 1,258 | -13/13 | 2,518 |