Text-to-Video
Built with a unified multimodal audio-video joint generation architecture, Seedance 2.0 supports four input modalities: text, image, audio, and video. Compared with Version 1.5, Seedance 2.0 delivers a substantial leap in generation quality. It achieves a higher usability rate for complex interaction and motion scenes, with significant improvements in physical accuracy, visual realism, and controllability, making it well-suited for high-quality creation scenarios.
1,271
Mar 2026
video
$7/1M
Catalog
Category rows come directly from Artificial Analysis when the endpoint exposes category-level Elo scores.
| Category | Elo | 95% CI | Appearances |
|---|---|---|---|
| Fantasy | 1,393 | -31/31 | 720 |
| Multi-scene | 1,365 | -31/31 | 784 |
| Action | 1,359 | -23/23 | 1,172 |
| 3D animation | 1,335 | -39/39 | 328 |
| Specific location or era | 1,321 | -21/21 | 1,440 |
| Sports | 1,320 | -24/24 | 1,034 |
| Long prompt | 1,318 | -23/23 | 1,327 |
| Cartoon and anime | 1,317 | -24/24 | 1,059 |
| People |
| 1,311 |
| -11/11 |
| 5,251 |
| Buildings | 1,307 | -13/13 | 3,677 |
| Sci Fi | 1,305 | -20/20 | 1,806 |
| Text | 1,304 | -27/27 | 999 |