Image-to-Video
Built with a unified multimodal audio-video joint generation architecture, Seedance 2.0 supports four input modalities: text, image, audio, and video. Compared with Version 1.5, Seedance 2.0 delivers a substantial leap in generation quality. It achieves a higher usability rate for complex interaction and motion scenes, with significant improvements in physical accuracy, visual realism, and controllability, making it well-suited for high-quality creation scenarios.
1,347
Mar 2026
video
$7/1M
Catalog
Category rows come directly from Artificial Analysis when the endpoint exposes category-level Elo scores.
| Category | Elo | 95% CI | Appearances |
|---|---|---|---|
| Action | 1,469 | -30/30 | 773 |
| Abstract | 1,428 | -41/41 | 340 |
| Buildings | 1,393 | -16/16 | 2,337 |
| Screens | 1,393 | -38/38 | 401 |
| Fantasy | 1,392 | -36/36 | 423 |
| Transport | 1,379 | -21/21 | 1,327 |
| Moving camera | 1,378 | -16/16 | 2,414 |
| Physics | 1,367 | -18/18 | 1,806 |
| Sports | 1,366 |
| -24/24 |
| 971 |
| Cartoon and anime | 1,365 | -40/40 | 366 |
| People | 1,362 | -14/14 | 3,086 |
| Short prompt | 1,354 | -11/11 | 4,639 |