Generate videos from text prompts with Seedance 2 — advanced camera control, native audio-video sync, and high-resolution output. Describe your scene and get a cinematic clip. Try free — pay per generation, no subscription.
About this model
SD 2.0 Text-to-Video is ByteDance's most advanced text-driven video generation model. Describe any scene in natural language and the model produces a cinematic clip with director-level camera control, native audio-video sync, and up to 2K resolution output. It understands complex prompts — lighting, motion physics, mood, and multi-shot storytelling — turning words into high-fidelity video sequences up to 15 seconds long.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $0.60 per video | muapiapp offers SD 2.0 Text-to-Video starting at $0.60 per video (5s, basic quality), scaling at $0.12/sec for basic and $0.25/sec for high quality across 5–15 second durations. |
| Fal.ai | $0.3024/sec (high) / $0.2419/sec (basic) | Fal.ai charges $0.3024/sec for high quality and $0.2419/sec for basic. muapiapp is 17% cheaper on high quality ($0.25/sec) and 50% cheaper on basic quality ($0.12/sec). |
| Replicate | $0.3024/sec (high) / $0.2419/sec (basic) | Replicate charges the same as Fal.ai — $0.3024/sec (high), $0.2419/sec (basic). muapiapp saves you 17–50% depending on quality tier. |
muapiapp offers SD 2.0 Text-to-Video starting at $0.60 per video (5s, basic quality), scaling at $0.12/sec for basic and $0.25/sec for high quality across 5–15 second durations.
Fal.ai charges $0.3024/sec for high quality and $0.2419/sec for basic. muapiapp is 17% cheaper on high quality ($0.25/sec) and 50% cheaper on basic quality ($0.12/sec).
Replicate charges the same as Fal.ai — $0.3024/sec (high), $0.2419/sec (basic). muapiapp saves you 17–50% depending on quality tier.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Text prompt describing the video. To use a fictional character, reference it inline with @character:<id> (the request_id from a completed Seedance 2 Character generation). Multiple characters are supported. Example: '@character:ab539e5f walks on the beach at sunset'. | A determined penguin straps itself into a homemade rocket sled on an icy mountain. The rocket ignites with a massive burst and launches the penguin across the frozen landscape at insane speed, blasting through snowdrifts and leaving a fiery trail behind. |
| Aspect Ratio | Enum (4 options) | - | 16:9 |
| Duration | Enum (3 options) | - | 5 |
| Quality | Enum (2 options) | - | basic |
Text prompt describing the video. To use a fictional character, reference it inline with @character:<id> (the request_id from a completed Seedance 2 Character generation). Multiple characters are supported. Example: '@character:ab539e5f walks on the beach at sunset'.
A determined penguin straps itself into a homemade rocket sled on an icy mountain. The rocket ignites with a massive burst and launches the penguin across the frozen landscape at insane speed, blasting through snowdrifts and leaving a fiery trail behind.-
16:9-
5-
basicDeveloper documentation
Write a Detailed Prompt: Describe the scene, subjects, lighting, mood, and camera movement. Be specific — 'slow dolly zoom into a neon-lit street at night' will outperform 'city street'.
Choose Quality: Select basic ($0.12/sec) for fast drafts or high ($0.25/sec) for final cinematic output.
Set Duration: Choose 5, 10, or 15 seconds. Longer durations allow richer storytelling.
Pick Aspect Ratio: Use 16:9 for widescreen, 9:16 for mobile/social, 4:3 or 3:4 for other formats.
Submit and Poll: You'll receive a request_id immediately. Poll the result endpoint until status is completed.
Frequently asked
It's ByteDance's state-of-the-art text-to-video model that generates cinematic clips from natural language prompts, with support for complex camera movements, native audio, and up to 2K resolution.
Basic quality uses the fast-t2v model at $0.12/sec — ideal for drafts and iteration. High quality uses the standard-t2v model at $0.25/sec for final, cinema-grade output with richer detail and smoother motion.
Yes, SD 2.0 generates audio natively alongside video, ensuring cinema-grade sound synchronized with the visual content.
SD 2.0 supports up to 2K resolution output.