SD 2 Text-to-Video (Pro) by ByteDance. Generates high-quality cinematic video from a text prompt with native audio-visual sync, up to 2K resolution, and 4–15 second duration.
About this model
SD 2 Text-to-Video (Pro) generates high-quality cinematic video from a text prompt. Powered by ByteDance's latest SD 2 model, it produces videos with native audio-visual synchronization, up to 2K resolution, and durations from 4 to 15 seconds. The Pro model prioritizes quality over speed.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $0.25 per second (e.g. $1.25 for 5 seconds) | Pay per second of video generated. |
| Fal.ai | $0.3024/sec (high) / $0.2419/sec (basic) | Fal.ai charges $0.3024/sec for high quality and $0.2419/sec for basic. muapiapp is 17% cheaper on high ($0.25/sec) and 50% cheaper on basic ($0.12/sec). |
| Replicate | $0.3024/sec (high) / $0.2419/sec (basic) | Replicate charges the same as Fal.ai — $0.3024/sec (high), $0.2419/sec (basic). muapiapp saves you 17–50% depending on quality tier. |
Pay per second of video generated.
Fal.ai charges $0.3024/sec for high quality and $0.2419/sec for basic. muapiapp is 17% cheaper on high ($0.25/sec) and 50% cheaper on basic ($0.12/sec).
Replicate charges the same as Fal.ai — $0.3024/sec (high), $0.2419/sec (basic). muapiapp saves you 17–50% depending on quality tier.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Text description of the video to generate. Use @character:<id> to anchor the video to a Seedance 2 character — automatically switches to image-to-video mode. | A cinematic shot of a futuristic city at night with neon lights reflecting on wet streets. |
| Aspect Ratio | Enum (6 options) | Output video aspect ratio. | 16:9 |
| Duration (seconds) | int | Video duration in seconds. | 5 |
Text description of the video to generate. Use @character:<id> to anchor the video to a Seedance 2 character — automatically switches to image-to-video mode.
A cinematic shot of a futuristic city at night with neon lights reflecting on wet streets.Output video aspect ratio.
16:9Video duration in seconds.
5Developer documentation
Write your prompt: Describe the scene, motion, lighting, and mood in detail. More descriptive prompts yield better results.
Set aspect ratio: Choose from 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16 to match your target platform.
Set duration: Choose between 4 and 15 seconds. Longer durations cost more.
Submit and poll: Submit the request to receive a request_id, then poll /predictions/{id}/result until status is completed.
Frequently asked
The Pro model (sd-2-text-to-video) prioritizes output quality and runs the full SD 2 model. The Fast variant (sd-2-text-to-video-fast) uses a lighter inference pass for quicker results at lower cost.
Six ratios are supported: 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. The default is 16:9.
You can generate videos between 4 and 15 seconds long. The default is 5 seconds.