SD 2 Text-to-Video (Fast) by ByteDance. Generates video from text at faster speeds with 4–15 second duration and 2K resolution.
About this model
SD 2 Text-to-Video (Fast) generates video from text prompts using ByteDance's SD 2 model at faster speeds and lower cost. It produces videos with audio-visual sync, up to 2K resolution, and 4–15 second duration.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $0.15 per second (e.g. $0.75 for 5 seconds) | Pay per second of video generated. |
| Fal.ai | $0.3024/sec (high) / $0.2419/sec (basic) | Fal.ai charges $0.3024/sec for high quality and $0.2419/sec for basic. muapiapp is 50% cheaper — Fast model ($0.15/sec) vs Fal.ai basic ($0.2419/sec). |
| Replicate | $0.3024/sec (high) / $0.2419/sec (basic) | Replicate charges the same as Fal.ai — $0.3024/sec (high), $0.2419/sec (basic). muapiapp Fast at $0.15/sec is 50% cheaper than Replicate's basic rate. |
Pay per second of video generated.
Fal.ai charges $0.3024/sec for high quality and $0.2419/sec for basic. muapiapp is 50% cheaper — Fast model ($0.15/sec) vs Fal.ai basic ($0.2419/sec).
Replicate charges the same as Fal.ai — $0.3024/sec (high), $0.2419/sec (basic). muapiapp Fast at $0.15/sec is 50% cheaper than Replicate's basic rate.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Text description of the video to generate. Use @character:<id> to anchor the video to a Seedance 2 character — automatically switches to image-to-video mode. | A cinematic shot of a futuristic city at night with neon lights reflecting on wet streets. |
| Aspect Ratio | Enum (6 options) | Output video aspect ratio. | 16:9 |
| Duration (seconds) | int | Video duration in seconds. | 5 |
Text description of the video to generate. Use @character:<id> to anchor the video to a Seedance 2 character — automatically switches to image-to-video mode.
A cinematic shot of a futuristic city at night with neon lights reflecting on wet streets.Output video aspect ratio.
16:9Video duration in seconds.
5Developer documentation
Write your prompt: Describe the scene clearly. The fast model responds well to concise, vivid descriptions.
Set aspect ratio: Choose from 21:9, 16:9, 4:3, 1:1, 3:4, or 9:16.
Set duration: Between 4 and 15 seconds.
Submit and poll: Use the returned request_id to poll for results.
Frequently asked
The Fast model uses a lighter inference pass. Results are typically slightly lower in detail than Pro but are generated more quickly and at lower cost.
Six ratios: 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16. Default is 16:9.
Between 4 and 15 seconds. Default is 5 seconds.