SD 2 Text-to-Video VIP 1080p by ByteDance. Generates cinematic 1080p video from a text prompt with priority routing, native audio-visual sync, and 4–15 second duration.
About this model
Seedance 2 Text-to-Video VIP 1080p generates cinematic full-HD video directly from a text prompt, with no image input required. Describe the visual scene, motion, mood, and style you want in natural language—the model synthesizes everything from scratch, outputting a coherent 1080p video with native audio-visual synchronization. Duration is flexible (4 to 15 seconds), and you can specify aspect ratios from cinema letterbox to mobile vertical or square formats.
This is the fastest route from idea to video for creators, marketers, and developers who want instant video from pure text. No need to hunt for stock footage or craft image references—just write what you envision, pick your aspect ratio, and let the model build the video. Muapi's unified API makes it accessible to builders of all kinds, with transparent per-generation pricing and no subscriptions.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $3.38 per generation | Muapi's transparent per-generation pricing for text-to-video makes pure prompt-based video generation affordable and predictable, with no subscription overhead. |
Muapi's transparent per-generation pricing for text-to-video makes pure prompt-based video generation affordable and predictable, with no subscription overhead.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Text description of the video to generate. | A cinematic shot of a futuristic city at night with neon lights reflecting on wet streets. |
| Aspect Ratio | Enum (6 options) | Output video aspect ratio. | 16:9 |
| Duration (seconds) | int | Video duration in seconds. | 5 |
Text description of the video to generate.
A cinematic shot of a futuristic city at night with neon lights reflecting on wet streets.Output video aspect ratio.
16:9Video duration in seconds.
5Developer documentation
Frequently asked
More specific is better. Include visual details (colors, lighting, composition), camera movement (pans, zooms, pulls), and mood or style references (cinematic, energetic, minimalist). Vague prompts produce generic results; rich prompts produce distinctive, compelling videos.
You can describe branded environments, generic people, or objects in detail. The model respects scene descriptions but may stylize certain elements. For brand-critical content, test with variations to see how closely the output matches your vision.
This model costs $3.38 per generation. You pay a flat fee for any video length between 4 and 15 seconds—no additional charges based on prompt complexity or duration within that range.