Seedance 2.5 Text to Video
Flagship Seedance 2 model with 4K support. Cinema-grade quality and premium motion control for production-ready content.
Seedance 2 is ByteDance's unified audio-video model for text-to-video, image-to-video, keyframe transitions, reference-driven scenes, editing, and character workflows.Muapi exposes 54 variants through one asynchronous API, with native audio, phoneme-aware lip-sync in 8+ languages, multi-shot director scripting, and pay-as-you-go pricing from $0.09/sec.
Flagship Seedance 2 model with 4K support. Cinema-grade quality and premium motion control for production-ready content.
Flagship image-to-video at up to 4K resolution. Maximum photorealism and motion precision in the Seedance family.
Next-generation text-to-video at up to 1080p. Enhanced motion quality and photorealism over Seedance 2.0.
Next-generation image-to-video at up to 1080p. Superior motion fidelity and visual consistency over Seedance 2.0.
Spicy variant — fast queue and a bolder visual style. Priority compute for text-to-video generation.
Spicy fast variant — quickest T2V with a bolder visual style and priority queue.
Spicy variant — fast queue and a bolder visual style. Priority image-to-video for high-volume workflows.
Spicy fast variant — fastest image-to-video available with a bolder visual style.
Mini Spicy variant — the fastest, lowest-cost text-to-video with a bolder visual style.
Mini Spicy variant — the fastest, lowest-cost image-to-video with a bolder visual style.
VIP variant — fast queue and low censorship. Priority compute for text-to-video generation.
VIP fast variant — quickest T2V with low censorship and priority queue.
VIP variant — fast queue and low censorship. Priority image-to-video for high-volume workflows.
VIP fast variant — fastest image-to-video available with low censorship.
VIP variant — fast queue and low censorship. Priority first/last-frame interpolation.
VIP fast variant — fastest first/last-frame interpolation with low censorship.
VIP variant — fast queue and low censorship. Priority reference-guided video generation.
VIP fast variant — quickest reference-guided option with low censorship.
VIP 1080p variant — full HD text-to-video with priority queue and low censorship.
VIP 1080p fast variant — high-speed full HD text-to-video with priority queue and low censorship.
VIP 1080p variant — animate images to full HD video with priority queue and low censorship.
VIP 1080p fast variant — fastest full HD image animation with priority queue and low censorship.
VIP 1080p variant — full HD reference-guided generation with up to 9 images, 3 videos, and 3 audio clips. Priority queue and low censorship.
VIP 1080p fast variant — fastest full HD reference-guided generation with priority queue and low censorship.
VIP 1080p variant — full HD first/last-frame interpolation with priority queue and low censorship.
VIP 4K variant — ultra-high-resolution 4K video from a text prompt with priority queue and low censorship.
VIP 4K variant — animate images to ultra-high-resolution 4K video with priority queue and low censorship.
VIP 4K variant — ultra-high-resolution 4K first/last-frame interpolation with priority queue and low censorship.
VIP 4K variant — ultra-high-resolution 4K reference-guided generation with up to 9 images, 3 videos, and 3 audio clips. Priority queue and low censorship.
VIP extend variant — fast queue and low censorship. Continues an existing Seedance 2.0 video at 720p while preserving visual style, motion, characters, and audio consistency.
VIP 1080p extend variant — full HD continuation of an existing Seedance 2.0 video with priority queue and low censorship. Preserves style, motion, and audio consistency.
Global variant. Generate cinematic videos from text prompts with precise motion and photorealistic quality.
Global fast variant. Reduced latency text-to-video, ideal for rapid iteration.
Global variant. Animate any image into a smooth, realistic video clip with full motion control.
Global fast variant. High-throughput image animation without sacrificing visual fidelity.
Global variant. Provide a start and end frame; seamlessly interpolates a fluid video between them.
Global fast variant. Same smooth first-to-last transitions at significantly lower latency.
Global variant. Use a reference image to guide character appearance across the generated video. No source video required.
Global fast variant. Maintain consistent character identity across scenes at high speed.
Remove SD 2.0 watermarks via AI inpainting. Flat $0.025 for clips up to 5s, then $0.005/sec beyond that.
Pro watermark removal for Seedance 2.0 videos. Flat $0.065 for clips up to 5s, then $0.013/sec beyond that.
Fastest and most affordable Seedance 2 variant. Ideal for rapid iteration, e-commerce assets, and high-volume batch workflows at 480p or 720p.
Animate images to video at the lowest cost in the Seedance 2 family. 1 image = start frame; up to 9 images = multi-reference via @image1–@image9.
Reference-guided mini-tier generation. Combine up to 9 images, 3 video clips, and 3 audio files with your prompt for character consistency at minimal cost.
Chinese Seedance 2 text-to-video at 480p resolution. Lower cost with low censorship — ideal for draft generation and rapid iteration.
Chinese Seedance 2 image-to-video at 480p resolution. Animate images at lower cost with low censorship.
Chinese Seedance 2 omni reference at 480p. Reference-guided generation at lower cost with low censorship.
Chinese Seedance 2 text-to-video with low censorship. Generate cinematic videos from text prompts.
Chinese Seedance 2 image-to-video with low censorship. Animate any image into a smooth, realistic video clip.
Chinese variant with low censorship. Reference-guided generation using a source video + image for strong character consistency.
Seamlessly continue an existing Seedance 2.0 video. Preserves visual style, motion, characters, and audio consistency across the new segment.
Edit existing videos using text prompts and optional reference images. Billing based on input video duration (max 15s).
Generate a reusable character sheet from reference images. Use the returned character ID in any Seedance 2 Omni Reference prompt.
Train a custom omni reference model on your subject. Returns a character ID for consistent identity across any scene or prompt.
Seedance 2 is ByteDance's unified audio-video generation model. It can create scenes from text, animate images, transition between first and last frames, follow multiple references, and apply edits or extensions to existing footage.
The family combines video generation with native dialogue, ambient sound, foley, and phoneme-aware lip-sync in 8+ languages. Director-mode scripting helps creators plan multi-shot sequences instead of generating isolated clips.
Muapi groups the current 54 variants behind one REST contract, so applications can switch between Chinese, Global, VIP, Mini, Spicy, Face Training, 2.1, and 2.5 options without another provider integration.
Guide a generation with reference images and supported source assets to keep characters, scenes, and style coherent.
Generate dialogue, ambience, foley, and music in the same audio-video pass.
Define the opening and closing images for a controlled visual transition.
Describe multi-shot sequences, camera movement, timing, and scene changes in one prompt.
Use premium variants when you need higher-resolution output and a priority queue.
Adapt the same generation workflow for landscape, portrait, square, and social-first formats.
Create talking-head and character videos with synchronized speech across supported languages.
Combine product references, camera directions, sound design, and multiple shots in a single workflow.
Turn scripts and keyframes into coherent clips for pitches, previsualization, and social content.
Use omni reference or face training to keep a recurring person or character recognizable.
Edit footage, extend a shot, or remove a watermark with a text instruction.
| Variant group | Resolution | Highlights | Best for |
|---|---|---|---|
| Chinese | 480p–720p | Core audio-video generation | General regional workflows |
| Global | 720p | Reference and fast international variants | Multilingual production |
| VIP | 720p–4K | Priority queue and low-censorship variants | Premium, high-volume generation |
| 2.1 / 2.5 | 720p | Newer Seedance family variants | Current-generation experiments |
Submit a prompt and generation settings, receive a request ID immediately, and poll the standard prediction result endpoint for the completed video URL.
curl -X POST https://api.muapi.ai/api/v1/seedance-2-text-to-video \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"A filmmaker walks through a rain-soaked neon market","resolution":"720p","duration":8,"aspect_ratio":"16:9"}'
# Response: {"request_id":"REQUEST_ID"}Start with the text-to-video endpoint; image, keyframe, reference, and VIP variants use the same Muapi task lifecycle.
curl https://api.muapi.ai/api/v1/predictions/REQUEST_ID/result \ -H "x-api-key: YOUR_API_KEY" # Poll until status is completed, then read the generated video URL.
Poll until the task is completed, then read the generated video URL and any synchronized audio output.
Seedance 2 is ByteDance's unified audio-video generation model on Muapi. It supports text-to-video, image-to-video, first/last-frame control, omni reference, multi-shot scripting, editing, and synchronized audio.
Muapi currently exposes 54 Seedance 2 model variants across Chinese, Global, VIP, Mini, Spicy, Face Training, 2.1, and 2.5 groups.
Omni Reference lets you supply multiple reference images so the model can maintain character appearance, scene composition, and style across a generation without a separate fine-tuning workflow.
Yes. Seedance 2 includes native audio and phoneme-aware lip-sync for 8+ languages, making it useful for dubbing, localization, and multilingual video.
Yes. The catalog includes video edit, extend, and watermark-removal variants that accept a source video and a text instruction.
Create a Muapi account, generate an API key, and call the endpoint matching your workflow. Access is available without a separate waitlist.
Use one API key for audio-video generation, references, keyframes, and production variants.