High-quality video generation with rich lighting, realistic physics, and cinematic texture controls.
Vidu's model ecosystem creates highly cinematic and physics-consistent video. MuAPI presents Vidu models next to other providers so teams can compare capability, pricing and model fit before choosing an integration path.


Vidu AI API models
Explore Vidu models for chat, code, image and video generation, including Gemini, Nano Banana and Veo-style workflows available through MuAPI.
$0.300 / second
Vidu's 2.0 model delivers advanced image-based video generation with enhanced lighting, emotion dynamics, and automatic frame interpolation for polished visual content.
$0.065 / second
Vidu Q2 Reference Video generates breathtaking cinematic clips from text prompts guided by multiple reference images. Each image refines the model’s understanding of subject, environment, and visual tone — ensuring perfect consistency in appearance and motion across every frame.
$0.200 / second
Vidu Q2 Pro Text-to-Video generates cinematic, prompt-faithful clips from text alone with strong temporal consistency and rich detail at up to 1080p. Pick this when you need polished output without a reference frame.
$0.300 / second
Vidu Q3 Turbo Text-to-Video is the fast, affordable tier of Vidu Q3 — same prompt understanding and motion quality, optimised for rapid iteration. Use it for storyboards, social cuts, and high-volume generation where speed and cost matter as much as polish.
$0.130 / second
Vidu Q2 Turbo Text-to-Video is the fast, affordable Q2 tier for prompt-only generation. Use it for storyboards, social cuts, and high-volume work where speed and cost matter.
$0.130 / second
Vidu Q2 Turbo Image-to-Video animates a starting image into a fast, prompt-guided clip while preserving subject identity. Built for speed and cost efficiency.
$0.750 / second
Vidu Q3 Pro Image-to-Video animates a single starting image into a smooth, prompt-guided clip up to 1080p. It preserves character identity, lighting, and composition while introducing natural motion, camera moves, and atmosphere — ideal for bringing concept art, product shots, and stills to life.
$0.300 / second
Vidu's 2.0 model offers enhanced visual quality and comprehensive workflow support across multiple resolution options for versatile content creation.
$0.400 / second
Vidu Q1 enables you to generate cinematic 1080p videos using multiple visual references—up to seven images—and text prompts. Designed for consistency, it preserves character appearance, props, and backgrounds across scenes while adding new motion and narrative elements.
$0.060 / second
Vidu Q2 Turbo Start–End Video creates highly detailed cinematic sequences by interpolating between two visual states — your start frame and end frame. Built for story moments, cinematic transformations, product reveals, and artistic transitions, it captures smooth motion, realistic lighting shifts, and dynamic camera movements while maintaining fidelity and emotional tone.
$0.130 / second
Vidu Q2 Pro Start–End Video is a professional-grade model built for cinematic transformation storytelling. It evolves a scene, subject, or concept from one moment to another through smooth visual interpolation, natural lighting transitions, and dynamic motion.
$0.750 / second
Vidu Q3 Pro Text-to-Video generates cinematic, prompt-faithful clips with strong temporal consistency, accurate motion, and rich detail across resolutions up to 1080p. Pick this when you want the highest visual fidelity Vidu Q3 can produce — great for hero shots, narrative beats, and stylized sequences driven purely from text.
$0.750 / second
Vidu Q3 Pro First-Last Frames interpolates a smooth, cinematic transition between two key images — your start frame and end frame — guided by a text prompt. Perfect for transformation reveals, scene transitions, product morphs, and storytelling beats that need a clean, controlled arc from A to B.
$0.300 / second
Vidu Q3 Turbo Image-to-Video animates a starting image into a fast, prompt-guided clip while keeping subject identity and composition intact. Built for speed and cost efficiency — perfect for batch animation, social content, and quick creative exploration.
$0.200 / second
Vidu Q2 Pro Image-to-Video animates a single starting image into a smooth, prompt-guided clip up to 1080p while preserving subject identity, lighting, and composition.
$0.300 / second
Vidu Q3 Turbo First-Last Frames interpolates a quick, cost-efficient transition between two key images — your start frame and end frame — guided by a text prompt. Great for transformation reveals, transitions, and short-form storytelling at scale.

$0.032 / generation
VIDU Reference-to-Image Q2 generates new high-quality images based on one or more reference images. It preserves the key identity, structure, or style of the reference while creating a new scene, variation, or enhanced composition. Ideal for character consistency, object re-interpretation, stylized redesigns, and cinematic recreations guided by reference inputs.

$0.040 / 1K tokens
VIDU Text-to-Image Q2 is a high-quality generative model focused on producing vivid, dynamic, and cinematic still images using natural language prompts. It excels at atmospheric depth, expressive lighting, surreal concepts, and motion-infused compositions typical of VIDU’s visual identity.