Advanced video generation models (Hailuo) specializing in realistic physics, dynamic camera movements, and human expressions.
MiniMax's model ecosystem focuses on photorealistic video dynamics. MuAPI presents MiniMax models next to other providers so teams can compare capability, pricing and model fit before choosing an integration path.


MiniMax AI API models
Explore MiniMax models for chat, code, image and video generation, including Gemini, Nano Banana and Veo-style workflows available through MuAPI.
$0.300 / second
Fast and lightweight text-to-video generation. Ideal for quick drafts, previews, or playful content where speed matters more than cinematic quality.
$0.600 / second
High-fidelity text-to-video with cinematic rendering. Best for storytelling, cinematic clips, or realistic visuals with depth, atmosphere, and detail.
$0.630 / second
Hailuo 2.3 Pro I2V breathes life into still images with stunning motion synthesis and cinematic camera control. Using deep motion understanding, it predicts realistic subject movement, depth, and environmental motion from a single input frame — delivering smooth, film-grade clips.
$0.630 / second
Hailuo 2.3 Pro T2V turns your imagination into motion-picture realism. It interprets natural language prompts and generates visually stunning cinematic sequences that capture depth, atmosphere, and authentic motion.
$0.150 / second
Transforms an image into video with light, natural motion. Great for social media, quick animations, and previews.
$0.360 / second
Hailuo 2.3 Standard I2V converts still images into visually immersive motion clips with stable dynamics and realistic movement. It provides a balanced mix of quality, speed, and coherence. In 768p video generation.
$0.240 / second
Minimax Hailuo 2.3 Fast is the lightweight, high-speed version of the Hailuo 2.3 family — designed for creators who need instant video generation with cinematic motion and scene consistency. In 768p video generation.
$0.600 / second
Advanced image-to-video with cinematic realism. Adds dynamic camera motion, realistic physics, and atmospheric detail for storytelling.

$1.000 / second
MiniMax H3 Text to Video creates video from a written prompt through the Muapi API.

$1.000 / second
MiniMax H3 Reference to Video creates a 2K video from a prompt plus image, video, and optional audio references through the Muapi API.

$1.000 / second
MiniMax H3 Image to Video animates a source image with a motion prompt through the Muapi API.
$0.360 / second
Hailuo 2.3 Standard T2V transforms pure imagination into moving cinematic visuals. Simply describe a scene, and this model generates a coherent, high-quality video that captures the prompt’s tone, environment, and emotion. In 768p video generation.

$0.260 / second
MiniMax H3 Open Text to Video generates coherent videos with native audio from text prompts, supporting 480p/768p resolution and 5-15s duration.

$0.260 / second
MiniMax H3 Open Image to Video animates a first-frame image (with optional last-frame guidance) into coherent video with native audio, supporting 480p/768p resolution and 5-15s duration.

$0.330 / second
MiniMax H3 Open Reference to Video generates coherent videos from prompts and multimodal references, guided by up to 9 reference images, 3 reference videos, and 3 reference audios with native stereo audio.

$0.650 / 1K tokens
Minimax Voice Clone creates a high-fidelity digital clone of a speaker’s voice from a short reference audio sample. It reproduces the speaker’s tone, emotion, accent, rhythm, and speaking style, then generates new speech from any text input.

$0.650 / 1K tokens
Speech-2.6-hd is Minimax’s high-definition text-to-speech model that turns written text into natural, human-like audio. It produces studio-quality speech with clear pronunciation, smooth pacing, realistic emotion, and no background noise.

$0.650 / 1K tokens
Speech-2.6-turbo is Minimax’s fast, lightweight text-to-speech model designed for quick audio generation while maintaining good natural voice quality. It produces clear speech with smooth pacing and minimal delay.

$0.200 / 1K tokens
Generate a full song with vocals or an instrumental-only track from a text prompt and structured lyrics with MiniMax Music 3.0.

$0.150 / generation
MiniMax H3 Text to Video LoRA generates coherent 480P / 768P videos with native stereo audio from text prompts and custom LoRA adapters.

$0.150 / generation
MiniMax H3 Image to Video LoRA animates a first-frame image, optionally with last-frame guidance, into coherent 480P / 768P videos with native stereo audio and LoRA support.

$0.150 / generation
MiniMax H3 Reference to Video LoRA generates coherent 480P / 768P videos from text prompts and multimodal reference inputs (up to 9 images, 3 videos, 3 audios) with native stereo audio and custom LoRA adapters.