Advanced video generation models (Hailuo) specializing in realistic physics, dynamic camera movements, and human expressions.

MiniMax's model ecosystem focuses on photorealistic video dynamics. MuAPI presents MiniMax models next to other providers so teams can compare capability, pricing and model fit before choosing an integration path.

MiniMax AI models on MuAPI
Back to Providers
Explore/MiniMax Models
MiniMax

MiniMax AI API models

MiniMax AI models on MuAPI

Explore MiniMax models for chat, code, image and video generation, including Gemini, Nano Banana and Veo-style workflows available through MuAPI.

All models

23 Models

Video Generation Models

Video

$0.300 / second

minimax-hailuo-02-standard-t2v

Fast and lightweight text-to-video generation. Ideal for quick drafts, previews, or playful content where speed matters more than cinematic quality.

Video

$0.600 / second

minimax-hailuo-02-pro-t2v

High-fidelity text-to-video with cinematic rendering. Best for storytelling, cinematic clips, or realistic visuals with depth, atmosphere, and detail.

Video

$0.630 / second

minimax-hailuo-2.3-pro-i2v

Hailuo 2.3 Pro I2V breathes life into still images with stunning motion synthesis and cinematic camera control. Using deep motion understanding, it predicts realistic subject movement, depth, and environmental motion from a single input frame — delivering smooth, film-grade clips.

Video

$0.630 / second

minimax-hailuo-2.3-pro-t2v

Hailuo 2.3 Pro T2V turns your imagination into motion-picture realism. It interprets natural language prompts and generates visually stunning cinematic sequences that capture depth, atmosphere, and authentic motion.

Video

$0.150 / second

minimax-hailuo-02-standard-i2v

Transforms an image into video with light, natural motion. Great for social media, quick animations, and previews.

Video

$0.360 / second

minimax-hailuo-2.3-standard-i2v

Hailuo 2.3 Standard I2V converts still images into visually immersive motion clips with stable dynamics and realistic movement. It provides a balanced mix of quality, speed, and coherence. In 768p video generation.

Video

$0.240 / second

minimax-hailuo-2.3-fast

Minimax Hailuo 2.3 Fast is the lightweight, high-speed version of the Hailuo 2.3 family — designed for creators who need instant video generation with cinematic motion and scene consistency. In 768p video generation.

Video

$0.600 / second

minimax-hailuo-02-pro-i2v

Advanced image-to-video with cinematic realism. Adds dynamic camera motion, realistic physics, and atmospheric detail for storytelling.

minimax-h3-text-to-video
Video

$1.000 / second

minimax-h3-text-to-video

MiniMax H3 Text to Video creates video from a written prompt through the Muapi API.

minimax-h3-reference-to-video
Video

$1.000 / second

minimax-h3-reference-to-video

MiniMax H3 Reference to Video creates a 2K video from a prompt plus image, video, and optional audio references through the Muapi API.

minimax-h3-image-to-video
Video

$1.000 / second

minimax-h3-image-to-video

MiniMax H3 Image to Video animates a source image with a motion prompt through the Muapi API.

Video

$0.360 / second

minimax-hailuo-2.3-standard-t2v

Hailuo 2.3 Standard T2V transforms pure imagination into moving cinematic visuals. Simply describe a scene, and this model generates a coherent, high-quality video that captures the prompt’s tone, environment, and emotion. In 768p video generation.

minimax-h3-open-text-to-video
Video

$0.260 / second

minimax-h3-open-text-to-video

MiniMax H3 Open Text to Video generates coherent videos with native audio from text prompts, supporting 480p/768p resolution and 5-15s duration.

minimax-h3-open-image-to-video
Video

$0.260 / second

minimax-h3-open-image-to-video

MiniMax H3 Open Image to Video animates a first-frame image (with optional last-frame guidance) into coherent video with native audio, supporting 480p/768p resolution and 5-15s duration.

minimax-h3-open-reference-to-video
Video

$0.330 / second

minimax-h3-open-reference-to-video

MiniMax H3 Open Reference to Video generates coherent videos from prompts and multimodal references, guided by up to 9 reference images, 3 reference videos, and 3 reference audios with native stereo audio.