用于聊天、推理和编码的多模态 Gemini 模型,以及专用的图像(Nano Banana)和视频(Veo)API。
Google 的模型生态覆盖对话式 AI、编码支持、图像生成和视频生成。MuAPI 将 Google 模型与其他提供商并列,帮助团队比较能力、价格和适合的集成方式。

$2.500 / 秒
VEO3 I2V animates static images into expressive video sequences, adding lifelike movement while preserving the original composition.
$0.600 / 秒
Get the ultra-high-definition 4K version of a Veo3.1 video generation task. This model is optimized for producing crisp, detailed videos suitable for professional and cinematic applications. It enhances visual fidelity while maintaining temporal coherence and realistic motion.
$0.600 / 秒
VEO3 Fast T2V creates short videos from text instantly, balancing speed and quality for quick content generation and prototyping.
$0.800 / 秒
Gemini Omni 1.1 Flash — text-to-video generation with native synchronized audio, supporting optional reference images, a reference video clip, voice profile, and character reference inputs in one call.
$0.800 / 秒
Gemini Omni 1.1 Flash Image to Video — animate a starting keyframe image into a video, with an optional ending keyframe image for controlled first/last-frame transitions, and native synchronized audio.
$0.800 / 秒
Gemini Omni 1.1 Flash Video Edit — restyle or edit a source video with a single text instruction, powered by Gemini Omni's natively multimodal any-to-any model.
$1.280 / 秒
Gemini Omni 1.1 Flash Reference to Video — generate video referencing up to 7 images and up to 3 short reference video clips (each up to 3 seconds), addressed by tag directly in the prompt.
$0.600 / 秒
Veo 3.1 Fast T2V is a high-speed AI video model that transforms text prompts into realistic 8-second videos. It emphasizes rapid generation while maintaining visual quality, accurate scene representation, and smooth motion. Ideal for social media, creative storytelling, or rapid concept visualization, it supports cinematic framing, dynamic lighting, and natural object movements.
$0.600 / 秒
Veo 3.1 Fast is an optimized version of Google’s Veo 3.1 AI that transforms static images into dynamic 8-second videos at higher speed. It preserves visual fidelity while enabling rapid generation, making it ideal for social media clips, storyboards, and quick creative previews.
$0.600 / 秒
Veo 3.1’s Extend Video mode lets you continue or expand an existing video clip seamlessly. Starting from a short generated video, you can prompt the model to extend the scene—keeping visual style, characters, motion, and audio consistent. This model needs original task_id of the video.
$0.300 / 秒
Veo 3.1 Lite is a lightweight variant of Google's Veo 3.1 model designed for faster, more accessible video generation from images.
$0.300 / 秒
Veo 3.1 Lite is a lightweight variant of Google's Veo 3.1 model designed for faster, more accessible video generation.
$3.000 / 秒
Veo 4 Image to Video — animate any still image with Veo 4's motion synthesis engine, supporting fine-grained camera control and realistic physics at up to 1080p.
$1.500 / 秒
Gemini Omni Image to Video — animate one or more reference images with a text prompt. Unified reasoning across modalities preserves subject identity and generates synchronized audio natively.
$2.500 / 秒
VEO3 T2V generates cinematic videos from text prompts, capturing dynamic motion, rich scenes, and storytelling visuals in stunning detail.
$0.600 / 秒
Quickly transform static images into short, motion-rich video clips with fast rendering and impressive quality — powered by Google's VEO3 on MuAPI.
$2.500 / 秒
Veo 3.1 is Google's advanced AI video generation model that allows users to create high-quality, 8-second videos from static images. This feature is particularly useful for transforming concept art, storyboards, or static visuals into dynamic video clips with synchronized audio.
$2.500 / 秒
Veo 3.1 is Google's advanced AI video generation model that transforms text prompts into high-quality videos. This model offers enhanced realism, richer audio, and improved narrative control, making it suitable for creators seeking cinematic-quality content.
$0.600 / 秒
Veo 3.1 R2V allows creators to generate dynamic videos using up to three reference images. The model maintains visual consistency of characters, objects, and style throughout the video, producing cinematic-quality 8-second clips. It’s perfect for turning concept art, storyboards, or character designs into short, animated sequences while preserving original aesthetics.
$3.000 / 秒
Veo 4 Text to Video — Google DeepMind's fourth-generation model delivering photorealistic, high-fidelity 1080p videos with exceptional prompt adherence and cinematic camera control.
$1.500 / 秒
Gemini Omni — natively multimodal any-to-any model. Generates high-fidelity video with synchronized audio directly from text prompts, with unified reasoning across modalities for more coherent scenes and fewer pipeline artifacts.
$2.400 / 秒
Gemini Omni Video Edit — natively multimodal video-to-video editing. Restyle, relight, swap subjects, or rewrite scenes from a source clip with a single prompt. Unified reasoning across modalities preserves motion and audio continuity while applying the edit.

$0.300 / 次生成
Generate a pack of high-quality, professional portraits in various styles (LinkedIn, CEO, Tinder, etc.) while preserving your facial features.

$0.030 / 次生成
Nano Banana is a mysterious, high-performance image model. It excels at precise, language-driven edits and consistent character preservation, allowing users to modify images with natural text commands.

$0.030 / 次生成
Nano Banana Effects is a creative visual effects model designed to transform ordinary images into fun, stylized, and eye-catching results. It applies artistic filters, 3D styles, cartoon transformations, and trending viral looks with a single click.

$0.030 / 1K tokens
Nano Banana is an advanced AI model excelling in natural language-driven image generation and editing. It produces hyper-realistic, physics-aware visuals with seamless style transformations.

$0.030 / 1K tokens
Google Imagen 4 is the latest text-to-image AI model from DeepMind, designed to produce stunningly photorealistic images with crisp detail, accurate text rendering, and creative flexibility. It supports high-resolution output (up to 2K), generates visuals in seconds, and embeds SynthID watermarks for authenticity.

$0.020 / 1K tokens
Imagen 4 Fast is optimized for speed and accessibility, allowing you to generate high-quality images in seconds. While slightly less detailed than the Ultra version, it excels at rapid ideation, drafts, storyboarding, and casual creativity.

$0.120 / 次生成
Nano Banana 2 Edit is the next-generation image editing model developed by Google DeepMind, following the original Nano Banana (also known as Gemini 2.5 Flash Image). It offers advanced image-edit capabilitie with improved resolution.

$0.060 / 1K tokens
Imagen 4 Ultra is Google’s flagship model, designed for photorealism, rich textures, and production-level imagery. It produces crisp, high-resolution visuals with advanced detail, lighting precision, and natural compositions.

$0.060 / 1K tokens
Nano Banana 2 (Gemini 3.1 Flash Image) is Google's most advanced image generation model, combining speed with high-fidelity 4K output and revolutionary character consistency.
$0.000 / 次生成
Generate a reusable character from a single reference image and a text description. Optionally attach a voice profile created with Gemini Omni Audio to give the character a consistent voice in future video generations.

$0.030 / 1K tokens
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest and most cost-efficient text-to-image model, delivering 4-second generation with exceptional prompt adherence, character consistency, and legible in-image text rendering.

$0.030 / 次生成
Nano Banana 2 Lite Edit (Gemini 3.1 Flash Lite Image) is Google's fastest and most cost-efficient image editing model, blending up to 14 reference images with exceptional prompt adherence and character consistency.

$0.120 / 1K tokens
Nano Banana 2 is the next-generation image generation developed by Google DeepMind, following the original Nano Banana (also known as Gemini 2.5 Flash Image). It offers advanced text-to-image capabilitie with improved resolution.

$0.060 / 次生成
Nano Banana 2 (Gemini 3.1 Flash Image) is Google's most advanced image generation model, combining speed with high-fidelity 4K output and revolutionary character consistency.

$0.004 / 1K tokens
Gemini Audio Vision uses Google Gemini's native audio understanding to analyze and describe audio content in detail — speech, tone, background sounds, speaker changes, and more. Upload an audio URL and a prompt, and Gemini returns a detailed text analysis. Token-based pricing.

$0.000 / 1K tokens
Gemini 3.7 Flash (OpenAI-compatible) is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: $0.60/M input tokens and $3.60/M output tokens. Two endpoints: standard async (/gemini-3-7-flash-openai) and live streaming (/gemini-3-7-flash-openai/stream) via SSE.

$0.001 / 1K tokens
Gemini 3 Flash is a fast, multimodal language model for real-time text generation. Supports text and image inputs, function calling, and Google Search grounding. Token-based pricing: $0.30/M input tokens and $1.80/M output tokens. Two endpoints: standard async (/gemini-3-flash) and live streaming (/gemini-3-flash/stream) via SSE.

$0.000 / 1K tokens
Gemini 3.5 Flash is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: $0.60/M input tokens and $3.60/M output tokens. Two endpoints: standard async (/gemini-3-5-flash) and live streaming (/gemini-3-5-flash/stream) via SSE.

$0.000 / 1K tokens
Gemini 3.5 Flash (OpenAI-compatible) is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: $0.60/M input tokens and $3.60/M output tokens. Two endpoints: standard async (/gemini-3-5-flash-openai) and live streaming (/gemini-3-5-flash-openai/stream) via SSE.

$0.001 / 1K tokens
Gemini 3 Pro is Google's powerful multimodal reasoning model, designed for complex problem solving, coding, and logical tasks. Supports text and image inputs. Token-based pricing: $4.00/M input tokens, $24.00/M output tokens. Two endpoints: standard async (/gemini-3-pro) and live streaming (/gemini-3-pro/stream) via SSE.

$0.001 / 1K tokens
Gemini 3.1 Pro is Google's next-generation multimodal model, optimized for complex reasoning, planning, coding, and multi-turn conversation. Supports text and image inputs. Token-based pricing: $4.00/M input tokens, $24.00/M output tokens. Two endpoints: standard async (/gemini-3-1-pro) and live streaming (/gemini-3-1-pro/stream) via SSE.

$0.000 / 1K tokens
Gemini 2.5 Pro is Google's advanced multimodal reasoning model, optimized for complex coding, logical tasks, and deep analysis. Supports text and image inputs. Token-based pricing: $1.25/M input tokens, $10.00/M output tokens. Two endpoints: standard async (/gemini-2-5-pro) and live streaming (/gemini-2-5-pro/stream) via SSE.

$0.000 / 1K tokens
Gemini 2.5 Flash is Google's high-speed multimodal language model, optimized for rapid text generation, real-time image understanding, and high-frequency tasks. Supports text and image inputs. Token-based pricing: $0.30/M input tokens, $2.50/M output tokens. Two endpoints: standard async (/gemini-2-5-flash) and live streaming (/gemini-2-5-flash/stream) via SSE.

$0.004 / 1K tokens
Gemini Video Vision uses Google Gemini's native video understanding to analyze and describe video content in detail — motion, composition, subjects, on-screen text, and more. Upload a video URL and a prompt, and Gemini returns a detailed text analysis. Token-based pricing.

$0.035 / 1K tokens
Gemini 3.1 Flash TTS turns written dialogue into expressive, natural multi-speaker speech with fine-grained control over voice, accent, emotional style, and pace. Ideal for fast, affordable voiceovers, character dialogue, and narration.

$0.035 / 1K tokens
Gemini 2.5 Pro TTS is Google's premium text-to-speech model for studio-quality, high-fidelity multi-speaker audio with expressive control over voice, accent, emotional style, and pace.

$0.000 / 1K tokens
Gemini 3.6 Flash is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: .60/M input tokens and .60/M output tokens. Two endpoints: standard async (/gemini-3-6-flash) and live streaming (/gemini-3-6-flash/stream) via SSE.

$0.000 / 1K tokens
Gemini 3.6 Flash (OpenAI-compatible) is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: .60/M input tokens and .60/M output tokens. Two endpoints: standard async (/gemini-3-6-flash-openai) and live streaming (/gemini-3-6-flash-openai/stream) via SSE.

$0.000 / 1K tokens
Gemini 3.7 Flash is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: $0.60/M input tokens and $3.60/M output tokens. Two endpoints: standard async (/gemini-3-7-flash) and live streaming (/gemini-3-7-flash/stream) via SSE.