Models/Google AI/Gemini API
Live6 featured endpoints

Google Gemini API — AI Video, Image & Text-to-Speech

Use Google Gemini media models on Muapi for AI image generation, cinematic video, and multi-speaker speech synthesis. One Muapi API key gives your application a consistent submit-and-poll workflow across these endpoints.This page focuses on media generation workflows. Gemini model families use different request fields and pricing; follow each family link for the full lineup, exact parameters, and current availability.

Nano Banana

Google image generation and editing family. Browse the hub for all variants.1 featured endpoints
ImageNewFeatured
Nano Banana

Nano Banana 2

Google image generation endpoint; the family hub includes additional generation and editing variants.

Image generation
$0.06/request
Try Model

Imagen 4

Google text-to-image models with standard, fast, and ultra tiers.1 featured endpoints
ImageLive
Imagen 4

Imagen 4 Fast

Fast text-to-image generation in Google’s Imagen 4 family.

Image generationFast
$0.02/generation
Try Model

Veo 3.1

Video generation, image-to-video, reference, extension, fast, and lite variants.2 featured endpoints
VideoLive
Veo 3.1

Veo 3.1 Text to Video

Generate a video from a prompt with the Veo 3.1 family.

Video generation
$2.50/request
Try Model
VideoLive
Veo 3.1

Veo 3.1 Fast Text to Video

Faster Veo 3.1 prompt-to-video generation.

Video generationFast
$0.60/request
Try Model

Gemini Omni

Video generation and editing workflows with synchronized audio.1 featured endpoints
VideoLive
Gemini Omni

Gemini Omni Text to Video

Generate video with synchronized audio using Gemini Omni.

Video + synchronized audio
From $0.16
Try Model

Gemini Text-to-Speech

Generate expressive multi-speaker dialogue audio from text.1 featured endpoints
Text to SpeechNew
Gemini TTS

Gemini 3.8 Flash TTS

Create expressive multi-speaker dialogue audio from structured text.

Dialogue to audioFast
$0.015/1,000 chars
Try Model

Google Gemini media API families

Google’s generative media lineup on Muapi spans image generation and editing, video creation, and text-to-speech. Nano Banana and Imagen cover images; Veo and Gemini Omni cover video; Gemini TTS generates spoken dialogue.

This landing page highlights representative endpoints so you can compare the families without repeating every model row from their dedicated hubs. Use the family pages for full catalogs, parameter details, and current availability.

What you can build

Create and edit images

Use Nano Banana for prompt-based image generation and reference-based edits, or Imagen 4 for high-quality text-to-image work.

Generate video

Call Veo 3.1 or Gemini Omni for prompt-to-video, image-driven motion, and related video workflows.

Synthesize speech

Send structured speakers and dialogue turns to Gemini TTS to generate multi-speaker audio.

Google AI API endpoints

FamilyInputOutputStarting priceFull lineup
Nano BananaPrompt and optional referencesImageFrom $0.03/request/nano-banana-api
Imagen 4PromptImageFrom $0.02/generation/imagen-4
Veo 3.1Prompt, image, or video referencesVideoFrom $0.30/request/veo-3.1
Gemini OmniPrompt, image, or video depending on endpointVideo, edits, or audioFrom $0.048/sec on selected endpoints/gemini-omni
Gemini TTSSpeaker and dialogue turnsAudioFrom $0.01/1,000 characters/gemini-tts

Pricing varies by model, duration, resolution, and endpoint. Confirm exact costs on each family page before submitting.

How to call the Google Gemini API

Submit a model-specific JSON request using your Muapi API key, save the request ID, and poll the shared result endpoint. Image, video, and speech endpoints use different request fields.

1. Generate an image with Nano Banana

curl -X POST https://api.muapi.ai/api/v1/nano-banana \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"A small glass greenhouse on a mossy hillside at sunrise","aspect_ratio":"1:1"}'

# Response: {"request_id":"REQUEST_ID"}

Send a prompt to a Nano Banana image endpoint.

2. Generate an image with Imagen 4

curl -X POST https://api.muapi.ai/api/v1/google-imagen4 \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"A richly detailed glasshouse at sunrise, dew on tropical leaves","aspect_ratio":"16:9","num_images":1}'

# Response: {"request_id":"REQUEST_ID"}

Submit a prompt and aspect ratio to Imagen 4.

3. Generate video with Veo 3.1

curl -X POST https://api.muapi.ai/api/v1/veo3.1-text-to-video \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"A cinematic sunrise over a misty mountain valley","duration":8}'

# Response: {"request_id":"REQUEST_ID"}

Call Veo with a prompt and duration.

4. Generate video with Gemini Omni

curl -X POST https://api.muapi.ai/api/v1/gemini-omni-text-to-video \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"A chef explains a recipe in a warm studio kitchen","resolution":"1080p","duration":8,"aspect_ratio":"16:9"}'

# Response: {"request_id":"REQUEST_ID"}

Create video with synchronized audio.

5. Generate speech with Gemini TTS

curl -X POST https://api.muapi.ai/api/v1/gemini-3-8-flash-tts \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Kore","accent":"Neutral","style":"Empathetic","pace":"Natural"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"Welcome to the show."}]}'

# Response: {"request_id":"REQUEST_ID"}

Provide structured speaker and dialogue turns.

6. Poll for the completed output

curl https://api.muapi.ai/api/v1/predictions/REQUEST_ID/result \
  -H "x-api-key: YOUR_API_KEY"

# Poll until status is completed, then read the generated asset URL from outputs.

Use the returned request ID with the shared result endpoint.

Google Gemini API FAQ

Which Google AI models can I call through Muapi?

The featured media families include Nano Banana, Imagen 4, Veo, Gemini Omni, and Gemini Text-to-Speech. Follow each family page for its complete endpoint lineup.

Does this page cover Gemini text chat models?

This catalog focuses on Google media-generation endpoints: images, video, and speech. Check Muapi’s live API schema for the currently available text-generation models and their request formats.

How do I get a Google Gemini API key?

Create a Muapi API key from Access Keys, then send it in the x-api-key header to the model endpoint.

Do all Google models use the same request format?

No. Image, video, and speech endpoints take different fields. Use the model-specific family pages and API examples for exact request schemas.

Build with Google AI media models

Use one Muapi API key to integrate Google image, video, and text-to-speech models into your application.