Models/Audio/Gemini TTS
Live4 models

Gemini TTS API — Google Gemini Text to Speech

Muapi's Gemini TTS API gives you Google's full Gemini text-to-speech lineup through one unified REST integration: Gemini 3.8 Flash TTS, Gemini 3.8 Flash Lite TTS, Gemini 3.1 Flash TTS, and Gemini 2.5 Pro TTS.Every model shares the same multi-speaker dialogue format — define speaker voices, accents, emotional styles, and pace, then script ordered dialogue turns with inline tone tags like [whispers] or [shouting]. Pricing runs from $0.01 per 1,000 characters with the same asynchronous submit-and-poll pattern across every tier.

i

The full Gemini TTS lineup is live on Muapi

Choose a Gemini TTS tier by fidelity, latency, and cost, define your speakers, script the dialogue, and poll the returned request ID for a downloadable audio file.

Gemini 3.8 Flash & Flash Lite TTS

Google's newest Gemini TTS generation — studio-grade fidelity or high-throughput, low-latency conversational speech2 models
TTSNew
3.8 Flash

Gemini 3.8 Flash TTS

Google's newest Flash-tier Gemini TTS model with studio-grade voice fidelity, expressive acting, and long-form stability at $0.015/1,000 characters.

Dialogue → audioFast
$0.015/1,000 chars
Try Model
TTSNew
3.8 Flash

Gemini 3.8 Flash Lite TTS

The most affordable Gemini TTS tier — high-throughput, low-latency conversational speech at $0.01/1,000 characters.

Dialogue → audioFast
$0.01/1,000 chars
Try Model

Gemini 3.1 Flash TTS

Fast, cost-efficient multi-speaker dialogue1 models
TTSLive
3.1 Flash

Gemini 3.1 Flash TTS

Cost-efficient multi-speaker speech with distinct voices, accents, emotional styles, and inline tone tags such as [whispers].

Dialogue → audioFast
$0.035/1,000 chars
Try Model

Gemini 2.5 Pro TTS

Premium studio-quality multi-speaker dialogue1 models
TTSLive
2.5 Pro

Gemini 2.5 Pro TTS

Premium multi-speaker speech for audiobooks, cinematic dialogue, and long-form narration with high-fidelity delivery.

Dialogue → audio
$0.035/1,000 chars
Try Model

What is the Gemini TTS API?

Google's Gemini TTS models turn a written script into expressive, multi-speaker audio. Every tier shares the same request shape: define one or more speakers with a voice, accent, emotional style, and pace, then write ordered dialogue turns that reference those speakers, with inline tone tags such as [whispers], [shouting], or [determination] for line-level emotional direction.

Muapi exposes all four Gemini TTS tiers — Gemini 3.8 Flash TTS, Gemini 3.8 Flash Lite TTS, Gemini 3.1 Flash TTS, and Gemini 2.5 Pro TTS — through one API key, one asynchronous submit-and-poll flow, and transparent pay-as-you-go billing that matches Google's own published rates.

Gemini TTS capabilities

Multi-speaker dialogue

Every Gemini TTS tier assigns distinct voices, accents, and delivery styles to scripted conversation turns via a shared `speakers` / `dialogue_turns` format.

Inline emotion and tone

Embed tags such as [whispers], [shouting], and [determination] directly in dialogue text for expressive, line-level control.

Four tiers, one shape

Pick Gemini 3.8 Flash TTS for studio-grade fidelity, Gemini 3.8 Flash Lite TTS for the lowest cost and latency, Gemini 3.1 Flash TTS for a fast prior-generation balance, or Gemini 2.5 Pro TTS for premium narration — all with the same request schema.

30 prebuilt voices, 8 accents

Choose from 30 Gemini voice names (Zephyr, Fenrir, Puck, Kore, and more) and 8 accents including Neutral, American variants, British (RP), British (Brixton), Transatlantic, and Australian.

Scene and tone controls

Set an optional `scene` description for acoustic setting and `sample_context` for overall narration tone, plus a `temperature` control for delivery variation.

Official Google pricing

Every tier is billed at Google's published Gemini TTS audio-output rate per 1,000 characters, with no markup and no subscription.

Gemini TTS use cases

Character dialogue & games

Voice NPCs and branching conversations with distinct, emotionally-directed characters using Gemini 3.8 Flash or Flash Lite TTS.

Conversational AI agents

Power real-time voice agents and IVR systems with Gemini 3.8 Flash Lite TTS's low-latency, high-throughput generation.

Podcasts & audiobooks

Produce multi-host conversational audio or long-form narration with Gemini 2.5 Pro TTS's studio-quality fidelity.

Prototyping

Iterate quickly on tone, accent, and pacing on Gemini 3.1 Flash TTS before a final high-fidelity render.

Gemini TTS model comparison

ModelProviderPriceBest For
Gemini 3.8 Flash TTSGoogle$0.015/1,000 charactersStudio-grade fidelity, expressive acting, long-form stability
Gemini 3.8 Flash Lite TTSGoogle$0.01/1,000 charactersHigh-throughput, low-latency conversational speech
Gemini 3.1 Flash TTSGoogle$0.035/1,000 charactersFast, cost-efficient multi-speaker dialogue
Gemini 2.5 Pro TTSGoogle$0.035/1,000 charactersPremium studio-quality multi-speaker dialogue

Gemini TTS API examples

Submit speakers and dialogue turns with your Muapi API key, save the request ID, and poll the standard prediction result endpoint for the audio URL.

Generate with Gemini 3.8 Flash TTS

curl -X POST https://api.muapi.ai/api/v1/gemini-3-8-flash-tts \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Kore","accent":"Neutral","style":"Empathetic","pace":"Natural"},{"speaker_id":"Speaker 2","voice_name":"Puck","accent":"British (RP)","style":"Deadpan","pace":"Natural"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"[warmly] Welcome to the show."},{"speaker_id":"Speaker 2","text":"[whispers] Let us begin."}]}'

# Response: {"request_id":"REQUEST_ID"}

Studio-grade voice fidelity and long-form stability for character dialogue and narration.

Generate with Gemini 3.8 Flash Lite TTS

curl -X POST https://api.muapi.ai/api/v1/gemini-3-8-flash-lite-tts \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Zephyr","accent":"American (Gen)","style":"Newscaster","pace":"Rapid Fire"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"Breaking news, delivered fast and clear."}]}'

The lowest-cost, lowest-latency Gemini TTS tier for high-throughput conversational speech.

Generate with Gemini 3.1 Flash TTS

curl -X POST https://api.muapi.ai/api/v1/gemini-3-1-flash-tts \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Fenrir","accent":"British (RP)","style":"Deadpan","pace":"Natural"},{"speaker_id":"Speaker 2","voice_name":"Puck","accent":"American (Gen)","style":"Empathetic","pace":"Staccato"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"[shouting] Halt, traveler!"},{"speaker_id":"Speaker 2","text":"[determination] I carry a message for the elder."}]}'

Fast, cost-efficient multi-speaker dialogue with the prior-generation Flash tier.

Generate with Gemini 2.5 Pro TTS

curl -X POST https://api.muapi.ai/api/v1/gemini-2-5-pro-tts \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Rasalgethi","accent":"Transatlantic","style":"Vocal Smile","pace":"Natural","audio_profile":"A warm, seasoned audiobook narrator"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"[gently] Once upon a time, in a quiet valley hidden away..."}]}'

Premium studio-quality multi-speaker dialogue for audiobooks and cinematic narration.

Poll the audio result

curl https://api.muapi.ai/api/v1/predictions/REQUEST_ID/result \
  -H "x-api-key: YOUR_API_KEY"

# Read the generated audio URL after status becomes completed.

Poll until status is completed, then read the downloadable audio URL from the response.

Gemini TTS API FAQ

What is the Gemini TTS API?

Muapi's Gemini TTS API exposes Google's full Gemini text-to-speech lineup — Gemini 3.8 Flash TTS, Gemini 3.8 Flash Lite TTS, Gemini 3.1 Flash TTS, and Gemini 2.5 Pro TTS — through one unified REST endpoint per model, all sharing the same multi-speaker dialogue request format.

Which Gemini TTS model should I use?

Use Gemini 3.8 Flash TTS for studio-grade fidelity at low cost, Gemini 3.8 Flash Lite TTS for the lowest latency and cost on high-volume or real-time workloads, Gemini 3.1 Flash TTS for a fast prior-generation option, or Gemini 2.5 Pro TTS for the highest-fidelity premium narration.

How much does the Gemini TTS API cost?

Pricing matches Google's official published rates: $0.015 per 1,000 characters for Gemini 3.8 Flash TTS, $0.01 for Gemini 3.8 Flash Lite TTS, and $0.035 for both Gemini 3.1 Flash TTS and Gemini 2.5 Pro TTS.

How do I create a multi-speaker conversation?

Define each voice in the `speakers` list with a unique `speaker_id` (in "Speaker N" format), then reference those IDs in your `dialogue_turns`. The model renders each line in that speaker's assigned voice, accent, style, and pace.

What voices and accents are available?

All four Gemini TTS tiers share 30 prebuilt voices (Zephyr, Fenrir, Puck, Kore, and more) and 8 accents including Neutral, American variants, British (RP), British (Brixton), Transatlantic, and Australian.

Can I control emotion and pacing?

Yes. Each speaker has `style` (Vocal Smile, Newscaster, Whisper, Empathetic, Promo/Hype, Deadpan) and `pace` (Natural, Rapid Fire, The Drift, Staccato) settings, plus inline tone tags like [whispers] or [shouting] embedded directly in the dialogue text.

What audio format is returned?

Each model returns a downloadable audio file URL from the standard prediction result endpoint, ready to embed, download, or pass into a video or lipsync pipeline.

How do I access the Gemini TTS API?

Create a Muapi API key, call the endpoint that matches the Gemini TTS tier you want, and poll the returned request ID. No waitlist is required — all four models are live.

Ready to generate Gemini TTS audio?

Use one Muapi API key across Gemini 3.8 Flash TTS, Gemini 3.8 Flash Lite TTS, Gemini 3.1 Flash TTS, and Gemini 2.5 Pro TTS.