Gemini 3.8 Flash TTS
Google's newest Flash-tier Gemini TTS model with studio-grade voice fidelity, expressive acting, and long-form stability at $0.015/1,000 characters.
Muapi's Gemini TTS API gives you Google's full Gemini text-to-speech lineup through one unified REST integration: Gemini 3.8 Flash TTS, Gemini 3.8 Flash Lite TTS, Gemini 3.1 Flash TTS, and Gemini 2.5 Pro TTS.Every model shares the same multi-speaker dialogue format — define speaker voices, accents, emotional styles, and pace, then script ordered dialogue turns with inline tone tags like [whispers] or [shouting]. Pricing runs from $0.01 per 1,000 characters with the same asynchronous submit-and-poll pattern across every tier.
The full Gemini TTS lineup is live on Muapi
Choose a Gemini TTS tier by fidelity, latency, and cost, define your speakers, script the dialogue, and poll the returned request ID for a downloadable audio file.
Google's newest Flash-tier Gemini TTS model with studio-grade voice fidelity, expressive acting, and long-form stability at $0.015/1,000 characters.
The most affordable Gemini TTS tier — high-throughput, low-latency conversational speech at $0.01/1,000 characters.
Cost-efficient multi-speaker speech with distinct voices, accents, emotional styles, and inline tone tags such as [whispers].
Premium multi-speaker speech for audiobooks, cinematic dialogue, and long-form narration with high-fidelity delivery.
Google's Gemini TTS models turn a written script into expressive, multi-speaker audio. Every tier shares the same request shape: define one or more speakers with a voice, accent, emotional style, and pace, then write ordered dialogue turns that reference those speakers, with inline tone tags such as [whispers], [shouting], or [determination] for line-level emotional direction.
Muapi exposes all four Gemini TTS tiers — Gemini 3.8 Flash TTS, Gemini 3.8 Flash Lite TTS, Gemini 3.1 Flash TTS, and Gemini 2.5 Pro TTS — through one API key, one asynchronous submit-and-poll flow, and transparent pay-as-you-go billing that matches Google's own published rates.
Every Gemini TTS tier assigns distinct voices, accents, and delivery styles to scripted conversation turns via a shared `speakers` / `dialogue_turns` format.
Embed tags such as [whispers], [shouting], and [determination] directly in dialogue text for expressive, line-level control.
Pick Gemini 3.8 Flash TTS for studio-grade fidelity, Gemini 3.8 Flash Lite TTS for the lowest cost and latency, Gemini 3.1 Flash TTS for a fast prior-generation balance, or Gemini 2.5 Pro TTS for premium narration — all with the same request schema.
Choose from 30 Gemini voice names (Zephyr, Fenrir, Puck, Kore, and more) and 8 accents including Neutral, American variants, British (RP), British (Brixton), Transatlantic, and Australian.
Set an optional `scene` description for acoustic setting and `sample_context` for overall narration tone, plus a `temperature` control for delivery variation.
Every tier is billed at Google's published Gemini TTS audio-output rate per 1,000 characters, with no markup and no subscription.
Voice NPCs and branching conversations with distinct, emotionally-directed characters using Gemini 3.8 Flash or Flash Lite TTS.
Power real-time voice agents and IVR systems with Gemini 3.8 Flash Lite TTS's low-latency, high-throughput generation.
Produce multi-host conversational audio or long-form narration with Gemini 2.5 Pro TTS's studio-quality fidelity.
Iterate quickly on tone, accent, and pacing on Gemini 3.1 Flash TTS before a final high-fidelity render.
| Model | Provider | Price | Best For |
|---|---|---|---|
| Gemini 3.8 Flash TTS | $0.015/1,000 characters | Studio-grade fidelity, expressive acting, long-form stability | |
| Gemini 3.8 Flash Lite TTS | $0.01/1,000 characters | High-throughput, low-latency conversational speech | |
| Gemini 3.1 Flash TTS | $0.035/1,000 characters | Fast, cost-efficient multi-speaker dialogue | |
| Gemini 2.5 Pro TTS | $0.035/1,000 characters | Premium studio-quality multi-speaker dialogue |
Submit speakers and dialogue turns with your Muapi API key, save the request ID, and poll the standard prediction result endpoint for the audio URL.
curl -X POST https://api.muapi.ai/api/v1/gemini-3-8-flash-tts \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Kore","accent":"Neutral","style":"Empathetic","pace":"Natural"},{"speaker_id":"Speaker 2","voice_name":"Puck","accent":"British (RP)","style":"Deadpan","pace":"Natural"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"[warmly] Welcome to the show."},{"speaker_id":"Speaker 2","text":"[whispers] Let us begin."}]}'
# Response: {"request_id":"REQUEST_ID"}Studio-grade voice fidelity and long-form stability for character dialogue and narration.
curl -X POST https://api.muapi.ai/api/v1/gemini-3-8-flash-lite-tts \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Zephyr","accent":"American (Gen)","style":"Newscaster","pace":"Rapid Fire"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"Breaking news, delivered fast and clear."}]}'The lowest-cost, lowest-latency Gemini TTS tier for high-throughput conversational speech.
curl -X POST https://api.muapi.ai/api/v1/gemini-3-1-flash-tts \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Fenrir","accent":"British (RP)","style":"Deadpan","pace":"Natural"},{"speaker_id":"Speaker 2","voice_name":"Puck","accent":"American (Gen)","style":"Empathetic","pace":"Staccato"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"[shouting] Halt, traveler!"},{"speaker_id":"Speaker 2","text":"[determination] I carry a message for the elder."}]}'Fast, cost-efficient multi-speaker dialogue with the prior-generation Flash tier.
curl -X POST https://api.muapi.ai/api/v1/gemini-2-5-pro-tts \
-H "x-api-key: YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Rasalgethi","accent":"Transatlantic","style":"Vocal Smile","pace":"Natural","audio_profile":"A warm, seasoned audiobook narrator"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"[gently] Once upon a time, in a quiet valley hidden away..."}]}'Premium studio-quality multi-speaker dialogue for audiobooks and cinematic narration.
curl https://api.muapi.ai/api/v1/predictions/REQUEST_ID/result \ -H "x-api-key: YOUR_API_KEY" # Read the generated audio URL after status becomes completed.
Poll until status is completed, then read the downloadable audio URL from the response.
Muapi's Gemini TTS API exposes Google's full Gemini text-to-speech lineup — Gemini 3.8 Flash TTS, Gemini 3.8 Flash Lite TTS, Gemini 3.1 Flash TTS, and Gemini 2.5 Pro TTS — through one unified REST endpoint per model, all sharing the same multi-speaker dialogue request format.
Use Gemini 3.8 Flash TTS for studio-grade fidelity at low cost, Gemini 3.8 Flash Lite TTS for the lowest latency and cost on high-volume or real-time workloads, Gemini 3.1 Flash TTS for a fast prior-generation option, or Gemini 2.5 Pro TTS for the highest-fidelity premium narration.
Pricing matches Google's official published rates: $0.015 per 1,000 characters for Gemini 3.8 Flash TTS, $0.01 for Gemini 3.8 Flash Lite TTS, and $0.035 for both Gemini 3.1 Flash TTS and Gemini 2.5 Pro TTS.
Define each voice in the `speakers` list with a unique `speaker_id` (in "Speaker N" format), then reference those IDs in your `dialogue_turns`. The model renders each line in that speaker's assigned voice, accent, style, and pace.
All four Gemini TTS tiers share 30 prebuilt voices (Zephyr, Fenrir, Puck, Kore, and more) and 8 accents including Neutral, American variants, British (RP), British (Brixton), Transatlantic, and Australian.
Yes. Each speaker has `style` (Vocal Smile, Newscaster, Whisper, Empathetic, Promo/Hype, Deadpan) and `pace` (Natural, Rapid Fire, The Drift, Staccato) settings, plus inline tone tags like [whispers] or [shouting] embedded directly in the dialogue text.
Each model returns a downloadable audio file URL from the standard prediction result endpoint, ready to embed, download, or pass into a video or lipsync pipeline.
Create a Muapi API key, call the endpoint that matches the Gemini TTS tier you want, and poll the returned request ID. No waitlist is required — all four models are live.
Use one Muapi API key across Gemini 3.8 Flash TTS, Gemini 3.8 Flash Lite TTS, Gemini 3.1 Flash TTS, and Gemini 2.5 Pro TTS.