Gemini 3.8 Flash Lite TTS: AI 音频与音乐

Generate low-latency, high-throughput multi-speaker speech with Gemini 3.8 Flash Lite TTS. Try free, pay per use, no subscription. Gemini 3.8 Flash Lite TTS is Google's high-throughput, low-latency, cost-efficient text-to-speech model for expressive multi-speaker conversational audio. Control voice, accent, emotional style, and pace per speaker for high-volume voiceovers, real-time dialogue, and budget-conscious narration.

Interactive model controls

Gemini 3.8 Flash Lite TTS is Google's high-throughput, low-latency, cost-efficient text-to-speech model for expressive multi-speaker conversational audio. Control voice, accent, emotional style, and pace per speaker for high-volume voiceovers, real-time dialogue, and budget-conscious narration.

📝

概览

关于此模型

Gemini 3.8 Flash Lite TTS is Google's high-throughput, low-latency, cost-efficient text-to-speech model, optimized for conversational speech at scale. Define any number of speakers, each with a distinct voice, accent, emotional style, and pace, then script a conversation across ordered dialogue turns with inline tone tags like [whispers] or [shouting]. It is ideal for high-volume voiceovers, real-time conversational agents, and budget-conscious narration where speed and cost matter most. For higher voice fidelity, see Gemini 3.8 Flash TTS; for premium studio-quality output, see Gemini 2.5 Pro TTS.

1Conversational AI: Power real-time voice agents and IVR systems with low-latency multi-speaker speech.
2High-volume production: Generate large batches of voiceovers and character lines at the lowest cost per line.
3Podcasts: Produce multi-host conversational audio from a written script.
4Prototyping: Iterate quickly on tone, accent, and pacing before a final high-fidelity render.
💰

价格与价值

成本分析

muapiapp$0.01 每 1,000 个字符

按次生成, 无需订阅. Matches Google's official Gemini 3.8 Flash Lite TTS audio-output rate — the most affordable tier in the Gemini TTS family.

Fal.ai暂不可用

不提供此模型。

Replicate暂不可用

不提供此模型。

** 竞品价格根据相似模型架构和使用层级估算。

⚙️

技术详情

配置参数

说话人array

说话人声音配置列表。每一轮对话通过 ID 引用一个说话人。

默认值[object Object],[object Object]
对话轮次array

按顺序排列的对话行列表。每一轮的 speaker_id 必须匹配上面定义的说话人。文本可以包含 [shouting] 或 [whispers] 等语气标签。

默认值[object Object],[object Object],[object Object],[object Object]
场景string

可选的 scene description that sets the acoustic setting, e.g. "A quiet, warm room with a fireplace crackling softly."

默认值
示例上下文string

可选的 overall tone/style, e.g. "Audiobook style narration. Tone is gentle and inviting."

默认值
温度number

采样温度(0-2)。数值越高,表达方式越丰富。

默认值1
📖

实施指南

开发者文档

How to Use Gemini 3.8 Flash Lite TTS

  1. Define your speakers: Add one entry per voice to the speakers list. Give each a speaker_id in "Speaker N" format, pick a voice_name, and set the accent, style, and pace. Optionally add an audio_profile describing the persona.

  2. Write the dialogue: Fill dialogue_turns with ordered lines. Each turn's speaker_id must match a speaker you defined. You can embed tone tags inline such as [whispers], [shouting], or [determination].

  3. Set the mood (optional): Use scene to describe the acoustic setting and sample_context to set the overall narration tone. Adjust temperature (0-2) for more or less varied delivery.

  4. Submit and poll: POST the request, then poll /predictions/{request_id}/result until the status is completed to get the audio URL.

curl -X POST https://api.muapi.ai/api/v1/gemini-3-8-flash-lite-tts \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "speakers": [
      {"speaker_id": "Speaker 1", "voice_name": "Fenrir", "accent": "British (RP)", "style": "Deadpan", "pace": "Natural"},
      {"speaker_id": "Speaker 2", "voice_name": "Puck", "accent": "American (Gen)", "style": "Empathetic", "pace": "Staccato"}
    ],
    "dialogue_turns": [
      {"speaker_id": "Speaker 1", "text": "[shouting] Halt, traveler!"},
      {"speaker_id": "Speaker 2", "text": "[determination] I carry a message for the elder."}
    ],
    "scene": "A dark, crumbling dungeon.",
    "temperature": 1
  }'
❓

常见问题

常见问答

How do I create a multi-speaker conversation?

Define each voice in the `speakers` list with a unique `speaker_id` (in "Speaker N" format), then reference those IDs in your `dialogue_turns`. The model renders each line in that speaker's assigned voice, accent, style, and pace.

Can I control emotion and pacing?

Yes. Each speaker has `style` (Vocal Smile, Newscaster, Whisper, Empathetic, Promo/Hype, Deadpan) and `pace` (Natural, Rapid Fire, The Drift, Staccato) settings. You can also embed inline tone tags like [whispers], [shouting], or [determination] directly in the dialogue text.

What voices and accents are available?

There are 30 prebuilt voices (Zephyr, Fenrir, Puck, Kore, and more) and 8 accents including Neutral, American variants, British (RP), British (Brixton), Transatlantic, and Australian.

How is pricing calculated?

Pricing is based on the total number of characters across all dialogue turns, at $0.01 per 1,000 characters — matching Google's official Gemini 3.8 Flash Lite TTS audio-output rate, the most affordable tier in the Gemini TTS family. You pay per generation with no subscription.

What is the maximum length per line?

Each dialogue turn's text can be up to 10,000 characters. You can include multiple turns to build longer conversations.

When should I use the Lite version instead of Flash or Pro?

Use Gemini 3.8 Flash Lite TTS for high-throughput, low-latency, cost-sensitive workloads like real-time conversational agents or large batch voiceover jobs. For higher voice fidelity, use Gemini 3.8 Flash TTS; for premium studio-quality narration, use Gemini 2.5 Pro TTS.