مباشر4 نموذجاً

واجهة برمجة تطبيقات Gemini TTS — تحويل النص إلى كلام متعدد المتحدثين بنبرات تعبيرية طبيعية

توفر واجهة برمجة تطبيقات Gemini TTS على MuAPI تحويل النص إلى كلام بجودة فائقة وواقعية بشرية مع دعم المحادثات متعددة المتحدثين والتحكم الدقيق في النبرة والمشاعر.أنشئ تعليقات صوتية للبودكاست، ومقاطع الفيديو، ومساعدي الصوت، وتطبيقات الألعاب بأصوات مذهلة تبدو طبيعية تماماً.

i

The full Gemini TTS lineup is live on Muapi

Choose a Gemini TTS tier by fidelity, latency, and cost, define your speakers, script the dialogue, and poll the returned request ID for a downloadable audio file.

Gemini 3.8 Flash & Flash Lite TTS

Google's newest Gemini TTS generation — studio-grade fidelity or high-throughput, low-latency conversational speech2 models
صوتNew
3.8 Flash

Gemini 3.8 Flash TTS

Gemini 3.8 Flash TTS — تحويل النص إلى كلام طبيعي ومعبر مع دعم الحوار متعدد المتحدثين.

Dialogue → audioFast
$0.015/1,000 chars
تجربة الصوت الآن
صوتNew
3.8 Flash

Gemini 3.8 Flash Lite TTS

Gemini 3.8 Flash Lite TTS — تحويل النص إلى كلام طبيعي ومعبر مع دعم الحوار متعدد المتحدثين.

Dialogue → audioFast
$0.01/1,000 chars
تجربة الصوت الآن

Gemini 3.1 Flash TTS

Fast, cost-efficient multi-speaker dialogue1 models
صوتLive
3.1 Flash

Gemini 3.1 Flash TTS

Gemini 3.1 Flash TTS — تحويل النص إلى كلام طبيعي ومعبر مع دعم الحوار متعدد المتحدثين.

Dialogue → audioFast
$0.035/1,000 chars
تجربة الصوت الآن

Gemini 2.5 Pro TTS

Premium studio-quality multi-speaker dialogue1 models
صوتLive
2.5 Pro

Gemini 2.5 Pro TTS

Gemini 2.5 Pro TTS — تحويل النص إلى كلام طبيعي ومعبر مع دعم الحوار متعدد المتحدثين.

Dialogue → audio
$0.035/1,000 chars
تجربة الصوت الآن

أصوات فائقة الواقعية مع Google Gemini TTS

يمثل Gemini TTS طفرة في تكنولوجيا توليد الصوت البشري، حيث يتفوق في محاكاة التوقفات الطبيعية، والنبرات التعبيرية، والتنغيم اللغوي الصحيح لمختلف اللغات واللهجات.

يدعم النموذج إخراج سيناريوهات حوارية كاملة تضم متحدثين متعددين في ملف صوتي واحد مع الحفاظ التام على هوية كل صوت.

قدرات Gemini TTS الصوتية

حوارات متعددة المتحدثين

توليد محادثات طبيعية بين شخصيتين أو أكثر في استدعاء برمجي واحد.

تحكم بالمشاعر والنبرات

توجيه نبرة الصوت لتكون حماسية أو هادئة أو احترافية أو درامية.

دعم متعدد اللغات

نطق طبيعي ودقيق للعربية والإنجليزية وعشرات اللغات العالمية.

زمن استجابة فائق السرعة

توليد سريع وبث صوتي فوري لتطبيقات المساعدين الذكيين.

Gemini TTS use cases

Character dialogue & games

Voice NPCs and branching conversations with distinct, emotionally-directed characters using Gemini 3.8 Flash or Flash Lite TTS.

Conversational AI agents

Power real-time voice agents and IVR systems with Gemini 3.8 Flash Lite TTS's low-latency, high-throughput generation.

Podcasts & audiobooks

Produce multi-host conversational audio or long-form narration with Gemini 2.5 Pro TTS's studio-quality fidelity.

Prototyping

Iterate quickly on tone, accent, and pacing on Gemini 3.1 Flash TTS before a final high-fidelity render.

Gemini TTS model comparison

ModelProviderPriceBest For
Gemini 3.8 Flash TTSGoogle$0.015/1,000 charactersStudio-grade fidelity, expressive acting, long-form stability
Gemini 3.8 Flash Lite TTSGoogle$0.01/1,000 charactersHigh-throughput, low-latency conversational speech
Gemini 3.1 Flash TTSGoogle$0.035/1,000 charactersFast, cost-efficient multi-speaker dialogue
Gemini 2.5 Pro TTSGoogle$0.035/1,000 charactersPremium studio-quality multi-speaker dialogue

Gemini TTS API examples

Submit speakers and dialogue turns with your Muapi API key, save the request ID, and poll the standard prediction result endpoint for the audio URL.

Generate with Gemini 3.8 Flash TTS

curl -X POST https://api.muapi.ai/api/v1/gemini-3-8-flash-tts \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Kore","accent":"Neutral","style":"Empathetic","pace":"Natural"},{"speaker_id":"Speaker 2","voice_name":"Puck","accent":"British (RP)","style":"Deadpan","pace":"Natural"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"[warmly] Welcome to the show."},{"speaker_id":"Speaker 2","text":"[whispers] Let us begin."}]}'

# Response: {"request_id":"REQUEST_ID"}

Studio-grade voice fidelity and long-form stability for character dialogue and narration.

Generate with Gemini 3.8 Flash Lite TTS

curl -X POST https://api.muapi.ai/api/v1/gemini-3-8-flash-lite-tts \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Zephyr","accent":"American (Gen)","style":"Newscaster","pace":"Rapid Fire"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"Breaking news, delivered fast and clear."}]}'

The lowest-cost, lowest-latency Gemini TTS tier for high-throughput conversational speech.

Generate with Gemini 3.1 Flash TTS

curl -X POST https://api.muapi.ai/api/v1/gemini-3-1-flash-tts \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Fenrir","accent":"British (RP)","style":"Deadpan","pace":"Natural"},{"speaker_id":"Speaker 2","voice_name":"Puck","accent":"American (Gen)","style":"Empathetic","pace":"Staccato"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"[shouting] Halt, traveler!"},{"speaker_id":"Speaker 2","text":"[determination] I carry a message for the elder."}]}'

Fast, cost-efficient multi-speaker dialogue with the prior-generation Flash tier.

Generate with Gemini 2.5 Pro TTS

curl -X POST https://api.muapi.ai/api/v1/gemini-2-5-pro-tts \
  -H "x-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"speakers":[{"speaker_id":"Speaker 1","voice_name":"Rasalgethi","accent":"Transatlantic","style":"Vocal Smile","pace":"Natural","audio_profile":"A warm, seasoned audiobook narrator"}],"dialogue_turns":[{"speaker_id":"Speaker 1","text":"[gently] Once upon a time, in a quiet valley hidden away..."}]}'

Premium studio-quality multi-speaker dialogue for audiobooks and cinematic narration.

Poll the audio result

curl https://api.muapi.ai/api/v1/predictions/REQUEST_ID/result \
  -H "x-api-key: YOUR_API_KEY"

# Read the generated audio URL after status becomes completed.

Poll until status is completed, then read the downloadable audio URL from the response.

الأسئلة الشائعة حول Gemini TTS API

هل يمكن توليد حوار بين متحدثين في نفس الطلب؟

نعم، يدعم Gemini TTS تحديد أصوار متعددة لنصوص حوارية مختلفة في نفس الملف الصوتي.

ما هي صيغ الصوت الناتجة؟

يدعم إخراج ملفات MP3 و WAV و OGG بجودة بث احترافية (24kHz و 48kHz).

هل يتوفر دعم للغة العربية؟

نعم، يقدم Gemini TTS دعماً ممتازاً للغة العربية الفصحى بنطق سليم ونبرة طبيعية.

كيف يتم حساب التكلفة؟

تتم الفوترة لكل 1,000 حرف من النص المحول إلى صوت.

امنح تطبيقاتك صوتاً بشرياً حقيقياً

ادمج Google Gemini TTS في منصتك البرمجية اليوم عبر واجهة MuAPI الموحدة.