Muapi's text-to-speech API converts written text into natural, human-like audio through a single unified REST endpoint. Choose from six models across three providers — MiniMax, Google Gemini, and ElevenLabs — spanning single-voice narration, multi-speaker dialogue with per-speaker voice and emotion control, and studio-grade or turbo-fast generation. Same submit-and-poll pattern as every other model on the platform, pay-per-generation, from $0.035 per 1,000 characters.
Studio-quality text-to-speech with crystal-clear pronunciation, smooth pacing, and realistic emotion. Best for audiobooks, podcasts, and video narration where fidelity matters most.
Fast, lightweight text-to-speech with minimal delay and natural voice quality. Adjustable speed, volume, pitch, and emotion for quick-turnaround audio.
Single-voice speech optimized for speed without sacrificing quality. Tune stability, similarity, and speaking rate — ideal for voiceovers, IVR prompts, and high-throughput narration.
Fast, cost-efficient multi-speaker TTS. Define any number of speakers with distinct voices, accents, and emotional styles, then script dialogue with inline tone tags like [whispers].
Premium studio-quality multi-speaker TTS. Assign each speaker a distinct voice, accent, and emotional style for audiobooks, cinematic dialogue, and long-form narration.
Multilingual multi-speaker dialogue from a structured script. Assign each speaker a voice ID and get lifelike intonation, turn-taking, and pacing handled automatically.
A text-to-speech (TTS) API converts written text into spoken audio via a single HTTP call — no audio engineering or voice-recording pipeline required. Submit a script, pick a voice and tone, and get back a downloadable audio file.
Muapi exposes six TTS endpoints through the same unified REST pattern used across every model on the platform — one API key, one request/poll flow, pay-as-you-go pricing. Pick single-voice narration for voiceovers and IVR, or multi-speaker dialogue for character conversations, podcasts, and audiobooks.
Gemini 3.1 Flash/2.5 Pro TTS and ElevenLabs Text to Dialogue V3 assign each speaker a distinct voice, accent, and emotional style across scripted conversation turns.
Gemini TTS models support inline tags like [whispers], [shouting], or [determination] for fine-grained expressive control within a single line.
Pick MiniMax Speech 2.6 HD or Gemini 2.5 Pro TTS for studio-grade fidelity, or Speech 2.6 Turbo / Gemini 3.1 Flash TTS when speed and cost matter more than peak quality.
Adjust speed, volume, pitch, emotion, stability, and similarity depending on the model — dial in the exact delivery your script needs.
ElevenLabs endpoints accept a custom ElevenLabs voice ID. Need to clone a voice first? See Muapi's voice cloning models.
MiniMax, Google Gemini, and ElevenLabs TTS models all run through the same submit-and-poll REST pattern — swap providers without rewriting integration code.
| Model | Provider | Price | Best For |
|---|---|---|---|
| MiniMax Speech 2.6 HD | MiniMax | $0.65 / generation | Studio-quality single-voice narration |
| MiniMax Speech 2.6 Turbo | MiniMax | $0.65 / generation | Fast single-voice generation, minimal delay |
| ElevenLabs TTS Turbo 2.5 | ElevenLabs | $0.05 / 1,000 characters | High-throughput single-voice narration, IVR |
| Gemini 3.1 Flash TTS | $0.035 / 1,000 characters | Fast, affordable multi-speaker dialogue | |
| Gemini 2.5 Pro TTS | $0.035 / 1,000 characters | Studio-quality multi-speaker dialogue | |
| ElevenLabs Text to Dialogue V3 | ElevenLabs | $0.10 / generation | Multilingual multi-speaker conversations |
POST /api/v1/{model-slug} with your text, voice, and optional speed/pitch/emotion parameters.GET /api/v1/predictions/{request_id}/result until status is completed, then download the audio output.Muapi's text-to-speech API converts written text into natural, human-like audio via a single REST endpoint. Six models across MiniMax, Google Gemini, and ElevenLabs cover single-voice narration and multi-speaker dialogue with tunable emotion, pace, and pitch.
Gemini 3.1 Flash TTS or Gemini 2.5 Pro TTS for inline emotion/tone tags across any number of speakers, or ElevenLabs Text to Dialogue V3 for multilingual conversations with per-speaker voice IDs.
From $0.035 per 1,000 characters on Gemini 3.1 Flash/2.5 Pro TTS, $0.05 per 1,000 characters on ElevenLabs TTS Turbo 2.5, $0.10 per generation on ElevenLabs Text to Dialogue V3, and $0.65 per generation on MiniMax Speech 2.6 HD/Turbo.
Yes — ElevenLabs endpoints accept a custom ElevenLabs voice ID. To create one first, see Muapi's voice cloning models (MiniMax Voice Clone, Suno Custom Voice Cloning).
Each model returns a downloadable audio file URL from the standard predictions/result endpoint — ready to embed, download, or pipe into a video-editing or lipsync pipeline.
Yes. Sign up at muapi.ai, create an API key from your dashboard, and start generating speech immediately — no waitlist required.