Muapi's video-to-audio API analyzes a silent video and generates a matching ambient track, foley, or sound effect synchronized to the on-screen action — no manual sound design required. The same MMAudio model family also supports generating a standalone sound effect from a text prompt alone. Both exposed through the same unified REST endpoint, billed at $0.001 per second of output audio.
Analyzes a silent video and generates a matching ambient track, foley, or sound effect synchronized to the on-screen action — describe the sound you want with a text prompt.
Generates a standalone sound effect or ambient audio clip from a text prompt alone — no video input required. Useful for building a sound library independent of any footage.
A video-to-audio API takes a silent video clip and generates a soundtrack that matches what's happening on screen — footsteps, wind, machinery, ambience — driven by a text prompt describing the sound you want, rather than a manual foley or sound-design pass.
Muapi exposes MMAudio Video to Audio for that exact workflow, plus MMAudio Text to Audio for generating a standalone sound effect or ambient clip with no video input at all. Both run through the same submit-and-poll REST pattern used across every model on the platform, billed per second of generated audio.
MMAudio Video to Audio watches the input video and generates audio timed to the visible action, guided by your text prompt describing the desired sound.
MMAudio Text to Audio generates a sound effect or ambience clip from text alone, useful for building a reusable sound library independent of any specific footage.
Both models accept a duration parameter (default 8 seconds) so the generated audio clip matches your video length or desired clip length.
Ideal for AI-generated video clips that come out silent — feed the clip straight into MMAudio Video to Audio to add a matching soundtrack in one API call.
Both MMAudio models bill at $0.001 per second of output audio — a full 8-second default clip costs $0.008.
Combine with Text to Speech for dialogue or Suno for music, then layer in MMAudio-generated ambience or effects for a fully-scored video.
| Model | Input | Price | Best For |
|---|---|---|---|
| MMAudio Video to Audio | Video + text prompt | $0.001 / sec ($0.008 for 8s default) | Adding a synchronized soundtrack to a silent video clip |
| MMAudio Text to Audio | Text prompt only | $0.001 / sec ($0.008 for 8s default) | Standalone sound effects and ambience, no video required |
POST /api/v1/{model-slug} with your prompt, video URL (video-to-audio only), and optional duration.GET /api/v1/predictions/{request_id}/result until status is completed, then download the audio.Muapi's video-to-audio API analyzes a silent video and generates a matching soundtrack via MMAudio Video to Audio, guided by a text prompt describing the desired sound.
Yes — MMAudio Text to Audio generates a standalone sound effect or ambient clip from a text prompt alone, no video input required.
Both MMAudio models cost $0.001 per second of generated audio — an 8-second clip (the default duration) costs $0.008.
Duration is a request parameter, defaulting to 8 seconds. Set it to match your video length or desired clip length.
No — MMAudio is built for ambience, foley, and sound effects. For speech and dialogue, use Muapi's Text to Speech API instead.
Yes. Sign up at muapi.ai, create an API key from your dashboard, and start generating audio immediately — no waitlist required.