Gemini Audio Vision API Pricing
Gemini Audio Vision uses Google Gemini's native audio understanding to analyze and describe audio content in detail — speech, tone, background sounds, speaker changes, and more. Upload an audio URL and a prompt, and Gemini returns a detailed text analysis. Token-based pricing.
About Gemini Audio Vision
Gemini Audio Vision brings Google Gemini's native audio understanding to the Muapi API. Instead of relying on a separate transcription pass, Gemini ingests the full audio file and reasons about it end-to-end — spoken content, speaker changes, tone and emotion, background sounds, music, and silence. Give it an audio URL and a prompt (or a structured system prompt) and it returns a detailed text analysis, making it useful for podcast summarization, call QA, content moderation, and sound-event detection. Pricing is token-based, similar to [Gemini Video Vision](/playground/gemini-video-vision) and [Gemini 2.5 Flash](/playground/gemini-2-5-flash), with a per-run minimum that reflects Google's higher per-token rate for audio input.
Interactive Savings Calculator
Estimate monthly API spend and compare absolute developer savings.
$20000.00
$2.00/M input tokens, $5.00/M output tokens (higher per-run minimum for audio)$300.00
Not availableDetailed Pricing Breakdown
| Provider | Estimated Rate | Notes |
|---|---|---|
| muapiapp | $2.00/M input tokens, $5.00/M output tokens (higher per-run minimum for audio) | Token-based billing with no subscription — pay only for what you use, including the higher token cost of audio input. |
| Fal.ai | Not available | Fal.ai's vision endpoints are image-only; they do not offer native audio-understanding via Gemini's file API. |
| Replicate | Not available | Replicate does not currently offer a hosted Gemini audio-understanding endpoint. |
Developer Integration Snippets
Model FAQ
Compare similar models
Ready to scale your production?
Get instant access to developer keys. Integrate high-speed dynamic models in minutes with our robust SDKs.

