gemini-audio-vision API — MuAPI: AI Large Language Models

Zintegruj model AI gemini-audio-vision za pomocą wydajnych interfejsów API z niskimi opóźnieniami i płatnością za rzeczywiste zużycie.

📝

Overview

About this model

Uzyskaj dostęp programistyczny do modelu gemini-audio-vision na platformie MuAPI z natychmiastową skalowalnością i niskimi opóźnieniami.

1Generowanie treści multimedialnych
2Integracja produktów AI
3Zautomatyzowane przepływy pracy
💰

Pricing & Value

Cost analysis

muapiapp$2.00/M 输入 Token, $5.00/M 输出 Token (higher per-run minimum for 音频)

Token-based billing with 无需订阅 — pay only for what you use, including the higher token cost of audio input.

Fal.aiNiedostępne

Fal.ai's vision endpoints are image-only; they do not offer native audio-understanding via Gemini's file API.

ReplicateNiedostępne

Replicate does not currently offer a hosted Gemini audio-understanding endpoint.

** Competitor pricing is estimated based on similar model architectures and usage tiers.

⚙️

Technical Details

Configuration schema

Promptstring

The question or instruction describing what to analyze in the audio.

Default ValueDescribe what is said and any background sounds in this audio, including speaker changes and tone.
Adres URL audiostring

URL of the audio to analyze.

Default Valuehttps://d3adwkbyhxyrtq.cloudfront.net/webassets/audiomodels/sample-audio.mp3
System Promptstring

Optional system-level instruction to guide the model's analysis style.

Default ValueRespond with a structured JSON analysis, not prose.
ModelEnum (1 options)

Gemini model to use for audio understanding.

Default Valuegemini-2.5-flash
📖

Implementation Guide

Developer documentation

Wyślij żądanie POST do punktu końcowego z wymaganymi parametrami i kluczem API MuAPI.

Common Questions

Frequently asked