gemini-audio-vision API — MuAPI: AI Large Taal Models

Integreer het gemini-audio-vision AI-model via krachtige API's met lage latentie en betalen naar verbruik.

📝

Overview

About this model

Krijg programmatische toegang tot het gemini-audio-vision-model op MuAPI met directe schaalbaarheid en lage latentie.

1Generatie van multimedia-inhoud
2AI-productintegratie
3Geautomatiseerde werkstromen
💰

Pricing & Value

Cost analysis

muapiapp$2.00/M 输入 Token, $5.00/M 输出 Token (higher per-run minimum for 音频)

Token-based billing with 无需订阅 — pay only for what you use, including the higher token cost of audio input.

Fal.aiNiet beschikbaar

Fal.ai 的视觉端点仅支持图像;它们无法通过 Gemini 文件 API 原生理解音频。

ReplicateNiet beschikbaar

Replicate 目前不提供托管的 Gemini 音频理解端点。

** Competitor pricing is estimated based on similar model architectures and usage tiers.

⚙️

Technical Details

Configuration schema

Promptstring

描述要分析的音频内容的问题或指令。

Default ValueDescribe what is said and any background sounds in this audio, including speaker changes and tone.
Audio URLstring

要分析的音频 URL。

Default Valuehttps://d3adwkbyhxyrtq.cloudfront.net/webassets/audiomodels/sample-audio.mp3
系统Promptstring

可选的 system-level instruction to guide the model's analysis style.

Default ValueRespond with a structured JSON analysis, not prose.
模型Enum (1 options)

用于音频理解的 Gemini 模型。

Default Valuegemini-2.5-flash
📖

Implementation Guide

Developer documentation

Stuur een POST-verzoek naar het endpoint met uw gewenste parameters en MuAPI-sleutel.

Common Questions

Frequently asked