gemini-audio-vision API — MuAPI: AI Large Language Models

รวมโมเดล AI gemini-audio-vision ผ่าน API ประสิทธิภาพสูงที่มีความหน่วงต่ำและจ่ายตามการใช้งานจริง

📝

Overview

About this model

เข้าถึงโมเดล gemini-audio-vision ผ่านการเขียนโปรแกรมบน MuAPI พร้อมการปรับขนาดได้ทันทีและความหน่วงต่ำ

1การสร้างเนื้อหามัลติมีเดีย
2การผสานรวมผลิตภัณฑ์ AI
3เวิร์กโฟลว์อัตโนมัติ
💰

Pricing & Value

Cost analysis

muapiapp$2.00/M 输入 Token, $5.00/M 输出 Token (higher per-run minimum for 音频)

Token-based billing with 无需订阅 — pay only for what you use, including the higher token cost of audio input.

Fal.aiไม่พร้อมใช้งาน

Fal.ai's vision endpoints are image-only; they do not offer native audio-understanding via Gemini's file API.

Replicateไม่พร้อมใช้งาน

Replicate does not currently offer a hosted Gemini audio-understanding endpoint.

** Competitor pricing is estimated based on similar model architectures and usage tiers.

⚙️

Technical Details

Configuration schema

พรอมต์ (Prompt)string

The question or instruction describing what to analyze in the audio.

Default ValueDescribe what is said and any background sounds in this audio, including speaker changes and tone.
URL เสียงstring

URL of the audio to analyze.

Default Valuehttps://d3adwkbyhxyrtq.cloudfront.net/webassets/audiomodels/sample-audio.mp3
System Promptstring

Optional system-level instruction to guide the model's analysis style.

Default ValueRespond with a structured JSON analysis, not prose.
โมเดลEnum (1 options)

Gemini model to use for audio understanding.

Default Valuegemini-2.5-flash
📖

Implementation Guide

Developer documentation

ส่งคำขอ POST ไปยังจุดสิ้นสุดพร้อมพารามิเตอร์ที่คุณต้องการและคีย์ MuAPI

Common Questions

Frequently asked