AI 大语言模型(LLM)

Muapi 为主流大语言模型提供统一查询接口。发送消息提示词,获取完成结果或 SSE 流式响应,并在同一个积分余额中按次使用。

  • 通过一个端点访问 Gemini、Claude、GPT、Grok、Qwen、Llama 和 DeepSeek
  • 支持异步执行和实时流式响应
  • 按次付费,无强制月度订阅
  • 统一的聊天、推理和文档解析 JSON 结构
模型指南:Grok 4.7 & 4.6 API

快速开始

此类别中的所有模型都使用相同的提交后轮询 API。将 gpt-5-5 替换为下方列表中的任意模型端点。

# 1. Submit
curl -X POST https://api.muapi.ai/api/v1/gpt-5-5 \
  -H "x-api-key: $MUAPI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"prompt":"Explain quantum computing in simple terms."}'
# → {"request_id":"abc123","status":"processing"}

# 2. Poll until completed
curl https://api.muapi.ai/api/v1/predictions/abc123/result \
  -H "x-api-key: $MUAPI_API_KEY"

排名前 5 的AI 大语言模型(LLM)模型

模型提供商成本适用场景
claude-sonnet-5$0.000Claude Sonnet 5 is Anthropic's flagship balanced model, offering the optimal combination of cost and state-of-the-art performance for coding, complex reasoning, and multimodal analysis. Supports text and image inputs. Token-based pricing: $3.00/M input tokens, $15.00/M output tokens. Two endpoints: standard async (/claude-sonnet-5) and live streaming (/claude-sonnet-5/stream) via SSE.
claude-haiku-4-5$0.000Claude Haiku 4.5 is Anthropic's fastest and most cost-effective model, designed for high-frequency queries, simple tasks, and near-instant response times. Supports text and image inputs. Token-based pricing: $0.60/M input tokens, $3.00/M output tokens. Two endpoints: standard async (/claude-haiku-4-5) and live streaming (/claude-haiku-4-5/stream) via SSE.
gemini-3-5-flash$0.000Gemini 3.5 Flash is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: $0.60/M input tokens and $3.60/M output tokens. Two endpoints: standard async (/gemini-3-5-flash) and live streaming (/gemini-3-5-flash/stream) via SSE.
gemini-2-5-flash$0.000Gemini 2.5 Flash is Google's high-speed multimodal language model, optimized for rapid text generation, real-time image understanding, and high-frequency tasks. Supports text and image inputs. Token-based pricing: $0.30/M input tokens, $2.50/M output tokens. Two endpoints: standard async (/gemini-2-5-flash) and live streaming (/gemini-2-5-flash/stream) via SSE.
gpt-5-nano$0.000GPT-5 Nano is a lightweight, high-speed language model from the GPT-5 family designed for instant text generation. It delivers intelligent, context-aware responses for creative writing, summarization, dialogue, code generation, and automation — all at low latency and cost. Perfect for chatbots, assistants, content tools, and real-time applications that need fast, reliable text output.

全部 38 个模型

claude-opus-5
-42%
文本生成
$0.0007$0.001

claude-opus-5

Claude Opus 5 is Anthropic's flagship most capable model, offering state-of-the-art performance for complex coding, reasoning, and multimodal analysis. Supports text and image inputs. Token-based pricing: $3.00/M input tokens, $15.00/M output tokens. Two endpoints: standard async (/claude-opus-5) and live streaming (/claude-opus-5/stream) via SSE.

deepseek-v4-pro
100%
文本生成
$0.0002$0.000

deepseek-v4-pro

DeepSeek V4 Pro is an advanced flagship multimodal reasoning model designed for complex coding, mathematical, and multi-step reasoning tasks.

claude-opus-4-7
-42%
文本生成
$0.0007$0.001

claude-opus-4-7

Claude Opus 4.7 is Anthropic's highly capable model for complex coding, long-context reasoning, and agentic workflows. Supports text and image inputs. Token-based pricing: $3.00/M input tokens, $15.00/M output tokens. Two endpoints: standard async (/claude-opus-4-7) and live streaming (/claude-opus-4-7/stream) via SSE.

gpt-5-nano
100%
文本生成
$0.0001$0.000

gpt-5-nano

GPT-5 Nano is a lightweight, high-speed language model from the GPT-5 family designed for instant text generation. It delivers intelligent, context-aware responses for creative writing, summarization, dialogue, code generation, and automation — all at low latency and cost. Perfect for chatbots, assistants, content tools, and real-time applications that need fast, reliable text output.

claude-sonnet-4-6
10%
文本生成
$1.1111$1.000

claude-sonnet-4-6

Claude Sonnet 4.6 delivers strong reasoning, advanced coding, and native computer-use functionality. Supports text and image inputs with up to 1M token context. Token-based pricing: $1.80/M input tokens, $9.00/M output tokens. Two endpoints: standard async (/claude-sonnet-4-6) and live streaming (/claude-sonnet-4-6/stream) via SSE.

gemini-3-flash
10%
文本生成
$0.0011$0.001

gemini-3-flash

Gemini 3 Flash is a fast, multimodal language model for real-time text generation. Supports text and image inputs, function calling, and Google Search grounding. Token-based pricing: $0.30/M input tokens and $1.80/M output tokens. Two endpoints: standard async (/gemini-3-flash) and live streaming (/gemini-3-flash/stream) via SSE.

claude-opus-4-6
-42%
文本生成
$0.0007$0.001

claude-opus-4-6

Claude Opus 4.6 is Anthropic's most capable model for complex coding, long-context reasoning, and agentic workflows. Supports text and image inputs. Token-based pricing: $3.00/M input tokens, $15.00/M output tokens. Two endpoints: standard async (/claude-opus-4-6) and live streaming (/claude-opus-4-6/stream) via SSE.

gpt-codex
100%
文本生成
$0.0003$0.000

gpt-codex

OpenAI GPT Codex delivers advanced coding capabilities with scalable reasoning depth. Supports multiple model variants (gpt-5-codex through gpt-5.4-codex) and multimodal inputs. Token-based pricing: $1.25/M input tokens, $9.00/M output tokens. Two endpoints: standard async (/gpt-codex) and live streaming (/gpt-codex/stream) via SSE.

gemini-3-pro
-42%
文本生成
$0.0007$0.001

gemini-3-pro

Gemini 3 Pro is Google's powerful multimodal reasoning model, designed for complex problem solving, coding, and logical tasks. Supports text and image inputs. Token-based pricing: $4.00/M input tokens, $24.00/M output tokens. Two endpoints: standard async (/gemini-3-pro) and live streaming (/gemini-3-pro/stream) via SSE.

gemini-3-5-flash
100%
文本生成
$0.0001$0.000

gemini-3-5-flash

Gemini 3.5 Flash is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: $0.60/M input tokens and $3.60/M output tokens. Two endpoints: standard async (/gemini-3-5-flash) and live streaming (/gemini-3-5-flash/stream) via SSE.

gpt-5-6-sol
-11%
文本生成
$0.0009$0.001

gpt-5-6-sol

GPT 5.6 Sol is OpenAI's flagship reasoning model, optimized for complex math, programming, and scientific research. Supports mixed text and image inputs, adjustable reasoning effort, and integrated web search tools. Pricing: $10.00/M input tokens, $60.00/M output tokens.

claude-haiku-4-5
100%
文本生成
$0.0001$0.000

claude-haiku-4-5

Claude Haiku 4.5 is Anthropic's fastest and most cost-effective model, designed for high-frequency queries, simple tasks, and near-instant response times. Supports text and image inputs. Token-based pricing: $0.60/M input tokens, $3.00/M output tokens. Two endpoints: standard async (/claude-haiku-4-5) and live streaming (/claude-haiku-4-5/stream) via SSE.

gpt-5-2
100%
文本生成
$0.0003$0.000

gpt-5-2

GPT 5.2 is a lightweight reasoning model with fast response times and deep coding capabilities. Supports image inputs, system prompts, web search capabilities, and reasoning effort control. Pricing: $1.25/M input tokens, $9.00/M output tokens.

claude-opus-4-8
-42%
文本生成
$0.0007$0.001

claude-opus-4-8

Claude Opus 4.8 is Anthropic's most capable model for complex coding, long-context reasoning, and agentic workflows. Supports text and image inputs. Token-based pricing: $3.00/M input tokens, $15.00/M output tokens. Two endpoints: standard async (/claude-opus-4-8) and live streaming (/claude-opus-4-8/stream) via SSE.

gemini-3-5-flash-openai
100%
文本生成
$0.0001$0.000

gemini-3-5-flash-openai

Gemini 3.5 Flash (OpenAI-compatible) is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: $0.60/M input tokens and $3.60/M output tokens. Two endpoints: standard async (/gemini-3-5-flash-openai) and live streaming (/gemini-3-5-flash-openai/stream) via SSE.

gemini-3-1-pro
-42%
文本生成
$0.0007$0.001

gemini-3-1-pro

Gemini 3.1 Pro is Google's next-generation multimodal model, optimized for complex reasoning, planning, coding, and multi-turn conversation. Supports text and image inputs. Token-based pricing: $4.00/M input tokens, $24.00/M output tokens. Two endpoints: standard async (/gemini-3-1-pro) and live streaming (/gemini-3-1-pro/stream) via SSE.

claude-opus-4-5
-42%
文本生成
$0.0007$0.001

claude-opus-4-5

Claude Opus 4.5 is Anthropic's highly capable model for complex coding, long-context reasoning, and agentic workflows. Supports text and image inputs. Token-based pricing: $3.00/M input tokens, $15.00/M output tokens. Two endpoints: standard async (/claude-opus-4-5) and live streaming (/claude-opus-4-5/stream) via SSE.

claude-sonnet-4-5
100%
文本生成
$0.0004$0.000

claude-sonnet-4-5

Claude Sonnet 4.5 is Anthropic's state-of-the-art model offering high intelligence, speed, and efficiency for code generation, writing, and logical analysis. Supports text and image inputs. Token-based pricing: $1.80/M input tokens, $9.00/M output tokens. Two endpoints: standard async (/claude-sonnet-4-5) and live streaming (/claude-sonnet-4-5/stream) via SSE.

gemini-2-5-pro
100%
文本生成
$0.0003$0.000

gemini-2-5-pro

Gemini 2.5 Pro is Google's advanced multimodal reasoning model, optimized for complex coding, logical tasks, and deep analysis. Supports text and image inputs. Token-based pricing: $1.25/M input tokens, $10.00/M output tokens. Two endpoints: standard async (/gemini-2-5-pro) and live streaming (/gemini-2-5-pro/stream) via SSE.

gemini-2-5-flash
100%
文本生成
$0.0001$0.000

gemini-2-5-flash

Gemini 2.5 Flash is Google's high-speed multimodal language model, optimized for rapid text generation, real-time image understanding, and high-frequency tasks. Supports text and image inputs. Token-based pricing: $0.30/M input tokens, $2.50/M output tokens. Two endpoints: standard async (/gemini-2-5-flash) and live streaming (/gemini-2-5-flash/stream) via SSE.

gpt-5-5
100%
文本生成
$0.0005$0.000

gpt-5-5

GPT 5.5 is OpenAI's state-of-the-art flagship reasoning model for high-complexity problems. Supports image and file uploads, system prompts, web search capabilities, and reasoning effort control. Pricing: $2.40/M input tokens, $16.00/M output tokens.

claude-fable-5
10%
文本生成
$0.0011$0.001

claude-fable-5

Claude Fable 5 is the latest flagship model from Anthropic. Supports text and image inputs with advanced reasoning and creative capabilities. Token-based pricing: $8.00/M input tokens, $40.00/M output tokens. Two endpoints: standard async (/claude-fable-5) and live streaming (/claude-fable-5/stream) via SSE.

generate-social-video-script
10%
文本生成
$0.1111$0.100

generate-social-video-script

Generate viral short-form video scripts for social media based on a topic and niche.

claude-sonnet-5
100%
文本生成
$0.0004$0.000

claude-sonnet-5

Claude Sonnet 5 is Anthropic's flagship balanced model, offering the optimal combination of cost and state-of-the-art performance for coding, complex reasoning, and multimodal analysis. Supports text and image inputs. Token-based pricing: $3.00/M input tokens, $15.00/M output tokens. Two endpoints: standard async (/claude-sonnet-5) and live streaming (/claude-sonnet-5/stream) via SSE.

grok-4-3
100%
文本生成
$0.0001$0.000

grok-4-3

Grok 4.3 is a highly capable multimodal reasoning model, supporting mixed text/image inputs, adjustable reasoning effort, and integrated web search tools. Token-based pricing: $2.50/M input tokens, $5.00/M output tokens. Two endpoints: standard async (/grok-4-3) and live streaming (/grok-4-3/stream) via SSE.

grok-4-5
100%
文本生成
$0.0001$0.000

grok-4-5

Grok 4.5 is a highly capable multimodal reasoning model, supporting mixed text/image inputs, adjustable reasoning effort, and integrated web search tools. Token-based pricing: $1.60/M input tokens, $4.80/M output tokens. Two endpoints: standard async (/grok-4-5) and live streaming (/grok-4-5/stream) via SSE.

gpt-5-6-luna
100%
文本生成
$0.0002$0.000

gpt-5-6-luna

GPT 5.6 Luna is OpenAI's fast, cost-effective multimodal reasoning model. Supports mixed text and image inputs, adjustable reasoning effort, and integrated web search tools. Pricing: $2.00/M input tokens, $12.00/M output tokens.

gpt-5-6-terra
100%
文本生成
$0.0004$0.000

gpt-5-6-terra

GPT 5.6 Terra is OpenAI's balanced multimodal reasoning model for general business and analytical tasks. Supports mixed text and image inputs, adjustable reasoning effort, and integrated web search tools. Pricing: $5.00/M input tokens, $30.00/M output tokens.

gemini-3-6-flash
100%
文本生成
$0.0001$0.000

gemini-3-6-flash

Gemini 3.6 Flash is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: .60/M input tokens and .60/M output tokens. Two endpoints: standard async (/gemini-3-6-flash) and live streaming (/gemini-3-6-flash/stream) via SSE.

gemini-3-6-flash-openai
100%
文本生成
$0.0001$0.000

gemini-3-6-flash-openai

Gemini 3.6 Flash (OpenAI-compatible) is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: .60/M input tokens and .60/M output tokens. Two endpoints: standard async (/gemini-3-6-flash-openai) and live streaming (/gemini-3-6-flash-openai/stream) via SSE.

grok-4-6
100%
文本生成
$0.0001$0.000

grok-4-6

Grok 4.6 is xAI’s next-generation multimodal reasoning model, supporting mixed text/image inputs, adjustable reasoning effort, and integrated web search tools. Token-based pricing: $1.50/M input tokens, $4.50/M output tokens. Two endpoints: standard async (/grok-4-6) and live streaming (/grok-4-6/stream) via SSE.

grok-4-7
100%
文本生成
$0.0001$0.000

grok-4-7

Grok 4.7 is xAI’s next-generation multimodal reasoning model, supporting mixed text/image inputs, adjustable reasoning effort, and integrated web search tools. Token-based pricing: $1.40/M input tokens, $4.20/M output tokens. Two endpoints: standard async (/grok-4-7) and live streaming (/grok-4-7/stream) via SSE. Coming soon.

gpt-5-mini
100%
文本生成
$0.0001$0.000

gpt-5-mini

GPT‑5 Mini is a compact yet powerful AI that converts plain text ideas into detailed, structured prompts suitable for use in text-to-image, text-to-video, and other generative AI models. It’s perfect for creators who want to quickly craft high-quality prompts without manually thinking about style, composition, and descriptive details. The model helps accelerate workflows for artists, video producers, and designers.

kimi-k3
100%
文本生成
$0.0001$0.000

kimi-k3

Moonshot Kimi K3 is a flagship 2.8T Mixture-of-Experts (MoE) LLM with a 1M token context window, designed for long-context reasoning, coding, and complex agent workflows. Token-based pricing: $0.70/M input tokens and $2.80/M output tokens. Two endpoints: standard async (/kimi-k3) and live streaming (/kimi-k3/stream) via SSE.

deepseek-v4-flash
100%
文本生成
$0.0001$0.000

deepseek-v4-flash

DeepSeek V4 Flash is an ultra-fast multimodal reasoning model optimized for low-latency text and image understanding tasks.

gpt-5-4
100%
文本生成
$0.0003$0.000

gpt-5-4

GPT-5.4 delivers powerful reasoning, coding, and professional knowledge work. Supports multimodal inputs (text and image) with adjustable reasoning depth. Token-based pricing: $1.25/M input tokens, $9.00/M output tokens. Two endpoints: standard async (/gpt-5-4) and live streaming (/gpt-5-4/stream) via SSE.

gemini-3-7-flash-openai
100%
文本生成
$0.0001$0.000

gemini-3-7-flash-openai

Gemini 3.7 Flash (OpenAI-compatible) is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: $0.60/M input tokens and $3.60/M output tokens. Two endpoints: standard async (/gemini-3-7-flash-openai) and live streaming (/gemini-3-7-flash-openai/stream) via SSE.

gemini-3-7-flash
100%
文本生成
$0.0001$0.000

gemini-3-7-flash

Gemini 3.7 Flash is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: $0.60/M input tokens and $3.60/M output tokens. Two endpoints: standard async (/gemini-3-7-flash) and live streaming (/gemini-3-7-flash/stream) via SSE.

常见问题

Which LLM model is best?

GPT 5.5 and Claude Opus 4.8 are the best for deep logic, complex coding, and multi-step reasoning. Gemini 3.5 Flash is highly optimized for rapid, real-time, low-latency applications. MuApi hosts all of them so you can benchmark quality vs. cost dynamically.

Do you support streaming?

Yes! Every LLM model has a streaming endpoint (e.g. `POST /api/v1/{model}/stream`) that pushes Server-Sent Events (SSE) for near-instant rendering in chat bubbles.