Real-time information synthesis models like Grok, alongside highly aesthetic image generation.
xAI's model ecosystem spans Grok's conversational reasoning and Grok Imagine. MuAPI presents xAI models next to other providers so teams can compare capability, pricing and model fit before choosing an integration path.


xAI AI API models
Explore xAI models for chat, code, image and video generation, including Gemini, Nano Banana and Veo-style workflows available through MuAPI.
$0.150 / second
Grok Imagine is xAI’s fast, creative text-to-video model that generates cinematic clips from 6 to 30 seconds with smooth motion, expressive lighting, and ambient audio. It turns a written idea into a visually rich video.
$0.150 / second
Grok Imagine is xAI’s multimodal image-to-video model, capable of animating still images into cinematic videos from 6 to 30 seconds with synchronized ambient audio. It focuses on realism, fluid motion, and expressive lighting transitions while maintaining high generation speed.
$0.640 / second
Generate videos from images using the Grok Imagine Video 1.5 Preview model with support for multiple aspect ratios, resolutions, and durations up to 15 seconds.
$0.050 / second
Grok Imagine Extend lets you continue and expand existing Grok Imagine video generations seamlessly. Starting from a previously generated video, you can extend the scene while maintaining visual style, characters, motion, and audio consistency. Requires the original task_id from the initial video generation.

$0.050 / generation
Grok Imagine Image 2.0 Edit applies a targeted, natural-language edit to a prior Grok Imagine Image 2.0 generation, changing only the described region while preserving the rest of the composition, style, and subject.

$0.050 / 1K tokens
Grok Imagine is xAI’s high-quality image generation model that transforms text prompts into detailed, stylish, and visually expressive images. It excels at creating vivid scenes, characters, environments, and concept art with strong lighting, depth, and artistic clarity. Get 6 images each time.

$0.050 / generation
Grok Imagine Image-to-Image transforms an existing image using natural language instructions while preserving scene structure, perspective, and lighting. It is ideal for object replacement, environment evolution, concept re-imagining, and creative edits that feel grounded and visually coherent rather than over-stylized.

$0.050 / 1K tokens
Grok Imagine Image 2.0 is xAI's next-generation image model, built for precise instruction-following, sharp typography, and detailed layout control. It follows instructions closely down to fine detail and plans layout so dense, multi-part visuals hold together, turning a text prompt into a polished, commercial-ready image in one request.

$0.050 / 1K tokens
Grok Imagine Quality is xAI's high-fidelity text-to-image mode that prioritizes accuracy and detail over speed. It produces sharper, more visually accurate images with stronger lighting, depth, and artistic clarity. Get 6 images each time.

$0.000 / 1K tokens
Grok 4.3 is a highly capable multimodal reasoning model, supporting mixed text/image inputs, adjustable reasoning effort, and integrated web search tools. Token-based pricing: $2.50/M input tokens, $5.00/M output tokens. Two endpoints: standard async (/grok-4-3) and live streaming (/grok-4-3/stream) via SSE.

$0.000 / 1K tokens
Grok 4.5 is a highly capable multimodal reasoning model, supporting mixed text/image inputs, adjustable reasoning effort, and integrated web search tools. Token-based pricing: $1.60/M input tokens, $4.80/M output tokens. Two endpoints: standard async (/grok-4-5) and live streaming (/grok-4-5/stream) via SSE.

$0.000 / 1K tokens
Grok 4.6 is xAI’s next-generation multimodal reasoning model, supporting mixed text/image inputs, adjustable reasoning effort, and integrated web search tools. Token-based pricing: $1.50/M input tokens, $4.50/M output tokens. Two endpoints: standard async (/grok-4-6) and live streaming (/grok-4-6/stream) via SSE.

$0.000 / 1K tokens
Grok 4.7 is xAI’s next-generation multimodal reasoning model, supporting mixed text/image inputs, adjustable reasoning effort, and integrated web search tools. Token-based pricing: $1.40/M input tokens, $4.20/M output tokens. Two endpoints: standard async (/grok-4-7) and live streaming (/grok-4-7/stream) via SSE. Coming soon.