Explore and test the latest AI models for image generation, video creation, audio synthesis, and 3D modelling. Try FLUX, Kling, Veo, Seedance, Nano Banana, Wan, Suno, and 500+ other models directly in your browser — pay per generation, no subscription required.
All models are accessible via the MuAPI REST API with a single API key. No per-provider accounts needed. Pay per generation and switch models freely.
Seedance 2.5 Omni Reference 480p is the early-access preview build of the Seedance 2.5 family at 480p resolution, blending multiple reference images, video clips, and audio into a single guided generation. Faster and more cost-effective than the 720p variant.
Seedance 2.5 Omni Reference is the early-access preview build of the Seedance 2.5 family, blending multiple reference images, video clips, and audio into a single guided generation. Supports clips up to 30 seconds and 480p/720p output (1080p is not yet supported by this preview).
Seedance 2.5 First & Last Frame 480p is the early-access preview build of the Seedance 2.5 family at 480p resolution, generating a smooth video transition between a start and end image. Faster and more cost-effective than the 720p variant.
Seedance 2.5 First & Last Frame is the early-access preview build of the Seedance 2.5 family, generating a smooth video transition between a start and end image. Supports clips up to 30 seconds and 480p/720p output (1080p is not yet supported by this preview).
Seedance 2.5 Image-to-Video 480p is the early-access preview build of the Seedance 2.5 family at 480p resolution, animating a single image into video. Faster and more cost-effective than the 720p variant. Supports clips up to 30 seconds.
Seedance 2.5 Text-to-Video 480p is the early-access preview build of the Seedance 2.5 family at 480p resolution — faster and more cost-effective than the 720p variant, ideal for previews and drafts. Supports clips up to 30 seconds.

MiniMax H3 Reference to Video creates a 2K video from a prompt plus image, video, and optional audio references through the Muapi API.

MiniMax H3 Image to Video animates a source image with a motion prompt through the Muapi API.

MiniMax H3 Text to Video creates video from a written prompt through the Muapi API.

Grok 4.7 is xAI’s next-generation multimodal reasoning model, supporting mixed text/image inputs, adjustable reasoning effort, and integrated web search tools. Token-based pricing: $1.40/M input tokens, $4.20/M output tokens. Two endpoints: standard async (/grok-4-7) and live streaming (/grok-4-7/stream) via SSE. Coming soon.

Grok 4.6 is xAI’s next-generation multimodal reasoning model, supporting mixed text/image inputs, adjustable reasoning effort, and integrated web search tools. Token-based pricing: $1.50/M input tokens, $4.50/M output tokens. Two endpoints: standard async (/grok-4-6) and live streaming (/grok-4-6/stream) via SSE. Coming soon.

Claude Opus 5 is Anthropic's flagship most capable model, offering state-of-the-art performance for complex coding, reasoning, and multimodal analysis. Supports text and image inputs. Token-based pricing: $3.00/M input tokens, $15.00/M output tokens. Two endpoints: standard async (/claude-opus-5) and live streaming (/claude-opus-5/stream) via SSE.

Wan 3.0 Image Edit is an upcoming AI image-editing model. Confirmed API controls, output specifications, and pricing will be published at launch.

Wan 3.0 Text to Image is an upcoming AI image model. Confirmed API controls, output specifications, and pricing will be published at launch.

Wan 3.0 Image to Video is an upcoming AI video model. Confirmed API controls, output specifications, and pricing will be published at launch.

Wan 3.0 Text to Video is an upcoming AI video model. Confirmed API controls, output specifications, and pricing will be published at launch.

Gemini 3.6 Flash (OpenAI-compatible) is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: .60/M input tokens and .60/M output tokens. Two endpoints: standard async (/gemini-3-6-flash-openai) and live streaming (/gemini-3-6-flash-openai/stream) via SSE.

Gemini 3.6 Flash is a high-speed, multimodal language model built for real-time text generation, supporting text and image inputs natively. Token-based pricing: .60/M input tokens and .60/M output tokens. Two endpoints: standard async (/gemini-3-6-flash) and live streaming (/gemini-3-6-flash/stream) via SSE.
FLUX 3 Image-to-Video animates a still image into a cinematic clip with optional native synchronized audio, using Black Forest Labs' unified image/video/audio architecture. Motion stays physically grounded and consistent with the source frame, making it suited for product animation, portrait bring-to-life effects, and scene extension.
FLUX 3 Text-to-Video generates cinematic video clips with optional native synchronized audio from a single unified model — the same architecture Black Forest Labs uses for FLUX 3's action-prediction research. Expect coherent motion, strong physical plausibility, and scene-appropriate ambient sound baked directly into generation.

FLUX 3 Dev is the faster, lower-cost variant of Black Forest Labs' FLUX 3 frontier model, planned for open-weight release. It trades a small amount of peak fidelity for significantly reduced latency and price, making it well suited for rapid iteration, prototyping, and high-volume text-to-image generation.

FLUX 3 Image-to-Image edits and restyles existing images using a text instruction plus up to several reference images. Built on Black Forest Labs' unified multimodal architecture, it preserves subject identity and scene structure while applying precise, prompt-driven edits — ideal for product retouching, style transfer, and character-consistent edits.

FLUX 3 Text-to-Image is Black Forest Labs' next-generation multimodal frontier model, jointly trained across image, video, and audio for outputs that are truer to life in every style. It generates highly photorealistic and stylistically flexible images from text prompts, with sharper detail, more coherent composition, and stronger prompt adherence than the FLUX.2 generation.
Video Background Remover automatically removes the background from any video, producing a clean cutout of the subject with a transparent or solid-color backdrop. It handles hair, edges, and fine detail with frame-accurate matting, supports videos up to 60 seconds, and can output transparent WebM/MOV, standard MP4, or animated GIF while optionally preserving the original audio.