Custom workflows and utility models optimized for performance, cost, and high scalability across image and video tasks.
MuAPI's custom utility and workflow models provide optimized performance and routing. MuAPI presents custom models next to other providers so teams can compare capability, pricing and model fit before choosing an integration path.


MuAPI AI API models
Explore MuAPI models for chat, code, image and video generation, including Gemini, Nano Banana and Veo-style workflows available through MuAPI.
$0.040 / second
Realistic lipsync video - optimized for speed, quality, and consistency.
$0.200 / second
InfiniteTalk Image-to-Video brings still portraits and character photos to life by generating natural, realistic talking videos. You provide a single face image and a dialogue script, and the model animates lip movement, facial expressions, and subtle head gestures to match the speech.
$0.200 / second
InfiniteTalk Video-to-Video enhances or transforms existing videos by syncing the subject’s lip movements and facial expressions with new dialogue or speech. Instead of starting from a still image, you provide a video clip, and the model seamlessly reanimates the speaker’s mouth and expressions to match the script.
$0.010 / second
Video Background Remover automatically removes the background from any video, producing a clean cutout of the subject with a transparent or solid-color backdrop. It handles hair, edges, and fine detail with frame-accurate matting, supports videos up to 60 seconds, and can output transparent WebM/MOV, standard MP4, or animated GIF while optionally preserving the original audio.
$0.040 / second
LatentSync is a video-to-video model that generates lip sync animations from audio using advanced algorithms for high-quality synchronization.
$0.040 / second
Generate realistic lipsync from any audio using VEED's latest model
$0.300 / second
Bring your characters and worlds to life with AI Dance Effects — a creative video effect that adds playful, dynamic, and cinematic motion to your generations. AI Dance Effects lets you guide how characters move, react, and express themselves.
$0.000 / second
Add AI-generated animated captions to any video using Vadoo's caption engine. Supports multiple languages and viral caption themes like Hormozi style. Perfect for social media creators, marketers, and content producers.
$0.250 / second
Convert any video into 175+ languages with synchronized voice translation, AI-voice cloning, and accurate lip sync. Just upload your video (or provide a link), select a target language, and HeyGen recreates the speech in that language. 0.05$ per second.
$0.240 / second
The AI Video Upscaler is a powerful tool designed to enhance the resolution and quality of videos. Whether you're working with low-resolution videos that need a boost or aiming to improve the clarity of existing footage, this upscaler leverages advanced machine learning models to deliver high-quality, upscaled videos.
$0.200 / second
Ovi is a unified audio–video generation model that can transform a static image plus a descriptive prompt into a short video with synchronized audio. It supports both text-to-video and image-conditioned video inputs. With built-in lip sync, background audio / sound effects, and dialogue support, Ovi brings still visuals to life in cinematic fashion. Videos are generated in 540p resolution.
$0.200 / second
Ovi is a unified model that generates synchronized video and audio from textual input. You write a scene description, including dialogue and ambient sounds, and Ovi produces a short video clip (typically ~5 seconds) where visuals and sound align naturally. Videos are generated in 540p resolution.
$0.000 / second
Add custom watermark to videos with adjustable position, opacity, and size. Free local processing using FFmpeg.
$0.100 / second
Drive a video's lip movements to match a target audio track, producing a lip-synced video output.

$0.630 / second
Generate animated motion graphics videos from a text prompt using AI-generated React/Remotion code rendered on Modal.

$0.010 / second
MMAudio-v2 generates high-quality, synchronized audio from video or text inputs. Seamlessly integrate it with AI video models to create fully-voiced, expressive video content.
$0.040 / second
Generate realistic lipsync animations from audio using advanced algorithms for high-quality synchronization.

$0.100 / second
Replace faces in videos with stunning realism. Our AI ensures accurate expression transfer, lighting consistency, and smooth frame-by-frame blending.
$0.300 / second
AI Video Effects applies advanced visual transformations, color grading, and cinematic filters to create stunning videos from images.
$0.030 / second
The AI Video Upscaler is a powerful tool designed to enhance the resolution and quality of videos. Whether you're working with low-resolution videos that need a boost or aiming to improve the clarity of existing footage, this upscaler leverages advanced machine learning models to deliver high-quality, upscaled videos.
$0.065 / second
The AI Video Watermark Remover is our flagship model designed to remove Sora 2 watermarks, logos, captions, and unwanted text from videos without compromising quality. Supporting a wide range of formats, it's fast, efficient, and processes with the highest quality.

$0.025 / second
Transform and resize your videos effortlessly with remix video tool.
$0.500 / second
Convert long-form videos into engaging short clips using AI clipping.
$0.050 / second
Combine multiple short video clips (5s, 10s, etc.) into a single seamless full-length video. Upload your clips in order and choose the final output aspect ratio. 'Auto' preserves the aspect ratio of your first clip.
$0.050 / second
Automatically crop and reframe a specific video segment to your chosen aspect ratio using AI subject tracking.
$0.630 / second
Edit and modify a previously generated motion graphics animation using a text instruction.
$0.300 / second
AI Video Effects applies advanced visual transformations, color grading, and cinematic filters to create stunning videos from images.
$0.300 / second
Motion Controls adds dynamic camera movements, speed ramps, and zoom effects to bring your images to life as smooth, engaging videos.
$0.300 / second
VFX delivers high-impact visual effects like explosions, particles, and cinematic overlays to transform static images into action-packed videos.

$0.020 / generation
Advanced facial recognition and blending algorithms enable precise face swaps while preserving skin tone, lighting, and facial geometry.

$0.010 / generation
Instantly remove image backgrounds with pixel-perfect precision. Ideal for product photos, profile pictures, and creative projects.

$0.060 / generation
Instantly generate studio-quality product images with AI. Upload your item photo and get clean, stylized shots perfect for e-commerce, ads, and catalogs.

$0.010 / generation
Smooth skin, reduce blemishes, and enhance complexion with natural-looking results. Perfect for portraits, selfies, and professional photo retouching.

$0.050 / generation
Create professional-grade product photos using AI. Upload your item image and describe it with a prompt, and get studio-style, lifestyle, or creative backgrounds in seconds

$0.030 / 1K tokens
Create stunning anime-style artwork instantly with our AI Anime Generator. Customize characters, scenes, and styles effortlessly in seconds!

$0.030 / generation
Expand the edges of any image with AI. This model continues your original photo or artwork beyond its borders while matching style, lighting, and content.

$0.050 / generation
Easily remove unwanted objects, people, or text from any image using AI. Just select the area you want to erase, and the model will intelligently fill the space with realistic background matching the surrounding environment. No Photoshop skills needed.

$0.002 / generation
The SDXL LoRA image model enhances Stable Diffusion XL with specialized fine-tuning, letting you generate images in unique styles, characters, or themes. By applying LoRA weights, you can create visuals that match a specific aesthetic, celebrity look, anime style, or custom-trained subject.

$0.020 / 1K tokens
Pony XL is a high-quality image generation model based on Stable Diffusion XL architecture. It specializes in character art, hybrid styles, and producing detailed, polished visuals even with simpler prompts.

$0.030 / generation
AI Image Effects applies advanced visual transformations, color grading, and cinematic filters to create stunning images from a image.

$0.000 / generation
Add custom watermark to images with adjustable position, opacity, and size. Free local processing using PIL.

$0.028 / 1K tokens
AI TikTok Carousel Generator — create viral TikTok carousel posts from a single text prompt. Choose a proven storytelling format (Problem-Solution, Listicle, Tutorial, Before & After), set your slide count (3-10), and get stunning AI-generated images at 1080x1920 portrait resolution, ready to upload to TikTok.

$0.010 / generation
Professional AI portrait styles including hair, makeup, style, and fashion transformations.

$0.100 / generation
Instantly change outfits in images using AI. Visualize different clothing styles without the need for physical trials—perfect for fashion, e-commerce, and virtual try-ons.

$0.010 / generation
Automatically add lifelike colors to black-and-white images. Our AI brings history to life with natural tones, accurate shading, and context-aware colorization.

$0.050 / generation
Bring your imagination to life with art inspired by the enchanting world of Studio Ghibli. This AI model generates dreamy, hand-drawn visuals with soft colors, whimsical characters, and painterly backgrounds

$0.020 / generation
Transform blurry or pixelated images into high-definition visuals. Our AI Image Upscaler uses deep learning to reconstruct details and bring your visuals to life.

$0.004 / 1K tokens
SDXL is a high-quality, large Stable Diffusion model for creating photorealistic and stylized images from text. It excels at fine detail, realistic lighting, and complex scenes.

$0.020 / 1K tokens
Neta Lumina is a powerful anime-style text-to-image model developed by Neta.art Lab. It’s built on Lumina-Image-2.0, fine-tuned with over 13 million high-quality anime images. It offers strong understanding of multilingual prompts, excellent detail fidelity, support for Danbooru tags, and leaning into niche styles like furry, Guofeng, pets, scenic backgrounds, etc.

$0.020 / 1K tokens
Croma Image is an advanced text-to-image generation model designed for high-quality, creative, and versatile visuals. It can produce anything from photorealistic portraits and products to imaginative concept art, fantasy illustrations, and cinematic scenes.

$0.020 / generation
SeedVR2 is a one-step diffusion-transformer model designed for image restoration, super-resolution, deblurring, and artifact removal. It enhances low-quality or compressed images into clean, sharp, high-resolution results while preserving natural colors and fine details.
$0.100 / 1K tokens
Generate viral short-form video scripts for social media based on a topic and niche.

$0.010 / 1K tokens
Convert text into natural-sounding speech using mmAudio-v2. Ideal for voiceovers, virtual assistants, and content narration with lifelike clarity and tone.

$0.100 / 1K tokens
Generate expressive, multilingual text-to-dialogue content using the ElevenLabs Text To Dialogue V3 model.

$0.010 / 1K tokens
Any LLM is a versatile large language model for text generation, comprehension, and diverse NLP tasks such as chat and summarization. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.025 / 1K tokens
Any LLM is a versatile large language model for text generation, comprehension, and diverse NLP tasks such as chat and summarization. Ready-to-use REST inference API, best performance, no coldstarts, affordable pricing.

$0.001 / 1K tokens
Classify any text snippet across OpenAI's standard moderation categories (sexual, hate, harassment, self-harm, violence, and more). Returns a boolean flag plus per-category booleans and confidence scores — a drop-in safety gate for chat inputs, generated text, and free-form prompts.
$0.010 / generation
Fetch the latest Reels metadata and metrics for an Instagram creator.

$0.010 / generation
Set or replace the thumbnail image of a video on a connected YouTube account.
$0.010 / generation
Fetch the latest Shorts from a YouTube channel by ID.
$0.010 / generation
Fetch the latest posts and performance metrics for a Twitter/X user.

$0.020 / generation
Post a video or image to a connected X (Twitter) account.

$0.010 / generation
Upload and publish a video to a connected YouTube account.
$0.010 / generation
Fetch the latest video posts and view metrics for a TikTok creator.

$0.020 / generation
Publish a video or image to a connected Instagram Business account.

$0.020 / generation
Publish a video or image to a connected Threads account.

$0.010 / generation
Detect unsafe or policy-violating content in any image. Supply an image URL (and optional context text) and receive structured flags for adult, violent, hateful, or otherwise restricted content — ideal for pre-screening user uploads before they reach generation pipelines.

$0.020 / generation
Upload and publish a video to a connected TikTok account.
$0.010 / generation
Fetch the latest Reels metadata and metrics for a Facebook page.

$0.020 / generation
Publish a video or image Pin to a connected Pinterest board.

$0.020 / generation
Publish a video or image to a connected Facebook Page.

$0.020 / generation
Publish a video or image to a connected LinkedIn profile or page.
$0.010 / generation
Retrieve profile details, stats, and metadata for a TikTok user.

$0.020 / generation
Detect and extract text fragments and their positions from an image using local OCR.

$0.010 / generation
Update the title, description, tags, category, privacy, or made-for-kids status of a video on a connected YouTube account.
$0.010 / generation
Download videos from YouTube in your chosen resolution or audio format.

$0.020 / generation
Scan a video for unsafe or policy-violating content. Supports MP4, MOV, and WebM URLs and returns structured safety classifications across harassment, hate, sexual, sexual-minors, and violence categories — useful for moderating user-generated or AI-generated video before publishing.