Send an image and an instruction ("change the background to a beach", "swap the jacket for a leather one", "remove the man on the left") and get an edited image back. Flux Kontext Pro/Max are the new standard for instruction-following edits; GPT-Image handles compositional rewrites; Seededit and Qwen Edit are cost-efficient alternatives. All exposed through the same JSON API.
Every model in this category uses the same submit-then-poll API. Replace flux-kontext-pro-i2i with any model endpoint from the list below.
# 1. Submit
curl -X POST https://api.muapi.ai/api/v1/flux-kontext-pro-i2i \
-H "x-api-key: $MUAPI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"prompt":"change the sky to a sunset"}'
# → {"request_id":"abc123","status":"processing"}
# 2. Poll until completed
curl https://api.muapi.ai/api/v1/predictions/abc123/result \
-H "x-api-key: $MUAPI_API_KEY"| Model | Provider | Cost | Best For |
|---|---|---|---|
| gpt-image-2-image-to-image | — | $0.090 | Transform and edit existing images using GPT Image 2 with text instructions. Supports up to 16 input images for precise style transfer, editing, and image transformation. |
| bytedance-seedream-5.0-pro-edit | — | $0.045 | Seedream 5.0 Pro Edit is ByteDance's flagship image editing model, extending Seedream 5.0 Lite Edit with higher-fidelity rendering, deeper visual reasoning, and finer typography control. It accepts one or more reference images and a natural-language instruction, targeting up to 2K output with stronger scene composition and richer prompt adherence for professional creative workflows. |
| nano-banana-pro-edit | — | $0.120 | Nano Banana 2 Edit is the next-generation image editing model developed by Google DeepMind, following the original Nano Banana (also known as Gemini 2.5 Flash Image). It offers advanced image-edit capabilitie with improved resolution. |
| nano-banana-2-edit | — | $0.060 | Nano Banana 2 (Gemini 3.1 Flash Image) is Google's most advanced image generation model, combining speed with high-fidelity 4K output and revolutionary character consistency. |
| bytedance-seedream-v4.5-edit | — | $0.050 | Seedream-v4.5 Edit allows you to transform an existing image using natural-language instructions. It preserves the core composition, lighting, and style of the original while modifying only the requested elements — perfect for object replacement, environment changes, stylistic adjustments, and high-detail creative reworks. |

Generate a pack of high-quality, professional portraits in various styles (LinkedIn, CEO, Tinder, etc.) while preserving your facial features.

Flux Kontext Max I2I in Max mode allows precise image enhancement and visual transformations while retaining the source layout. It’s powerful for retouching, photo-to-art workflows, concept refinement.

Decompose a composite image into a base image plus clean, separate layer assets with Seedream 5.0 Flash Layer Decomposition. Optionally target elements with prompts or <bbox> regions.

Transform blurry or pixelated images into high-definition visuals. Our AI Image Upscaler uses deep learning to reconstruct details and bring your visuals to life.

Advanced facial recognition and blending algorithms enable precise face swaps while preserving skin tone, lighting, and facial geometry.

Instantly change outfits in images using AI. Visualize different clothing styles without the need for physical trials—perfect for fashion, e-commerce, and virtual try-ons.

Instantly remove image backgrounds with pixel-perfect precision. Ideal for product photos, profile pictures, and creative projects.

Instantly generate studio-quality product images with AI. Upload your item photo and get clean, stylized shots perfect for e-commerce, ads, and catalogs.

Smooth skin, reduce blemishes, and enhance complexion with natural-looking results. Perfect for portraits, selfies, and professional photo retouching.

Takes an input images and transforms it based on a new prompt. Keeps structure or pose while changing style, appearance, or details.

Create professional-grade product photos using AI. Upload your item image and describe it with a prompt, and get studio-style, lifestyle, or creative backgrounds in seconds

Bring your imagination to life with art inspired by the enchanting world of Studio Ghibli. This AI model generates dreamy, hand-drawn visuals with soft colors, whimsical characters, and painterly backgrounds

Expand the edges of any image with AI. This model continues your original photo or artwork beyond its borders while matching style, lighting, and content.

Easily remove unwanted objects, people, or text from any image using AI. Just select the area you want to erase, and the model will intelligently fill the space with realistic background matching the surrounding environment. No Photoshop skills needed.

Flux Kontext Pro I2I variant enables transforming base images into refined artwork while keeping structure intact. It’s useful for sketch refinement, visual style changes, and creative edits such as re-dressing, relighting, or re-theming with prompt guidance.

Transform an input image based on a new prompt — like changing style, lighting, or composition. Useful for reinterpreting visuals while keeping structure.

Seededit allows precise edits to images using masks and prompt guidance. Whether you're replacing backgrounds, changing clothing, or inpainting missing areas, Seededit ensures realistic, high-quality results with semantic control.

Ideogram’s Character Reference model enables consistent character generation using just one reference image. Upload a clear character portrait—and you can place that character in unlimited scenes, styles, poses, or narratives with visual fidelity maintained across all outputs.

The Qwen Edit Image Model allows you to modify existing images using text-based editing prompts. Instead of generating from scratch, you can upload a base image and describe the desired changes (e.g., replacing objects, altering colors, adding new elements).

AI Image Effects applies advanced visual transformations, color grading, and cinematic filters to create stunning images from a image.

Topaz Image Upscale is a high-quality image-to-image enhancement model that increases resolution, sharpness, and detail using AI super-resolution. It improves clarity, restores texture, reduces noise, and produces crisp, high-res output while preserving natural look and fine edges.

Nano Banana is a mysterious, high-performance image model. It excels at precise, language-driven edits and consistent character preservation, allowing users to modify images with natural text commands.

Flux Kontext Effects is a creative image and video model that applies stylized transformations, cinematic filters, and artistic reinterpretations to your inputs. Instead of generating new content from scratch, it enhances or reimagines existing images and videos with unique looks — ranging from surreal effects to realistic cinematic moods.

Minimax’s I2I “Subject Reference” model enables you to transform images while preserving the appearance of a subject using a single reference image. Ideal for maintaining character likeness—features, clothing, or expression—across different styles or settings.

WAN 2.6 Image Edit applies targeted, instruction-based edits to an existing image while preserving composition, perspective, and lighting. It’s ideal for object replacement, material changes, environment tweaks, and style adjustments with clean integration and minimal artifacts—keeping the original scene coherent and cinematic.

Flux PuLID is an innovative image-to-image model that enables consistent face rendering across different styles or scenes—without needing any model fine-tuning. By providing a reference image (e.g., a portrait), the model generates new visuals while maintaining your subject’s identity with high fidelity.

Flux Redux is a transformation model that reimagines or enhances your input images while preserving their main structure and subject. It’s built for creative refinement — whether you want style transfer, artistic reinterpretation, cinematic polish, or mood transformation.

[Beta] Turn fictional character references into reusable video characters. Upload reference images and describe the outfit to get a character_id you can use in SD 2.0 Omni Reference.

Ideogram V3 Reframe is a specialized image-to-image model built on Ideogram 3.0, designed to intelligently extend and adapt images across diverse aspect ratios and resolutions. Leveraging advanced AI outpainting, it preserves visual consistency while enabling creative reframing for digital, print, and video content.

Nano Banana Effects is a creative visual effects model designed to transform ordinary images into fun, stylized, and eye-catching results. It applies artistic filters, 3D styles, cartoon transformations, and trending viral looks with a single click.

Qwen Image Edit Plus is an upgraded image-editing model that supports multiple image references and superior text editing. Powered by the 20B-parameter Qwen architecture, it allows changes like background swap, style transfer, object removal/addition, and precise text edits (bilingual: English/Chinese) while maintaining visual consistency and preserving details of the original images.

Seedream 5.0 Lite Edit is an advanced image transformation model by ByteDance, enabling precise, controllable edits using natural language. It specializes in high-fidelity style transfer (Anime, Cyberpunk, Fantasy), background swaps, and object modification while preserving original lighting, color tones, and character consistency for professional-grade creative reworks.

ReVE Edit is a next-generation image editing model that allows users to apply detailed visual transformations through natural language. Whether you want to restyle portraits, modify backgrounds, or create artistic reinterpretations, ReVE Edit delivers realistic and coherent results while preserving structure and identity.

Nano Banana 2 Edit is the next-generation image editing model developed by Google DeepMind, following the original Nano Banana (also known as Gemini 2.5 Flash Image). It offers advanced image-edit capabilitie with improved resolution.

Kling O1 Image Edit applies targeted transformations to an existing image while preserving composition, lighting, and visual consistency. Use it to replace objects, retouch elements, change materials, or apply stylistic shifts with high fidelity and minimal artifacts.

Flux 2 Dev Edit takes an existing image and applies transformations, replacements, or style changes based on a text instruction. It preserves composition, lighting, and the overall scene while modifying only what the edit prompt specifies. Ideal for creative replacements, stylistic adjustments, object swaps, and environment changes while keeping the original artistic integrity.

Flux-2-Pro Edit enables precise, high-fidelity modifications to an existing image while preserving its lighting, style, mood, and composition. It’s ideal for replacing objects, altering materials, adjusting environmental elements, or performing stylistic transformations without damaging the original scene’s quality. Flux-2-Pro maintains ultra-detailed textures and cinematic realism during edits.

VIDU Reference-to-Image Q2 generates new high-quality images based on one or more reference images. It preserves the key identity, structure, or style of the reference while creating a new scene, variation, or enhanced composition. Ideal for character consistency, object re-interpretation, stylized redesigns, and cinematic recreations guided by reference inputs.

Qwen Image Edit 2511 performs precise, instruction-driven edits on an existing image while preserving composition, lighting, and overall style. It’s well-suited for object replacement, material changes, localized edits, and subtle scene adjustments with strong visual consistency and minimal artifacts.

Qwen Image Text-to-Image 2512 generates high-resolution, visually consistent images from text prompts. It focuses on strong scene structure, clean composition, and atmospheric lighting, making it well-suited for cinematic environments, surreal concepts, fantasy and sci-fi worlds.

GPT-Image-1.5 Edit applies precise, instruction-based modifications to an existing image while preserving composition, lighting, perspective, and visual coherence. It’s well-suited for object replacement, concept evolution, symbolic edits, and creative transformations that feel natural and intentional rather than destructive.

Professional AI portrait styles including hair, makeup, style, and fashion transformations.

Add custom watermark to images with adjustable position, opacity, and size. Free local processing using PIL.

Wan 3.0 Image Edit is an upcoming AI image-editing model. Confirmed API controls, output specifications, and pricing will be published at launch.

Qwen 2.0 Pro Image Edit model with maximum precision and modifications.

Flux-2-Klein-4B Turbo Edit provides ultra-fast, instruction-based image editing. This high-efficiency variant of Klein 4B Edit is optimized for near-instant swaps and tweaks while preserving layout and lighting. Ideal for real-time design tools and quick creative adjustments.

Alibaba WAN 2.7 Image Edit performs prompt-driven image editing with support for multiple-image references.

Qwen 3.0 Image to Image transforms and edits existing reference images according to text prompts with precise style and content preservation.

Edit and transform existing images using Kling O3 with natural language instructions. Supports up to 10 reference images, 1K/2K/4K resolutions, and up to 9 outputs per request.

Generate a reusable character from a single reference image and a text description. Optionally attach a voice profile created with Gemini Omni Audio to give the character a consistent voice in future video generations.

Nano Banana 2 Lite Edit (Gemini 3.1 Flash Lite Image) is Google's fastest and most cost-efficient image editing model, blending up to 14 reference images with exceptional prompt adherence and character consistency.

Automatically add lifelike colors to black-and-white images. Our AI brings history to life with natural tones, accurate shading, and context-aware colorization.

Edit a specific part of an image using natural language. Ideal for object removal, replacement, or content-aware filling.

Seedream 5.0 Pro Edit is ByteDance's flagship image editing model, extending Seedream 5.0 Lite Edit with higher-fidelity rendering, deeper visual reasoning, and finer typography control. It accepts one or more reference images and a natural-language instruction, targeting up to 2K output with stronger scene composition and richer prompt adherence for professional creative workflows.

The Wan2.5 Edit Image model allows you to transform existing images with precision and creativity. By providing an image along with an edit prompt, you can make realistic changes, enhancements, or stylistic adjustments—whether it’s altering objects, changing backgrounds, adding details, or applying an entirely new artistic style.

Seedream v4 Edit refines or transforms existing images based on a new prompt and a reference. Instead of masking, you provide a source image and describe how it should be altered — adjusting style, details, or replacing elements while keeping the subject consistent.

SeedVR2 is a one-step diffusion-transformer model designed for image restoration, super-resolution, deblurring, and artifact removal. It enhances low-quality or compressed images into clean, sharp, high-resolution results while preserving natural colors and fine details.

FLUX 3 Image-to-Image edits and restyles existing images using a text instruction plus up to several reference images. Built on Black Forest Labs' unified multimodal architecture, it preserves subject identity and scene structure while applying precise, prompt-driven edits — ideal for product retouching, style transfer, and character-consistent edits.

Flux-2-Flex Edit allows flexible transformation of an existing image: object replacement, material changes, lighting adjustments, style shifts, or localized edits. It preserves the original scene’s geometry, perspective, and lighting while modifying only what the edit prompt specifies.

Seedream-v4.5 Edit allows you to transform an existing image using natural-language instructions. It preserves the core composition, lighting, and style of the original while modifying only the requested elements — perfect for object replacement, environment changes, stylistic adjustments, and high-detail creative reworks.

Grok Imagine Image-to-Image transforms an existing image using natural language instructions while preserving scene structure, perspective, and lighting. It is ideal for object replacement, environment evolution, concept re-imagining, and creative edits that feel grounded and visually coherent rather than over-stylized.

Flux-2-Klein-4B Edit applies lightweight, instruction-based edits to an existing image. It’s best for clear object swaps, small visual changes, and cute enhancements while preserving the original scene’s layout and lighting. Ideal for fast edits, UI demos, and simple creative tweaks.

Flux-2-Klein-9B Edit performs higher-quality image edits with better detail retention, lighting consistency, and texture handling compared to smaller variants. It’s well-suited for cute character edits, object additions, and visual refinements that need to look natural and polished while keeping the original scene intact.

Qwen 2.0 Image Edit model with precise background modification and enhancements.

Nano Banana 2 (Gemini 3.1 Flash Image) is Google's most advanced image generation model, combining speed with high-fidelity 4K output and revolutionary character consistency.

Flux-2-Klein-9B Turbo Edit offers high-quality, ultra-fast image editing with superior detail retention. This high-efficiency version of Klein 9B Edit handles lighting and textures with precision while delivering edits much faster than the standard variant. Best for polished character edits and professional refinements where speed is critical.

Qwen 3.0 Pro Image to Image transforms and edits existing reference images according to text prompts with professional-grade precision.

Alibaba WAN 2.7 Image Edit Pro performs prompt-driven image editing with multi-image reference support and up to 2K output.

Change face expressions and camera angles using Qwen Image Edit with predefined angle and expression LoRA weights.

Grok Imagine Image 2.0 Edit applies a targeted, natural-language edit to a prior Grok Imagine Image 2.0 generation, changing only the described region while preserving the rest of the composition, style, and subject.

Transform and edit existing images using GPT Image 2 with text instructions. Supports up to 16 input images for precise style transfer, editing, and image transformation.

Decompose complex images into clean, editable multi-layer assets with ByteDance Seedream 5 Pro Layer Decomposition.

Meta Muse Image Edit applies a text instruction to up to 10 reference images at once, for instruction-driven edits, composites, and style transforms.

Transform and edit existing images using GPT-Image-2.5 Flare, OpenAI's default GPT Image 2.5 tier. Precision editing changes only what you ask for while preserving the rest of the scene, and multi-turn consistency keeps earlier edits intact across a chain of changes.

Transform and edit existing images using GPT-Image-2.5 Sunburst, OpenAI's slower, higher-precision GPT Image 2.5 tier, tuned for maximum edit precision on a single critical edit.

Upscale an image faithfully with Topaz's precision models (Standard/High Fidelity/CGI/Text Refine/Faces), preserving the original look while sharpening fine detail.

Qwen 2.1 Image to Image edits and recombines up to 10 reference images from a text prompt, with an optional inpainting mask for precise local edits.

Upscale an image with Topaz's creative/generative-detail model, adding plausible new fine detail while enlarging up to 4x.

Upscale an image with Topaz's Wonder/Recover generative models, reconstructing realistic fine detail, textures, and faces at high resolution.

Seedream 5.0 Flash Edit is ByteDance's fast, low-cost image-to-image tier, editing or combining up to 10 reference images from a text instruction at 1K or 2K.
Image editing models take an existing image plus an instruction and preserve composition where possible. Text-to-image regenerates from scratch. For consistent edits across a series, image-edit models are usually the right tool.
No — Flux Kontext, GPT-Image, and Reve all do mask-free instruction edits. Mask-based inpainting is also available if you need pixel-precise control.
Flux Kontext Pro for high-fidelity background swaps and lighting changes; the dedicated `product-shot` endpoint for standardized e-commerce flows.