Transform and edit existing images using GPT Image 2 with text instructions. Supports up to 16 input images for precise style transfer, editing, and image transformation.
About this model
GPT Image 2 Image to Image transforms and edits existing images using natural language instructions, supporting up to 16 input images for precise editing, style transfer, and creative transformation. It follows detailed instructions to modify composition, style, and content while preserving important elements, with selectable quality (low / medium / high) and 1K / 2K / 4K resolutions.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $0.025 – $0.150 per image | Pay only for what you use, regardless of the number of input images. Low: $0.025 / $0.040 / $0.075 (1K / 2K / 4K). Medium: $0.030 / $0.045 / $0.090. High: $0.060 / $0.090 / $0.150. |
| Fal.ai | $0.040 – $0.110 per image | Low and medium quality only. Low/Medium: $0.040 / $0.060 / $0.110 (1K / 2K / 4K). |
| Replicate | Not available | GPT Image 2 is not available on Replicate. |
Pay only for what you use, regardless of the number of input images. Low: $0.025 / $0.040 / $0.075 (1K / 2K / 4K). Medium: $0.030 / $0.045 / $0.090. High: $0.060 / $0.090 / $0.150.
Low and medium quality only. Low/Medium: $0.040 / $0.060 / $0.110 (1K / 2K / 4K).
GPT Image 2 is not available on Replicate.
** Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Text instructions describing the desired transformation. Maximum 20,000 characters. | Preserve the original subject identity, facial structure, hairstyle, pose, body proportions, camera framing, and overall composition from the source image. Transform the environment into a flooded underwater railway terminal with giant glass ceilings showing whales and ocean water above. Add realistic cinematic water reflections, soft underwater caustic lighting, floating atmospheric particles, subtle wetness on clothing and skin, and highly detailed environmental storytelling elements such as abandoned luggage, vines, cracked marble, and cinematic fog depth. Maintain natural realism and believable anatomy while enhancing texture fidelity, lighting realism, color harmony, and cinematic atmosphere. Keep the face sharp and recognizable with authentic skin detail and emotionally grounded expression. Avoid over-stylization, distorted anatomy, plastic skin, blurry textures, extra limbs, cartoon aesthetics, oversaturated colors, low-detail backgrounds, or artificial-looking lighting. |
| Image URLs | array | Upload or provide input images to transform. Up to 16 images supported. | https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/gpt-image-2-image-to-image-in.jpg |
| Aspect Ratio | Enum (6 options) | Output image aspect ratio. Note: images with aspect ratio 'auto' (or unspecified) will only be converted to 1K; 1:1 cannot be converted to 4K — otherwise the task will fail to create. | auto |
| Resolution | Enum (3 options) | Image resolution. Note: images with a 1:1 aspect ratio cannot be converted to 4K. Images with aspect ratio 'auto' (or unspecified) will only be converted to 1K; otherwise the task will fail to create. | 2K |
| Quality | Enum (3 options) | Generation quality. 'low' and 'medium' use a faster, cheaper backend; 'high' uses the higher-fidelity backend. | high |
Text instructions describing the desired transformation. Maximum 20,000 characters.
Preserve the original subject identity, facial structure, hairstyle, pose, body proportions, camera framing, and overall composition from the source image. Transform the environment into a flooded underwater railway terminal with giant glass ceilings showing whales and ocean water above. Add realistic cinematic water reflections, soft underwater caustic lighting, floating atmospheric particles, subtle wetness on clothing and skin, and highly detailed environmental storytelling elements such as abandoned luggage, vines, cracked marble, and cinematic fog depth. Maintain natural realism and believable anatomy while enhancing texture fidelity, lighting realism, color harmony, and cinematic atmosphere. Keep the face sharp and recognizable with authentic skin detail and emotionally grounded expression. Avoid over-stylization, distorted anatomy, plastic skin, blurry textures, extra limbs, cartoon aesthetics, oversaturated colors, low-detail backgrounds, or artificial-looking lighting.Upload or provide input images to transform. Up to 16 images supported.
https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/gpt-image-2-image-to-image-in.jpgOutput image aspect ratio. Note: images with aspect ratio 'auto' (or unspecified) will only be converted to 1K; 1:1 cannot be converted to 4K — otherwise the task will fail to create.
autoImage resolution. Note: images with a 1:1 aspect ratio cannot be converted to 4K. Images with aspect ratio 'auto' (or unspecified) will only be converted to 1K; otherwise the task will fail to create.
2KGeneration quality. 'low' and 'medium' use a faster, cheaper backend; 'high' uses the higher-fidelity backend.
highDeveloper documentation
Upload your images: Provide up to 16 input images via the images_list field. These are the images to be transformed.
Write your prompt: Describe the transformation you want. Be specific about the target style, changes, or effects.
Pick a quality: Choose low or medium for fast, lower-cost edits, or high for the highest fidelity output. high is the default.
Pick aspect ratio and resolution: Select an aspect ratio (e.g. 1:1, 16:9, 9:16) and a resolution (1K, 2K, or 4K). Note that 1:1 cannot be rendered at 4K, and auto aspect ratio is locked to 1K.
Submit and review: Click Generate. The model applies your instructions to the input images and returns the transformed result.
Frequently asked
You can provide up to 16 images in the `images_list` field. The model will use all of them as reference when generating the output.
`low` and `medium` are faster and cheaper and are well suited to drafts and bulk edits. `high` runs the full-fidelity model and is best for finals. Pricing scales with both quality and resolution.
Style transfer, background changes, object modifications, composition adjustments, and creative reinterpretations are all supported via natural language instructions in the prompt.
Yes, you can instruct the model to preserve specific elements (like a product shape or a person) while changing other aspects such as background or style.