Grok Imagine Image 2.0 is xAI's next-generation image model for precise generation and editing. It follows detailed instructions, renders sharp typography, and supports multi-reference editing with up to 5 input images.
About this model
Grok Imagine Image 2.0 is xAI's next-generation image model, built for precise generation and editing rather than one-shot novelty. It follows instructions closely down to fine detail, plans typography and layout the way a designer would so dense, multi-part visuals hold together, and renders small text sharply. The model preserves what you put in — subjects, style, and composition — across successive generations and edits.
Editing is treated as a first-class capability, not an add-on. Optional reference images let you target precise edits or combine up to 5 inputs in a single generation, removing the need for manual compositing when building a character, a location, and its props from separate sources while holding one consistent style. The model supports a wide range of aspect ratios, from tall banners to wide cinematic frames, so a single composition can be recomposed to fit any placement.
Cost analysis
| Provider | Cost | Notes |
|---|---|---|
| muapiapp | $0.05 | muapiapp offers cost-effective access to Grok Imagine Image 2.0 once live, in line with muapiapp's typical 20-50% savings over calling providers directly. |
| OpenAI (gpt-image-2) | Varies | OpenAI's competing image model ranks closely with Image 2.0 on public leaderboards but is typically priced higher per generation for comparable quality. |
| ByteDance (Seedream 5.0 Pro) | Varies | Seedream 5.0 Pro is a comparable high-end image model; muapiapp aims to undercut direct provider pricing while matching output quality. |
muapiapp offers cost-effective access to Grok Imagine Image 2.0 once live, in line with muapiapp's typical 20-50% savings over calling providers directly.
OpenAI's competing image model ranks closely with Image 2.0 on public leaderboards but is typically priced higher per generation for comparable quality.
Seedream 5.0 Pro is a comparable high-end image model; muapiapp aims to undercut direct provider pricing while matching output quality.
* Competitor pricing is estimated based on similar model architectures and usage tiers.
Configuration schema
| Parameter | Type | Description | Default |
|---|---|---|---|
| Prompt | string | Text prompt describing the image, or the edit to apply when reference images are provided. | A high-contrast halftone portrait of a young man rendered entirely in fine white dots on a black background, sharp facial detail, editorial poster style. |
| Image URLs | array | Optional list of reference image URLs for multi-reference editing (e.g. combining a subject with a location or prop). Up to 5 images. | https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/grok-imagine1.png |
| Aspect Ratio | Enum (9 options) | Aspect ratio of the output image. | 1:1 |
Text prompt describing the image, or the edit to apply when reference images are provided.
A high-contrast halftone portrait of a young man rendered entirely in fine white dots on a black background, sharp facial detail, editorial poster style.Optional list of reference image URLs for multi-reference editing (e.g. combining a subject with a location or prop). Up to 5 images.
https://d3adwkbyhxyrtq.cloudfront.net/webassets/videomodels/grok-imagine1.pngAspect ratio of the output image.
1:1Developer documentation
Prepare Your Prompt
Add Reference Images (Optional)
images_list. Omit this field for pure text-to-image generation.Choose an Aspect Ratio
1:2) to wide cinematic frames (2:1). The default is 1:1.Submit and Generate
grok-imagine-image-2 endpoint. The model returns a generated or edited image consistent with your prompt and any reference inputs.Iterate
Frequently asked
Image 2.0 focuses on precise instruction-following, sharp typography and layout for dense visuals, and multi-reference editing that accepts up to 5 input images in one generation — capabilities aimed at real creative work rather than single-shot images.
Yes. Provide one or more reference image URLs in `images_list` along with a prompt describing the edit, and the model will apply targeted changes while preserving the rest of the image.
Up to 5 reference images can be supplied in a single request, useful for combining a subject, a location, and props into one consistent scene.
Image 2.0 supports a wide range of ratios including 1:1, 1:2, 2:1, 9:16, 16:9, 2:3, 3:2, 3:4, and 4:3, so the same composition can be recomposed for different placements.
This model page is live ahead of API availability. You can request access now and it will be enabled as soon as the endpoint goes live.