How To Use Grok Imagine Image 2.0: Prompts & Workflows [2026]

A prompting guide for xAI's Grok Imagine Image 2.0 on fal, covering its text-to-image and edit endpoints, in-frame typography, the settings that change the output, and pricing.

John OzuysalSep 10, 202613 min read
How To Use Grok Imagine Image 2.0: Prompts & Workflows [2026]

Grok Imagine Image 2.0 is xAI's image generation and editing model, run on fal across two endpoints: xai/grok-imagine-image/v2.0/text-to-image and xai/grok-imagine-image/v2.0/edit. Typography is the axis this generation was built on, with wording you spell out coming back set into the scene. A quality of low or medium pairs with 1K or 2K resolution, giving four billing combinations from $0.04 up to $0.08 per image. Editing accepts up to three input images alongside an instruction, and each input adds $0.01.

In this guide, I'll cover how to prompt both of Grok Imagine Image 2.0's endpoints on fal, from a one-line brief through to a storefront scene carrying five separate blocks of type, with prompts you can paste into fal's playground or drop straight into an API call.

TL;DR

Grok Imagine Image 2.0 is xAI's image generation and editing model, and fal runs it across two endpoints: xai/grok-imagine-image/v2.0/text-to-image and xai/grok-imagine-image/v2.0/edit.

Typography is the axis this generation was built on: wording you spell out in the prompt comes back set into the scene, with the hierarchy holding the way you set it out.

A quality setting of low or medium pairs with 1K or 2K resolution, giving four billing combinations from $0.04 up to $0.08 per image.

Editing accepts up to three input images alongside a written instruction, with nothing to mask or select first, and each input adds $0.01 to the bill.

Where is the best place to access Grok Imagine Image 2.0?

The best place to access Grok Imagine Image 2.0 for developers and marketers is on fal, as it runs through two endpoints: xai/grok-imagine-image/v2.0/text-to-image for generation and xai/grok-imagine-image/v2.0/edit for editing.

Each has its own browser playground for testing.

The pair answers to the same @fal-ai/client package, so one integration covers both.

There's no subscription that you need to manage and no seat count.

You add credit from $1, and every request bills per output image. Commercial use applies to both endpoints on fal.

Here's how you submit a generation request:

javascript
import { fal } from "@fal-ai/client";

const result = await fal.subscribe("xai/grok-imagine-image/v2.0/edit", {
  input: {
    prompt: "Make this scene more realistic but still keep the game vibes"
  },
  logs: true,
  onQueueUpdate: (update) => {
    if (update.status === "IN_PROGRESS") {
      update.logs.map((log) => log.message).forEach(console.log);
    }
  },
});

console.log(result.data);
console.log(result.requestId);

How do you write a prompt for Grok Imagine Image 2.0?

You want to write the composition out region by region, spell every word that has to appear inside the frame, and close the gaps you care about before the model closes them for you.

The thing about Grok Imagine Image 2.0 is that it doesn't leave decisions unmade. Anything you skip still gets an answer, picked by the model.

xAI says it built this generation to work out type and page layout up front, in a designer's order.

A prompt that arrives already art-directed gives it much less to invent.

These four gaps close most of the distance:

  • The subject as a physical object: material, finish, the wear on it, and how light behaves across that surface.
  • The camera: where it stands in relation to the subject, and whether the lens flattens the scene or stretches it.
  • The regions: what fills the top of the frame and what stays empty for copy.
  • The wording: spelled exactly, in the order it appears, with the line breaks you want kept.

I'd say that three or four sentences will cover most images. Here's a worked example:

Prompt (Settings: 4:3, 2K, medium quality): Overhead campaign still of a cream leather handbag laid open on a bleached oak table, one hard window light from the left throwing a clean edge shadow, the contents arranged in a loose row beside it, a folded silk scarf and a brass keyring and a small printed card reading HOUSE OF MERIDIAN in narrow black capitals with EDITION 04 beneath it. Muted palette of bone, camel and warm gray, nothing in the top quarter of the frame.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

What does the revised_prompt field tell you?

Responses from Grok Imagine Image 2.0 include a revised_prompt field alongside the image, and it's essentially the enhanced prompt that was used to generate the image.

You can treat it as the model's own notes on your brief. Let's try it with something deliberately underwritten:

Prompt (Settings: 1:1, 1K, low quality): A luxury perfume bottle product shot, cinematic lighting, highly detailed, 8k.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

That line names a bottle and stops.

The model still needs a material, a surface underneath it, a light direction and a background, and whatever it settles on shows up in revised_prompt.

You can read the field as a list of what to take back. Let's run it a second time:

Prompt (Settings: 1:1, 2K, medium quality): A heavy square glass perfume bottle on a slab of gray travertine, shot at bottle height from slightly off center, a single hard light from the back left throwing a bright edge down one corner of the glass and a hard shadow across the stone to the right, the liquid inside a deep amber, ORRIS ABSOLUTE etched into the front of the glass in small serif capitals, flat dark wall filling the band above the bottle.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

The second version hands over four fewer decisions. What comes back lands close to the picture you had before you started typing.

There's a ceiling on this, though.

Once a prompt carries more than about eight hard requirements, some of them lose, and the ones that drop are rarely the ones you'd have picked.

The way I typically generate photorealistic images is that I rank the requirements before writing, keep the four that decide the image, then push the remainder into an edit pass on the frame that wins.

How does Grok Imagine Image 2.0 render text inside an image?

You need to spell the words in the prompt exactly as they should appear, broken where the finished lockup breaks.

A single generation will hold around twelve separate text blocks, across a range of sizes and spread over more than one surface.

Type also takes on the material it's printed on, which is why the surface belongs in the prompt alongside the wording.

Lettering laid flat over a photograph tends to read as a mockup, and that difference shows up most on packaging and garment work.

When you write the type, name the treatment and the surface it belongs to.

The treatment: heavy grotesque, narrow serif capitals, hand lettered, foil stamped, screen printed, embossed.

The surface: the front face of a carton, a cotton chest panel, a painted brick wall, a paper band around a jar.

The hierarchy returns the way you wrote it.

Set it out in the prompt starting with the largest line and working down.

Prompt (Settings: 3:4, 2K, medium quality): A charcoal cotton t-shirt laid flat on a light gray studio sweep, shot straight down under soft even light, the chest panel carrying a screen printed graphic reading CASCADE ATHLETIC in wide condensed capitals across the top, EST 2019 in small type below a thin horizontal rule, and NORTH DIVISION set in a narrow column down the left side of the print, cracked vintage ink texture on every letter, the weave of the fabric showing through the print.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

Dense layouts put that ordering rule under the most pressure:

Prompt (Settings: 2:3, 2K, medium quality): A cinema poster for a cold war thriller, a lone figure in a long coat crossing an empty airfield at dawn, printed on matte stock with visible offset dot texture, THE HELSINKI FILE set very large in heavy sans capitals across the lower third, one credit line in small type beneath it reading DIRECTED BY ELENA VOROS, a block of six tiny distributor credits along the very bottom edge, and IN CINEMAS MARCH running in a thin band across the top.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

Which prompt patterns work best on Grok Imagine Image 2.0?

Here are four setups tailored to what Grok Imagine Image 2.0 does well:

Packaging with a full type lockup

Prompt (Settings: 3:4, 2K, medium quality): A slim cardboard coffee box standing on a raw plaster shelf in low warm light, matte uncoated stock in deep green, the front panel printed in cream ink reading NORTHBOUND across the top in wide capitals, SINGLE ORIGIN ETHIOPIA on the second line, a thin rule under that, then WASHED, 250 G in small type at the base, a second box lying on its side behind it slightly out of focus.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

Ultrawide key art at 2:1

The ratio list includes 2:1 and 20:9, so a title card gets generated at the shape it will actually run at.

Prompt (Settings: 2:1, 2K, medium quality): Ultrawide key art for a streaming thriller, a woman in a wet raincoat standing small at the far right of a flooded underground parking garage, sodium lights reflecting off the water, the rest of the frame given over to darkness and pooled light, SIXTEEN NIGHTS set in thin widely spaced capitals low on the left side, deep teal and amber, heavy film grain.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

A campaign frame with copy space held open

You want to say where the copy block goes and say that it stays clear, as unclaimed space gets filled.

Prompt (Settings: 16:9, 2K, medium quality): A dark gray electric sedan photographed at three quarter front on a wet salt flat at dusk, low camera close to the ground, a hard rim light running along the roof line from a low sun behind, the reflection of the car breaking across the wet surface below, the entire left third of the frame held as flat empty sky with no cloud detail for a copy block.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

Type across several surfaces at once

This one tests the multi-block claim in a single pass:

Prompt (Settings: 9:16, 2K, medium quality): A narrow bakery storefront on a cobbled street shot straight on in flat morning light, a painted timber fascia reading GOLDEN HOUR in tall serif capitals, a smaller window decal reading BREAD, PASTRY, COFFEE, a chalkboard on the sidewalk listing four items in hand lettered script, a brass plate beside the door engraved NO. 14, and a paper sign taped inside the glass reading BACK AT 2.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

falMODEL APIs

The fastest, cheapest and most reliable way to run genAI models. 1 API, 100s of models

falSERVERLESS

Scale custom models and apps to thousands of GPUs instantly

falCOMPUTE

A fully controlled GPU cloud for enterprise AI training + research

How do you edit images with Grok Imagine Image 2.0?

Editing runs on xai/grok-imagine-image/v2.0/edit, an endpoint that takes up to three image URLs alongside a written instruction, then hands back the frame with your change applied, and everything around it held in place, with no mask to paint and no region to select first.

On this endpoint, aspect_ratio defaults to auto, holding the shape of the first image you pass in unless one of the 13 fixed ratios overrides it, and every input image adds $0.01 on top of whatever the output costs.

The wording carries the whole burden of aim, so an instruction that names only the change leaves its reach open to interpretation, and naming the fixed parts alongside it closes that gap, with two changes in one instruction better split across two passes.

Rewriting type on a product

Take the coffee box from the pattern above.

Prompt: Only the printed panel changes. Set it to read NORTHBOUND, then DECAF COLOMBIA, then 500 G across the same three lines in the same cream ink.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

Combining a subject with a setting

The three-image ceiling covers jobs like dropping a product into a scene that was shot separately.

The coffee box shelf goes in first, keeping auto on its 3:4 frame, with the perfume bottle second.

Prompt: The plaster shelf, the wall behind it and the warm light in the first image are fixed. Swap the coffee boxes for the perfume bottle from the second image, matching the light direction and the shadow softness already in the frame.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

Moving a subject to a new location

This one runs on the sedan from the campaign frame above.

Prompt: The car is fixed, along with its angle, its paint and its reflections. Rebuild everything behind and around it as a parking garage at night under strip lighting, holding the left third of the frame clear.

Generated using Grok Imagine Image 2.0 on fal, an AI model from xAI.

What are Grok Imagine Image 2.0's settings?

Both endpoints take the same short parameter set, with image_urls the only field unique to editing:

  • aspect_ratio: 13 fixed shapes on the generation endpoint, covering ultrawide through square to tall vertical, with 20:9 the widest, 9:20 the tallest and 1:1 the default. The edit endpoint adds auto on top of those.
  • num_images: how many images one request returns, each billed separately. fal caps this at four per request.
  • output_format: jpeg by default, with png and webp available. Reach for png when something will be laid over the image later.
  • resolution: 1k or 2k, defaulting to 1k.
  • quality: low or medium, defaulting to medium.
  • image_urls: edit endpoint only, three images maximum.
  • sync_mode: hands the image back as a data URI, and the output then stays out of your request history.

How much does Grok Imagine Image 2.0 cost on fal?

fal bills per output image, at a rate depending on which quality mode and which resolution you asked for.

Low quality at 1K: $0.04 per image.

Medium quality at 1K: $0.06 per image.

Low quality at 2K: $0.06 per image.

Medium quality at 2K: $0.08 per image.

Editing: the same rates, with $0.01 added for every input image.

Against real work, a campaign exploration of 32 drafts at low quality and 1K comes to $1.28, and eight finals at medium quality and 2K add $0.64, for $1.92 across 40 images.

Six edit passes at medium quality and 2K with one input image each run $0.54.

➡️ There's one billing condition to know before you scale a batch: a request judged to be in violation of xAI's terms is still charged.

Recently Added

Start generating with Grok Imagine Image 2.0 on fal

Grok Imagine Image 2.0's playground opens without a charge, and billing starts at the first generation.

My loop with this model begins with four low-quality 1K variations to settle the wording.

The prompt that wins gets one medium 2K pass.

Anything needing changes after that goes to the edit endpoint, from a recolor to a new background.

Signing up for fal is free, and you top up from $1 when you're ready to generate.

Over 1,000 models run on the same fal API, and a finished frame can move to an upscaler or a video model without a new integration.

Frequently asked questions about Grok Imagine Image 2.0

Can I use Grok Imagine Image 2.0 output commercially?

Yes. Both Grok Imagine Image 2.0 endpoints carry the commercial use label, so images can go into paid campaigns and client deliverables.

What is the largest size Grok Imagine Image 2.0 outputs?

2K. Resolution takes 1k or 2k on both endpoints and defaults to 1k, so ask for 2k when you want the larger file.

How many images can Grok Imagine Image 2.0 edit at once?

Three. The edit endpoint accepts up to three input images in one request, and each of them adds $0.01 to the cost of the output.

What is new in Grok Imagine Image 2.0?

The quality parameter is new to this generation, with low or medium selectable per request.

The generation endpoint also carries 13 aspect ratios, from 20:9 at the widest to 9:20 at the tallest.

Does Grok Imagine Image 2.0 generate video?

No. Grok Imagine Image 2.0 generates and edits still images, and xAI's video models run on fal under separate endpoints.

About the author
John Ozuysal

Founder of House of Growth. 2x entrepreneur, 1x exit, mentor at 500, Plug and Play, and Techstars.

Build with generative media on fal

Hundreds of production-ready image, video, and audio models behind one API.