Guides

Grok Imagine Prompts: A Working Guide for Images and Video

Grok Imagine prompts that actually work: the photographer formula, copy-ready examples, the reference aspect-ratio trap, and the video dialect.

By Aditya Bawankule10 min read

Good Grok Imagine prompts are shorter than you think. Grok rewards camera language over adjective piles, front-loads whatever you put in the first twenty words, and holds one subject in one scene far better than it holds three. Most bad results are not a missing detail, they are a buried subject under a paragraph of style words. This guide gives you the formula we use in production, copy-ready prompts you can paste and adapt, the reference-image trap that quietly ruins framing, and the separate dialect Grok speaks for video.

The Grok Imagine prompt formula

Write like a photographer briefing a shoot, not like a designer writing a mood board. Grok responds to lens, light, and framing vocabulary much more strongly than it responds to stacked quality adjectives. The skeleton, in this order:

  1. Shot type and subject, first. "Full body vertical photograph of a..." The opening words carry the most weight, so the thing you actually want goes there, before any style language.
  2. One or two identity traits. Hair, build, age range. Two is plenty. A list of six competing descriptors produces an average of six people.
  3. Wardrobe and setting, plainly. "Wearing a cream linen shirt, on a sunlit balcony." Simple beats specific here; Grok invents good detail on its own and fights over-specified prop lists.
  4. Light and lens. "Golden hour, 85mm, shallow depth of field." This is the single highest-leverage clause in a Grok prompt.
  5. Finish. "Photorealistic, sharp focus, professional photography." Three words, not thirty.

That is the whole thing. A prompt built this way usually lands between twenty-five and fifty words, which is roughly a third of what people write when they first try the model.

Twenty-nine words, built on the shot-type-first formula. The lens and light clause is doing most of the work.

The failure mode on the other side is what we call the laundry list: "ultra detailed, 8k, hyperrealistic, no plastic skin, no airbrushing, perfect anatomy, award winning." Negative phrasing is unreliable on every current image model, and Grok is no exception. Saying "no plastic skin" does not remove plastic skin, it adds the tokens for plastic skin. Describe the state you want instead: "natural skin texture, visible freckles, soft window light."

Copy-ready Grok Imagine prompts

These are written out in full so you can paste them as they are. Each one is followed by the parts to swap; the clause order and the light-and-lens clause stay put.

Portrait and lifestyle

"Full body vertical photograph of a woman in her late twenties with dark curly hair, wearing a cream linen shirt, standing on a sunlit balcony, golden hour backlight, 85mm, shallow depth of field, photorealistic, sharp focus."

Swap the age range, the one hair trait, the outfit, and the balcony. Leave the lens, the light, and the three finish words alone. This is the workhorse. It produces a clean, well-lit figure with believable light falloff on the face, and it is the base for persona work: see AI Instagram models for what changes when the same face has to appear a hundred times.

Cinematic and editorial

"Cinematic film still of a lone fisherman on a wet stone jetty at dawn, anamorphic framing, cool teal grade, volumetric haze, strong silhouette, shallow depth of field."

Swap the subject, the location, the hour, and the color grade. Everything else holds. Grok is genuinely good at atmosphere. Name one grade and one weather condition; naming three fights itself. For a broader vocabulary of looks to drop into the grade slot, AI art styles covers the style words that reliably land.

Grok image prompts for products and objects

"Product photograph of a matte black ceramic kettle on a travertine slab, single softbox from the left, seamless sand-colored backdrop, macro detail on the brushed steel handle, studio lighting, sharp focus."

Swap the object, the surface, the backdrop color, and the material the macro detail lands on. Objects are the easiest category on this model because there is no anatomy to get wrong. The one thing to avoid: do not ask for the label text. Grok does not render legible type reliably, and a mangled brand name on an otherwise perfect bottle is a wasted generation. Generate the object clean, then add real typography in a layout tool, or route the job to a model built for text (more on that below).

Character and concept art

"Full body character portrait of a desert scavenger, patched canvas duster, tall angular shoulders, neutral standing pose, even studio light, plain slate background, sharp focus."

Swap the archetype, the one costume detail, and the one silhouette detail. Neutral pose plus flat background is deliberate. That is a reference frame, not a hero image, and it is the input you want if the character is going to be reused. The mechanics of reuse are in consistent AI characters.

A frame you plan to animate

"Medium shot of a mechanic in an open workshop bay, clean separation between subject and background, even light, mechanic centered, no motion blur, sharp focus."

Swap the subject and the setting, and change both mentions of the subject so they match. If a still is destined for image-to-video, compose it for stability: readable silhouette, uncluttered background, subject not cropped at an awkward joint. A busy frame turns into a morphing frame the moment motion is applied.

What Grok image generation gets right, and where it breaks

Knowing where this model breaks decides which jobs you send it and which you route elsewhere.

It is fast and cheap, and it is the loosest of the major models on content. That combination is why it dominates glam, pin-up, and character work where stricter models refuse or sanitize. It is also loosest on anatomy: hands, fingers, and limb counts fail at a noticeable rate. That is a regeneration problem, not a prompt problem. Do not write "perfect hands, five fingers" into the prompt; just generate again.

It is weak at text. Signs, labels, book spines, and logos come back as convincing gibberish. Never build a concept that depends on legible words in the image.

It is weak at precise spatial instructions. "Her left hand on the railing, the dog to her right, a bicycle leaning against the far wall" is three spatial constraints, and Grok will satisfy roughly one of them. Cut to one subject and one action beat. Multi-element layout work belongs on Nano Banana, which parses declarative sentences literally, or on GPT Image, which is the strongest of the three at compositional intent and coherent typography.

Generated media is reviewed and not used for training. xAI states that "generated media is subject to content policy review and is not used for training." Loose does not mean unmoderated, and a refusal on a given prompt is final for that call.

Reference images: match the aspect ratio or Grok stretches them

This is the detail that costs people the most time, because the output looks plausible until you notice everyone in it is slightly too wide.

xAI's image guide confirms that "multi-image editing supports up to 3 source images in a single request for combining subjects, transferring styles, and composing scenes." What the docs do not tell you is what happens when a source image and the requested output shape disagree. In our testing, Grok does not letterbox or crop the reference to fit. It stretches it to fill the output canvas. Feed a 16:9 photo into a 9:16 render and every face in the reference arrives at the model already squeezed narrow, and the model faithfully reproduces a squeezed face.

Left: a 16:9 reference stretched into a 9:16 canvas. Right: the same reference letterboxed first, proportions intact.

The fix is mechanical. Before you upload, pad or crop the reference to the exact aspect ratio you are going to request. A letterboxed reference with neutral bars beats a stretched one every time, and cropping to the output shape beats both when the crop does not lose the subject. Dream Pixel Forge does this for you: references are letterboxed to the output aspect ratio automatically before they reach the provider, so the distortion never happens.

Two more reference rules that hold across our runs: give each reference one job (subject, or style, or object, never all three), and use fewer references rather than more. Three competing sources produce a compromise; one clear source produces the thing you asked for.

Grok Imagine video prompts: prompt the motion, not the look

Grok Imagine 1.0 shipped on February 2, 2026 with clips up to 10 seconds at 720p and native synchronized audio generated in the same pass. On the API side, xAI's video guide describes configurable duration up to 15 seconds, and splits the models: grok-imagine-video-1.5 handles text-to-video and image-to-video, while reference-to-video stays on grok-imagine-video. So the honest answer to "how long is a Grok clip" depends on the door: the consumer app ships 10 seconds, the API lets you configure up to 15, and Dream Pixel Forge runs 6-second clips. Image-to-video takes exactly one source image, which becomes the starting frame; there is no token syntax and no way to index a second image into the shot.

That last point drives the whole prompting approach. When a source image exists, the look is already decided. Spending your prompt re-describing the scene fights the frame you already approved. Write the motion instead:

  • One camera move, named once. "Slow push-in" or "static locked-off shot." Two competing motion words make the model over-direct.
  • One action beat. A clip holds exactly one, whether it runs six seconds or fifteen. "He wipes his hands and looks at the camera" works. "He crosses the workshop, picks up a wrench, kneels, and starts loosening a bolt" produces a rushed morph.
  • Dialogue as an exact quoted line plus a tone, never a topic. He says flatly: "that bolt has been seized since the spring" beats "he explains the repair."
  • Name the ambience. Audio is promptable. "Distant traffic through an open roller door, no music" is a real instruction, not decoration.

A complete Grok Imagine video prompt for an animated still is often one sentence: "Slow push-in, he sets the wrench down and looks at the camera, low workshop hum." The same aspect-ratio rule applies here as for image references, and for the same reason: a mismatched source frame gets stretched to the output canvas. For the wider picture of animating stills, see image to video AI, and for the models that lead on identity lock and audio fidelity, Seedance and Veo 3 respectively.

How to use Grok Imagine

There are three doors, and they behave differently enough that the prompt advice above shifts slightly between them.

The consumer apps. Grok Imagine lives inside Grok and on X, as a chat-driven grok AI image generator with a generate-and-remix loop. It is the fastest way to feel out the model. The catch is that generation sits behind a paid X or SuperGrok subscription with daily caps that xAI does not publish in its developer docs, so the number you can produce today is whatever the app tells you it is. If your work depends on a known volume, this is the wrong door.

The Imagine API. Direct, metered, scriptable. Covered in the next section.

A studio that routes to it. Dream Pixel Forge runs Grok Imagine as one adapter behind a provider router, which means you get the model without the account, the caps, or the SDK. The same prompts on this page work in the freeform image generator, and the router falls back to another model automatically when a call fails rather than handing you an error.

The Grok Imagine API: what it costs

xAI publishes per-unit pricing on its models page. As of this writing:

ModelPriceNotes
grok-imagine-image$0.02 per imageThe standard image model. 5 requests per second.
grok-imagine-image-quality$0.05 per imageHigher-quality tier, also aliased grok-imagine-image-pro.
grok-imagine-video$0.050 per secondHandles reference-to-video, which 1.5 does not.
grok-imagine-video-1.5$0.080 per secondText-to-video and image-to-video.

Two API-only capabilities worth knowing: you can "configure output count (up to 10 images per request)" along with aspect ratio and resolution, and image edits bill for both the input and the output image. Batching ten variations of one prompt for twenty cents is the cheapest exploration loop in this model family, and it is the single best argument for using the API over the app.

In Dream Pixel Forge, Grok Imagine images cost 2 credits each and support seven aspect ratios (1:1, 2:3, 3:2, 3:4, 4:3, 9:16, 16:9) with up to three input images. Grok video runs 6-second clips at 480p or 720p in 16:9 or 9:16. Because it is cheap and fast, it is the default lane for iteration: find the winning composition on Grok, then re-render the keeper on a premium model if the brief needs it.

Fixing a bad Grok generation

Change one thing, matched to the symptom. Rewriting the whole prompt destroys the parts that were already working.

What went wrongThe one fix
Pretty image, wrong contentMove the shot type and subject to the front, cut style words
Correct but blandAdd exactly one light anchor, one lens anchor, one texture anchor
Subjects merged or scene is chaoticCut to one subject and one action beat
Face or proportions changed between runsSwitch to a reference image; stop adding identity adjectives
Reference subject looks stretchedMatch the reference aspect ratio to the output before uploading
Text or logo is gibberishDo not fix it here. Route the job to Nano Banana or GPT Image
Hands or fingers are wrongRegenerate. This is not promptable on this model

Run these prompts

Run the same subject twice, once as a keyword pile and once as the photographer skeleton, and compare. Paste any prompt from this page into the freeform generator below and change one clause at a time.

Try it now

Free to try, no account needed

Reference images

Select up to 8 images to guide the result.

0/8
Example output for Freeform Image Generator
Example output

Your generated image will replace this example.

The short version

Shot type first, two identity traits, plain wardrobe and setting, one light and lens clause, three words of finish. One subject, one scene. Never depend on legible text. Match reference aspect ratios to the output or accept the stretch. For video, prompt the motion and leave the look to the source frame. Grok will not follow a paragraph of instructions. Brief it like a photograph and it renders fast and cheap.

Tools for this guide

Frequently asked questions

How do you write a good Grok Imagine prompt?

Write like a photographer, not a designer. Put the shot type and subject first ("Full body vertical photograph of a..."), add one or two identity traits, then wardrobe and setting in plain words, then one light and lens clause ("golden hour, 85mm, shallow depth of field"), then three words of finish. Good Grok prompts usually land between 25 and 50 words. The first ~20 words carry the most weight, so anything buried behind a paragraph of style adjectives gets diluted. Avoid negative phrasing: "no plastic skin" adds the tokens for plastic skin rather than removing it.

What is Grok Imagine bad at?

Three things, reliably. Text: signs, labels, and logos come back as convincing gibberish, so never build a concept that needs legible words. Precise spatial layout: give it three positional constraints and it satisfies roughly one, so cut to one subject and one action beat. Anatomy: hands and proportions fail at a noticeable rate, and that is a regeneration problem, not a prompt problem. Route text-heavy and multi-element layout work to Nano Banana or GPT Image instead.

Why do my Grok Imagine reference images look stretched?

Because Grok fills the output canvas with the reference rather than letterboxing or cropping it. Feed a 16:9 photo into a 9:16 render and every face in that reference arrives at the model already squeezed narrow, and the model faithfully reproduces the distortion. Pad or crop your reference to the exact output aspect ratio before uploading. Dream Pixel Forge letterboxes references to the output aspect ratio automatically, so the stretch never reaches the provider.

How much does the Grok Imagine API cost?

Per xAI's models page: grok-imagine-image is $0.02 per image, grok-imagine-image-quality (aliased grok-imagine-image-pro) is $0.05 per image, grok-imagine-video is $0.050 per second, and grok-imagine-video-1.5 is $0.080 per second. Image models are rate limited to 5 requests per second, you can request up to 10 images per call, and edits bill for both the input and the output image. In Dream Pixel Forge, Grok Imagine images cost 2 credits each and 6-second Grok video clips cost 24 credits at 480p.

How do you prompt Grok Imagine for video?

Prompt the motion, not the look. When a source image exists the frame is already decided, so re-describing the scene fights it. Name one camera move ("slow push-in" or "static locked-off shot"), one action beat, dialogue as an exact quoted line plus a tone, and the ambient audio you want. Six to ten seconds holds exactly one beat; chained actions produce rushed morphs. One sentence is usually the whole prompt: "Slow push-in, she looks up from the book and smiles, ambient cafe noise."