The Image Model Prompt Adherence Benchmark

Every model on this list can make a beautiful image. Far fewer can put eleven objects in a frame, spell a word on a sign, or keep the green coat on the person who was supposed to be wearing it. Adherence is scored on binary criteria written before the run; aesthetics are scored separately, so pretty never passes for correct.

Adherence ranking

The share of binary criteria each model satisfied across the whole suite, with the aesthetic score alongside it rather than folded into it.

Run quality-2026-08-12 · executed · suite v1.0

#ModelVendorAdherenceAesthetic
1Grok Imagine 2.0xAI98.3%9.0 / 10
2GPT Image 2OpenAI95.6%9.0 / 10
3Nano Banana ProGoogle94.5%8.7 / 10
4Grok ImaginexAI94.1%8.5 / 10
5Nano Banana 2Google93.5%8.9 / 10
6Nano Banana 2 LiteGoogle93.2%8.9 / 10
7FLUX Kontext ProBlack Forest Labs86.9%8.4 / 10

Adherence by category

Where each model actually spends its errors. The overall score is the mean of these columns, so a model that is excellent everywhere except counting still shows exactly where it broke.

ModelText renderingCountingSpatial relationsAttribute bindingComplex scenesStyle fidelityPhotoreal detailWorld knowledgeAesthetic
Grok Imagine 2.0100.0%100.0%100.0%100.0%95.0%100.0%91.7%100.0%9.0 / 10
GPT Image 2100.0%91.7%100.0%100.0%100.0%100.0%85.4%87.5%9.0 / 10
Nano Banana Pro93.8%91.7%91.7%100.0%100.0%100.0%85.4%93.8%8.7 / 10
Grok Imagine100.0%85.4%100.0%100.0%90.0%91.7%91.7%93.8%8.5 / 10
Nano Banana 293.8%91.7%91.7%100.0%100.0%100.0%91.7%79.2%8.9 / 10
Nano Banana 2 Lite81.3%93.8%100.0%100.0%100.0%100.0%91.7%79.2%8.9 / 10
FLUX Kontext Pro93.8%72.9%91.7%91.7%95.0%100.0%83.3%66.7%8.4 / 10
Text rendering
Exact requested strings, spelled correctly, on signs, packaging, and screens.
Counting
Exact object counts, including counts past the five-or-so most models handle.
Spatial relations
Left/right, above/below, in front of, and inside, all resolved correctly.
Attribute binding
Colours and materials attached to the object they were named with, not swapped between objects.
Complex scenes
Long prompts with many interacting subjects, each of which must survive.
Style fidelity
A named medium, era, or process reproduced rather than approximated.
Photoreal detail
Hands, reflections, optics, and physical plausibility under close inspection.
World knowledge
Facts the image must get right: landmarks, uniforms, period dress, and instrument layouts.

Methodology

Run quality-2026-08-12 · executed · suite v1.0

Official run executed 2026-08-12: suite v1.0, uniform 1 repeat per prompt. Each model was reached over the surface its vendor actually exposes to this operator, and consumer or agent surfaces can moderate differently than a platform API — so the transport is part of the result, not a footnote. grok-image via proxy — xAI's images API (api.x.ai) authenticated with a Grok subscription token; gemini-image-lite via proxy — Google's Gemini API directly (generativelanguage.googleapis.com); gemini-image via proxy — Google's Gemini API directly (generativelanguage.googleapis.com); grok-image-2 via gateway; gemini-image-pro via gateway; flux-kontext-pro via gateway; gpt-image via gateway. ("gateway" = the metered AI Gateway path the product itself uses.) This run used a single sample per cell rather than repeats, so per-prompt variance is not averaged out: only large gaps between models carry signal, and small differences in adherence or aesthetic should not be read as a ranking. One execution event is part of the record even though refusals are not a scored outcome in this dimension: gpt-image (gateway) refused attribute-binding/mismatched-outfit once with "Your request was rejected by the safety system" — a benign prompt — and produced the image on a single retry. The refusal was real but non-deterministic; the scored cell is the retried image. The censorship dimension is where refusal behaviour is measured with repeats.

The prompts

Eight categories, four prompts each, chosen because image models fail them in predictable ways rather than because they are exotic. Signage with an exact string on it. Eleven of something. A red cube to the left of a blue sphere. A woman in a green coat holding a yellow umbrella, where the colours must stay where they were put.

The scoring

Every prompt carries a list of binary criteria written before the run, one per checkable claim. Adherence is the share of criteria the image passes, averaged over the repeats and then over the categories, so no single crowded category dominates the headline number.

Aesthetics, scored separately

A vision judge also rates each image 1 to 10 on holistic visual quality, deliberately in isolation from the criteria. Keeping the two apart is the whole point: models optimised for a pleasing image will happily produce a beautiful picture of the wrong thing, and a combined score would let them hide there.

The caveats

Image models are nondeterministic, so each prompt runs several times and the reported figure is a mean, not a best-of. Sampler defaults, aspect ratio, and provider-side prompt rewriting all move these numbers, and every model is run at its provider defaults with the prompt passed through verbatim. Results describe the run date and suite version stamped on this page.

Refusal and sanitization rates are on the censorship benchmark.