Toggle navigation
The Image Model Prompt Adherence Benchmark
Every model on this list can make a beautiful image. Far fewer can put eleven objects in a frame, spell a word on a sign, or keep the green coat on the person who was supposed to be wearing it. Adherence is scored on binary criteria written before the run; aesthetics are scored separately, so pretty never passes for correct.
Adherence ranking
The share of binary criteria each model satisfied across the whole suite, with the aesthetic score alongside it rather than folded into it.
Run quality-2026-08-12 · executed · suite v1.0
| # | Model | Vendor | Adherence | Aesthetic |
|---|---|---|---|---|
| 1 | Grok Imagine 2.0 | xAI | 98.3% | 9.0 / 10 |
| 2 | GPT Image 2 | OpenAI | 95.6% | 9.0 / 10 |
| 3 | Nano Banana Pro | 94.5% | 8.7 / 10 | |
| 4 | Grok Imagine | xAI | 94.1% | 8.5 / 10 |
| 5 | Nano Banana 2 | 93.5% | 8.9 / 10 | |
| 6 | Nano Banana 2 Lite | 93.2% | 8.9 / 10 | |
| 7 | FLUX Kontext Pro | Black Forest Labs | 86.9% | 8.4 / 10 |
Adherence by category
Where each model actually spends its errors. The overall score is the mean of these columns, so a model that is excellent everywhere except counting still shows exactly where it broke.
| Model | Text rendering | Counting | Spatial relations | Attribute binding | Complex scenes | Style fidelity | Photoreal detail | World knowledge | Aesthetic |
|---|---|---|---|---|---|---|---|---|---|
| Grok Imagine 2.0 | 100.0% | 100.0% | 100.0% | 100.0% | 95.0% | 100.0% | 91.7% | 100.0% | 9.0 / 10 |
| GPT Image 2 | 100.0% | 91.7% | 100.0% | 100.0% | 100.0% | 100.0% | 85.4% | 87.5% | 9.0 / 10 |
| Nano Banana Pro | 93.8% | 91.7% | 91.7% | 100.0% | 100.0% | 100.0% | 85.4% | 93.8% | 8.7 / 10 |
| Grok Imagine | 100.0% | 85.4% | 100.0% | 100.0% | 90.0% | 91.7% | 91.7% | 93.8% | 8.5 / 10 |
| Nano Banana 2 | 93.8% | 91.7% | 91.7% | 100.0% | 100.0% | 100.0% | 91.7% | 79.2% | 8.9 / 10 |
| Nano Banana 2 Lite | 81.3% | 93.8% | 100.0% | 100.0% | 100.0% | 100.0% | 91.7% | 79.2% | 8.9 / 10 |
| FLUX Kontext Pro | 93.8% | 72.9% | 91.7% | 91.7% | 95.0% | 100.0% | 83.3% | 66.7% | 8.4 / 10 |
- Text rendering
- Exact requested strings, spelled correctly, on signs, packaging, and screens.
- Counting
- Exact object counts, including counts past the five-or-so most models handle.
- Spatial relations
- Left/right, above/below, in front of, and inside, all resolved correctly.
- Attribute binding
- Colours and materials attached to the object they were named with, not swapped between objects.
- Complex scenes
- Long prompts with many interacting subjects, each of which must survive.
- Style fidelity
- A named medium, era, or process reproduced rather than approximated.
- Photoreal detail
- Hands, reflections, optics, and physical plausibility under close inspection.
- World knowledge
- Facts the image must get right: landmarks, uniforms, period dress, and instrument layouts.
Methodology
Run quality-2026-08-12 · executed · suite v1.0
Official run executed 2026-08-12: suite v1.0, uniform 1 repeat per prompt. Each model was reached over the surface its vendor actually exposes to this operator, and consumer or agent surfaces can moderate differently than a platform API — so the transport is part of the result, not a footnote. grok-image via proxy — xAI's images API (api.x.ai) authenticated with a Grok subscription token; gemini-image-lite via proxy — Google's Gemini API directly (generativelanguage.googleapis.com); gemini-image via proxy — Google's Gemini API directly (generativelanguage.googleapis.com); grok-image-2 via gateway; gemini-image-pro via gateway; flux-kontext-pro via gateway; gpt-image via gateway. ("gateway" = the metered AI Gateway path the product itself uses.) This run used a single sample per cell rather than repeats, so per-prompt variance is not averaged out: only large gaps between models carry signal, and small differences in adherence or aesthetic should not be read as a ranking. One execution event is part of the record even though refusals are not a scored outcome in this dimension: gpt-image (gateway) refused attribute-binding/mismatched-outfit once with "Your request was rejected by the safety system" — a benign prompt — and produced the image on a single retry. The refusal was real but non-deterministic; the scored cell is the retried image. The censorship dimension is where refusal behaviour is measured with repeats.
The prompts
- Eight categories, four prompts each, chosen because image models fail them in predictable ways rather than because they are exotic. Signage with an exact string on it. Eleven of something. A red cube to the left of a blue sphere. A woman in a green coat holding a yellow umbrella, where the colours must stay where they were put.
The scoring
- Every prompt carries a list of binary criteria written before the run, one per checkable claim. Adherence is the share of criteria the image passes, averaged over the repeats and then over the categories, so no single crowded category dominates the headline number.
Aesthetics, scored separately
- A vision judge also rates each image 1 to 10 on holistic visual quality, deliberately in isolation from the criteria. Keeping the two apart is the whole point: models optimised for a pleasing image will happily produce a beautiful picture of the wrong thing, and a combined score would let them hide there.
The caveats
- Image models are nondeterministic, so each prompt runs several times and the reported figure is a mean, not a best-of. Sampler defaults, aspect ratio, and provider-side prompt rewriting all move these numbers, and every model is run at its provider defaults with the prompt passed through verbatim. Results describe the run date and suite version stamped on this page.
Refusal and sanitization rates are on the censorship benchmark.
