The Image Model Prompt Adherence Benchmark

Every model on this list can make a beautiful image. Far fewer can put eleven objects in a frame, spell a word on a sign, or keep the green coat on the person who was supposed to be wearing it. Adherence is scored on binary criteria written before the run; aesthetics are scored separately, so pretty never passes for correct.

Sample data

Every number on this page is placeholder data. The first official benchmark run is pending, and nothing here should be cited, quoted, or used to choose a model.

Adherence ranking

The share of binary criteria each model satisfied across the whole suite, with the aesthetic score alongside it rather than folded into it.

Run 2026-08-sample · executed August 1, 2026 · suite v0.1.0-sample

#ModelVendorAdherenceAesthetic
1GPT Image 2OpenAI85.1%8.1 / 10
2Nano Banana 2Google83.7%8.3 / 10
3Nano Banana 2 LiteGoogle73.3%7.2 / 10
4Grok ImaginexAI68.4%7.6 / 10

Adherence by category

Where each model actually spends its errors. The overall score is the mean of these columns, so a model that is excellent everywhere except counting still shows exactly where it broke.

ModelText renderingCountingSpatial relationsAttribute bindingComplex scenesStyle fidelityPhotoreal detailWorld knowledgeAesthetic
GPT Image 291.7%83.3%86.1%88.9%80.6%77.8%86.1%86.1%8.1 / 10
Nano Banana 288.9%80.6%83.3%86.1%77.8%80.6%88.9%83.3%8.3 / 10
Nano Banana 2 Lite72.2%69.4%75.0%77.8%66.7%72.2%77.8%75.0%7.2 / 10
Grok Imagine58.3%61.1%69.4%72.2%63.9%75.0%80.6%66.7%7.6 / 10
Text rendering
Exact requested strings, spelled correctly, on signs, packaging, and screens.
Counting
Exact object counts, including counts past the five-or-so most models handle.
Spatial relations
Left/right, above/below, in front of, and inside, all resolved correctly.
Attribute binding
Colours and materials attached to the object they were named with, not swapped between objects.
Complex scenes
Long prompts with many interacting subjects, each of which must survive.
Style fidelity
A named medium, era, or process reproduced rather than approximated.
Photoreal detail
Hands, reflections, optics, and physical plausibility under close inspection.
World knowledge
Facts the image must get right: landmarks, uniforms, period dress, and instrument layouts.

Methodology

Run 2026-08-sample · executed August 1, 2026 · suite v0.1.0-sample

Placeholder data used to build and review the leaderboard surface before the first official run. Shape, category ids, and model roster match what the real runner emits: 9 censorship categories and 8 quality categories, 4 prompts each, 4 repeats per prompt — 144 censorship generations and 128 quality generations per model. Every number here is invented and must not be cited.

The prompts

Eight categories, four prompts each, chosen because image models fail them in predictable ways rather than because they are exotic. Signage with an exact string on it. Eleven of something. A red cube to the left of a blue sphere. A woman in a green coat holding a yellow umbrella, where the colours must stay where they were put.

The scoring

Every prompt carries a list of binary criteria written before the run, one per checkable claim. Adherence is the share of criteria the image passes, averaged over the repeats and then over the categories, so no single crowded category dominates the headline number.

Aesthetics, scored separately

A vision judge also rates each image 1 to 10 on holistic visual quality, deliberately in isolation from the criteria. Keeping the two apart is the whole point: models optimised for a pleasing image will happily produce a beautiful picture of the wrong thing, and a combined score would let them hide there.

The caveats

Image models are nondeterministic, so each prompt runs several times and the reported figure is a mean, not a best-of. Sampler defaults, aspect ratio, and provider-side prompt rewriting all move these numbers, and every model is run at its provider defaults with the prompt passed through verbatim. Results describe the run date and suite version stamped on this page.

Refusal and sanitization rates are on the censorship benchmark.