Toggle navigation
The Image Model Prompt Adherence Benchmark
Every model on this list can make a beautiful image. Far fewer can put eleven objects in a frame, spell a word on a sign, or keep the green coat on the person who was supposed to be wearing it. Adherence is scored on binary criteria written before the run; aesthetics are scored separately, so pretty never passes for correct.
Sample data
Every number on this page is placeholder data. The first official benchmark run is pending, and nothing here should be cited, quoted, or used to choose a model.
Adherence ranking
The share of binary criteria each model satisfied across the whole suite, with the aesthetic score alongside it rather than folded into it.
Run 2026-08-sample · executed August 1, 2026 · suite v0.1.0-sample
| # | Model | Vendor | Adherence | Aesthetic |
|---|---|---|---|---|
| 1 | GPT Image 2 | OpenAI | 85.1% | 8.1 / 10 |
| 2 | Nano Banana 2 | 83.7% | 8.3 / 10 | |
| 3 | Nano Banana 2 Lite | 73.3% | 7.2 / 10 | |
| 4 | Grok Imagine | xAI | 68.4% | 7.6 / 10 |
Adherence by category
Where each model actually spends its errors. The overall score is the mean of these columns, so a model that is excellent everywhere except counting still shows exactly where it broke.
| Model | Text rendering | Counting | Spatial relations | Attribute binding | Complex scenes | Style fidelity | Photoreal detail | World knowledge | Aesthetic |
|---|---|---|---|---|---|---|---|---|---|
| GPT Image 2 | 91.7% | 83.3% | 86.1% | 88.9% | 80.6% | 77.8% | 86.1% | 86.1% | 8.1 / 10 |
| Nano Banana 2 | 88.9% | 80.6% | 83.3% | 86.1% | 77.8% | 80.6% | 88.9% | 83.3% | 8.3 / 10 |
| Nano Banana 2 Lite | 72.2% | 69.4% | 75.0% | 77.8% | 66.7% | 72.2% | 77.8% | 75.0% | 7.2 / 10 |
| Grok Imagine | 58.3% | 61.1% | 69.4% | 72.2% | 63.9% | 75.0% | 80.6% | 66.7% | 7.6 / 10 |
- Text rendering
- Exact requested strings, spelled correctly, on signs, packaging, and screens.
- Counting
- Exact object counts, including counts past the five-or-so most models handle.
- Spatial relations
- Left/right, above/below, in front of, and inside, all resolved correctly.
- Attribute binding
- Colours and materials attached to the object they were named with, not swapped between objects.
- Complex scenes
- Long prompts with many interacting subjects, each of which must survive.
- Style fidelity
- A named medium, era, or process reproduced rather than approximated.
- Photoreal detail
- Hands, reflections, optics, and physical plausibility under close inspection.
- World knowledge
- Facts the image must get right: landmarks, uniforms, period dress, and instrument layouts.
Methodology
Run 2026-08-sample · executed August 1, 2026 · suite v0.1.0-sample
Placeholder data used to build and review the leaderboard surface before the first official run. Shape, category ids, and model roster match what the real runner emits: 9 censorship categories and 8 quality categories, 4 prompts each, 4 repeats per prompt — 144 censorship generations and 128 quality generations per model. Every number here is invented and must not be cited.
The prompts
- Eight categories, four prompts each, chosen because image models fail them in predictable ways rather than because they are exotic. Signage with an exact string on it. Eleven of something. A red cube to the left of a blue sphere. A woman in a green coat holding a yellow umbrella, where the colours must stay where they were put.
The scoring
- Every prompt carries a list of binary criteria written before the run, one per checkable claim. Adherence is the share of criteria the image passes, averaged over the repeats and then over the categories, so no single crowded category dominates the headline number.
Aesthetics, scored separately
- A vision judge also rates each image 1 to 10 on holistic visual quality, deliberately in isolation from the criteria. Keeping the two apart is the whole point: models optimised for a pleasing image will happily produce a beautiful picture of the wrong thing, and a combined score would let them hide there.
The caveats
- Image models are nondeterministic, so each prompt runs several times and the reported figure is a mean, not a best-of. Sampler defaults, aspect ratio, and provider-side prompt rewriting all move these numbers, and every model is run at its provider defaults with the prompt passed through verbatim. Results describe the run date and suite version stamped on this page.
Refusal and sanitization rates are on the censorship benchmark.
