Research · Original model evidence
Which image modelactually does what you asked?
We run these models in production every day. So we freeze the prompts, run every model through the same test, inspect the image that came back, and publish the evidence alongside the score — including run dates, suite versions, category breakdowns, and raw outcome counts.
Full size- models across the programme
- 7
- fixed, versioned test suites
- 3
- battery cells returned an image
- 91/91
- published protocol for every suite
- v1.0
The leaderboard
Both dimensions in one table, ordered by permissiveness. Higher is better in every column except sanitization, where a high number means the model changed your prompt without saying so.
Run 2026-08 · executed · suite v1.0
| Model | Vendor | Permissiveness | Sanitization | Adherence | Aesthetic |
|---|---|---|---|---|---|
| Nano Banana 2 Lite | 92.0% | 0.5% | 93.2% | 8.9 / 10 | |
| Nano Banana 2 | 91.0% | 2.4% | 93.5% | 8.9 / 10 | |
| Grok Imagine | xAI | 82.9% | 3.9% | 94.1% | 8.5 / 10 |
| GPT Image 2 | OpenAI | 69.0% | 1.9% | 95.6% | 9.0 / 10 |
| Grok Imagine 2.0 | xAI | — | — | 98.3% | 9.0 / 10 |
| Nano Banana Pro | — | — | 94.5% | 8.7 / 10 | |
| FLUX Kontext Pro | Black Forest Labs | — | — | 86.9% | 8.4 / 10 |
Three ways to fail
The Image Model Benchmarks
A returned image is not the same as a correct one. Each programme isolates a different failure: policy, instruction following, or the defaults a model falls back to when the prompt gets hard.
Weapons
Brands
Suggestive01 · Policy behaviour · Actual run output
The Image Model Censorship Benchmark
The same lawful request can be answered, refused, or quietly rewritten. We separate all four outcomes across public figures, IP, brands, weapons, satire, medical imagery, and more.
See the censorship leaderboard- 01Grok Imagine 2.098.3%
- 02GPT Image 295.6%
- 03Nano Banana Pro94.5%
02 · Instruction following · Measured scores
The Image Model Prompt Adherence Benchmark
Exact text, object counts, spatial relations, attribute binding, dense scenes, style fidelity, detail, and world knowledge — scored against binary criteria so pretty never passes as correct.
See the adherence leaderboard
03 · Launch-day defaults · Actual battery output
The Standard Image Model Battery
Thirteen prompts frozen in the repo: people, skin, hands, text, adherence, named styles, and one edit on a fixed source. Every new model lands beside the old ones without moving the goalposts.
See the battery resultsWhere the public arenas stand
Our benchmarks measure what a model does with a fixed prompt. The arenas measure which output people prefer, over millions of head-to-head votes, which is a different question and a useful one. This section is external data, quoted with its source and its as-of date, and refreshed here monthly.
Artificial Analysis — text to image
Retrieved 11 August 2026. The page states no update date.
- Rank 1GPT Image 2 (high)OpenAI
- 1,371
- Rank 2Reve 2.1Reve
- 1,324
- Rank 3Nano Banana 2Google
- 1,322
- Rank 4GPT Image 1.5 (high)OpenAI
- 1,313
- Rank 5MAI-Image-2.5Microsoft AI
- 1,309
- Rank 6Nano Banana ProGoogle
- 1,300
Ideogram 4.0 (Quality) leads the open-weights models at 1,214. The board flags five models added in the last month, all Ideogram or Cosmos variants.
Source: Elo leaderboardLMArena — text to image
Updated 10 August 2026, 5,919,946 votes across 77 models.
- Rank 1gpt-image-2 (medium)OpenAI
- 1381 ±5
- Rank 2mai-image-2.6-previewMicrosoft AI
- 1336 ±11
- Rank 3grok-imagine-image-2.0 (low)SpaceXAI
- 1316 ±12
- Rank 4reve-2.1Reve
- 1302 ±8
- Rank 5muse-imageMeta
- 1282 ±7
- Rank 6gemini-3.1-flash-imageGoogle
- 1264 ±5
LMArena marks ranks 2 to 5 here as preliminary: grok-imagine-image-2.0 sits third on 2,676 votes against 70,065 for the model above it, so that placement is thin and should be read as provisional.
Source: Score leaderboardArtificial Analysis — text to video, with audio
Retrieved 11 August 2026. The page states no update date.
- Rank 1Gemini Omni FlashGoogle
- 1,242
- Rank 2MiniMax H3MiniMax
- 1,237
- Rank 3Dreamina Seedance 2.0 720pByteDance Seed
- 1,223
- Rank 4Wan2.7-260612Alibaba
- 1,160
- Rank 5HappyHorse-1.1Alibaba-ATH
- 1,147
- Rank 6Kling 3.0 1080p (Pro)KlingAI
- 1,109
This is the with-audio board, which is what the site shows by default; the no-audio board ranks differently. On the same site's image-to-video board, Dreamina Seedance 2.0 720p leads at 1,198 with MiniMax H3 second at 1,193.
Source: Elo leaderboardDo not read these boards against each other. The two image arenas use different prompt sets, different voter pools, and different variants of the same models, and their scales are not comparable: the model at the top of one is not scored on the axis that put it at the top of the other. They are also not comparable with anything on our own pages, which score a fixed prompt against a checklist rather than counting votes.
How it is measured
The full protocol, including repeat counts and the nondeterminism caveat, is on each benchmark page.
Step 01
Fixed prompt suites
Three versioned suites of prompts, written once and reused verbatim across every model. The censorship suite covers contested-but-legitimate requests; the quality suite covers prompts models are known to fail; the standard battery is frozen and rerun on every model launch so a new model lands next to the old ones.
Step 02
Provider APIs, declared repeat counts
Every prompt runs through the provider's own API, at default safety settings. Some suites repeat every cell to measure nondeterminism; single-sample runs say so explicitly. Repeat counts matter because one refusal, or one unusually good image, proves very little.
Step 03
A vision model judges the image
Each returned image is scored against a per-prompt checklist by a vision model, not by whether the API returned a 200. That is the only way to catch a model that answers and quietly changes the request. The battery is the one exception, and it says so: how a body is rendered has no objective checklist, so those readings are made by eye and labelled as such.
Step 04
Everything is published
Prompt ids, category breakdowns, outcome counts, run date, and suite version all ship with the results, so a run can be argued with rather than taken on faith.
Questions
What does the permissiveness score measure?
- The percentage of generations across the whole censorship suite where the model produced exactly what the prompt asked for. Anything less than full compliance, whether that is an outright refusal or a quietly altered image, counts against it. It is a measure of whether the model does what you asked, not a judgement about whether it should.
What is silent sanitization?
- A model silently sanitizes when it never refuses but does not comply either: it swaps a named public figure for a generic lookalike, replaces a specific character with an original design, blurs a logo, or adds clothing nobody asked for. The API returns success, the image looks fine, and the request was not honoured. It is the single hardest failure mode to notice in production, which is why it gets its own headline metric.
Are the prompts trying to break the models?
- No. The censorship suite deliberately avoids anything genuinely prohibited. Every prompt in it is something a working designer, journalist, illustrator, or teacher would legitimately request: a named politician at a podium, a period-accurate rifle, a surgical illustration, a satirical cartoon. The question is what happens to lawful work, not whether guardrails can be defeated.
Do these results change over time?
- Yes, and that is the point of stamping every table with a run date and a suite version. Providers update safety filters far more often than they update model weights, and a model that complied last quarter can start sanitizing without any announcement. Results are only claims about the model as it behaved on the run date.
Which model should I use?
- It depends on which failure hurts more. If your work involves real people, real brands, or named characters, read the censorship table first — the models are far apart there, and the gap is not where most people expect it. If it involves text in the image, exact counts, or complex layouts, read the prompt adherence benchmark. Compare a model across dimensions only where it appears in both runs, and keep the protocol differences in view: the suites use different prompts, scoring rules, and repeat counts.