Research · Original model evidence

Which image modelactually does what you asked?

We run these models in production every day. So we freeze the prompts, run every model through the same test, inspect the image that came back, and publish the evidence alongside the score — including run dates, suite versions, category breakdowns, and raw outcome counts.

Seven AI image models answering the same neutral portrait prompt, shown side by sideFull size
One prompt, seven models, first render each. The frozen battery exposes each model's default person before any preference score gets involved.
models across the programme
7
fixed, versioned test suites
3
battery cells returned an image
91/91
published protocol for every suite
v1.0

The leaderboard

Both dimensions in one table, ordered by permissiveness. Higher is better in every column except sanitization, where a high number means the model changed your prompt without saying so.

Run 2026-08 · executed · suite v1.0

ModelVendorPermissivenessSanitizationAdherenceAesthetic
Nano Banana 2 LiteGoogle92.0%0.5%93.2%8.9 / 10
Nano Banana 2Google91.0%2.4%93.5%8.9 / 10
Grok ImaginexAI82.9%3.9%94.1%8.5 / 10
GPT Image 2OpenAI69.0%1.9%95.6%9.0 / 10
Grok Imagine 2.0xAI98.3%9.0 / 10
Nano Banana ProGoogle94.5%8.7 / 10
FLUX Kontext ProBlack Forest Labs86.9%8.4 / 10

Three ways to fail

The Image Model Benchmarks

A returned image is not the same as a correct one. Each programme isolates a different failure: policy, instruction following, or the defaults a model falls back to when the prompt gets hard.

Where the public arenas stand

Our benchmarks measure what a model does with a fixed prompt. The arenas measure which output people prefer, over millions of head-to-head votes, which is a different question and a useful one. This section is external data, quoted with its source and its as-of date, and refreshed here monthly.

Artificial Analysis — text to image

Retrieved 11 August 2026. The page states no update date.

Rank 1GPT Image 2 (high)OpenAI
1,371
Rank 2Reve 2.1Reve
1,324
Rank 3Nano Banana 2Google
1,322
Rank 4GPT Image 1.5 (high)OpenAI
1,313
Rank 5MAI-Image-2.5Microsoft AI
1,309
Rank 6Nano Banana ProGoogle
1,300

Ideogram 4.0 (Quality) leads the open-weights models at 1,214. The board flags five models added in the last month, all Ideogram or Cosmos variants.

Source: Elo leaderboard

LMArena — text to image

Updated 10 August 2026, 5,919,946 votes across 77 models.

Rank 1gpt-image-2 (medium)OpenAI
1381 ±5
Rank 2mai-image-2.6-previewMicrosoft AI
1336 ±11
Rank 3grok-imagine-image-2.0 (low)SpaceXAI
1316 ±12
Rank 4reve-2.1Reve
1302 ±8
Rank 5muse-imageMeta
1282 ±7
Rank 6gemini-3.1-flash-imageGoogle
1264 ±5

LMArena marks ranks 2 to 5 here as preliminary: grok-imagine-image-2.0 sits third on 2,676 votes against 70,065 for the model above it, so that placement is thin and should be read as provisional.

Source: Score leaderboard

Artificial Analysis — text to video, with audio

Retrieved 11 August 2026. The page states no update date.

Rank 1Gemini Omni FlashGoogle
1,242
Rank 2MiniMax H3MiniMax
1,237
Rank 3Dreamina Seedance 2.0 720pByteDance Seed
1,223
Rank 4Wan2.7-260612Alibaba
1,160
Rank 5HappyHorse-1.1Alibaba-ATH
1,147
Rank 6Kling 3.0 1080p (Pro)KlingAI
1,109

This is the with-audio board, which is what the site shows by default; the no-audio board ranks differently. On the same site's image-to-video board, Dreamina Seedance 2.0 720p leads at 1,198 with MiniMax H3 second at 1,193.

Source: Elo leaderboard

Do not read these boards against each other. The two image arenas use different prompt sets, different voter pools, and different variants of the same models, and their scales are not comparable: the model at the top of one is not scored on the axis that put it at the top of the other. They are also not comparable with anything on our own pages, which score a fixed prompt against a checklist rather than counting votes.

How it is measured

The full protocol, including repeat counts and the nondeterminism caveat, is on each benchmark page.

  1. Step 01

    Fixed prompt suites

    Three versioned suites of prompts, written once and reused verbatim across every model. The censorship suite covers contested-but-legitimate requests; the quality suite covers prompts models are known to fail; the standard battery is frozen and rerun on every model launch so a new model lands next to the old ones.

  2. Step 02

    Provider APIs, declared repeat counts

    Every prompt runs through the provider's own API, at default safety settings. Some suites repeat every cell to measure nondeterminism; single-sample runs say so explicitly. Repeat counts matter because one refusal, or one unusually good image, proves very little.

  3. Step 03

    A vision model judges the image

    Each returned image is scored against a per-prompt checklist by a vision model, not by whether the API returned a 200. That is the only way to catch a model that answers and quietly changes the request. The battery is the one exception, and it says so: how a body is rendered has no objective checklist, so those readings are made by eye and labelled as such.

  4. Step 04

    Everything is published

    Prompt ids, category breakdowns, outcome counts, run date, and suite version all ship with the results, so a run can be argued with rather than taken on faith.

Questions

What does the permissiveness score measure?

The percentage of generations across the whole censorship suite where the model produced exactly what the prompt asked for. Anything less than full compliance, whether that is an outright refusal or a quietly altered image, counts against it. It is a measure of whether the model does what you asked, not a judgement about whether it should.

What is silent sanitization?

A model silently sanitizes when it never refuses but does not comply either: it swaps a named public figure for a generic lookalike, replaces a specific character with an original design, blurs a logo, or adds clothing nobody asked for. The API returns success, the image looks fine, and the request was not honoured. It is the single hardest failure mode to notice in production, which is why it gets its own headline metric.

Are the prompts trying to break the models?

No. The censorship suite deliberately avoids anything genuinely prohibited. Every prompt in it is something a working designer, journalist, illustrator, or teacher would legitimately request: a named politician at a podium, a period-accurate rifle, a surgical illustration, a satirical cartoon. The question is what happens to lawful work, not whether guardrails can be defeated.

Do these results change over time?

Yes, and that is the point of stamping every table with a run date and a suite version. Providers update safety filters far more often than they update model weights, and a model that complied last quarter can start sanitizing without any announcement. Results are only claims about the model as it behaved on the run date.

Which model should I use?

It depends on which failure hurts more. If your work involves real people, real brands, or named characters, read the censorship table first — the models are far apart there, and the gap is not where most people expect it. If it involves text in the image, exact counts, or complex layouts, read the prompt adherence benchmark. Compare a model across dimensions only where it appears in both runs, and keep the protocol differences in view: the suites use different prompts, scoring rules, and repeat counts.