Toggle navigation
Research
Which image modelactually does what you asked?
We run these models in production every day, so we test them the way we need them tested and publish what comes back. The first programme is The Image Model Benchmarks, measuring two things nobody reports honestly: how often a model refuses or quietly rewrites a legitimate prompt, and how closely it follows a hard one. Same prompts, same repeats, every model, judged on the image that came back rather than on the status code.
Everything published here ships with its data: run dates, suite versions, category breakdowns, and raw outcome counts, so a result can be argued with rather than taken on faith.
Sample data
Every number on this page is placeholder data. The first official benchmark run is pending, and nothing here should be cited, quoted, or used to choose a model.
The leaderboard
Both dimensions in one table, ordered by permissiveness. Higher is better in every column except sanitization, where a high number means the model changed your prompt without saying so.
Run 2026-08-sample · executed August 1, 2026 · suite v0.1.0-sample
| Model | Vendor | Permissiveness | Sanitization | Adherence | Aesthetic |
|---|---|---|---|---|---|
| Grok Imagine | xAI | 79.2% | 16.0% | 68.4% | 7.6 / 10 |
| Nano Banana 2 | 47.2% | 34.0% | 83.7% | 8.3 / 10 | |
| Nano Banana 2 Lite | 38.2% | 36.1% | 73.3% | 7.2 / 10 | |
| GPT Image 2 | OpenAI | 19.4% | 29.2% | 85.1% | 8.1 / 10 |
The Image Model Benchmarks
Our first research programme, in two halves. Each has its own suite, its own leaderboard, and its own methodology.
The Image Model Censorship Benchmark
Nine categories of contested-but-legitimate prompts: public figures, IP characters, brands, weapons, violence, political satire, suggestive, medical, and religious. Four possible outcomes per generation, including the one nobody reports — the model complies in form and rewrites your prompt in substance.
See the censorship leaderboardThe Image Model Prompt Adherence Benchmark
Eight categories of prompts that image models reliably fail: rendering exact text, counting objects, spatial relations, attribute binding, dense scenes, style fidelity, photoreal detail, and world knowledge. Scored on binary criteria, with a separate aesthetic score so pretty never passes for correct.
See the adherence leaderboardHow it is measured
The full protocol, including repeat counts and the nondeterminism caveat, is on each benchmark page.
- 01
Fixed prompt suites
Two versioned suites of prompts, written once and reused verbatim across every model. The censorship suite covers contested-but-legitimate requests; the quality suite covers prompts models are known to fail.
- 02
Provider APIs, repeated runs
Every prompt runs several times per model through the provider's own API, at default safety settings. Repeats matter: these models are nondeterministic, and a single refusal proves very little.
- 03
A vision model judges the image
Each returned image is scored against a per-prompt checklist by a vision model, not by whether the API returned a 200. That is the only way to catch a model that answers and quietly changes the request.
- 04
Everything is published
Prompt ids, category breakdowns, outcome counts, run date, and suite version all ship with the results, so a run can be argued with rather than taken on faith.
Questions
What does the permissiveness score measure?
- The percentage of generations across the whole censorship suite where the model produced exactly what the prompt asked for. Anything less than full compliance, whether that is an outright refusal or a quietly altered image, counts against it. It is a measure of whether the model does what you asked, not a judgement about whether it should.
What is silent sanitization?
- A model silently sanitizes when it never refuses but does not comply either: it swaps a named public figure for a generic lookalike, replaces a specific character with an original design, blurs a logo, or adds clothing nobody asked for. The API returns success, the image looks fine, and the request was not honoured. It is the single hardest failure mode to notice in production, which is why it gets its own headline metric.
Are the prompts trying to break the models?
- No. The censorship suite deliberately avoids anything genuinely prohibited. Every prompt in it is something a working designer, journalist, illustrator, or teacher would legitimately request: a named politician at a podium, a period-accurate rifle, a surgical illustration, a satirical cartoon. The question is what happens to lawful work, not whether guardrails can be defeated.
Do these results change over time?
- Yes, and that is the point of stamping every table with a run date and a suite version. Providers update safety filters far more often than they update model weights, and a model that complied last quarter can start sanitizing without any announcement. Results are only claims about the model as it behaved on the run date.
Which model should I use?
- It depends on which failure hurts more. The two dimensions pull against each other: the most permissive model on this list is also the weakest at following a hard prompt, and the model with the best prompt adherence is the most restrictive one here by a wide margin. If your work involves real people, real brands, or real products, read the censorship table first; if it involves text in the image, exact counts, or complex layouts, read the adherence table first.
