Research

Which image modelactually does what you asked?

We run these models in production every day, so we test them the way we need them tested and publish what comes back. The first programme is The Image Model Benchmarks, measuring two things nobody reports honestly: how often a model refuses or quietly rewrites a legitimate prompt, and how closely it follows a hard one. Same prompts, same repeats, every model, judged on the image that came back rather than on the status code.

Everything published here ships with its data: run dates, suite versions, category breakdowns, and raw outcome counts, so a result can be argued with rather than taken on faith.

Sample data

Every number on this page is placeholder data. The first official benchmark run is pending, and nothing here should be cited, quoted, or used to choose a model.

The leaderboard

Both dimensions in one table, ordered by permissiveness. Higher is better in every column except sanitization, where a high number means the model changed your prompt without saying so.

Run 2026-08-sample · executed August 1, 2026 · suite v0.1.0-sample

ModelVendorPermissivenessSanitizationAdherenceAesthetic
Grok ImaginexAI79.2%16.0%68.4%7.6 / 10
Nano Banana 2Google47.2%34.0%83.7%8.3 / 10
Nano Banana 2 LiteGoogle38.2%36.1%73.3%7.2 / 10
GPT Image 2OpenAI19.4%29.2%85.1%8.1 / 10

The Image Model Benchmarks

Our first research programme, in two halves. Each has its own suite, its own leaderboard, and its own methodology.

How it is measured

The full protocol, including repeat counts and the nondeterminism caveat, is on each benchmark page.

  1. 01

    Fixed prompt suites

    Two versioned suites of prompts, written once and reused verbatim across every model. The censorship suite covers contested-but-legitimate requests; the quality suite covers prompts models are known to fail.

  2. 02

    Provider APIs, repeated runs

    Every prompt runs several times per model through the provider's own API, at default safety settings. Repeats matter: these models are nondeterministic, and a single refusal proves very little.

  3. 03

    A vision model judges the image

    Each returned image is scored against a per-prompt checklist by a vision model, not by whether the API returned a 200. That is the only way to catch a model that answers and quietly changes the request.

  4. 04

    Everything is published

    Prompt ids, category breakdowns, outcome counts, run date, and suite version all ship with the results, so a run can be argued with rather than taken on faith.

Questions

What does the permissiveness score measure?

The percentage of generations across the whole censorship suite where the model produced exactly what the prompt asked for. Anything less than full compliance, whether that is an outright refusal or a quietly altered image, counts against it. It is a measure of whether the model does what you asked, not a judgement about whether it should.

What is silent sanitization?

A model silently sanitizes when it never refuses but does not comply either: it swaps a named public figure for a generic lookalike, replaces a specific character with an original design, blurs a logo, or adds clothing nobody asked for. The API returns success, the image looks fine, and the request was not honoured. It is the single hardest failure mode to notice in production, which is why it gets its own headline metric.

Are the prompts trying to break the models?

No. The censorship suite deliberately avoids anything genuinely prohibited. Every prompt in it is something a working designer, journalist, illustrator, or teacher would legitimately request: a named politician at a podium, a period-accurate rifle, a surgical illustration, a satirical cartoon. The question is what happens to lawful work, not whether guardrails can be defeated.

Do these results change over time?

Yes, and that is the point of stamping every table with a run date and a suite version. Providers update safety filters far more often than they update model weights, and a model that complied last quarter can start sanitizing without any announcement. Results are only claims about the model as it behaved on the run date.

Which model should I use?

It depends on which failure hurts more. The two dimensions pull against each other: the most permissive model on this list is also the weakest at following a hard prompt, and the model with the best prompt adherence is the most restrictive one here by a wide margin. If your work involves real people, real brands, or real products, read the censorship table first; if it involves text in the image, exact counts, or complex layouts, read the adherence table first.