Guides
Nano Banana Prompts: The Formula That Actually Works
Nano Banana prompts work when you write a scene, not a keyword pile. The formula, worked examples, editing rules, and which of the three models to route a brief to.
Nano Banana prompts work when you write a scene in plain sentences, not a pile of keywords. Google's own guidance is blunt about it: "A simple list of keywords won't cut it; you need to describe the scene narratively." Start the prompt with a strong verb naming the operation, describe what you want rather than what you do not want, and put the subject before the style language. That single habit fixes most bad results.
What follows: which of the three Nano Banana models to send a brief to, the prompt shape that survives contact with the model, what to change when a render fails, and the jobs where Nano Banana is the wrong model no matter how good the prompt is.
What Nano Banana actually is in 2026
Nano Banana is the nickname that stuck to Google's Gemini image models, a throwaway LMArena codename a DeepMind product manager made out of her own two nicknames and never expected to go public (Google tells the story in how Nano Banana got its name). There are now three current models under the name, and they take prompts differently enough that "a Nano Banana prompt" is not one thing.
| Model | API id | Output | Google list price, 1K image |
|---|---|---|---|
| Nano Banana 2 Lite | gemini-3.1-flash-lite-image | 1K | $0.0336 |
| Nano Banana 2 | gemini-3.1-flash-image | 1K | $0.067 |
| Nano Banana Pro | gemini-3-pro-image | 1K, 2K, 4K | $0.134 (4K: $0.24) |
Nano Banana 2 (Gemini 3.1 Flash Image) launched on February 26, 2026, and Google describes it as combining "the advanced world knowledge, quality and reasoning you love in Nano Banana Pro, at lightning-fast speed." Nano Banana 2 Lite followed on June 30, 2026, generating "an image in as little as four seconds" and pitched by Google at "rapid-firing ideas, A/B testing ad variations, or powering social apps for millions of users." Nano Banana Pro is Gemini 3 Pro Image, the quality ceiling, and the one Google points at posters, diagrams, and multilingual text.
Two limits worth knowing before you write anything: reference-image caps differ by model, running to 14 images on Lite but splitting into per-role budgets on Nano Banana 2 and Pro, and Nano Banana Pro maintains, in DeepMind's words, "consistency and resemblance of up to five characters." Every output carries a SynthID watermark.
The Nano Banana prompt formula
The model is an instruction follower rather than a keyword matcher, so the prompt wants a skeleton. Google's prompting guide gives the order as subject, action, location and context, composition, style. One thing is worth putting in front of it:
- Artifact type and goal. "A die-cut sticker of", "a product hero photograph of", "an infographic diagram showing". Naming the thing you are making first is the highest-leverage word in the prompt.
- Subject and scene in a complete sentence.
- Composition: shot type, framing, angle.
- Style, light, and palette: one style anchor, one lighting anchor, one texture anchor. Not five of each.
- Constraints and references last: background, aspect ratio, what must stay unchanged.
Three rules apply on top of the skeleton, and breaking any one of them causes a recognizable failure.
Describe the desired state positively. Google's guide says it plainly: use "empty street" instead of "no cars". Negatives are unreliable across every image model we run. Nano Banana honors a short "no text, no faces" better than Grok or GPT Image do, but a positively-described scene still beats a list of prohibitions. If you want to understand why negative prompting works so poorly on modern models, we wrote that up in negative prompts.
Say what the background is. Left unspecified, Nano Banana invents framing, borders, and drop shadows. "Plain white background, edge to edge" is a phrase worth memorizing.
One subject, one scene, unless you are on Pro doing deliberate multi-element layout work. Chaotic renders are almost always two subjects fighting, not a missing adjective.
Nano Banana prompt examples that hold up
The best Nano Banana prompts all look the same on the page: two to four sentences, subject first, one lighting idea, one lens idea, an explicit background. Here are four we actually use, with the failure each one is designed to avoid.
Product mockup with legible label text
Text is where Nano Banana beats its competitors, and where most prompts throw the advantage away by not stating the exact string. Put the words in quotes and describe the typography.
A high-end commercial product photograph of a matte black cold brew
coffee can standing on a wet slate slab. Studio three-point softbox
lighting, shallow depth of field at f/2.0, 85mm lens. On the can, render
the label text "MIDNIGHT ROAST" in a heavy condensed sans-serif font,
with the smaller line "cold brew, no sugar" beneath it in a thin
uppercase sans-serif. Plain deep charcoal background, edge to edge.
Cinematic color grading with muted teal tones, photorealistic,
sharp focus.Google's guide lists the rules that make this work: "Enclose your desired words in quotes", describe the font ("a bold, white, sans-serif font"), and use its text-first trick when the copy is not settled yet, because "Gemini Image models work best if you first converse with it to generate the text concepts, and then ask for an image with that text". Deciding the wording before you prompt the visual is the whole game. "Come up with a slogan and put it on the can" renders worse than a confirmed literal string, every time.
Lifestyle photorealism without the beauty filter
Nano Banana 2 is the model we reach for when a person needs to look like a person: freckles, real skin texture, ordinary light. It beat Grok Imagine on exactly that axis in our own bakeoffs. Write it as a photograph, not as a character description.
A wide editorial photograph of a woman in her early thirties sitting at
a scratched wooden cafe table by a rain-streaked window, holding a
ceramic mug in both hands and looking out at the street. She has visible
freckles, untouched skin texture, and dark hair tied back loosely. Soft
overcast daylight from the window, warm interior lamps behind her, 35mm
lens, shallow depth of field, photorealistic, sharp focus, natural color
grading.The temptation is to add "no plastic skin, hyper detailed pores, not airbrushed". Resist it. That habit came from older Stable Diffusion-era models, and it makes Nano Banana renders worse: those words push the model toward the exaggerated texture look they are meant to prevent.
An infographic, which is the job Pro was built for
Multi-element layouts with multiple text labels are the one place where the Flash models flatten and Pro earns its price. Name the artifact type, then dictate every label as an exact string.
An infographic diagram titled "How A Pour Over Works" showing a glass
pour-over dripper on a kitchen scale, with four numbered callout labels
reading "1. Rinse the filter", "2. Bloom for 30s", "3. Pour in slow
spirals", and "4. Draw down". Clean editorial layout on a warm off-white
background, thin dark gray line art with a single amber accent color,
all label text in a bold geometric sans-serif font, generous white
space, no photographic elements.A sticker, which is the job Lite was built for
Flat, single-subject graphic work does not need Pro. Lite renders this in about four seconds, at a quarter of Pro's price.
A die-cut sticker of a smiling red fox wearing a tiny yellow raincoat,
bold black outlines, cel-shaded flat color, subtle drop shadow, plain
white background, edge to edge.If stickers are the actual job rather than the example, the sticker design tool wraps this prompt shape in a form, and how to make stickers to sell covers the print side.
Editing prompts are surgical, not descriptive
This is the single biggest difference between people who get good results from Nano Banana and people who do not. An edit prompt should name the change and explicitly protect everything else. Google's guide calls the technique semantic masking: "You can define a mask through text to edit a specific part of an image while leaving the rest untouched", and it tells you to "Be explicit about what to keep exactly the same."
In practice:
Change only the jacket to burgundy corduroy. Keep the face, pose,
lighting, background, and composition exactly the same.Re-describing the whole frame is not an edit. It is a regeneration with extra steps, and it degrades everything that was already right.
Reference images: one role each, and fewer than you think
Nano Banana is the strongest of the mainstream models at reference-driven generation, and Google's multimodal formula is worth internalizing: references, then a relationship instruction, then the new scenario. Their example is a good template. "Using the attached napkin sketch as the structure and the attached fabric sample as the texture, transform this into a high-fidelity 3D armchair render. Place it in a sun-drenched, minimalist living room."
The operating rule we hold to: give each reference exactly one role, object, character, or style, and say which. Stacking six references that each partially describe the subject produces a blend of all of them. Fewer, clearly-roled references beat more, every time. The published reference caps are a capability, not a target, and they are lower on Pro than on Lite, so a pipeline that assumes a flat 14 will break when it routes a brief to Pro.
Reference images also do the job that prompt adjectives cannot. Identity does not survive re-describing. If the same face has to appear across twenty images, no amount of "heart-shaped face, green eyes, 28 years old" will hold it, and the fix is structural rather than lexical. We cover why in consistent AI characters.
Nano Banana Pro prompts: when to change models, not words
Nano Banana Pro prompts are not written differently from Flash prompts. What changes is which briefs you send there. Switch to Pro when:
- The composition has many elements. Dashboards, exploded diagrams, multi-panel layouts, anything with more than two or three text strings. Flash flattens these; Pro holds them.
- The output is going to print. Pro is the only Nano Banana model where 2K and 4K are native rather than an upscale. Roughly, 2K is a good 6-inch print at 300 DPI and 4K is a good 12-inch one.
- Several characters have to stay recognizable across a set, up to DeepMind's stated five.
- Flash refused or flattened. Pro is less prone to sanitizing a brief.
Do not switch to Pro because a single-subject render came out bland. That is a prompt problem: add exactly one style anchor, one lighting anchor, and one texture anchor, and re-run on Flash. Paying double for the same weak prompt gets you a sharper version of the same weak image.
Google is also candid about where Pro still fails, listing "small faces, accurate spelling, and fine details" among its weak points, and warning that it can struggle with "grammar, spelling, cultural nuances, or idiomatic phrases". Proofread rendered text. It is right far more often than any previous model and it is not right always.
When Nano Banana is the wrong model
Some briefs will not improve with better prompting, because the model is the wrong one.
- Glam, pin-up, and body-forward character work. Grok Imagine has a looser filter and a stronger instinct for this, and it wants a short photographer-style prompt rather than declarative sentences. That dialect is in Grok Imagine prompts.
- Strictly brand-safe commercial layout. GPT Image 2 is the strongest of the three at compositional intent, the "three items in a row, evenly spaced" kind of instruction, and it has the tightest content filter, which is a feature when the output is client work. See GPT Image prompts.
- Anything that moves. Nano Banana is stills only. For video, the prompt craft is a different discipline entirely: motion, one action beat, one camera move. Start with Veo 3 prompts for the cinematic end and Seedance prompts for multi-reference products, environments, and stylized characters.
- A specific illustration tradition. Style vocabulary matters more than model choice here. Our AI art styles guide has the terms that actually move a render.
How to use Nano Banana
There are more doors than most people realize, and they differ in cost, limits, and whether you get the raw file.
- The Gemini app, on web, Android, and iOS. Pick the Flash-Lite, Flash, or Pro model and you get Nano Banana 2 Lite, Nano Banana 2, or Nano Banana Pro respectively. This is the fastest way to try a prompt.
- Google Search, through AI Mode and Lens, plus Google Ads, Google Workspace in Slides and Vids, and Flow.
- Google AI Studio, which is the free-ish playground for the API models.
- The Gemini API and Vertex AI, for anything programmatic.
- Third-party studios and routers that carry the models, including this one.
On watermarking, Google is explicit: "All media generated by Google's tools are embedded with our imperceptible SynthID digital watermark." Free and Pro plan users in the Gemini app also get a visible Gemini sparkle in the corner, which is removed for Ultra subscribers and for developers using the API. If you are producing client work, that visible mark is the detail that decides which door you use.
Nano Banana API access and what it costs
The Nano Banana API is the Gemini API, and the model ids are the ones in the table above. Google's pricing page lists no free tier for the image models, so API use is paid from the first image: about $0.0336 per 1K image on Lite, $0.067 on Nano Banana 2, and $0.134 on Nano Banana Pro, rising to $0.24 for 4K. Batch processing is half price. Those are the numbers to hold any wrapper service against. The full per-resolution price table, where the key comes from, and the call shape in code are in the Nano Banana API guide.
In Dream Pixel Forge the same models are priced in credits: Nano Banana 2 Lite is 2 credits an image, Nano Banana 2 is 4, and Nano Banana Pro is 8 at 1K, 10 at 2K, and 16 at 4K, on the same balance as Grok Imagine, GPT Image 2, FLUX Kontext Pro, and the video models. Run them in the freeform image generator, or see pricing for current credit packs.
If the images are being generated by code or by an agent rather than by you, the same models are reachable through an MCP server, a CLI, and an agent API, so a tool-calling assistant can generate, edit, and validate without a browser. That setup is in image generation MCP server.
Fixing a bad render: one change at a time
The most common mistake after a disappointing generation is rewriting the whole prompt. Match the symptom to a single fix instead, re-run, and only then change something else.
| Symptom | The single change |
|---|---|
| Label text came back as decorative squiggles | Put the exact string in quotes and name the font weight, rather than asking for a slogan |
| The strings are spelled right but land on the wrong callouts | Number them in the prompt in reading order; past three or four strings, move the brief to Pro |
| Flash flattened a multi-panel layout into one busy image | Name the artifact type in the first three words: infographic diagram showing, three-panel comparison chart of |
| An edit changed the face or the light along with the jacket | Name only the change, then list what must stay exactly the same. Do not re-describe the frame |
| Reference images blended into an average of themselves | Give each reference one role, object or character or style, and drop the rest |
| Skin went plastic after the anti-AI adjectives went in | Delete the negative list; keep two texture words plus one lighting and one lens clause |
| Invented border, matte, or drop shadow around the image | State the background explicitly: plain white background, edge to edge |
The bottom line
Nano Banana rewards writing like a director briefing a photographer: name the artifact, describe the scene in sentences, quote the exact text, state the background, and pick one anchor each for style, light, and texture. It punishes keyword piles, negative laundry lists, and edit prompts that re-describe the frame.
Get that shape right and the model choice becomes a cost decision rather than a quality one: Lite for volume, Nano Banana 2 for most real work, Pro for layout, text density, and print. And when the brief is glam, or strictly brand-safe, or moving, the honest answer is that a better Nano Banana prompt is not what you need.






