Guides

Veo 3 Prompts: A Working Guide to Veo 3.1 (2026)

Veo 3 prompts that work: one camera move, one action, audio written out loud. Plus which Veo version you are running, real per-second costs, and where Veo refuses.

By Aditya Bawankule12 min readUpdated July 29, 2026

Good Veo 3 prompts do four things: name one camera move, give the subject one action, describe the sound out loud, and stop. Everything else in a prompt is decoration, and past a point decoration makes the clip worse, not better.

One thing to know up front: "Veo 3" today means Veo 3.1. Google shut down the original veo-3.0-generate-001 and veo-3.0-fast-generate-001 models on June 30, 2026, and what runs behind the Gemini app, Flow, Vertex AI, and every reseller today is Veo 3.1, Veo 3.1 Fast, or Veo 3.1 Lite. We ship Veo 3.1 in production alongside two other video engines, and this guide covers what that experience teaches: how to prompt it, what it refuses, and what a clip actually costs.

Veo 3 vs Veo 3.1: which model you are running

The version history matters because a lot of prompt advice was written for the original 2025 model, and some of it no longer applies.

  • Veo 3 shipped in 2025 with the headline feature: native audio generated in the same pass as the picture. The API models were deprecated and shut down on June 30, 2026.
  • Veo 3.1 and Veo 3.1 Fast arrived on October 15, 2025, adding video extension and reference images (up to three) to guide appearance.
  • Veo 3.1 Lite landed on March 31, 2026 as the cheap iteration tier, and reached Vertex AI in early April.

Practically, the naming does not change how you write. Veo 3.1 follows the same prompt grammar as Veo 3 with better instruction-following, so a good Veo 3 prompt is a good Veo 3.1 prompt. What changed is what you can add around the prompt: reference images, first and last frames, and extending a clip you already like.

What is a Veo 3 prompt?

A Veo 3 prompt is a plain-language description of a single shot, including its sound. It is not a script, not a shot list, and not a config file. The API takes one string and returns one clip of 4, 6, or 8 seconds at 720p or 1080p, in 16:9 or 9:16, with a SynthID watermark baked in.

The mental model that produces the best results is that you are briefing a cinematographer for one take, not writing a story. Everything you say has to be executable inside eight seconds by one camera. "She walks in, sits down, opens her laptop, and starts typing" is four takes crammed into one, and the model will rush all four into a morph. "She sets the laptop down and exhales" is one take.

How to write good prompts for Veo 3

Google publishes two prompt frameworks, and they agree with each other. The DeepMind prompt guide lists seven elements: shot framing and motion, style, lighting, character description, location, action, and dialogue. The Google Cloud guide for Veo 3.1 compresses those into a five-part formula: cinematography, subject, action, context, style and ambiance.

Neither tells you what to leave out. Here is the version that holds up, in order.

  1. Name the camera once. One move per clip: "slow push-in", "handheld tracking shot", "static locked-off shot". Two competing camera instructions make the model over-direct, and you get a swooping mess. If you do not care about the camera, say "static shot" rather than saying nothing.
  2. Describe the subject specifically, once. "A woman in her late twenties with wavy brown hair and light freckles, in a gray wool coat" gives the model something to hold. "A woman" gives it a lottery ticket. Repeating the description later in the prompt does not reinforce it, it just competes.
  3. Give exactly one action. One beat per clip. Eight seconds holds a glance, a turn, a pour, a reach. It does not hold a sequence.
  4. Set the location and light in one sentence. "A rain-streaked diner window at night, sodium streetlight through the glass." Sensory, concrete, singular.
  5. Write the sound explicitly. This is the step most people skip, and it is the one Veo is actually differentiated on. Covered below.
  6. Add style last, if at all. "Shot on 35mm film, shallow depth of field" or "claymation" or "VHS texture". Put it at the end so it colors the shot rather than replacing it.

Length is not the variable people think it is. Adding detail helps until it starts contradicting itself, and a 200-word prompt full of adjectives is mostly self-contradiction. A tight 40-word prompt with one camera move, one action, and named audio beats a paragraph almost every time.

Negative prompts: describe the absence, do not name it

Veo handles negation badly, the same way most generative models do. Google's own guidance is to phrase exclusions as descriptions: instead of "no buildings", write "a desolate landscape with no buildings or roads" as part of the scene description. In practice the safest habit is to describe the thing you want and never mention the thing you do not. Say "an empty stretch of two-lane highway at dawn", not "a highway with no cars".

Prompting audio and dialogue, the part Veo is actually for

Veo generates speech, sound effects, and ambience in the same pass as the video. If your prompt says nothing about sound, you get whatever the model infers, which is usually generic ambience and, worse, sometimes invented background chatter or music you did not want. Naming the audio is free control.

Three rules cover almost everything:

  • Dialogue is an exact quoted line plus a tone, never a topic. Write she says warmly: "this is my morning routine". Do not write "she talks about her morning routine", which returns mouth movement attached to invented words. Google's guide uses the same pattern: A woman says, "We have to leave now."
  • Label sound effects and ambience separately. The Veo 3.1 guide uses explicit labels: SFX: thunder cracks in the distance and Ambient noise: the quiet hum of a starship bridge. Labeling works better than burying the sound in the scene sentence.
  • Say what you do not want to hear, once. "No music" is the one negative worth including, because unrequested soundtrack is the most common audio failure. "Soft vinyl crackle, no music" is a complete audio brief.

One cost note that follows directly from this: audio is not free. On the rates we pay for Veo 3.1, a silent clip costs half what the same clip costs with audio. If the clip is going to get a scripted voiceover in the edit anyway, generate it silent and halve the bill. In Dream Pixel Forge that is a toggle, and the credit counter updates before you spend.

Veo 3 prompt examples you can copy

These are structured the way described above: one camera move, one action, explicit audio. Swap the nouns, keep the shape.

  • Product hero (16:9). Slow push-in on a matte black espresso machine on a walnut counter, morning window light raking across the steel. A single curl of steam rises from the cup. Shot on 35mm, shallow depth of field. Ambient noise: low kitchen hum and the hiss of steam. No music.
  • Talking-head hook (9:16). Handheld phone selfie shot, a woman in her late twenties with dark shoulder-length hair in a sunlit kitchen, holding the phone at arm's length. She looks straight into the lens and says brightly: "okay, three weeks in and I have thoughts". Ambient noise: soft room tone, distant traffic. No music.
  • Atmosphere b-roll (16:9). Static locked-off shot of a rain-streaked diner window at night, sodium streetlight refracting through the water. Rain intensifies across the shot. SFX: rain on glass, a car passing on wet asphalt. No dialogue, no music.
  • Character moment (9:16). Slow tracking shot following a man in his sixties in a canvas work jacket walking a gravel path between vineyard rows at golden hour. He stops and runs a hand over the leaves. Ambient noise: gravel underfoot, cicadas, light wind. No music.
  • Stylized (16:9). Claymation, static wide shot, a small round robot on a workbench cluttered with tools. It tilts its head and its single eye blinks twice. SFX: soft servo whirr, a tiny metallic click. Warm tungsten lighting.

Notice what is absent: no shot numbers, no "cinematic masterpiece 8K", no stacked style tags. Quality adjectives without a referent are noise. "Shot on 35mm" tells the model something; "cinematic" tells it nothing it did not already assume.

Timestamp prompting for a multi-beat clip

If you genuinely need two beats in eight seconds, Google documents a timestamp syntax for it: [00:00-00:04] for the first shot, [00:04-00:08] for the second. It works, and it is the one exception to the one-action rule. It is also the fastest way to get a rushed clip if you ask for more than two beats, so treat it as a two-shot tool, not a storyboard engine.

Do you need a Veo 3 JSON prompt?

No. JSON prompting is a community convention, not an input format. The Veo API takes a prompt string; a JSON blob you paste into that field is just a string with braces in it. It is not parsed as structure and there is no documented schema.

What JSON does do, and why people swear by it, is force you to fill every slot: camera, subject, action, setting, lighting, audio. That discipline is real and it is exactly what the checklist above gives you without the syntax tax. If a JSON template helps you remember the audio field, use it. Just do not believe it is unlocking a hidden parser, and do not let it push you into a 300-key prompt where half the fields contradict each other.

Image to video: prompt the motion, not the look

The highest-leverage Veo workflow is not text to video at all. Generate the still first, get the composition and branding right for a couple of credits, then animate the exact frame you approved. The full argument for that order is in our guide to image to video AI, but the prompting rule is short: when a source frame exists, the look is already decided, so spend the entire prompt on movement.

A good image-to-video prompt for Veo is one line: "slow push-in, she looks up from the book and smiles, ambient cafe noise, no music". Re-describing the scene fights the source frame and produces drift. You can leave the prompt empty and let the model pick a default motion, but a one-line motion hint beats a blank prompt reliably enough that it is worth always writing one.

Reference images work differently on Veo than on its competitors, and this trips people up. Veo 3.1 accepts up to three reference images and has no token syntax: you do not write [Image 1] or <IMAGE_1>, you just describe the scene plainly and the references guide appearance implicitly. Seedance uses [Image N] tags to direct who does what across up to nine references; Grok animates a single source image with no tagging syntax at all, and its video model takes references separately rather than by name. Copying one model's syntax into another's prompt box is a common and invisible failure: the tokens are simply read as text.

Try it now

Free to try, no account needed

Reference images

Select up to 8 images to guide the result.

0/8
Example output for Freeform Image Generator
Example output

Your generated image will replace this example.

Where Veo refuses: the photoreal-person wall

Feed Veo an image-to-video request whose source frame contains a photoreal human, and it does not error usefully. It runs, bills nothing, and returns an empty result: the responsible-AI filter has caught it. Seedance is at least explicit and rejects the request with "input image may contain real person", whether the person arrives as the source frame or as a reference. We measured this across both engines on July 20, 2026. It is not a prompt problem, and no amount of rewording gets around it.

So if your work is photoreal people in motion, do not architect around Veo as the only engine. In Dream Pixel Forge the single-image animate path for a photoreal persona routes to Grok Imagine Video, which renders where the other two refuse.

For the persona side of that workflow, where identity has to hold across dozens of clips, our guide to AI UGC video covers the character-sheet approach that keeps the same face across a campaign.

How to use Veo 3: where you can actually run it

Google Veo 3 prompts are portable, but access is not. There are four practical routes, and they differ mostly in whether you are paying per month or per second.

  • Gemini app. The consumer path. With a personal account you need a Google AI Pro or Ultra subscription; Google's own help page is explicit that video generation is not on the free tier, and both paid tiers cap how many videos you get.
  • Flow. Google's filmmaking front-end, built around Veo, with scene extension and frame-to-frame transitions. Best if you are cutting a sequence rather than generating single clips.
  • Gemini API and Vertex AI. Pay per second of output. Published rates are 0.40 dollars per second for Veo 3.1 at 720p and 1080p, 0.10 to 0.12 dollars per second for Fast, and 0.05 to 0.08 dollars per second for Lite. There is no free tier on any of them. An 8-second full-quality clip is therefore about 3.20 dollars, which is the number to keep in your head when you are deciding whether to iterate at full quality.
  • Studios and resellers. Anything that puts Veo behind a credit system alongside other models. This is where you land if you want one interface, one bill, and the ability to fall back to another engine when Veo refuses.

Whichever route you take, the specs are the same: 4, 6, or 8 seconds, 720p or 1080p, 16:9 or 9:16, native audio, SynthID watermarking on every output. Vertex AI added an upscaling pass to 1080p and 4K in April 2026, in preview, which works on any video and not just Veo output.

What Veo 3.1 costs inside Dream Pixel Forge

We run Veo 3.1 and Veo 3.1 Fast in the same studio as Seedance 2.0 and Grok Imagine Video, on one credit balance, so the honest comparison is per clip.

ModelCredits per second (audio)Credits per second (silent)8-second clip
Veo 3.12613208 credits
Veo 3.1 Fast10780 credits

The premium Veo models require the Pro or Studio plan. Starter unlocks the standard video engines, Seedance 2.0 and Grok Imagine Video, which is deliberately the cheaper place to test: find the hook and the framing on a 20 to 40 credit clip, then spend Veo credits once on the version you are actually shipping. Do the arithmetic and the reason is obvious: at 208 credits, an 8-second Veo 3.1 clip eats a fifth of the Pro plan's 1,000 monthly credits, so a month of Pro is four or five full-quality Veo clips. The pricing page has current plans and per-model costs.

Common Veo prompt mistakes

Five failures account for most bad clips, and four of them are prompt-side.

  • Chaining actions. Four verbs in eight seconds produces a morph. One beat per clip, then extend or cut.
  • Two camera moves. "Sweeping aerial that pushes in and orbits" is three instructions. Pick one and the shot stabilizes.
  • Silent on audio. Leaving sound undescribed gets you invented music and background chatter. Two words ("no music") fixes most of it.
  • Dialogue as a topic. "He explains the offer" gets you a mouth moving over nonsense. Quote the line.
  • Re-describing the scene on an image-to-video run. The frame is already decided. Prompt only the motion.

The short version of this Veo 3 prompt guide

One camera move, one action, a specific subject, one sentence of setting, and the sound written out loud. Style at the end, negatives phrased as descriptions, dialogue in quotes. Iterate on a still or a fast model, then spend the premium credits once. And remember that "Veo 3" today means Veo 3.1.

If you want to compare Veo's motion character against the alternatives before committing, our companion guides cover Seedance prompts for multi-reference products, environments, and stylized characters and Grok Imagine prompts for the engine that animates one photoreal persona still. On the stills side, Nano Banana prompts and GPT Image prompting cover the frames you will be feeding into all of this. Generating the still first is still the cheapest way to make a good clip.

Tools for this guide

Frequently asked questions

What is a Veo 3 prompt?

A Veo 3 prompt is a plain-language description of a single shot, including its sound. It is not a script, a shot list, or a config file. The API takes one string and returns one clip of 4, 6, or 8 seconds at 720p or 1080p, in 16:9 or 9:16, with a SynthID watermark applied. The useful mental model is briefing a cinematographer for one take rather than writing a story: everything in the prompt has to be executable inside eight seconds by one camera. Note that the model answering to "Veo 3" today is Veo 3.1, since Google shut down the original Veo 3 API models on June 30, 2026.

How to write good prompts for Veo 3?

Name one camera move, describe the subject specifically once, give exactly one action, set location and lighting in a sentence, write the audio explicitly, and put style last. Google's DeepMind guide lists seven elements (shot framing and motion, style, lighting, character description, location, action, dialogue) and the Google Cloud guide for Veo 3.1 compresses them into cinematography plus subject plus action plus context plus style and ambiance. What neither says is what to leave out: two camera moves, chained actions, and stacked quality adjectives all make clips worse. A tight 40-word prompt with named audio beats a 200-word paragraph.

What are some good video prompts?

Good video prompts pair one camera instruction with one action and an explicit audio line. For example: "Slow push-in on a matte black espresso machine on a walnut counter, morning window light raking across the steel. A single curl of steam rises from the cup. Shot on 35mm, shallow depth of field. Ambient noise: low kitchen hum and the hiss of steam. No music." Or for vertical: "Handheld phone selfie shot, a woman in her late twenties in a sunlit kitchen. She looks into the lens and says brightly: 'okay, three weeks in and I have thoughts.' Ambient noise: soft room tone. No music." Dialogue is always an exact quoted line plus a tone, never a topic.

Is Google's Veo 3 free?

No. Google's help page for video generation in the Gemini app states that with a personal account you need a Google AI Pro or Google AI Ultra subscription, and both tiers cap how many videos you can generate. On the developer side there is no free tier for any Veo model on the Gemini API: published rates are 0.40 dollars per second for Veo 3.1 at 720p and 1080p, 0.10 to 0.12 dollars per second for Veo 3.1 Fast, and 0.05 to 0.08 dollars per second for Veo 3.1 Lite. An 8-second full-quality clip is therefore about 3.20 dollars.

How to get access of Veo 3?

There are four practical routes. The Gemini app is the consumer path and needs a Google AI Pro or Ultra subscription. Flow is Google's filmmaking front-end built around Veo, with scene extension and frame-to-frame transitions. The Gemini API and Vertex AI bill per second of output, with no free tier. Finally, studios and resellers put Veo behind a credit system alongside other models, which is what you want if you need one interface, one bill, and a fallback engine for the cases where Veo refuses. Dream Pixel Forge runs Veo 3.1 and Veo 3.1 Fast this way, next to Seedance 2.0 and Grok Imagine Video.

What does a Veo 3.1 clip cost on Dream Pixel Forge?

Veo 3.1 costs 26 credits per second with audio and 13 per second silent, so a standard 8-second clip is 208 credits. Veo 3.1 Fast is 10 credits per second with audio and 7 silent, so 80 credits for 8 seconds. The premium Veo models require the Pro plan (29 dollars a month) or Studio (79 dollars a month); Starter (12 dollars a month) unlocks the standard video engines, Seedance 2.0 and Grok Imagine Video, at roughly 20 to 40 credits a clip. The intended workflow is to find the hook and framing on a cheap clip, then spend Veo credits once on the version you ship.