Skip to content

AI

Veo 3.1 Prompt Guide: How to Write Cinematic AI Video Prompts

Subject, action, camera movement, lighting and sound — the structure behind prompts that actually look directed.

· 9 min read

Veo 3.1 Prompt Guide: How to Write Cinematic AI Video Prompts

Most people write video prompts the way they write image prompts: a pile of adjectives attached to a noun. That works when you only need one frame. Video is different. A generated clip has a beginning and an end, a camera that has to go somewhere, motion that has to obey some kind of physics, and — in Veo 3.1 — sound that arrives with the picture. Writing a good prompt is much closer to writing a shot list than to describing a photograph.

This guide covers the five-part prompting structure Google describes for Veo 3.1, the camera and motion language that actually changes the output, how to handle dialogue and ambience, and the mistakes that quietly ruin otherwise good prompts. Everything here is written to be used, not admired — and if you would rather assemble prompts with a form than from a blank page, our free Veo Prompt Builder does the structuring for you.

What Veo 3.1 actually gives you

Before the writing, the constraints. Veo 3.1 supports both text-to-video and image-to-video workflows. Google's Gemini API documentation describes eight-second generations at 720p, 1080p or 4K, with native audio in the relevant Veo 3.1 API model, in landscape 16:9 or portrait 9:16. The Gemini API workflow also exposes first-frame and last-frame control, video extension, and up to three reference images for keeping a character, object or style consistent across shots.

One caveat worth internalising: the available controls vary by product surface. A consumer app, a studio tool and the raw API do not necessarily expose the same knobs, and they change over time. Write your prompts against the model's behaviour, not against a feature list you saw in a screenshot.

The practical consequence of the eight-second window is the single most useful rule in this article: one shot, one idea. Eight seconds is a single beat of a scene. Ask for a chase, a reveal and an emotional close-up in the same prompt and you will get a muddled average of all three.

The Veo 3.1 prompt formula

Google's Veo 3.1 prompting guide presents a five-part formula: Cinematography + Subject + Action + Context + Style & Ambiance. Think of it as a useful checklist rather than a rigid order: covering each of these areas tends to produce a cleaner, more controllable result than leaving any of them out.

  1. Cinematography — shot size, angle, camera movement, lens and focus behaviour.
  2. Subject — who or what the shot is about, described concretely and only once.
  3. Action — the single thing that happens, with direction and speed.
  4. Context — location, time of day, weather, what fills the background.
  5. Style & Ambiance — film look, colour palette, lighting quality, mood.

For practical use, you can add audio and aspect-ratio instructions separately when the surface supports them. The discipline matters more than the sequence: each element gets its own clause, and nothing gets described twice.

Cinematography

Weak: "cinematic shot of a car". Strong: "low-angle tracking shot, 24mm lens, camera moving parallel to the car at speed". The second version tells the model where the camera is, what it is doing and how wide the world looks.

Subject

Describe the subject once, with two or three specific identifying details — clothing, age, material, condition. Repeating the description later in the prompt in slightly different words is a common cause of identity drift mid-clip.

Action

One verb, one direction. "She lifts the visor and steps forward" is fine. "She runs, stops, turns, laughs and walks away" is four shots pretending to be one.

Context

Context is where realism comes from. "A workshop" gives the model nothing to hold onto. "A concrete workshop at dusk, tools on a pegboard, rain on the skylight" gives it lighting, texture and depth for free.

Style & Ambiance

Name a look rather than stacking superlatives: anamorphic film look, documentary handheld, high-contrast noir, soft pastel daylight. Two or three style terms beat ten.

Camera language that actually helps

These are the terms that reliably change the output, and what each one is for:

  • Wide shot — establishes scale and location. Use when the environment is the point.
  • Medium shot — the default for a person doing something with their hands.
  • Close-up — emotion, texture, detail. Pairs badly with large movement.
  • Low angle — makes the subject dominant. High angle makes it vulnerable or schematic.
  • Tracking shot — camera travels alongside a moving subject.
  • Dolly / push-in — camera moves toward the subject. Slow push-ins read as tension or revelation.
  • Orbit — camera arcs around a mostly static subject. Specify a direction and roughly how far.
  • Crane / aerial — vertical or high-altitude movement, good for reveals.
  • Handheld — small imperfect movement. Add "subtle" unless you want visible shake.
  • Shallow depth of field — isolates the subject. Deep focus keeps the whole environment readable.

Pick one primary movement per clip. A slow push-in that also orbits, cranes and whip-pans in eight seconds is not a shot; it is a request for chaos.

Motion and physics

Video models infer physics from your description, so describe motion the way a camera operator would brief it. Three things are worth stating explicitly:

  • What moves, and in which direction — "steam drifts upward and left", "the drone banks right and climbs".
  • How fast — slow, steady, sudden. Speed words do real work.
  • What stays stable — naming what should not change is one of the cheapest ways to reduce warping and identity drift.

Environmental reactions sell the shot: dust lifting as a hatch opens, water rippling under rotor wash, fabric catching wind. And check for contradictions before you submit — "static locked-off tripod shot" plus "the camera sweeps across the room" is an instruction the model has to resolve by guessing.

On exclusions: rather than relying on a wall of "no X, no Y" wording, describe the scene you do want, positively and concretely. "Empty rain-slick street at dawn" works better than "no people, no cars, no signs".

Lighting and atmosphere

Lighting is the fastest route from "AI clip" to "shot". A short vocabulary covers most needs: natural daylight, overcast soft light, golden hour, practical lights (lamps and screens inside the frame), neon, moonlight, hard dramatic contrast, silhouette and rim light. Atmospherics — fog, haze, rain, dust, steam — give light something to travel through, which is why hazy scenes almost always look more filmic.

Name a key light direction when it matters: "warm amber key from the left, cool blue fill from the window behind". That single clause does more for the image than five aesthetic adjectives.

Dialogue, sound effects and ambience

Veo 3.1's relevant API model generates native audio, and prompting can specify dialogue, sound effects and ambient sound. Keep the three separate and label them:

  • Dialogue — put the exact words in quotes and attribute them: The engineer says, calmly, "It's holding." Keep it to one short line; keep dialogue short enough to fit comfortably within the clip.
  • Sound effects — name discrete events: SFX: a latch clicks, then a low hydraulic hiss.
  • Ambience — the continuous bed: Ambient: distant wind, faint fluorescent hum.

If you do not want speech, say so plainly — a line such as "no dialogue, ambient sound only" is clearer than silence on the subject. And unless you specifically want on-screen text, state "no on-screen text", because rendered captions and signage are a frequent unwanted extra.

Text-to-video prompt example

A complete prompt using the framework, written as a single eight-second shot:

Slow dolly push-in, low angle, 35mm lens, shallow depth of field.
A veteran flight engineer in an oil-stained jacket stands beneath the
open doors of a mountain hangar. She wipes her hands on a rag and
looks up as the hangar lights flicker on one bank at a time.
Late golden hour, dust suspended in the air, snow-capped ridge
visible outside. Anamorphic film look, warm amber key light against
cold blue shadows, gentle grain.
Ambient: low wind, distant metallic clank, hum of fluorescent tubes.
Dialogue: she says, quietly, "Alright. Let's see what you can do."
16:9, no on-screen text.

Every clause has a job: camera, subject, action, context, style, audio, format. No adjective appears twice.

Image-to-video prompt example

For image-to-video, the picture already answers "what does this look like". Your prompt should answer "what moves, and where does the camera go". Re-describing the supplied image in detail tends to fight the reference rather than support it.

Animate the supplied still. The camera performs a slow 15-degree
orbit to the right while pushing in slightly. The steam from the
mug drifts upward and bends left. The subject blinks once and turns
her head toward the window. Everything else in the frame stays
stable — no change to clothing, hairstyle or background layout.
Ambient: soft rain against glass, faint café chatter.
16:9, subtle handheld micro-movement only.

Note the explicit stability clause. Telling the model what must remain unchanged is the difference between an animated photograph and a slowly melting one.

9:16 prompts for Shorts and Reels

Vertical is not landscape rotated. The frame is tall and narrow, so composition advice changes:

  • State vertical 9:16 explicitly at the top of the prompt.
  • Favour medium and close-up framing; wide shots lose their subject in a vertical frame.
  • Keep important subjects and text away from the extreme edges so the shot remains adaptable across social platforms.
  • Prefer vertical motion (push-in, tilt, rise) over long horizontal pans, which run out of frame almost immediately.
  • Front-load the visual hook; the first second decides whether anyone sees the other seven.
Vertical 9:16. Medium close-up, eye level, static tripod with a
slow push-in. A watchmaker's hands lower a hairspring into a
movement, centred in the middle third of the frame with clean
headroom. Overhead practical lamp, dark workbench, shallow depth
of field so the background falls away.
SFX: metal tweezers ticking against brass, a single soft click.

More of our short-form work, and the clips these prompts feed into, live on the Hangar Works videos page.

Common Veo prompting mistakes

  • Vagueness. "Cinematic shot of a city" leaves every meaningful decision to the model.
  • Too many actions. Eight seconds holds one beat. Split the rest into separate shots, or use extension.
  • Contradictory camera directions. Static and sweeping, wide and close, locked-off and handheld — pick one.
  • Adjective stacking. "Epic, breathtaking, ultra-detailed, hyper-realistic, masterpiece" adds noise, not quality.
  • Ignoring sound. With native audio available, leaving the audio line blank wastes half the medium.
  • Redescribing the subject. Every restatement is an invitation for the model to change the face, the jacket or the paint colour.
  • Negative-only phrasing. Describe the scene you want; use exclusions sparingly and as a supplement.

A reusable Veo prompt template

Copy this, fill the braces, delete what you do not need:

[Cinematography] {shot size} + {angle} + {camera movement} + {lens / depth of field}
[Subject] {who or what, described once, consistently}
[Action] {one clear action, with direction and speed}
[Context] {location, time of day, weather, background detail}
[Style & Ambiance] {film look, colour palette, lighting, mood}
[Audio] {dialogue in quotes} / {sound effects} / {ambient bed}
[Format] {16:9 or 9:16}, {resolution}, 8-second single shot

Keep each line to one idea. If a line needs a comma-separated list longer than three items, it probably belongs in a different shot.

Build your next prompt in seconds

Free Hangar Works Veo Prompt Builder

Assemble cinematography, subject, action, context, style and audio into a clean, copyable prompt — no account, no API key, entirely in your browser.

Build your prompt with the free Veo Prompt Builder

For more on the systems behind these models, browse our AI coverage, including what happens if AI becomes smarter than us.

Frequently asked questions

How long can a Veo 3.1 clip be?
Google's Gemini API documentation describes eight-second generations, with video extension available in the API workflow to build longer sequences from successive shots. Treat each prompt as one eight-second beat.
Does Veo 3.1 generate sound?
The relevant Veo 3.1 API model generates native audio, and prompts can specify dialogue, sound effects and ambient sound. Available controls can vary by product or API surface, so check the surface you are using.
Should I use negative prompts?
Prefer describing the scene you want positively and concretely. A clear description of an empty rain-slick street works better than a list of things that should not appear; use exclusions only as a supplement.
What resolution and aspect ratios does Veo 3.1 support?
For the relevant current Gemini API Veo 3.1 model, documentation describes 720p, 1080p and 4K output in landscape 16:9 or portrait 9:16. Available controls can vary by product surface, so check the interface you are using. State the aspect ratio explicitly in your prompt, especially for vertical video.
How do image-to-video prompts differ from text-to-video prompts?
With image-to-video the still already defines the look, so the prompt should describe motion and camera movement rather than redescribing the image. Naming what must stay stable reduces drift and warping.

Related stories