Veo 3 Cinematic Product Ad Prompt
This prompt directs a 6-second cinematic product ad for Veo 3 as a three-shot sequence: an extreme macro detail opener, a slow orbital hero move, and a wide resolve with space for your tagline. Fill in the product, lighting, and palette — the shot-list structure is what makes it render like a real commercial instead of a slideshow.
A 6-second cinematic product spot for Veo 3: hero shot, detail macro, logo resolve — the commercial structure that converts.
Cinematic product advertisement, 6 seconds, single continuous shot sequence described as one prompt: SHOT 1 (0–2s): Extreme close-up macro of [PRODUCT DETAIL — e.g. water beading on a watch face], shallow depth of field, dramatic side light. SHOT 2 (2–4s): Slow orbital move around [PRODUCT] on [SURFACE], [LIGHTING — e.g. golden-hour window light], dust particles drifting. SHOT 3 (4–6s): Pull back to wide as [TAGLINE MOMENT — e.g. the product settles into its case with a soft click], clean background, space left for text overlay. Style: high-end commercial, 35mm film grain, [COLOR PALETTE], no readable text in frame, no human faces, photorealistic motion.
See what this prompt creates
Most AI video ads fail for the same reason most student films fail: they're one undifferentiated shot of a thing, slowly rotating, for eight seconds. Real commercials are edited — macro, hero, resolve — and this prompt bakes that edit into the generation by describing a three-shot sequence with timings. Veo 3 handles this structured approach far better than "make a cool product video," because each shot gives the model a concrete camera job instead of a vibe.
The load-bearing decision is no readable text in frame. Video models still garble text — your brand name will come out as alien typography. The prompt reserves clean space in shot 3 specifically so you can composite the tagline in editing, where typography belongs. Same for faces: "no human faces" keeps the render in the model's comfort zone (products, light, motion) and out of the uncanny valley.
Getting the best result
- Keep it to 6 seconds. That's the native strength zone — one idea, three shots. Longer generations wander; chain multiple 6s clips in editing instead.
- Describe motion per shot. "Slow orbital move," "dust particles drifting" — each shot needs its own motion direction. Static descriptions produce expensive stills.
- Pick one palette and commit. "Warm amber and deep teal" across all three shots is what makes separate generations cut together. Palette is the cheapest continuity tool you have.
- Generate the resolve with headroom. Shot 3's "space left for text overlay" only works if the product sits off-center — check the framing before you plan the title card.
3 variations to try
UGC-style ad: swap the whole aesthetic: "handheld phone footage, natural window light, authentic unboxing energy" — the TikTok-native version of the same 3-shot structure.
Luxury slow: "all three shots at half speed feel, deep shadows, single light source" for the premium fragrance-watch-and-wait pacing.
Loopable: "shot 3 composition matches shot 1, seamless loop" — the ambient background-loop variant for websites and displays.
How to use this prompt
- Copy it — hit the copy button above to grab the raw prompt text.
- Fill the brackets — replace anything in [BRACKETS] with your own details.
- Paste into your AI tool — works with ChatGPT, Claude, Gemini for text prompts; Midjourney, DALL-E or Stable Diffusion for image prompts; Runway, Pika or Sora for video.
- Iterate — tweak one bracket at a time until the output is exactly right.
Frequently asked questions
Why three shots instead of one continuous video?
Video models render single composed shots far better than multi-scene narratives. The 3-shot structure (macro → hero → resolve) mirrors real commercial editing and gives the model a concrete camera job per segment instead of a vague vibe.
Can I put my logo or tagline in the video?
Not in the generation — video models garble readable text. The prompt deliberately leaves clean space in shot 3; composite your logo and tagline in CapCut, Premiere, or DaVinci afterward, where typography stays crisp.
How long should each AI ad clip be?
Six seconds is the sweet spot for current models — long enough for a complete micro-story, short enough to stay coherent. Need 30 seconds? Generate five 6s shots and cut them together. That's the professional workflow.
Does this work in Kling or Pika too?
The shot-list language ports directly — camera moves, lighting, and timing are universal. Veo 3 currently renders the most cinematic product footage, but paste the same prompt into Kling or Pika and you'll get a usable variant of the same spot.