DEV Community

Sora2 Hub Team
Sora2 Hub Team

Posted on

How to Turn a Product Photo Into a Video Ad with AI (Step by Step)

A practical workflow for turning one packshot into a short, platform-ready video ad, from prepping the photo to picking a model, writing the motion prompt and checking the result.

Quick answer: Start with a sharp, well-lit photo of the real product. If you need a new setting, build a styled first frame with an image model such as Nano Banana Pro or GPT Image 2. Then animate that frame with an image-to-video model. Veo 3.1 suits single cinematic shots with native audio. Kling 3.0 suits multi-shot ads up to 15 seconds. Seedance 2.0 suits ads built from several references (extra images, a video for camera movement, a music track). Keep the camera move simple, generate a few variations, and check every frame for a warped logo or a changed product before you publish.


Why start from a photo instead of text

Text-to-video is fine for mood footage, but it invents the product. For an ad, the product on screen has to match what the customer receives. Image-to-video starts from your actual packshot, so shape, color and label are anchored from the first frame. All three models in this guide support image-to-video. Veo 3.1 and Kling 3.0 also let you set start and end frames.


Step 1: Prepare the product photo

The output is limited by the input. Before you generate anything:

  • Use the highest-resolution original you have. Don't use a compressed screenshot from your store page.
  • Show the product clearly. Keep hands, props and other packaging out of the way of the label and logo.
  • Light it evenly. Hard reflections on glossy packaging often turn into flicker once the image moves.
  • Pick the hero angle. The model will mostly keep the angle you give it, so choose the one you want viewers to remember.
  • Check your rights. Use photos you own or have licensed, especially if they show people.

A plain background is fine. You can build the scene in the next step.


Step 2 (optional): Build a styled first frame with an image model

A product on a white background works for a clean spin or zoom. For a lifestyle ad (a skincare bottle on a bathroom shelf, a sneaker on wet pavement), first make a still frame with the product placed in that scene.

  • Nano Banana Pro (Google's Gemini 3 Pro Image). Google says it can blend multiple reference images, keep up to six objects at high fidelity, output at 1K, 2K or 4K, and render legible text. That last point matters when the label has to survive.
  • GPT Image 2 (OpenAI). OpenAI describes it as a model for fast, high-quality generation and editing, with high-fidelity image inputs and inpainting. Inpainting is useful when you only want to change the background around an untouched product.

Whichever you use, zoom in and compare the label to the real product. Both vendors note that their models can still get small text and fine details wrong. A frame that looks right at thumbnail size can have a misspelled ingredient list.

Generate the frame in the ratio you plan to publish in, such as 9:16 for vertical feeds. That way you won't have to crop the product later.


Step 3: Pick the video model for the job

If your ad needs… Try Why (per the vendor's documentation)
One polished hero shot with sound Veo 3.1 Generates 4, 6 or 8 s clips with native audio at 720p, 1080p or 4K (1080p and 4K at 8 s only), 16:9 or 9:16. Accepts up to three reference images to preserve a product's appearance, plus first and last frames.
Several shots in one clip (close-up, then wide, then a pack shot) Kling 3.0 Multi-shot generation, 3–15 s durations, native audio, start frame plus element references. Kuaishou cites better text preservation for signage and logos, with e-commerce ads as an example.
Ads built from several assets (extra product angles, a reference video for the camera move, a music track) Seedance 2.0 Accepts text, images, video and audio together. ByteDance says one generation can use up to 9 images, 3 video clips and 3 audio clips, with multi-shot output up to 15 s.

If you aren't sure, run the same first frame through two models and compare them. Results depend heavily on the product. Reflective, transparent and heavily branded packaging all behave differently.


Step 4: Write a motion prompt, not a description

The image already shows what the product looks like. The prompt should describe what happens. A good structure:

  1. Camera: one move per shot (slow push-in, orbit, top-down tilt).
  2. Action: what moves in the scene (steam rising, water droplets running down the can, fabric shifting in a breeze).
  3. Product constraint: tell the model to keep the product unchanged.
  4. Light and mood: golden hour, soft studio light, neon reflections.
  5. Audio (for models with native audio): ambient sound, a sound effect, or a short spoken line.

Example (single shot, Veo 3.1 or Kling 3.0):

Slow push-in on the amber glass serum bottle on a marble bathroom shelf. Morning sunlight moves across the shelf and a few water droplets slide down the glass. The bottle, label and cap stay exactly as in the image. Soft, calm mood. Audio: quiet room tone and a light glass clink at the end.

Example (multi-shot, Kling 3.0):

Shot 1: macro close-up of the sneaker's sole stepping into a shallow puddle, splash in slow motion. Shot 2: wide shot of a runner on a wet city street at dusk, neon reflections. Shot 3: the sneaker on a black plinth, slow orbit, logo clearly visible. Keep the sneaker design and logo identical across all shots.

Avoid asking for the product to transform, open, or show features it doesn't have. Every extra instruction is another chance for the model to change the product.


Step 5: Generate variations and check every frame

Generate three or four versions of each shot, then review them at full size, frame by frame:

  • Logo and label: do the letters stay readable and stable, or do they melt or shift mid-clip?
  • Geometry: did the bottle gain a second cap, did the shoe change its stitching, did the strap disappear?
  • Color: compare the product color to the real item under neutral light.
  • Hands and physics: if a person handles the product, check fingers and contact points.
  • Audio: make sure any generated speech says nothing you can't legally claim.

If a clip is close but flawed, try a shorter duration or a simpler camera move before you rewrite the whole prompt. Count cost per usable clip, not per generation.


Step 6: Finish the ad in an editor

AI output is raw footage, not a finished ad. In your editor:

  • Add the offer and call to action as real text overlays. That's more reliable than asking the model to render them.
  • Add licensed music if you didn't use native audio, and set levels for sound-off viewing with captions.
  • Export each platform's ratio and check its current ad specs. They change, so check the platform's help pages rather than an old blog post.
  • Follow disclosure rules. Many ad platforms and marketplaces have policies on AI-generated or altered media. Check the ones you advertise on.

Most important: the ad must show the product as it really is. A video that makes a product look bigger, glossier or more capable than it is can mislead customers and lead to returns.


Doing it all in one place

This workflow often uses two or three models: an image model for the first frame and one or two video models to compare. If you'd rather not keep separate accounts and bills, Sora2 Hub is a credit-based studio that offers Nano Banana Pro, GPT Image 2, Veo 3.1, Kling 3.0, Seedance 2.0, Hailuo and Wan from one credit balance. You can make the first frame, animate it with two different models and compare the clips without switching tools.


Checklist

  1. High-resolution, well-lit photo of the real product, with rights cleared.
  2. Optional styled first frame, with the label checked at full zoom.
  3. Model chosen for the job: Veo 3.1 (hero shot), Kling 3.0 (multi-shot), Seedance 2.0 (multi-reference).
  4. Motion prompt: camera, action, "keep the product unchanged", light, audio.
  5. Three or four variations per shot, checked frame by frame.
  6. Text, CTA, music and captions added in the editor. Platform specs and disclosure rules checked.

Sources: Google Gemini API Veo 3.1 documentation; Google DeepMind Nano Banana Pro page and Gemini API image generation docs; OpenAI GPT-Image-2 model page and image generation guide; Kling VIDEO 3.0 model guide and Kling API capability map; ByteDance Seed "Seedance 2.0 Official Launch" post. Checked October 2026. Model capabilities change often, so confirm current limits before production.

Top comments (0)