Create AI Influencer Ad Videos with Nano Banana and Kling in ComfyUI

You need a UGC-style promotional video. A creator holding your product, talking about it on camera, ready to post as a Reel or Short.
No filming. No actors. No editing.
Model photo and product image in. Talking influencer video out.
Run it now on Floyo!
Why This Workflow
Single-pass text-to-video struggles with precise product placement. The product ends up in the wrong hand, at the wrong angle, or hallucinated entirely.
This workflow solves it by splitting the job in two. Stage 1 places the product accurately in a photo. You approve it before any video runs. Stage 2 animates only what's already correctly placed. Clean output every time.
product placed naturally in the model's hands or scene
you approve the photo before committing to video generation
Kling animates with natural movement and lip-synced dialogue
vertical output ready for Reels, Shorts, and social ads
no filming, editing, or animation skills required
How It Works
Stage 1: Nano Banana (Promotional Photo)
Upload your model image and product image. Write a prompt describing how the product should appear in the scene. Nano Banana combines both into a realistic ad photo, the model holding or presenting the product naturally.
Run it multiple times. Each generation produces a different pose, framing, and placement. When you find one that looks right, save that image. That's what feeds Stage 2.
Stage 2: Kling (Talking Video)
Upload your approved photo into the final image input. Write the dialogue prompt, what you want the influencer to say. Kling animates the image with subtle body movement, natural expressions, and lip-synced speech. Output is a short vertical video ready to post.
Key Inputs
Model Image
The person whose look carries through both stages. Clean, well-lit shot with a simple background.
Works well with:
studio or neutral background portraits
clear upper body or full body shots
forward-facing or slight angle poses
Product Image
The item you want featured. Clean product shot on a neutral background for the most accurate placement.
Works well with:
single product on white or simple background
clearly visible product with readable details
skincare, supplements, fashion, tech accessories, food products
Stage 1 Prompt
Describe how the product should appear in the scene.
Examples:
"holding the serum bottle at chest height, warm studio lighting, lifestyle photography, natural smile""presenting the product toward camera, clean background, beauty editorial style, soft light""product placed on table beside model, natural daylight, café setting, candid lifestyle shot"
Stage 2 Dialogue Prompt
Write exactly what you want the influencer to say. Keep it short, 1 to 3 sentences works best for social content length.
Examples:
"says: I've been using this every morning and my skin has never looked better. Seriously, try it. Link in bio.""says: This is the one product I actually notice a difference from. Three weeks in and I'm not stopping.""says: Found this and I'm obsessed. The texture is incredible — you need to try it."
What This Is Great For
E-commerce and product marketing: Generate promotional videos for product pages, social ads, and email campaigns without a production shoot.
Social content at scale: Each Stage 1 run produces a different pose and framing. Each Stage 2 run can carry different dialogue. Test multiple angles and scripts from the same two input images.
Small brands without production budgets: Professional UGC content requires model fees, a videographer, and post-production. This workflow produces the same format from two images and a prompt.
Multilingual campaigns: Change the Stage 2 dialogue prompt for different languages. Same visual, different speech.
What to Watch Out For
Get Stage 1 right before moving to Stage 2. The video only animates what's already in the photo. A weak product placement in Stage 1 produces a weak video. Regenerate Stage 1 freely, it's fast.
Keep dialogue short. 1 to 3 sentences is the sweet spot. Longer scripts can affect lip-sync accuracy and produce clips that run too long for the format.
Input image quality is the ceiling. Blurry, low-resolution, or poorly lit source images reduce quality in both stages. Clean, well-lit inputs produce the best output.
Products with complex geometry or very small packaging may need more Stage 1 regenerations to place accurately.