Prompts
How to Create AI Influencers in Hebrew with Nano Banana Pro, Google Flow, and CapCut
7 min readAmichai Shekel
A practical guide to creating AI influencers in Hebrew using Nano Banana Pro, Google Flow, Veo 3.1, and CapCut. Includes prompts, steps, product integration, video continuity, and tips for working with Hebrew.
This guide breaks down one very specific process: how to create a digital character that looks like a real influencer, make her hold a product, talk about it in Hebrew, and then connect several video clips into a continuous video. This is not just about an AI influencer. It is simply the easiest way to explain what is happening on screen: a realistic character, a real or invented product, natural speech, movement, and continuity between scenes. Once you understand the workflow, you can take it in many directions: product videos, UGC content, business presenters, short explainers, internal company content, or a character that appears repeatedly and speaks in the same style.
The technology is not yet perfect in Hebrew. There will be glitches, strange words, and hand gestures that require additional trial and error. But the very fact that you can create a character, a product, Hebrew speech, and video continuity in a few simple steps already opens a massive door for content creation.
What We Will Build in This Guide
- A realistic AI influencer character
- An initial image suited for Reels, TikTok, or Stories in a 9:16 aspect ratio
- A continuation image where the character changes pose or moves closer to the camera
- An image where the character is holding a real product
- A video where the character speaks Hebrew and presents the product
- A video longer than 8 seconds by connecting multiple clips
- Basic editing in CapCut to export the final product
Why It Matters
The Tools We Will Use
- Gemini: Working with Nano Banana Pro and generating images
- Dedicated Gem: A tool that analyzes a reference image and generates a similar prompt
- Google Flow: Generating images, video, and connecting frames
- Veo 3.1: The video model that animates the character inside Flow
- CapCut: Quick editing and combining clips into a single video
- ElevenLabs / HeyGen: An advanced option for dubbing and lip-syncing if you want cleaner Hebrew
Gemini
Generating Images with Nano Banana Pro
The first step in the process. Inside Gemini, activate Create Image / Nano Banana Pro, enter a prompt or a reference image, and generate the initial character. You can continue talking to the model like in Photoshop: change poses, swap products, fix the height ratio, or add and remove elements.
Tip for agents: If Gemini refuses to edit an image or gets confused, open a new chat and upload the latest image. Do not bang your head against the wall.
Google Flow
Turning Images into Video with Veo 3.1
The central tool in this guide. Flow allows you to generate images, turn a first frame and last frame into a video, continue scenes, and create a sequence of multiple clips. It is still experimental, but it offers very powerful capabilities for anyone who wants to work with Google video models.
Tip for agents: Before every generation, make sure the ratio is Portrait 9:16 if your goal is a Reel, TikTok, or Story.
CapCut
Editing, Cutting, and Combining Clips
After generating several video clips in Flow, download them and combine them in CapCut. This is where you shorten weak parts, connect scenes, add subtitles, music, and sound, and export a video ready for social media.
Tip for agents: Flow can connect scenes, but CapCut gives you much more control over the final edit.
Step 1: Finding a Good Reference
Step 2: Turning a Reference into a Prompt
Basic Prompt for Creating a Realistic UGC Influencer
Create a photorealistic vertical 9:16 UGC-style selfie photo of a young Israeli woman in her early 20s, casually filming herself on a phone in a real apartment kitchen. She has natural skin texture, realistic facial features, slightly imperfect lighting, and a casual everyday outfit. The background should feel lived-in and local, not like a studio: small kitchen details, soft daylight, slight visual imperfections, realistic phone-camera look. She is holding an iced coffee and looking directly at the camera as if she is about to recommend something to her followers. No text, no captions, no logos, no overly polished fashion-shoot look.
The secret to realism: Do not go for the prettiest look. An overly perfect image looks like AI. Imperfect lighting, a slightly messy background, a local feel, and a shot that feels like a phone camera all help it pass the eye test much better.
Step 3: Creating a Continuation Frame
Prompt for Creating a Continuation Frame
Using the same woman, same kitchen, same outfit, same phone-camera style and same lighting, create a second frame where she leans closer toward the camera and gently reaches forward as if she is cleaning the phone lens. Keep her identity consistent. Keep the background consistent. Make it feel like the next moment in the same casual selfie video. Vertical 9:16. No text, no captions, no logos.
Step 4: Integrating a Product in Hand
Prompt for Integrating a Product into the Image
Using the influencer image as the main reference and the product image as the product reference, create a new vertical 9:16 image where the same woman naturally holds the product in her hand, close to the camera, as if she is about to talk about it in a casual UGC video. Keep her identity, outfit, lighting and kitchen background consistent. Make the product proportions realistic. Make the hand grip natural. Preserve the exact product packaging as much as possible. No text overlays, no captions, no extra logos.
Pay attention to proportions. If you upload a very long product image, the model sometimes changes the aspect ratio of the entire image. Always add to the prompt: vertical 9:16. If the image turns out too long or too wide, ask for a correction immediately.
Step 5: Moving to Google Flow
AI Master Club opens the door for you: a live meeting every week, recordings, and a community. 47 NIS per month, cancel anytime.
Prompt for Flow: Influencer Speaking Hebrew About a Product
The woman is filming a casual selfie-style UGC video in Hebrew. She speaks naturally and emotionally, like she is telling her followers about a new product trend she just discovered. She starts with the first frame and smoothly transitions into holding the product shown in the final frame. Her facial expressions should be excited and believable, with natural hand movement and casual body language. She occasionally looks at the product and then back at the camera. Keep the same identity, same outfit, same kitchen background and same vertical 9:16 phone-video style. Make the Hebrew speech sound as natural as possible.
Step 6: Creating a Video Longer Than 8 Seconds
Prompt for Scene Continuation
Continue the same selfie video from the previous final frame. The woman keeps speaking in Hebrew with the same voice, same personality and same excited UGC tone. She now puts the product aside, picks up her phone, and reacts as if she is showing something she found online. Keep the movement smooth and natural. Keep the same kitchen, same outfit, same lighting and same vertical 9:16 style. Avoid sudden camera jumps or changes in identity.
Step 7: Editing in CapCut
What to Do When the Hebrew Breaks
Common mistakes to avoid
- Creating a character that is too pretty and too perfect. It looks less reliable
- Using a studio background when a UGC feel is needed
- Not maintaining a 9:16 ratio in all images
- Changing too many things between frame and frame
- Giving the model too much long Hebrew text
- Expecting every first attempt to be perfect
- Trying to create a long video instead of combining short segments
Ideas for business uses
- UGC video for a physical product
- AI presenter showcasing a service
- Short tutorial videos for employees
- A character explaining a new feature in the app
- TikTok or Reels videos without a production day
- A/B testing of several characters and several messages
- Product videos for customers, without coordinating a shoot and without hiring actors
Frequently asked questions
Yes, but it is usually preferable through tools like HeyGen that are designed for cloning a speaking character. You can also work with your images in image models, but you will see every small nuance that the model missed. For people who do not know you, it will sometimes pass better than you think.
Yes. First you create a simulation of the product in a clean image, and then you use it as a reference image for integration in the character hand. If you have only a video of the product, you can stop a frame, extract an image, clean the background, and use it.
Not always. Kling and Veo 3.1 are both strong models. The advantage of Flow is that it sits inside the Google environment and connects images, video, and continuity comfortably. In certain cases Kling will give better motion. It is worth testing the same input in several models.
Yes, but not as a single segment. You create several short videos, use the last frame of each segment as the first frame of the next segment, and then combine in Flow or CapCut.
Image quality. The more accurate the first image, product image, and last frame are, the better the video turns out. A weak image will usually lead to a weak video.
This is the foundation. From here you can continue to more advanced capabilities: digital doubles in HeyGen, working with Freepik, Kling, Midjourney, Google AI Studio, Hixfield, and automations that generate dozens of videos at scale.


