AI UGC
Machine.
This is the clean version of the build: what goes into each section, what comes out, the exact prompts and what to check before spending credits on the next stage.
Run the cheap layers first.
V1 is the workflow built during the training. V2 is the latest version to use after you understand the build.
- Clone Workflow V1 to build directly beside the lessons, or V2 for the latest version.
- Replace the website URL and product image.
- Run the text nodes and check the prompts.
- Generate the avatar and product composite.
- Check all three starting frames for drift.
- Check the three script splits and list selectors.
- Run one A-roll clip and one B-roll clip.
- Only then release the remaining video nodes.
Know what each branch receives.
The tool is replaceable. The routing is the skill.
Pletor is used in the workshop because the image and URL inputs remain connected, the canvas is built around workflows and the same system can be deployed as a private app. It is more expensive than Higgsfield Canvas and some alternatives. Use whichever canvas lets you build, inspect and repair the same routing.
- List the inputs before adding nodes.
- List the outputs before choosing models.
- Build one branch at a time.
- Use the AI assistant to accelerate work, not to hide the logic.
Turn two simple inputs into a reusable avatar generator.
This module builds the first working branch of the machine. A website URL gives the model the brand, product and likely customer context. A text node turns that context into a candid UGC avatar prompt. GPT Image 2 then generates the original character. Once that branch works, the same inputs and output can be exposed as a simple private app for a client or team member.
What you build
- Add a text input for the product website URL.
- Add an image input for one or more clean product references.
- Group and label both nodes as Inputs.
- Connect the website URL to a Generate Text node and label it Avatar Prompter.
- Give that node permanent instructions so it researches the product, chooses a plausible customer and outputs one image prompt with no preamble.
- Connect the prompt output to GPT Image 2 and generate a 9:16 original avatar.
- Test the branch one node at a time before using the global Run button.
- Expose only the inputs and avatar output when you deploy the branch as an app.
Decisions that prevent bad generations
Keep the product out of the original avatar
This first image establishes the person, room and phone-camera texture. Composite the real product afterward so the model does not invent packaging or place a random object in the frame.
Choose an easy product interaction
AI still struggles with complicated hinges, lids and mechanisms. Start with somebody holding the product while talking. Only ask them to use it when the interaction is visually simple and you have strong references.
Use a clean product image
Upload the clearest front-facing reference you have. Multiple references can help difficult packaging, but more images do not rescue unclear or conflicting product photography.
Use AI assistance carefully
Pletor can help build or repair nodes, but it can also produce confident nonsense. Learn the routing yourself so you can see what went wrong instead of repeatedly generating the whole workflow.
Recommended settings from the lesson
Why deploy the app version? Your client sees the URL and image inputs plus the finished avatar. The workflow stays behind the interface, so they cannot break the routing. You still retain access to the full canvas when something needs fixing.
Approve one product-holding frame, then branch from it.
This lesson connects the original avatar and product reference to Nano Banana Pro, creates the first product-holding frame and uses that approved result to generate three related locations. The finished branches keep the same person and product while varying the room, pose, expression and angle.
The build order
- Confirm the product URL was pasted as plain text and the avatar matches the likely customer.
- Connect the original avatar and product image to Nano Banana Pro.
- Use a short holding prompt first. Add a researched compositing node only when the product needs size or usage context.
- Generate and approve the product-holding starting frame.
- Duplicate the scene branch twice and connect the approved frame to all three.
- Give each branch a different room, pose, expression and camera angle.
- Run the full image workflow and check consistency before connecting video models.
Simple prompt or researched prompt?
Start simple: ask the avatar to hold the supplied product naturally while maintaining eye contact with the camera. If that works, keep it. A researched compositing node is useful when the product has unusual dimensions or needs a specific grip, but extra complexity is not automatically better.
Known trade-offs
Small packaging text
Nano Banana Pro preserves the original frame well, but tiny label text can warp when the product is far from camera. Bring it closer, strengthen the references or accept minor text variation in a moving shot.
Wrong customer from the URL
Paste without formatting. Hidden rich-text data can disrupt the website input and send the prompt branch in the wrong direction.
Tripods and phone overlays
Remove the exact failure with a direct negative instruction. Avoid broad rewrites that may damage the parts already working.
Testing cost
Run the text and image layers several times if this is for a client. Lock them before the expensive A-roll and B-roll nodes are released.
One performance, three safely routed clips.
Generate the complete 30-second story in one text node so the logic and accent stay connected. Use a caret as the delimiter, split once, then route List Selector 1, 2 and 3 into the matching starting frames. Never connect the unsorted list to all three video branches.
- Feed the URL and three frames into the script node.
- Ask for three sections short enough for roughly ten seconds each.
- Separate each complete prompt with ^.
- Use one Split Text node and three numbered selectors.
- Read and edit every prompt before paying for video.
- Run one A-roll branch first, then release the other two.
Make it sound like a creator, not a miniature sales page.
Use a recognisable problem, an honest reaction and a simple change in behaviour. Expressions and product glances belong beside the thought that caused them. Avoid placing the same gesture at the end of every clip.
Keep the workflow model-agnostic. When one generator improves, replace the final A-roll node instead of rewriting the prompt and routing.
Two useful product shots from the same split pattern.
The B-roll generator receives the product website and clean reference, then outputs two distinct prompts separated by a caret. Use one close detail or interaction shot and one wider day-in-the-life shot. Keep both visibly phone-shot rather than turning them into glossy product cinematography.
- Create the B-roll instruction in GPT-5.5 with web research enabled.
- Return only two prompts separated by ^.
- Split once and select item 1 and item 2.
- Generate five-second 9:16 clips while testing.
- Check product scale, logo, label and motion before expanding the branch.
Automate the assembly without pretending editing has disappeared.
Merge the three A-roll clips in order. A clean cut is safer than forcing a transition between generated rooms. Add captions through ZapCap, CapCut or your preferred caption tool. B-roll timing still benefits from editorial judgement unless you build a more advanced timing branch.
- Merge A-roll 1, 2 and 3.
- Use no transition or a very subtle zoom.
- Keep captions inside mobile safe areas.
- Review every subtitle before delivery.
- Add storage or delivery nodes only after the core output is reliable.
Prove the system with a completely new product.
Replace the URL and product image, then run the text and image layers. Once the avatar, composite and three scenes are approved, release one A-roll and one B-roll test. The finished run produces the remaining clips, merged A-roll and captions.
The completed workflow is the starting point. Add hook variations, placements, statics, delivery and new product branches when the core machine works reliably.
Create the original avatar.
Give this text node the website URL and product image for context. The product helps it choose a believable customer, but the instructions keep it out of the generated frame.
You will receive a website URL and a product image for context. Use the website and product image to understand the product and identify the kind of person most likely to use it. Choose a plausible age, gender, ethnicity, location, outfit and everyday environment for that customer. Output one complete starting-frame prompt for a candid UGC avatar generation. The avatar should be looking towards the lens, ready to talk to their audience. The final image prompt must include this visual direction: Shot on a mediocre phone camera, not a polished AI render: visible pores and realistic skin texture, real iPhone colour science, natural smartphone HDR, slight overexposure with blown highlights near the window, faint lens flare at the edge of frame, minor compression artefacts, faint noise, slight autofocus imperfection, framing slightly off-centre, one-handed grab-shot feel, no beauty filter and no skin smoothing. Everything should feel accidental and mundane, believable as a real phone photo. Avoid anything that reads as rendered or AI-generated. Critical: produce only a starting-frame avatar. Do not include the product, packaging, brand name or any product information in the image prompt. The product will be composited later. Output the final image prompt only. No explanation, headings or preamble.
Original image model: GPT Image 2. Suggested format: 9:16, social story, 2K where available.
Add the product without rebuilding the frame.
Start with the short prompt used in the recording. If the model does not understand the product's size or how it should be held, use the researched Compositing Prompter underneath it.
Have the avatar hold the supplied product naturally while maintaining eye contact with the camera. Keep the avatar's identity, skin, hair, clothing, background, lighting, camera quality and image texture unchanged.
Researched Compositing Prompter
Give this text node the original avatar image, clean product image and website URL. Its output feeds the image-editing node.
You will receive an image of a person, a product image and a website URL. Use the website to understand the product, including its real size, proportions and how it would naturally be held. Output one compositing prompt for Nano Banana Pro. Instruct it to composite the supplied product into the person's hand as if they are naturally holding it while talking to the camera and delivering a message to their audience. The person must remain facing the camera. The prompt must preserve the original person's identity, face, skin, hair, clothing, background, camera, lighting, colour, quality and texture exactly. Do not beautify, relight, denoise or regenerate the rest of the frame. Output the compositing prompt only. No explanation or preamble.
Strict local edit alternative
Use this when the character already holds an object and you want to replace only that object.
Perform a strictly local edit only. Replace only the object inside the person’s hands with the supplied product. Preserve the original image outside that area exactly: no changes to the face, skin, hair, clothing, background, lighting, colour grade, sharpness, noise, texture or resolution. Do not beautify, restyle, relight, denoise or regenerate the full frame. Preserve the supplied product’s exact proportions, logo, label text and colours. Do not have the person using the product. Have them simply holding it as if delivering a message to the audience on camera.
Make three consistent scenes.
Frame 2 uses the approved product-holding frame. Frame 3 receives both approved frames so it knows which rooms have already been used.
Starting frame 2
You will receive one approved starting-frame image and a product image for reference. Place the same avatar in a different room in the same house, looking towards the lens as if talking to the camera about the product. Use a slightly different natural expression, pose and camera angle. They may hold the product slightly differently, but it must remain natural. Use a medium shot from a stationary front-facing phone camera. Do not include a tripod, phone, interface or recording overlay in the image. Maintain the avatar's exact identity, face, hair, outfit, product, product label, lighting style, phone-camera quality, colour and texture.
Starting frame 3
You will receive two approved starting-frame images and a product image for reference. Place the same avatar somewhere else in the same house that is clearly different from both supplied rooms. Keep them looking towards the lens as if talking to the camera about the product. Use another slightly different natural expression, pose, product hold and camera angle. Use a medium shot from a stationary front-facing phone camera. Do not include a tripod, phone, interface or recording overlay in the image. Maintain the avatar's exact identity, face, hair, outfit, product, product label, lighting style, phone-camera quality, colour and texture.
Change the room, pose, expression and angle. Keep everything that makes the person and footage recognisable.
Write and split the performance.
You will receive a product website URL, three approved starting frames and the avatar context. Read the website for accurate product information. Write one candid TikTok or Instagram-style UGC script that can be delivered across three separate clips. The combined spoken duration must be about 30 seconds. Give each section enough dialogue to feel complete, but keep it short enough for a natural delivery in a roughly 10-second clip. Write from relatable real-world experience rather than listing product facts. Avoid a polished testimonial, generic brand language, a formal introduction or a hard sell. The speaker should sound as if they are working out what to say while talking to their audience. Use each corresponding starting frame only as visual context for that section. Do not invent a pose, prop or action that is not possible from the supplied frame. Add one or two small expressions or gestures in the place they would naturally occur. A glance at the product should happen while the speaker is thinking about or referring to it, not automatically at the end. Choose one plausible broad regional accent based on the character and context, using the nearest major city where appropriate. Do not use the accent to stereotype or change the writing. Use exactly the same accent in all three sections. For each section, output two or three simple performance lines followed by the dialogue in parentheses. Finish every section with the chosen accent, no background music and no camera movement. Separate the three final video prompts using this exact delimiter: ^ Output the three prompts only. No explanation, headings or preamble.
Create two B-roll shots that still look like UGC.
You will receive: 1. A product website URL 2. A product reference image 3. An optional A-roll frame showing the visual environment and camera style Read the website and create two distinct 9:16 B-roll video prompts for the product. The two shots should be meaningfully different: 1. One close product-detail or interaction shot 2. One wider day-in-the-life product shot Use the supplied product reference. Preserve the product’s proportions, colour, logo and label. Make the footage feel like candid cellphone UGC recorded at home. Use natural phone-camera quality, auto-exposure, slightly blown highlights, minor compression, realistic motion and an imperfect handheld or casually placed camera. Use the A-roll frame only as context for the environment, exposure, texture and general styling. Do not include the avatar in the B-roll unless the prompt explicitly asks for a hand or partial body interaction. Avoid glossy studio lighting and polished commercial product cinematography. Separate the two final prompts using this exact delimiter: ^ Output the two final prompts only. No explanation, headings or preamble.
Models and useful settings.
Do not rerun the whole machine.
The edit changes the whole photograph
Use the strict local-edit prompt and Nano Banana 2. Lock face, skin, hair, clothing, background, light, colour, noise, texture and resolution.
The character changes between scenes
Use the product-holding frame as the reference. Change only room position, pose, expression and camera angle.
All video nodes receive the same script
Use one split-text node followed by separate list selectors. Do not duplicate the unsorted split output.
The accent changes
Choose it once inside the script node and include it inside all three final prompts.
The B-roll looks like a glossy advert
Ask for cellphone footage at home, auto-exposure, imperfect placement, minor compression and softly blown highlights.
Credits disappear while debugging
Run text, image, routing and one short video test in that order. Add a human review gate before the remaining expensive nodes.
Now make one run yours.
Clone it, replace the inputs and post your first result or question inside RCV support.