RCV Library
Back to training
Written companion · Keep this beside the training

AI UGC
Machine.

This is the clean version of the build: what goes into each section, what comes out, the exact prompts and what to check before spending credits on the next stage.

Start here

Run the cheap layers first.

V1 is the workflow built during the training. V2 is the latest version to use after you understand the build.

  1. Clone Workflow V1 to build directly beside the lessons, or V2 for the latest version.
  2. Replace the website URL and product image.
  3. Run the text nodes and check the prompts.
  4. Generate the avatar and product composite.
  5. Check all three starting frames for drift.
  6. Check the three script splits and list selectors.
  7. Run one A-roll clip and one B-roll clip.
  8. Only then release the remaining video nodes.
Recorded full run: roughly 733 Pletor credits, around $14 to $15 at the pricing shown during the training. Model and platform pricing changes.
The workflow map

Know what each branch receives.

Product URLAvatar prompt builderGPT Image 2Original avatar
Avatar + product imageStrict local editNano Banana 2Product-holding frame
Product-holding frameThree scene branchesNano Banana 2Three starting frames
URL + avatar contextScript promptSplit text + selectorsThree A-roll prompts
URL + product imageB-roll promptSplit text + selectorsTwo B-roll prompts
Module 01 companion · Workflow options

The tool is replaceable. The routing is the skill.

Pletor is used in the workshop because the image and URL inputs remain connected, the canvas is built around workflows and the same system can be deployed as a private app. It is more expensive than Higgsfield Canvas and some alternatives. Use whichever canvas lets you build, inspect and repair the same routing.

  1. List the inputs before adding nodes.
  2. List the outputs before choosing models.
  3. Build one branch at a time.
  4. Use the AI assistant to accelerate work, not to hide the logic.
Do not rebuild the course because a model or platform changes. Swap the final node and preserve the system around it.
Module 02 companion · Avatar generation + app example

Turn two simple inputs into a reusable avatar generator.

This module builds the first working branch of the machine. A website URL gives the model the brand, product and likely customer context. A text node turns that context into a candid UGC avatar prompt. GPT Image 2 then generates the original character. Once that branch works, the same inputs and output can be exposed as a simple private app for a client or team member.

What you build

  1. Add a text input for the product website URL.
  2. Add an image input for one or more clean product references.
  3. Group and label both nodes as Inputs.
  4. Connect the website URL to a Generate Text node and label it Avatar Prompter.
  5. Give that node permanent instructions so it researches the product, chooses a plausible customer and outputs one image prompt with no preamble.
  6. Connect the prompt output to GPT Image 2 and generate a 9:16 original avatar.
  7. Test the branch one node at a time before using the global Run button.
  8. Expose only the inputs and avatar output when you deploy the branch as an app.
The reusable pattern: “You will receive this input. Use it to understand this context. Perform this task. Output only this result.” Put the instructions inside the node once, then replace the URL whenever you run the workflow.

Decisions that prevent bad generations

Keep the product out of the original avatar

This first image establishes the person, room and phone-camera texture. Composite the real product afterward so the model does not invent packaging or place a random object in the frame.

Choose an easy product interaction

AI still struggles with complicated hinges, lids and mechanisms. Start with somebody holding the product while talking. Only ask them to use it when the interaction is visually simple and you have strong references.

Use a clean product image

Upload the clearest front-facing reference you have. Multiple references can help difficult packaging, but more images do not rescue unclear or conflicting product photography.

Use AI assistance carefully

Pletor can help build or repair nodes, but it can also produce confident nonsense. Learn the routing yourself so you can see what went wrong instead of repeatedly generating the whole workflow.

Recommended settings from the lesson

Prompt modelGPT-5.5
Image modelGPT Image 2
Format9:16 social story
Resolution2K
QualityMedium for testing

Why deploy the app version? Your client sees the URL and image inputs plus the finished avatar. The workflow stays behind the interface, so they cannot break the routing. You still retain access to the full canvas when something needs fixing.

Module 03 companion · Multi-scene compositing

Approve one product-holding frame, then branch from it.

This lesson connects the original avatar and product reference to Nano Banana Pro, creates the first product-holding frame and uses that approved result to generate three related locations. The finished branches keep the same person and product while varying the room, pose, expression and angle.

The build order

  1. Confirm the product URL was pasted as plain text and the avatar matches the likely customer.
  2. Connect the original avatar and product image to Nano Banana Pro.
  3. Use a short holding prompt first. Add a researched compositing node only when the product needs size or usage context.
  4. Generate and approve the product-holding starting frame.
  5. Duplicate the scene branch twice and connect the approved frame to all three.
  6. Give each branch a different room, pose, expression and camera angle.
  7. Run the full image workflow and check consistency before connecting video models.
Do not regenerate everything to fix one error. If the product scale is wrong, fix the composite. If the room is wrong, fix that scene branch. Preserve the approved layers.

Simple prompt or researched prompt?

Start simple: ask the avatar to hold the supplied product naturally while maintaining eye contact with the camera. If that works, keep it. A researched compositing node is useful when the product has unusual dimensions or needs a specific grip, but extra complexity is not automatically better.

Known trade-offs

Small packaging text

Nano Banana Pro preserves the original frame well, but tiny label text can warp when the product is far from camera. Bring it closer, strengthen the references or accept minor text variation in a moving shot.

Wrong customer from the URL

Paste without formatting. Hidden rich-text data can disrupt the website input and send the prompt branch in the wrong direction.

Tripods and phone overlays

Remove the exact failure with a direct negative instruction. Avoid broad rewrites that may damage the parts already working.

Testing cost

Run the text and image layers several times if this is for a client. Lock them before the expensive A-roll and B-roll nodes are released.

Composite modelNano Banana Pro
Format9:16 social story
Resolution2K
Primary referenceApproved product-holding frame
Module 04 companion · Script, split and A-roll

One performance, three safely routed clips.

Generate the complete 30-second story in one text node so the logic and accent stay connected. Use a caret as the delimiter, split once, then route List Selector 1, 2 and 3 into the matching starting frames. Never connect the unsorted list to all three video branches.

  1. Feed the URL and three frames into the script node.
  2. Ask for three sections short enough for roughly ten seconds each.
  3. Separate each complete prompt with ^.
  4. Use one Split Text node and three numbered selectors.
  5. Read and edit every prompt before paying for video.
  6. Run one A-roll branch first, then release the other two.
Reusable pattern: the same split-and-select routing can create ten hooks, ten static variations, multiple placements or any other repeated output.
Module 05 companion · Natural A-roll

Make it sound like a creator, not a miniature sales page.

Use a recognisable problem, an honest reaction and a simple change in behaviour. Expressions and product glances belong beside the thought that caused them. Avoid placing the same gesture at the end of every clip.

Gemini Omni FlashFast, cheap and naturally delivered. Can skip words when the script is too long.
Seedance 2More expensive and slower. Often more reliable with exact speech.
Best testUse the same frame and prompt, then compare speech, teeth, hands and delivery.

Keep the workflow model-agnostic. When one generator improves, replace the final A-roll node instead of rewriting the prompt and routing.

Module 06 companion · B-roll

Two useful product shots from the same split pattern.

The B-roll generator receives the product website and clean reference, then outputs two distinct prompts separated by a caret. Use one close detail or interaction shot and one wider day-in-the-life shot. Keep both visibly phone-shot rather than turning them into glossy product cinematography.

  1. Create the B-roll instruction in GPT-5.5 with web research enabled.
  2. Return only two prompts separated by ^.
  3. Split once and select item 1 and item 2.
  4. Generate five-second 9:16 clips while testing.
  5. Check product scale, logo, label and motion before expanding the branch.
Module 07 companion · Merge and captions

Automate the assembly without pretending editing has disappeared.

Merge the three A-roll clips in order. A clean cut is safer than forcing a transition between generated rooms. Add captions through ZapCap, CapCut or your preferred caption tool. B-roll timing still benefits from editorial judgement unless you build a more advanced timing branch.

  1. Merge A-roll 1, 2 and 3.
  2. Use no transition or a very subtle zoom.
  3. Keep captions inside mobile safe areas.
  4. Review every subtitle before delivery.
  5. Add storage or delivery nodes only after the core output is reliable.
Module 08 companion · Final test

Prove the system with a completely new product.

Replace the URL and product image, then run the text and image layers. Once the avatar, composite and three scenes are approved, release one A-roll and one B-roll test. The finished run produces the remaining clips, merged A-roll and captions.

New URLNew imageSame routingNew output

The completed workflow is the starting point. Add hook variations, placements, statics, delivery and new product branches when the core machine works reliably.

Prompt 01

Create the original avatar.

Give this text node the website URL and product image for context. The product helps it choose a believable customer, but the instructions keep it out of the generated frame.

You will receive a website URL and a product image for context.

Use the website and product image to understand the product and identify the kind of person most likely to use it. Choose a plausible age, gender, ethnicity, location, outfit and everyday environment for that customer.

Output one complete starting-frame prompt for a candid UGC avatar generation. The avatar should be looking towards the lens, ready to talk to their audience.

The final image prompt must include this visual direction:

Shot on a mediocre phone camera, not a polished AI render: visible pores and realistic skin texture, real iPhone colour science, natural smartphone HDR, slight overexposure with blown highlights near the window, faint lens flare at the edge of frame, minor compression artefacts, faint noise, slight autofocus imperfection, framing slightly off-centre, one-handed grab-shot feel, no beauty filter and no skin smoothing. Everything should feel accidental and mundane, believable as a real phone photo. Avoid anything that reads as rendered or AI-generated.

Critical: produce only a starting-frame avatar. Do not include the product, packaging, brand name or any product information in the image prompt. The product will be composited later.

Output the final image prompt only. No explanation, headings or preamble.

Original image model: GPT Image 2. Suggested format: 9:16, social story, 2K where available.

Prompt 02

Add the product without rebuilding the frame.

Start with the short prompt used in the recording. If the model does not understand the product's size or how it should be held, use the researched Compositing Prompter underneath it.

Have the avatar hold the supplied product naturally while maintaining eye contact with the camera. Keep the avatar's identity, skin, hair, clothing, background, lighting, camera quality and image texture unchanged.

Researched Compositing Prompter

Give this text node the original avatar image, clean product image and website URL. Its output feeds the image-editing node.

You will receive an image of a person, a product image and a website URL.

Use the website to understand the product, including its real size, proportions and how it would naturally be held.

Output one compositing prompt for Nano Banana Pro. Instruct it to composite the supplied product into the person's hand as if they are naturally holding it while talking to the camera and delivering a message to their audience. The person must remain facing the camera.

The prompt must preserve the original person's identity, face, skin, hair, clothing, background, camera, lighting, colour, quality and texture exactly. Do not beautify, relight, denoise or regenerate the rest of the frame.

Output the compositing prompt only. No explanation or preamble.

Strict local edit alternative

Use this when the character already holds an object and you want to replace only that object.

Perform a strictly local edit only. Replace only the object inside the person’s hands with the supplied product. Preserve the original image outside that area exactly: no changes to the face, skin, hair, clothing, background, lighting, colour grade, sharpness, noise, texture or resolution. Do not beautify, restyle, relight, denoise or regenerate the full frame. Preserve the supplied product’s exact proportions, logo, label text and colours.

Do not have the person using the product. Have them simply holding it as if delivering a message to the audience on camera.
Face unchangedLabel readableHands make senseTexture preserved
Prompt 03

Make three consistent scenes.

Frame 2 uses the approved product-holding frame. Frame 3 receives both approved frames so it knows which rooms have already been used.

Starting frame 2

You will receive one approved starting-frame image and a product image for reference.

Place the same avatar in a different room in the same house, looking towards the lens as if talking to the camera about the product. Use a slightly different natural expression, pose and camera angle. They may hold the product slightly differently, but it must remain natural.

Use a medium shot from a stationary front-facing phone camera. Do not include a tripod, phone, interface or recording overlay in the image.

Maintain the avatar's exact identity, face, hair, outfit, product, product label, lighting style, phone-camera quality, colour and texture.

Starting frame 3

You will receive two approved starting-frame images and a product image for reference.

Place the same avatar somewhere else in the same house that is clearly different from both supplied rooms. Keep them looking towards the lens as if talking to the camera about the product. Use another slightly different natural expression, pose, product hold and camera angle.

Use a medium shot from a stationary front-facing phone camera. Do not include a tripod, phone, interface or recording overlay in the image.

Maintain the avatar's exact identity, face, hair, outfit, product, product label, lighting style, phone-camera quality, colour and texture.

Change the room, pose, expression and angle. Keep everything that makes the person and footage recognisable.

Prompt 04

Write and split the performance.

You will receive a product website URL, three approved starting frames and the avatar context.

Read the website for accurate product information. Write one candid TikTok or Instagram-style UGC script that can be delivered across three separate clips. The combined spoken duration must be about 30 seconds. Give each section enough dialogue to feel complete, but keep it short enough for a natural delivery in a roughly 10-second clip.

Write from relatable real-world experience rather than listing product facts. Avoid a polished testimonial, generic brand language, a formal introduction or a hard sell. The speaker should sound as if they are working out what to say while talking to their audience.

Use each corresponding starting frame only as visual context for that section. Do not invent a pose, prop or action that is not possible from the supplied frame. Add one or two small expressions or gestures in the place they would naturally occur. A glance at the product should happen while the speaker is thinking about or referring to it, not automatically at the end.

Choose one plausible broad regional accent based on the character and context, using the nearest major city where appropriate. Do not use the accent to stereotype or change the writing. Use exactly the same accent in all three sections.

For each section, output two or three simple performance lines followed by the dialogue in parentheses. Finish every section with the chosen accent, no background music and no camera movement.

Separate the three final video prompts using this exact delimiter:
^

Output the three prompts only. No explanation, headings or preamble.
One split-text node. Three list selectors. Match selector 1 to frame 1, selector 2 to frame 2 and selector 3 to frame 3.
Prompt 05

Create two B-roll shots that still look like UGC.

You will receive:
1. A product website URL
2. A product reference image
3. An optional A-roll frame showing the visual environment and camera style

Read the website and create two distinct 9:16 B-roll video prompts for the product.

The two shots should be meaningfully different:
1. One close product-detail or interaction shot
2. One wider day-in-the-life product shot

Use the supplied product reference. Preserve the product’s proportions, colour, logo and label.

Make the footage feel like candid cellphone UGC recorded at home. Use natural phone-camera quality, auto-exposure, slightly blown highlights, minor compression, realistic motion and an imperfect handheld or casually placed camera.

Use the A-roll frame only as context for the environment, exposure, texture and general styling. Do not include the avatar in the B-roll unless the prompt explicitly asks for a hand or partial body interaction.

Avoid glossy studio lighting and polished commercial product cinematography.

Separate the two final prompts using this exact delimiter:
^

Output the two final prompts only. No explanation, headings or preamble.
Checked against the final training

Models and useful settings.

Avatar promptGPT-5.5 with web research
Original avatarGPT Image 2 · 9:16 · 2K
Composite and scenesNano Banana 2 or Nano Banana Pro
Script and B-roll promptsGPT-5.5 with web research
A-rollGemini Omni Flash or Seedance 2 · roughly 9 to 10 seconds · audio on
B-rollGemini Omni Flash or Seedance 2 · 5 seconds while testing
AssemblyMerge node, then ZapCap or your preferred caption tool
Fix the broken layer

Do not rerun the whole machine.

The edit changes the whole photograph

Use the strict local-edit prompt and Nano Banana 2. Lock face, skin, hair, clothing, background, light, colour, noise, texture and resolution.

The character changes between scenes

Use the product-holding frame as the reference. Change only room position, pose, expression and camera angle.

All video nodes receive the same script

Use one split-text node followed by separate list selectors. Do not duplicate the unsorted split output.

The accent changes

Choose it once inside the script node and include it inside all three final prompts.

The B-roll looks like a glossy advert

Ask for cellphone footage at home, auto-exposure, imperfect placement, minor compression and softly blown highlights.

Credits disappear while debugging

Run text, image, routing and one short video test in that order. Add a human review gate before the remaining expensive nodes.

You have the machine

Now make one run yours.

Clone it, replace the inputs and post your first result or question inside RCV support.