A shot-by-shot production plan for rebuilding ChillFit's rank-06 creative (Library ID 1585698999664983, 31.44 s, 360×640@30) with exactly one variable changed: the woman in pink becomes Roxie. Same set, same motion, same Dutch captions, same audio, same edit. The point of the experiment is to measure how close current video models get to a real, high-spend UGC ad. The W1 pilot is generated and approved (motion transfer won); this v2 plan scales that locked recipe to the whole video.
A perfect replication is only meaningful if we change one thing. The controlled variable is the performer; every other element is either rebuilt to match the original exactly or reused verbatim from the source file.
| Element | Treatment |
|---|---|
| Performer | Swapped → Roxie (charsheet: content/_character/roxie-charsheet-sage-gymwear-bilateral.png), in her canonical dark-sage seamless unitard + white crew socks. This is the experiment. |
| Set | Rebuilt to match: pale wooden spindle-back chair, white paneled double closet doors with black lever handle at frame right, light grey carpet, soft daylight from front-left. |
| Camera | Rebuilt to match: locked tripod, portrait 9:16, full body centred, chair mid-frame, slight eye-level. Zero camera motion in every shot. |
| Motion / choreography | Rebuilt to match, shot by shot (specs in Acts A–B). |
| Edit | Identical cut points, to the frame (timeline below). |
| Captions & overlays | Rebuilt pixel-close: same Dutch copy, same font style, same colours, same muscle-diagram strip, same ChillFit watermark and icon bug. |
| Audio | Reused verbatim — demuxed from the source file (Dutch female VO 0.1–20.8 s + upbeat track). Keeps the comparison purely about video. |
| App-UI tail (20.7–31.4 s) | Two options — reuse verbatim vs. rebuild with a "before" Roxie in the challenge cards. Decision point in Cost & approvals. |
higgsfield generate cost quotes (MiniMax H3: 24 cr / 6 s,
40 cr / 10 s; Nano Banana 2: 2 cr / image).Measured with scene detection (threshold sweep at 0.10/0.25) plus frame sampling
at the boundaries, and an ElevenLabs transcription of the audio track (language: nld).
| Shot | In → Out | Dur | Content | Caption on screen |
|---|---|---|---|---|
| A | 0.00 → 4.20 | 4.20s | 3×3 grid — nine simultaneous chair-exercise clips | "Ik ben 65," → "ik doe deze workout elke dag." |
| W1 | 4.20 → 7.17 | 2.97s | Single: overhead claps, legs open–close wide | "gewoon 7 minuten per dag," |
| W2 | 7.17 → 10.73 | 3.56s | Single: cactus-arm pumps | "en na één week zul je een verschil opmerken," |
| W3 | 10.73 → 13.73 | 3.00s | Single: alternating arm raises | "na twee weken zullen anderen het opmerken," |
| W4 | 13.73 → 17.60 | 3.87s | Single: seated leg extension + cross-body arm swings | "en na drie weken," → "al je vrienden zullen vragen hoe je het hebt gedaan." |
| W5 | 17.60 → 20.73 | 3.13s | Single: wide double-leg raise → settles, smiles; app icon pops over chest | "het geheim?" → "download deze app" |
| C | 20.73 → ~28.3 | ~7.6s | 28-day chair-yoga challenge card UI, finger taps DAG-1 then DAG-11 | Card copy (Dutch), stats bar |
| D | ~28.3 → 31.44 | ~3.1s | End card: 3 phone mockups, ChillFit logo, store badges | "Planner voor thuistraining" |
"Ik ben 65. Ik doe deze work-out elke dag, gewoon 7 minuten per dag en na één week zul je een verschil opmerken. Na twee weken zullen anderen het opmerken en na drie weken zullen al je vrienden vragen hoe je het hebt gedaan. Het geheim? Download deze app."
Before assembly I'll re-run the transcription with --full to get word-level
timestamps, and snap every caption in/out to the spoken word — the original's captions track the
VO within ~0.1 s.
Nine simultaneous portrait panels, all the same set and framing, each a different chair move. Caption band sits between panel rows 1 and 2. This is the highest-effort shot to fake and the strongest test of character consistency: nine Roxies on screen at once.
| Panel | Move (as observed) | Sourced from |
|---|---|---|
| 1 · top-left | Arms clasp overhead in a diamond, then lower; knees open–close wide | Crop of clip G1 |
| 2 · top-mid | Both arms raised in cactus position, small pulses | Crop of W2 (offset start) |
| 3 · top-right | Alternating single-arm raises, opposite arm resting | Crop of W3 (offset start) |
| 4 · mid-left | Lateral arm sweeps crossing at the chest | Crop of clip G2 |
| 5 · centre | Knee hug — one foot lifts, hands pull the knee in | Crop of clip G3 |
| 6 · mid-right | Seated single-leg extension with side arm reach | Crop of W4 (offset start) |
| 7 · bottom-left | Arms open to a wide "W", chest opener pulses | Crop of clip G4 |
| 8 · bottom-mid | Overhead clap into torso side-bend | Crop of W1 (offset start) |
| 9 · bottom-right | Straight-leg raise with cross-body arm swing | Crop of W4 (mirrored, different offset) |
So the grid costs only four dedicated extra clips (G1–G4); five panels are centre-crops of the Act B masters at staggered start offsets (and one mirror) so no two panels read as the same footage. Each panel is a 1:1.78 portrait crop scaled to 120×213 in the 360×640 master — small enough to be forgiving of minor artifacts, which is why the grid clips are the right place to absorb the weaker takes.
One camera setup, five jump cuts. Every clip below is generated at 6 s and trimmed to its slot; generating a hair long gives trim room to pick the cleanest motion cycle. Shared prompt preamble for all five (and G1–G4):
Locked-off tripod shot, vertical 9:16, no camera movement. A fit woman in her
fifties with long wavy blonde hair, wearing a dark sage-green seamless workout
unitard and white crew socks, sits centred on a light wooden spindle-back chair
in front of white paneled closet doors with a black lever handle, light grey
carpet. Soft even daylight. She smiles naturally while exercising. Amateur
iPhone footage look, slight compression, true-to-life motion, no slow motion.
Parchment-textured page, header "28-DAAGSE STOELYOGA UITDAGING" in a heavy rounded serif (Cooper Black class), stats bar LENGTE 168CM · GEWICHT 100KG · DOEL 70KG (values red, goal green). Below, a 4×7 sketch calendar of seated exercises. A hand with red nail polish taps day 1 → the DAG-1 card zooms in: a live-action video of a heavier woman in a plum workout set doing seated knee lifts, three Dutch exercise lines, a cyan progress border animating around the card, four sketch thumbnails at the bottom. Back to the calendar, tap → DAG-11 card: the same woman visibly slimmer. The before/after arc lives inside these two card videos.
If rebuilt: the page is an HTML/CSS animation captured at 30 fps (parchment, calendar, zooms, cyan border), the tapping hand is a cut-out PNG animated in CSS, and the two inner card videos are generated — a softened, heavier "day-1 Roxie" variant and today's Roxie for day 11. That heavier variant is an identity edit of the charsheet (Nano Banana) used as the start frame.
Magenta-to-pink gradient, three fanned iPhone mockups of the ChillFit app (plan list, workout player at 00:08, calories screen), ChillFit icon + script wordmark, "Planner voor thuistraining", App Store / Google Play badges. Contains no human.
Recommended — reuse the source tail verbatim (20.73–31.44 s spliced from the original file). The experiment is about replicating human video; the UI tail is motion graphics we already know how to build and would dilute the credit budget. The remake then measures exactly the segment where generation is hard.
Alternative — full rebuild including both card videos with before/after Roxie (+48 cr generation, +~half a day of HTML animation work). Right choice later if this creative graduates from experiment to a runnable 50Queen ad, which would also mean swapping every ChillFit mark for 50Queen.
| Asset | Spec | Build method |
|---|---|---|
| Captions (video) | Chunky rounded sans (Baloo 2 / Fredoka SemiBold class), white fill, soft dark outline, centred ~62% down-frame; key tokens "65" and "7" in hot magenta (~#E93FD5) with white stroke. Grid captions invert: near-black fill, white stroke. | ASS subtitle track (libass): outline+shadow per style, word-timed from the --full transcript. |
| Muscle strip | Three semi-transparent white line-art female anatomy figures (front/back) across the lower quarter; the working muscle groups fill solid white per shot (mapping in Act B). | One SVG base drawn once (traced from a source frame's geometry), per-shot fills toggled, exported as transparent PNGs, overlay enable='between(t,…)'. |
| Watermark | "ChillFit" in a pink script face, ~25% opacity, two instances mid-frame, present on all live-action shots. | Text-on-transparent PNG, static overlay. |
| App icon bug | Rounded-square magenta→violet gradient icon, white script "ChillFit", pops over the chest at 19.0 s with a quick scale-in. | SVG → PNG, ffmpeg overlay with a 6-frame scale ease. |
| Grid frame | 3×3 layout, hairline white gutters (~2 px at 360w), panels flush to frame edges. | ffmpeg xstack over a white base. |
Every overlay is rebuilt at 720×1280 (2× the source) so the master looks clean, then the delivery encode downsamples to the source's exact 360×640 for honest A/B comparison.
Revised 2026-08-28. The original scope reused the Dutch track verbatim, but the direction has shifted toward a usable 50Queen asset: captions are now English and the ChillFit watermark is dropped, so a Dutch VO no longer fits. The new audio recipe:
qwen_audio_tts, preset "Roxie", f6448975…, English,
"cheerful upbeat morning coach"). Script = the English translation of the Dutch VO, one line
per caption: "I'm 65. I do this workout every day — just 7 minutes a day, and after one
week you'll notice a difference. After two weeks others will notice, and after three weeks all
your friends will ask how you did it. The secret? Download this app." Generated per-line
so each line can be nudged to fit its shot's slot.Purist fallback, still one command away: mux the untouched Dutch track for a strict A/B-comparison cut. Both cuts can share all video.
minimax_h3, 2K, ~4 cr/s: 24 cr per 6 s clip). Chosen per prior research: Seedance 2.0 reads too polished for UGC-style ads, and H3 was queued as the next candidate. Its param sheet exposes start_image/end_image, image_references and video_references — with one hard constraint: start-frames cannot be combined with reference media. That constraint forces the strategy split below.--video-references
= the exact source segment for that shot, --image-references = Roxie charsheet
+ the approved cand2 set still. The still is the pilot's one fix — S1 ran
charsheet-only and drifted (face thinner, flatter expression); S2 proved a pinned frame holds the
face. Adding cand2 to the reference set is the best-of-both, and it also anchors the chair, whose
hoop-back drifted in the W1 output. Two more prompt-level fixes ride along: "smiling warmly
throughout" (S1's expression went neutral) and an explicit cadence cue per shot ("claps once per
second" — S1 ran ~30% slow).higgsfield generate create minimax_h3 \
--prompt "Replicate the reference video exactly: same locked camera framing, same
rhythm, same choreography of <shot motion, with cadence> — performed by the woman
from the reference images: early fifties, long wavy blonde hair, dark sage-green
seamless workout unitard, white crew socks, smiling warmly throughout. Same room:
white paneled closet doors, black lever handle, light wooden spindle-back chair,
grey carpet. Amateur iPhone look. No text, no captions, no watermarks." \
--video-references <source segment upload> \
--image-references <charsheet> --image-references <cand2 still> \
--duration 6 --aspect-ratio 9:16
| Clip | Video reference (source cut) | Feeds | Credits |
|---|---|---|---|
| W1 (done) | 4.20 → 7.17 | Act B slot 1 + grid panel 8 · optional 24 cr redo with the three fixes | 0 (spent) |
| W2 | 7.17 → 10.73 | Act B slot 2 + grid panel 2 | 24 |
| W3 | 10.73 → 13.73 | Act B slot 3 + grid panel 3 | 24 |
| W4 | 13.73 → 17.60 | Act B slot 4 + grid panels 6, 9 (mirrored) | 24 |
| W5 | 17.60 → 20.73 | Act B slot 5 + the reveal | 24 |
| G1–G4 | Upscaled crops of grid panels 1, 4, 5, 7 (0.00 → 4.20) | Grid-only moves (overhead clasp, chest crosses, knee hug, W-openers) | 96 |
Grid-clip caveat: G1–G4's video references are 120×213 panel crops upscaled 4× — much weaker motion signal than the full-frame W refs. If transfer degrades, the fallback is prompted motion (S2-style) for those four: the panels render at 120×213 in the final grid, small enough to forgive tamer motion. Decide per clip at review, not up front.
xstack layout=3x3 on white, 4.20 s.remake_720.mp4 (master), remake_360.mp4 (source-matched), sxs_compare.mp4 (original | remake hstacked, shared audio).Spent so far: 56 cr (4 stills + the two-arm W1 pilot). Settled along the way: strategy = motion transfer, wardrobe = sage unitard, no ChillFit watermark, captions English. Remaining spend to a finished human section:
| Item | Model | Qty | Credits |
|---|---|---|---|
| Main clips W2–W5 (locked recipe) | MiniMax H3 · 6s | 4 | 96 |
| Grid clips G1–G4 | MiniMax H3 · 6s | 4 | 96 |
| Optional W1 redo (smile + cadence + chair fixes) | MiniMax H3 · 6s | 1 | 24 |
| VO — Roxie TTS preset, ~9 lines + retakes | Qwen Audio TTS | — | ~10 |
| Retry allowance (~25%) | — | — | ~55 |
| Total — full run, tail spliced from source | ~280 | ||
| Later, only on graduation to a runnable 50Queen ad: tail rebuild with 50Queen UI + before/after card videos | H3 + NB2 + HTML | — | +~130 |
Recommended — English VO in Roxie's own TTS voice over a demucs-separated music bed (rationale in Audio). Fallback cut with the original Dutch track costs one extra mux for a strict replication A/B.
Split. The end card (28.3→31.44s) is rebuilt for 50Queen
— done: three fanned device frames carrying real product screens (session player · Today · live
quiz), the icon + Instrument Serif wordmark, "Fifteen minutes, three days a week.", and the real
CTA pill "Get my plan now · 50queen.com" (web-only product; the repo's tests forbid store
badges). Rendered frame-by-frame in PIL at 720×1280·30fps mirroring the source's scale-in → fan
→ lockup timing; zero generation credits. Asset: endcard/endcard_50queen.mp4.
The challenge-card section (20.7→28.3s) stays spliced from the source until
graduation — it needs real card videos, not just graphics.