The pipeline
No two of these ads were made the same way, but every one of them moved through the same ten stages. The stages are the reusable part; the technique inside stage seven changed per beat, sometimes per shot.
-
Source the swipe. Pull an advertiser's top creatives from the
Facebook Ad Library sorted by
total_impressions, via thefb-ad-library-downloadskill, and archive them underads/_inspirations/page-<id>-top15/with a README recording rank, Library ID and start date. Two byte-identical files under different Library IDs is the strongest performance signal available — Meta publishes no spend, so duplication and longevity are the only proxies. -
Tear the source down to timecodes.
ffprobefor the shape,ffmpeg scdetfor the real cut points, contact sheets at different rates for different jobs (8fps across the hook, 1fps across the body), keyframe extraction, ElevenLabs transcription with word-level timestamps, and Gemini viadescribe-visualfor what the frames mean. Output is a beat table: every cut, every caption, every overlay, timestamped. - Publish a plan before spending anything. A beat-by-beat artifact mapping source beat → remake spec, with the prompts, per-beat credit cost and an adaptation table — published as a Claude Artifact and mirrored to Cloudflare. Approval happens on the plan, not on the footage.
- Rewrite the copy ladder onto our positioning. Structure and rhythm are kept; the claims are replaced. Never "bed pilates", never "28 days", never a weight-loss promise — always a 12-week at-home program and a capability claim.
-
Anchor the character. Roxie, from a wardrobe-specific bilateral
charsheet registered in
content/_character/roxie.json. The charsheet is an input to the image model; it is almost never handed to a video model directly. - Probe the one question that decides viability. Six credits, not a hundred and twenty. One alpha-blend of the strobe pair. One day-card. One three-second window. Nothing else gets funded until the probe comes back clean.
- Generate per beat, verify per beat. Still factory first (Nano Banana Pro, references = charsheet + the source frame for pose), then motion with whichever model suits that specific movement, then a same-cadence frame grid against the source before the beat is accepted.
- Assemble in code. ffmpeg and PIL. Captions, UI screens, chrome and end cards are drawn programmatically, and segments are pre-rendered so a re-cut costs nothing. Audio goes on last, placed from STT word timestamps so voice and captions cannot drift apart.
- Review against the source, side by side. A second artifact with the cut, the original, a synced comparison and every rejected iteration still playable. Eugene either approves or names the delta; the delta becomes the next version.
-
Ship into placements. 9:16 plus a 4:5 feed crop, wired through
asset_feed_specplacement customization, a UTM per creative, and a CPL kill rule against the current benchmark.
From the very first before/after build: an ad decomposes into generation problems and editing problems, and the editing problems are usually the ones that make it good. The strobe transition everybody liked was never generated — it was two clips and a cut list. What generation had to deliver was a matched pair of frames. Get the pair right and the effect is twenty lines of ffmpeg; get it wrong and no video model saves you.
Dancebit — "Walk 6,587 steps in 10 minutes"
- Source
- instagram.com/reels/DRPXnNyjMyA · 43.35s · one continuous locked-off take, zero cuts
- Ours
ads/workout_program/final.mp4· 70.27s · 1080×1920- Status
- Built — two known artifacts unfixed
- Teardown
.claude/ads-research/teardown-dancebit-DRPXnNyjMyA.md
The first remake, and the one that set the reading habit. The teardown's finding was that a funded brand with a real budget still shot in a kitchen in one take — all the polish lives in the overlay layer. So the remake didn't try to copy the footage. It copied the chrome.
Steps
- Keyframe the source at 16 points, build three contact sheets at different rates (8fps over the hook, 4fps over the intro, 1fps over the body), and write the beat sheet with the overlay state at every beat.
- Use our own filmed workout footage as the base rather than generating the performance —
workout-trimmed.mp4, 65.77s, eight movements. - Drive Kling 3.0 Motion Control from that footage with the Roxie charsheet as the character, converting real reference movement into our talent.
- Rebuild the overlay layer from scratch in headless Chrome (
render_overlays.py→ PNGs) and composite with ffmpeg — the title lockup, per-movement labels, the in-workout chrome, the end card. - Measure the chrome card rather than eyeballing it (244px), and set its bottom edge at y=1620 for a 300px margin that clears Meta Reels' ~269px UI zone.
- Lift the source's music bed, loop it with a 1.2s crossfade to cover 70.3s, fade under the end card — flagged as needing its own licence before any paid run.
What we changed
- Mulberry brand chrome throughout; no countdown; eight movements labelled.
- End card copy rewritten to remove a pricing claim: "Get your customized fitness plan on us / 12 MINUTES · NO EQUIPMENT · WOMEN 50+" — with the note that "on us" still reads as free and had better be true before the paywall.
- The source's before/after belly inset was not reproduced. The slide-in, timing and geometry were built and verified against the talent's pointing gesture, but the slot was left holding a stat card, with an app screenshot recommended as the fill. A fabricated shrinking-belly panel is manufactured evidence aimed at exactly the anxiety the product addresses, and it's also the format Meta bans outright.
@sue.giers — handstand up the G-Wagon
- Source
ads/sue-giers-7646356284442152224.mp4· 9.008s · 1080×1920- Ours
- Never shipped — three plan revisions, two probe generations
- Status
- Pipeline probe
- Plan
.claude/ads-research/genplan-sue-clone.md
Asked as "AI generate the same video as this". The value it delivered wasn't a creative — it was the model constraints that every later build ran on.
Steps
- 5fps contact sheet plus scene detection, which corrected the existing teardown: this is two shots with a hard cut at 7.70s, not one continuous take. That halved the difficulty — no handstand-into-seated transition ever needed generating.
- Write a frame-1 spec precise enough to be a prompt: camera at knee height tilted up, subject right-of-centre, SUV entering frame left, flat fluorescent light, no grade, mild sensor noise, "amateur social video — not cinematic".
- Spec the text overlay as a post step with in/out times — video models can't hold legible multi-line text across eight seconds.
- Three costed approaches so the decision was Eugene's: v1 Seedance with the source as a video reference (~150cr), v2 a single Kling Turbo prompt at general strokes (15cr), v3 Nano Banana charsheet → start and end frames → non-turbo Kling.
- Run v3b deliberately as a controlled experiment — keep the brand name, drop the video reference — to isolate which input had triggered the earlier unexplained moderation failure.
kling3_0_turbo accepts one reference and treats it as frame 1 — hand
it a multi-panel charsheet and it animates the sheet. The charsheet is an
input to the image model, never to the video model. Non-turbo Kling takes
a start and an end image, and pinning both ends is what stops arc drift.
Real-person prompts get ip_detected; vehicles get described by shape,
not brand.
Also recorded, and worth more than the pipeline: a faithful clone of this ad would have been the wrong creative. It has no product, no offer and no CTA, and it sells peak aspiration to women who may not see themselves in it. Run it to prove the pipeline; put money behind a different concept.
"Says who?" — the before/after strobe
- Sources
- @kangmiratips21 7606799992237444383 (the strobe mechanic) + @dianecaldwell_ 7620248665193499934 (the copy and the CTA)
- Ours
ads/before-after-strobe-cut/final.mp4· 14.79s · 1080×1920- Status
- Shipped and ran — became the account's benchmark creative at $0.85 CPL
- Spend
- ~122 cr per complete pass
The build that produced the house method. The plan opened by separating what had to be generated from what had to be edited, and then optimised almost entirely around one dependency: the strobe only works if the before frame and the after frame are composed to interleave.
Steps
- Generate the "after" character first, then derive the "before" from it. Never the reverse — models anchor far better on "make this person heavier and older" than the opposite, and the after is the shot the whole ad has to pay off.
- Generate the before scene frame (B1) from the before sheet.
- Generate the after scene frame (B2) with B1 passed as a second reference, instructed to use it for camera framing only — head size, head height, body centring, camera height, crop. This single step is what makes the strobe read as one person morphing rather than two clips fighting.
- Gate on a 50% alpha-blend of B1 and B2 before generating any video. Six credits answers both viability questions — can we get a convincing unflattering before, and will the pair interleave — at 5% of a full pass.
- Four MiniMax H3 clips from the start frames. Generate 6s, use 2–3s: video models drift toward flattering the longer they run, so the before clips are deliberately near-motionless and cut early.
- Pick the music first, analyse it (113.6 BPM, drop at 6.82s, bar lines at 8.93 / 11.05 / 13.16), then place every cut on its beat.
- Build the strobe as a frame pattern — 3 frames before, 2 frames after, ×5 — resolving onto the after exactly on the drop.
- Burn the captions in post: Avenir Black, white with a 3px black stroke, at 72% frame height.
"you're 56 / it's too late to change"dies into the strobe;says who?arrives on the other side.
The build diverged from the plan in one place worth keeping: instead of the
start-frame workflow, it used an 8-panel storyboard — alternating
pose and transition panels — fed to MiniMax H3 as image_references.
Two-reference identity hijack. Passing the charsheet together with a real frame from the source TikTok made Nano Banana take the room and the real creator's face from the TikTok frame, ignoring the sheet entirely; text instructions did not override it. Always two-step: generate the character in the target state on a plain background first, then pass that as reference 1 alongside the scene photo.
H3 drifts framing and grade. The after clip pulled wider than its start frame; the before clip warmed a cool marble bathroom to beige. Not prompt-fixable. Which means the strobe pair has to be re-aligned after generation, not just before.
"Says who?" — the late-40s variant
- Source
- Our own winner, re-aged. Spec in
dissection-a-says-who-age-variants.md - Ours
ads/before-after-40s/final-40s-FINAL.mp4+ a 4:5 feed crop- Status
- Launched — ran against the 56 creative in one 45–65 ad set, $50/day CBO
- Spend
- ~278 cr across three rebuilds
The first time we remade ourselves. Everything except the age variables was frozen: cut list, track, strobe pattern, caption style and timing, the three after contexts. Only the character, the ages in the copy and the hook line moved.
Steps
- Two candidate "after" characters generated as full sheets and put to Eugene as a pick — A (chestnut, 47) and B (Black American woman, 48) — with late-40s calibration written explicitly into the prompt (faint crow's feet, visible pores, first silver strands only).
- Derive both "before" sheets from the chosen afters. QA caught the predicted failure immediately: the +22lb barely registered on candidate A. Every image model is trained toward attractive, and it bites here first.
- Force the body delta at the scene-frame level, not the sheet level — the scene frames are what appear on screen, so the weight instruction goes there, described relative to clothing fit rather than as a number.
- Re-run the alpha-blend strobe gate on the new pair before any video spend. Passed.
- Four H3 clips (10s / 6s / 10s / 10s), start-image mode, motion prompts carried over from the original build.
- Byte-copy the audio bed out of the shipped winner so the 6.82s drop timing is inherited exactly and the cut list transplants without re-timing.
- Assemble with the same
build.sh; only the caption strings change.
Two direction changes, both on the after side
The v1 cut was rejected: the run and car clips read late-50s, not late-40s. The correction wasn't a smaller delta but a bigger one — push the after to late 30s, then again to early 30s (plump collagen, zero lines, sculpted jaw). The before kept its grey roots; the dull-roots-to-glossy-colour jump is part of the transformation. That preference is now standing policy: the after should read dramatically younger and prettier than the before.
What fixed the age drift in the rebuild: dual-anchoring every scene frame on the sheet plus an already-approved scene frame, moderate closed smiles instead of wide-open laughs, hair down, and explicit do-not-age guards in the clip prompts. All three v3 clips passed age QA first try.
Shipping decisions
- No age-split ad sets. One ad set, Women 45–65, two creatives inside, letting the auction pick the winner per impression.
- Entities recreated rather than edited whenever nothing had spend history.
- Placement customization via
asset_feed_spec: 9:16 to stories/reels, 4:5 to feeds. - A Leads campaign with
destination_type: UNDEFINEDmakes Ads Manager demand WhatsApp — set it toWEBSITE. Residual UI errors after that are stale drafts, invisible to the API; discard drafts clears them.
ChillFit rank06 — "I'm 55" chair workout
- Source
- Ad Library 1585698999664983 · ChillFit's #6 by impressions · "TENHO 65 ANOS"
- Ours
ads/50queen-im-55-chair-workout/final/…-720.mp4· 31.5s · audio on- Status
- Final — 28 Aug
- Spend
- ~374 cr + ~1.5k ElevenLabs characters
The first end-to-end remake where every frame is ours and every sound is ours — zero source assets survive in the cut. It also produced the shared 50Queen ending that two later ads splice in.
Steps
- Pull the advertiser's top 10 from the Ad Library by impressions, note the byte-identical duplicates (rank 1 and rank 5 were the same creative under two IDs), and pick the chair ad.
- Write a 16-beat storyboard — the source torn down beat by beat with the remake spec written against each one — then a shot-by-shot production plan with prompts and costs. Both published for approval.
- Race a pilot to settle the generation recipe, then apply it: motion transfer with the source segment as
--video-references, the charsheet and an approved still as--image-references. - Extract the source's muscle-diagram strips by median-diff matting and rebuild them as our own overlays.
- Audition the voice properly — Seed Audio candidates, then ElevenLabs, landing on "Ginger" at speed 0.90 / stability 35% / similarity 75%. Captions burned to the exact speech timings.
- Build the non-performance beats in code: the parchment calendar, the red-nail tapping hand, the progress borders, the end card with three real app screens in device frames.
- Pre-render each timeline segment so re-assembly and re-syncing never require regeneration — an audio-only change is a remix and a remux, nothing more.
What we changed
| ChillFit | 50Queen |
|---|---|
| "I'm 65" | "I'm 55", aimed at menopause |
| 7 minutes a day | 15 minutes a day |
| 28-day challenge | 12-week plan |
| App Store badges | Web CTA — 50Queen is web-only |
| Body transformation | Capability arc; stats bar reads AGE 55 · 15 MIN · GOAL: STRONGER |
That last row is a policy hedge, not a style choice: framing the before/after as capability and energy rather than weight keeps it clear of Meta's before/after rules. The music is generated, so there's no licensing exposure, and the voice is a licensed library voice.
Sweat rank03 — the overhead pilates ad
- Source
- Sweat / Kayla Itsines, rank 3 of the page's top 15 · 14.7s · running byte-identical under two ad IDs
- Ours
ads/50queen-overhead-guided-workouts/50queen_remake_v3.mp4· 20.73s- Status
- Final — audio must be swapped before a paid run
- Spend
- ~610 cr including the instructive failures
The hardest replication and the one that taught the most. One locked overhead camera, eleven exercise vignettes joined by jump cuts, a kinetic type ladder, and a phone raised to the lens. Three beat types, three completely different pipelines.
The segment recipe — exercise beats
- Find the hidden cuts. What looks like a continuous take is eleven ~1s pose vignettes under identical framing. They only appear at a low scene-detection threshold, and every boundary was verified frame-by-frame against a cut-verification grid. Round 2's "flying legs" were the model physically bridging poses that were never connected.
- Chop at those boundaries so every reference segment is genuinely continuous.
- Slow each segment to ~6.4s — H3 refuses durations under 6, and
duration 4fails silently (four wasted jobs before that was pinned down). - Run H3 with the slowed segment as the motion reference, at a matching duration, with the charsheet plus a per-beat pose still as identity references.
- Retime back to real speed with motion-blended resampling (
minterpolate mi_mode=blend), not plain frame decimation.
The phone bridge, and the end card
- Generate the start frame with a blank phone, animate with prompt-driven image-to-video (pose retargeting drops props, so motion references are for body-space movement only), then perspective-warp the real app splash onto the glass per frame. The tracker needed an ROI gate because the mat itself is a phone-shaped dark quad and hijacked the first pass.
- The end card is built, not generated — plum sampled from the app splash, the vector wordmark, and a phone bezel playing a real screen recording of the production workout player, captured off the dev harness by a headless-browser screenshot burst at ~12fps and blended to 30.
- The type ladder was reverse-engineered to a recipe: Arial Black at 168px, sheared 13.5° about the block's own centre, shadow at 38% alpha with a 7px blur — verified against the source frame side by side.
Reference length must match output length. Give H3 a 3s reference for a 6s output and it invents choreography to fill the gap.
A single shared anchor still over-anchors. Later beats collapse toward it. Per-beat pose stills fixed it.
Never drop image references to dodge moderation. With weak identity conditioning, H3 clones the reference video's subject — one beat came back as the Sweat athlete on a hot-pink mat, and shipped that way until it was caught. The fix is to swap in any modest approved still of the same character, plus an explicit "she is NOT the person in the reference video".
Retime gently, and only where it's earned. Compressing a 6.5s master into a 1.5s beat decimates 80% of the frames and reads as jitter. Pick a window of the slow master and compress ~2.5×. Acted and idle beats play at 1:1 with a window trim; only exercise-form beats get compressed at all. Every one of those pacing fixes cost zero credits — it's a re-render, not a regeneration.
Also settled here: Kling Motion Control cannot read bird's-eye footage. The pose tracker misreads the foreshortening and returns ghost duplicate arms and a standing profile glued to the mat.
Sweat rank01 — the $4 talking head
- Source
- Sweat's rank-1 creative · single-take UGC talking head pitching a price
- Ours
ads/four-dollars-talkinghead/final-v4.mp4· 25.3s- Status
- Shippable — pending two claim checks
- Spend
- ~1,400 cr
A format remake rather than a shot remake: same four-beat skeleton — price hook, absurdity anchor, finger-counted value stack, CTA — with our offer inside it. The interesting work was almost entirely in performance and audio.
Steps
- Measure the source voiceover word by word — pauses, stretched words, energy dynamics — and turn that measurement into explicit delivery direction.
- Write the script line and the delivery direction into the same text prompt, one MiniMax H3 generation per beat from an approved Nano Banana Pro still.
- Transcribe the outputs back and check the read landed verbatim, measured the same way as the source. All four beats passed first take.
- Audition the voice for real — 10+ takes — and choose H3's own voice over our Roxie TTS preset, which reads as announcer TTS the moment it's on camera.
- Finish the audio with a zero-credit chain: per-clip EQ match to the source's spectrum, near-dry convolution reverb at 80ms, the source ad's own room tone harvested from its single silent gap and spectrally resynthesised, mixed ~6 dB under its noise floor, then loudnorm to −26 LUFS. Music was A/B'd and dropped.
- Build captions from the output audio's word timestamps — auto-phrased into 2–5 word pills at real pause boundaries, swapped on phrase boundaries, in the source's own caption system.
- Cut 12 shots on exact word timestamps, alternating native crops, with one tighter punch-in reserved for the price line. Dead air trimmed at every join.
"$4 per week" is $49 ÷ 12 weeks — it moves if the price ladder moves. "Free to try for a week" is only true while the 7-day trial is live on the deployed paywall. Both are written into the file as blocking checks rather than remembered.
Cost note worth carrying: H3 pricing swung from 5 to 58 credits per second within a single day. Probe the rate before regenerating anything.
Sweat rank08 — at-home workouts for women over 50
- Source
- Sweat rank-8 running Meta ad, "At-home Pilates" · 14.2s · five continuous movement shots
- Ours
ads/50queen-rank08-athome-workouts/50queen-rank08-remake-final.mp4· 13.7s- Status
- Final — 29 Aug
- Spend
- ~160 cr (failed jobs auto-refunded)
The clearest demonstration that there is no single technique. Five movement beats needed five different generation strategies, and each was approved before the next was attempted.
| Beat | Movement | Technique |
|---|---|---|
| S1 | Wide squat, overhead reach | Carbon copy — Kling Motion Control driven by the source at its native 360×640, palindrome-padded past the ~4s minimum |
| S2 | Side plank → one-arm bridge | Slow driver — source retimed to 2× slow-mo, output retimed back, halving per-frame displacement so pose tracking survives floor work. Driver starts late to dodge moderation on a hip close-up |
| S3 | Same squat, other side | Mirrored S1 — driver h-flipped with a mirrored start frame; the room stays true because the background comes from the start frame |
| S4 | Down-dog → upward dog | Pinned endpoints — Kling image-to-video with start frame, generated end frame, and a prompt naming the vinyasa. Motion control mangled the inversion; prompt-only i2v just held the pose |
| S5 | Kneeling gate side bend | Seedance omni-reference — pose named in the prompt, source segment as motion reference, start frame for appearance |
| S6 | End card | Composited — pure ffmpeg and PIL, phone playing a real recording of the workout player |
What we changed
Structure preserved exactly — five shots, one caption each, brand card with the offer. The caption ladder rewritten onto our positioning, and a Roxie voiceover added as a sound-on layer over a mute-first cut where the captions already carry the whole pitch.
Four of the six segments are motion-driven by, or mirrored from, the source footage. Structure, timing and caption rhythm aren't protectable — but the README records that before scaling spend on this creative, those segments should be re-driven from our own filmed reference of the same movements so the motion provenance is fully ours.
LazyFit rank04 — the 12-week at-home challenge
- Source
- LazyFit "28-Day Bed Pilates Challenge", rank 4 of the page's top 15
- Ours
ads/50queen-12week-at-home-challenge/final/…-1080.mp4· 30.0s- Status
- Final — v9, 28 Aug
- Spend
- ~35 stills, 12 clips, 3 TTS reads, 1 music track
A painterly/écorché illustrated ad — which meant the generation problem was as much about style conversion order as about motion.
The recipe, validated across every movement beat
- Segment the reference to movement boundaries using scene detection plus dense timestamped frame grids. Keyframes must come from a segment's actual first frame, never from arbitrary-rate samples.
- Build pose masters in the painterly style first, and reach peak poses by limb-only edits. Never pose-edit inside a derived style — it grows limbs.
- Convert the verified painterly poses to écorché by style conversion, mat swap included.
- Animate with H3 image-to-video using the rest pose as both start and end frame, with the motion described in the prompt — clips then open at the true starting position and loop seamlessly.
- Composite in ffmpeg: paper ground, strip pan, PIL-built captions and UI and cards, keyed hand cutout — alpha-scan the cutout's edges before animating it, because this one exits bottom-right.
- Audio last: per-line VO placed from STT word timestamps, ducked music, loudnorm. Captions inherit the same anchor table, so VO and captions cannot drift.
What we changed
- The source's plus-size / ex-revenge hook replaced with a capability promise: "If you're out of shape, I beg you to try this."
- 28-day bed pilates → 12-week at-home program.
- The transformation beat is posture and energy, not body size — slouched to upright — as a deliberate ad-policy hedge.
- The UI demo shows the real product: real session names and durations pulled from
sessions.ts, real movement demo clips, and a Week 2 that actually unlocks after Week 1 is completed.
Nine versions, each one a named review note: audible VO → ducked music → the first hand → voice swap → natural pace → real workout cards → uncropped hand → tap-collapse-tap flow → quieter music and the Week-2 unlock. The version history is the log.
LazyFit rank05 — the printable day cards
- Source
- LazyFit's #5 creative — "Printable Bed Pilates", the printer-prints-day-cards format · 37.5s
- Ours
ads/50queen-beginner-friendly-at-home/final/…-1080.mp4· 36.6s- Status
- Final — v4.2
- Spend
- ~225 cr + ~1.2k ElevenLabs characters
Steps
- Frame-by-frame teardown: shot table, caption and VO transcript, motion grids.
- Probe exactly one card and publish it as its own artifact — that probe was the go/no-go gate for the whole build, and it settled the recipe before production started.
- Per card: 3fps frame-grid the source segment to write the beat description → composition-matched Roxie still (Nano Banana Pro + charsheet, pose copied from the source frame) → Seedance 2.0 motion copy with the source segment as video reference, the still as image reference, and a beat prompt at source-matched duration → same-cadence grid verification → composite into programmatic card chrome.
- Draw everything else in code — hook sheet, captions, app screens, chrome in PIL; printer, tumble, tap and collapse animations in ffmpeg.
- Wire the app demo to the real product: Week-1 movement clips with their real names and durations from the content model.
- Voice: ElevenLabs Jessica reading the caption ladder verbatim, over a licensed bed.
Seedance copies motion. MiniMax H3 interprets it. Kling Motion Control refuses stylized references. That one line decides the model per beat on every build since.
Two operational gotchas: the account caps at 6 concurrent jobs, so parallel sessions collide — run sequentially with backoff. And the ripped source's "PROTECTED" watermark ghosts contaminated any paper texture sampled from it, so all paper is now synthesized and the early chrome had to be rebuilt.
Adaptation worth noting: the source's "lose 30 lbs in 28 days" was carried verbatim through a deliberate carbon-copy pass — to isolate structure from copy — and then dropped for "feel healthier in 28 days" in the shipping cut. The claim class was flagged as the one Meta's health policies target, and the decision to drop it was made explicitly rather than by default.
BetterMe Men — the challenge statics
- Source
- A BetterMe Men static from the Ad Library, plus its sibling video from the same campaign
- Ours
- Four 1080×1920 statics in
ads/50queen-12week-challenge-statics/ - Status
- Shippable — four with distinct test roles
The only non-video remake, and the only structural replication — the mechanisms were copied, none of the pixels were.
Steps
- Archive both source creatives locally with frames and transcript, then break them down into a seven-mechanism analysis: art-object instead of a real body, date-anchored start, challenge framing, milestone ladder, placement-native CTA, age-gate quiz landing.
- Write an adaptation table — mechanism by mechanism, kept or replaced, with the reason.
- Generate the backgrounds with Nano Banana 2 at 9:16 2K, one pass each, as expressionist oil-paint scenes.
- Composite the type in HTML, not in the painting, and screenshot it at 1080×1920 through the headless browser. v1's dark-on-paint failed legibility; v2 put white Avenir Next over a soft dark wash.
- Ship four variants with explicitly different jobs: a primary, a B-test, a clone control deliberately outside our voice rules to isolate whether the "before-body" device beats the function scenes, and a hold.
BetterMe's transformation promise — "unrecognisable by October" — is outside our voice rules. So the milestone ladder became function milestones: mornings feel looser → stairs without stopping → up off the floor without hands. Which happen to be the program's own path names, so the ad and the product say the same thing.
Structurally useful: the dates live in the compositor HTML, not in the paintings. Re-dating the set is a text edit and four screenshots; the backgrounds are reusable indefinitely.
Wall-of-text UGC — the $1.03 format
- Sources
- uskin.app's AI UGC ad (via a tweet claiming a $1.03 result) + @walkingwithnat's persona account — 759K likes on 3.6K followers
- Ours
- Two creatives in
ads/50queen-wall-of-text-ugc/final/· 5s and 8s - Status
- Shippable — organic-first ship path
- Spend
- ~270 cr across ~17 rounds
A faceless, phone-rough vertical video that is read, not watched. A ~60-word overlay takes 18–24 seconds against a 5–8 second loop, which forces rewatches and manufactures the watch-time signal. The copy opens with a concession a brand would never make, and the product is named exactly once, in a parenthetical.
Steps
- Frames with Nano Banana Pro at 2K 9:16, charsheet as the identity reference.
- Motion per play. Play A used Kling 3.0 with start and end frames and the "slight movement" trick — the end frame is the same tightness as the start, so the only motion is the phone tilt settling, and the end frame is regenerated from the start frame so the grip matches. Play B was a race — Kling against H3 on a jog prompt with no motion reference available — and H3 won on stride quality.
- Crop ~10% off the top of both, because head-bob pulls the mouth into frame.
- Rough the plate deliberately: centre-crop to exact 1080×1920, bounce through 480p with sensor noise, come back to 1080.
- Overlay plain white Arial Bold 60px, per-line centred, no stroke, no shadow, block at 46% height — one ffmpeg call, so copy variants cost nothing and can be tested freely against the same plate.
At extreme close-up the image reference dominates skin texture — anchor close-ups on a frame whose skin already looks right. A reference image containing a text overlay will bleed that text into the output. And model outputs are not true 9:16 (716×1284, 716×1272) — always re-crop rather than assuming.
Ship path is different from every other build here: post organic on the page, boost small, then run winners as Reels-placement ads — and kill anything that doesn't beat the Says-Who $0.85 CPL benchmark after roughly $15 of spend.
Which model for which beat
| Model | Use it for | Constraints that bite |
|---|---|---|
| nano_banana_pro / 2 | Every still — charsheets, scene frames, pose masters, end-card art | Up to 14 references; holds identity off a charsheet. Two references both containing a person → identity hijack. Trained toward attractive; fights any "unflattering" brief. |
| minimax_h3 | Interpreting motion; talking heads; image-to-video from a pinned pose | duration 4 fails silently — use ≥6. Reference length must match output length. start_image can't be mixed with image_references. Drifts framing and grade from the start frame. Clones the reference video's subject if identity conditioning is weak. Outputs are not true 9:16. Moderation-rejects some wardrobe outright. |
| seedance_2_0 / 2_5 | Copying motion faithfully; stylized and illustrated references | Up to 9 image references plus video references. ~2s minimum reference — palindrome-pad shorter ones. Account caps at 6 concurrent jobs. |
| kling3_0 | Pinned-endpoint shots — start frame and end frame both supplied | The turbo variant takes exactly one reference and treats it as frame 1, so it can't take a charsheet at all. |
| kling3_0_motion_control | Carbon-copying a real movement onto our character | Cannot read bird's-eye footage. Refuses stylized references. ~4s minimum driver. Never upscale the driver — feed it at native resolution. Failed jobs auto-refund. |
| ElevenLabs | On-camera-adjacent voiceover (Ginger, Jessica) | Needs a TTS-scoped key — STT keys 401 on text_to_speech. Library voices may need a paid tier. |
| seed_audio / Sonilo | Original music beds — no licensing exposure | Source-ad audio is fine for an academic replication and must be replaced before any paid run. |
Rules earned the hard way
Sequencing
- Generate the aspirational state first and derive the diminished one from it — models anchor better downhill than up.
- Gate on the cheapest possible probe. One blend check, one card, one three-second window. Never fund a full pass on a hypothesis.
- Verify each beat against the source at the same cadence before accepting it, not at the end of the assembly.
- Get the beat approved before generating the next one. Six different techniques on one 14-second ad is a normal outcome.
Identity and moderation
- Never pass two references that both contain a person and expect to control which one wins.
- Never drop image references to slip past moderation — the model will clone the reference video's subject instead.
- When a pose still trips moderation, swap in any modest approved still of the same character and describe the scene rather than the body.
- Force body deltas at the scene-frame level; the scene frames are what appear on screen.
- Dual-anchor every scene frame on the charsheet plus an already-approved frame, and write explicit do-not-drift guards into clip prompts.
Post is where the quality is
- Pick the music first, analyse the BPM and the drop, then place every cut on the beat.
- Pre-render segments so re-cuts, re-timings and copy variants cost zero credits. Most review notes are re-renders, not regenerations.
- Place voice from STT word timestamps and let the captions inherit the same anchor table — then they can't drift.
- Burn all text in post. No video model holds legible multi-line text.
- Measure before positioning — card heights, safe areas, caption widths. Eyeballing costs a version.
Claims and provenance
- Every price or offer claim in a cut gets written into the README as a blocking pre-ship check against the deployed funnel.
- Before/after arcs are framed as capability and energy, never as weight. That's a policy position, not a style preference.
- No fabricated body-results imagery. It's manufactured evidence pointed at the exact anxiety the product addresses, and it's also the format Meta bans.
- Source audio and source-driven motion are fine for a replication spike and must be replaced or re-driven from our own footage before spend scales.
Cost
- Roughly 160–610 credits for a full video remake; the talking head at 1,400 was the outlier, and generation pricing moved 10× inside one day.
- Reconcile spend per job, never by account balance — concurrent sessions pollute the delta.
- Probe the current rate before regenerating anything you've generated before.