50Queen · ads · production reference

How we remake an ad

Twelve remakes between 5 and 30 August 2026, reconstructed from the session transcripts and the logs they left behind. What follows is the pipeline that emerged, then each ad we took apart — the source, the steps that produced our version, the adaptation we made to it, and the rule we only learned by getting it wrong.

12 remakes 9 shipped masters ~3,900 Higgsfield credits 1 character · Roxie
The shape of every build

The pipeline

No two of these ads were made the same way, but every one of them moved through the same ten stages. The stages are the reusable part; the technique inside stage seven changed per beat, sometimes per shot.

  1. Source the swipe. Pull an advertiser's top creatives from the Facebook Ad Library sorted by total_impressions, via the fb-ad-library-download skill, and archive them under ads/_inspirations/page-<id>-top15/ with a README recording rank, Library ID and start date. Two byte-identical files under different Library IDs is the strongest performance signal available — Meta publishes no spend, so duplication and longevity are the only proxies.
  2. Tear the source down to timecodes. ffprobe for the shape, ffmpeg scdet for the real cut points, contact sheets at different rates for different jobs (8fps across the hook, 1fps across the body), keyframe extraction, ElevenLabs transcription with word-level timestamps, and Gemini via describe-visual for what the frames mean. Output is a beat table: every cut, every caption, every overlay, timestamped.
  3. Publish a plan before spending anything. A beat-by-beat artifact mapping source beat → remake spec, with the prompts, per-beat credit cost and an adaptation table — published as a Claude Artifact and mirrored to Cloudflare. Approval happens on the plan, not on the footage.
  4. Rewrite the copy ladder onto our positioning. Structure and rhythm are kept; the claims are replaced. Never "bed pilates", never "28 days", never a weight-loss promise — always a 12-week at-home program and a capability claim.
  5. Anchor the character. Roxie, from a wardrobe-specific bilateral charsheet registered in content/_character/roxie.json. The charsheet is an input to the image model; it is almost never handed to a video model directly.
  6. Probe the one question that decides viability. Six credits, not a hundred and twenty. One alpha-blend of the strobe pair. One day-card. One three-second window. Nothing else gets funded until the probe comes back clean.
  7. Generate per beat, verify per beat. Still factory first (Nano Banana Pro, references = charsheet + the source frame for pose), then motion with whichever model suits that specific movement, then a same-cadence frame grid against the source before the beat is accepted.
  8. Assemble in code. ffmpeg and PIL. Captions, UI screens, chrome and end cards are drawn programmatically, and segments are pre-rendered so a re-cut costs nothing. Audio goes on last, placed from STT word timestamps so voice and captions cannot drift apart.
  9. Review against the source, side by side. A second artifact with the cut, the original, a synced comparison and every rejected iteration still playable. Eugene either approves or names the delta; the delta becomes the next version.
  10. Ship into placements. 9:16 plus a 4:5 feed crop, wired through asset_feed_spec placement customization, a UTM per creative, and a CPL kill rule against the current benchmark.
The one structural insight

From the very first before/after build: an ad decomposes into generation problems and editing problems, and the editing problems are usually the ones that make it good. The strobe transition everybody liked was never generated — it was two clips and a cut list. What generation had to deliver was a matched pair of frames. Get the pair right and the effect is twenty lines of ffmpeg; get it wrong and no video model saves you.

Remake 01 · 5–7 Aug

Dancebit — "Walk 6,587 steps in 10 minutes"

Source
instagram.com/reels/DRPXnNyjMyA · 43.35s · one continuous locked-off take, zero cuts
Ours
ads/workout_program/final.mp4 · 70.27s · 1080×1920
Status
Built — two known artifacts unfixed
Teardown
.claude/ads-research/teardown-dancebit-DRPXnNyjMyA.md

The first remake, and the one that set the reading habit. The teardown's finding was that a funded brand with a real budget still shot in a kitchen in one take — all the polish lives in the overlay layer. So the remake didn't try to copy the footage. It copied the chrome.

Steps

  1. Keyframe the source at 16 points, build three contact sheets at different rates (8fps over the hook, 4fps over the intro, 1fps over the body), and write the beat sheet with the overlay state at every beat.
  2. Use our own filmed workout footage as the base rather than generating the performance — workout-trimmed.mp4, 65.77s, eight movements.
  3. Drive Kling 3.0 Motion Control from that footage with the Roxie charsheet as the character, converting real reference movement into our talent.
  4. Rebuild the overlay layer from scratch in headless Chrome (render_overlays.py → PNGs) and composite with ffmpeg — the title lockup, per-movement labels, the in-workout chrome, the end card.
  5. Measure the chrome card rather than eyeballing it (244px), and set its bottom edge at y=1620 for a 300px margin that clears Meta Reels' ~269px UI zone.
  6. Lift the source's music bed, loop it with a 1.2s crossfade to cover 70.3s, fade under the end card — flagged as needing its own licence before any paid run.

What we changed

Remake 02 · 6–7 Aug

@sue.giers — handstand up the G-Wagon

Source
ads/sue-giers-7646356284442152224.mp4 · 9.008s · 1080×1920
Ours
Never shipped — three plan revisions, two probe generations
Status
Pipeline probe
Plan
.claude/ads-research/genplan-sue-clone.md

Asked as "AI generate the same video as this". The value it delivered wasn't a creative — it was the model constraints that every later build ran on.

Steps

  1. 5fps contact sheet plus scene detection, which corrected the existing teardown: this is two shots with a hard cut at 7.70s, not one continuous take. That halved the difficulty — no handstand-into-seated transition ever needed generating.
  2. Write a frame-1 spec precise enough to be a prompt: camera at knee height tilted up, subject right-of-centre, SUV entering frame left, flat fluorescent light, no grade, mild sensor noise, "amateur social video — not cinematic".
  3. Spec the text overlay as a post step with in/out times — video models can't hold legible multi-line text across eight seconds.
  4. Three costed approaches so the decision was Eugene's: v1 Seedance with the source as a video reference (~150cr), v2 a single Kling Turbo prompt at general strokes (15cr), v3 Nano Banana charsheet → start and end frames → non-turbo Kling.
  5. Run v3b deliberately as a controlled experiment — keep the brand name, drop the video reference — to isolate which input had triggered the earlier unexplained moderation failure.
Learned here, used everywhere after

kling3_0_turbo accepts one reference and treats it as frame 1 — hand it a multi-panel charsheet and it animates the sheet. The charsheet is an input to the image model, never to the video model. Non-turbo Kling takes a start and an end image, and pinning both ends is what stops arc drift. Real-person prompts get ip_detected; vehicles get described by shape, not brand.

Also recorded, and worth more than the pipeline: a faithful clone of this ad would have been the wrong creative. It has no product, no offer and no CTA, and it sells peak aspiration to women who may not see themselves in it. Run it to prove the pipeline; put money behind a different concept.

Remake 03 · 6 Aug

"Says who?" — the before/after strobe

Sources
@kangmiratips21 7606799992237444383 (the strobe mechanic) + @dianecaldwell_ 7620248665193499934 (the copy and the CTA)
Ours
ads/before-after-strobe-cut/final.mp4 · 14.79s · 1080×1920
Status
Shipped and ran — became the account's benchmark creative at $0.85 CPL
Spend
~122 cr per complete pass

The build that produced the house method. The plan opened by separating what had to be generated from what had to be edited, and then optimised almost entirely around one dependency: the strobe only works if the before frame and the after frame are composed to interleave.

Steps

  1. Generate the "after" character first, then derive the "before" from it. Never the reverse — models anchor far better on "make this person heavier and older" than the opposite, and the after is the shot the whole ad has to pay off.
  2. Generate the before scene frame (B1) from the before sheet.
  3. Generate the after scene frame (B2) with B1 passed as a second reference, instructed to use it for camera framing only — head size, head height, body centring, camera height, crop. This single step is what makes the strobe read as one person morphing rather than two clips fighting.
  4. Gate on a 50% alpha-blend of B1 and B2 before generating any video. Six credits answers both viability questions — can we get a convincing unflattering before, and will the pair interleave — at 5% of a full pass.
  5. Four MiniMax H3 clips from the start frames. Generate 6s, use 2–3s: video models drift toward flattering the longer they run, so the before clips are deliberately near-motionless and cut early.
  6. Pick the music first, analyse it (113.6 BPM, drop at 6.82s, bar lines at 8.93 / 11.05 / 13.16), then place every cut on its beat.
  7. Build the strobe as a frame pattern — 3 frames before, 2 frames after, ×5 — resolving onto the after exactly on the drop.
  8. Burn the captions in post: Avenir Black, white with a 3px black stroke, at 72% frame height. "you're 56 / it's too late to change" dies into the strobe; says who? arrives on the other side.

The build diverged from the plan in one place worth keeping: instead of the start-frame workflow, it used an 8-panel storyboard — alternating pose and transition panels — fed to MiniMax H3 as image_references.

Two failures confirmed at cost

Two-reference identity hijack. Passing the charsheet together with a real frame from the source TikTok made Nano Banana take the room and the real creator's face from the TikTok frame, ignoring the sheet entirely; text instructions did not override it. Always two-step: generate the character in the target state on a plain background first, then pass that as reference 1 alongside the scene photo.

H3 drifts framing and grade. The after clip pulled wider than its start frame; the before clip warmed a cool marble bathroom to beige. Not prompt-fixable. Which means the strobe pair has to be re-aligned after generation, not just before.

Remake 04 · 21–22 Aug

"Says who?" — the late-40s variant

Source
Our own winner, re-aged. Spec in dissection-a-says-who-age-variants.md
Ours
ads/before-after-40s/final-40s-FINAL.mp4 + a 4:5 feed crop
Status
Launched — ran against the 56 creative in one 45–65 ad set, $50/day CBO
Spend
~278 cr across three rebuilds

The first time we remade ourselves. Everything except the age variables was frozen: cut list, track, strobe pattern, caption style and timing, the three after contexts. Only the character, the ages in the copy and the hook line moved.

Steps

  1. Two candidate "after" characters generated as full sheets and put to Eugene as a pick — A (chestnut, 47) and B (Black American woman, 48) — with late-40s calibration written explicitly into the prompt (faint crow's feet, visible pores, first silver strands only).
  2. Derive both "before" sheets from the chosen afters. QA caught the predicted failure immediately: the +22lb barely registered on candidate A. Every image model is trained toward attractive, and it bites here first.
  3. Force the body delta at the scene-frame level, not the sheet level — the scene frames are what appear on screen, so the weight instruction goes there, described relative to clothing fit rather than as a number.
  4. Re-run the alpha-blend strobe gate on the new pair before any video spend. Passed.
  5. Four H3 clips (10s / 6s / 10s / 10s), start-image mode, motion prompts carried over from the original build.
  6. Byte-copy the audio bed out of the shipped winner so the 6.82s drop timing is inherited exactly and the cut list transplants without re-timing.
  7. Assemble with the same build.sh; only the caption strings change.

Two direction changes, both on the after side

The v1 cut was rejected: the run and car clips read late-50s, not late-40s. The correction wasn't a smaller delta but a bigger one — push the after to late 30s, then again to early 30s (plump collagen, zero lines, sculpted jaw). The before kept its grey roots; the dull-roots-to-glossy-colour jump is part of the transformation. That preference is now standing policy: the after should read dramatically younger and prettier than the before.

What fixed the age drift in the rebuild: dual-anchoring every scene frame on the sheet plus an already-approved scene frame, moderate closed smiles instead of wide-open laughs, hair down, and explicit do-not-age guards in the clip prompts. All three v3 clips passed age QA first try.

Shipping decisions

Remake 05 · 25–28 Aug

ChillFit rank06 — "I'm 55" chair workout

Source
Ad Library 1585698999664983 · ChillFit's #6 by impressions · "TENHO 65 ANOS"
Ours
ads/50queen-im-55-chair-workout/final/…-720.mp4 · 31.5s · audio on
Status
Final — 28 Aug
Spend
~374 cr + ~1.5k ElevenLabs characters

The first end-to-end remake where every frame is ours and every sound is ours — zero source assets survive in the cut. It also produced the shared 50Queen ending that two later ads splice in.

Steps

  1. Pull the advertiser's top 10 from the Ad Library by impressions, note the byte-identical duplicates (rank 1 and rank 5 were the same creative under two IDs), and pick the chair ad.
  2. Write a 16-beat storyboard — the source torn down beat by beat with the remake spec written against each one — then a shot-by-shot production plan with prompts and costs. Both published for approval.
  3. Race a pilot to settle the generation recipe, then apply it: motion transfer with the source segment as --video-references, the charsheet and an approved still as --image-references.
  4. Extract the source's muscle-diagram strips by median-diff matting and rebuild them as our own overlays.
  5. Audition the voice properly — Seed Audio candidates, then ElevenLabs, landing on "Ginger" at speed 0.90 / stability 35% / similarity 75%. Captions burned to the exact speech timings.
  6. Build the non-performance beats in code: the parchment calendar, the red-nail tapping hand, the progress borders, the end card with three real app screens in device frames.
  7. Pre-render each timeline segment so re-assembly and re-syncing never require regeneration — an audio-only change is a remix and a remux, nothing more.

What we changed

ChillFit50Queen
"I'm 65""I'm 55", aimed at menopause
7 minutes a day15 minutes a day
28-day challenge12-week plan
App Store badgesWeb CTA — 50Queen is web-only
Body transformationCapability arc; stats bar reads AGE 55 · 15 MIN · GOAL: STRONGER

That last row is a policy hedge, not a style choice: framing the before/after as capability and energy rather than weight keeps it clear of Meta's before/after rules. The music is generated, so there's no licensing exposure, and the voice is a licensed library voice.

Remake 06 · 26–29 Aug

Sweat rank03 — the overhead pilates ad

Source
Sweat / Kayla Itsines, rank 3 of the page's top 15 · 14.7s · running byte-identical under two ad IDs
Ours
ads/50queen-overhead-guided-workouts/50queen_remake_v3.mp4 · 20.73s
Status
Final — audio must be swapped before a paid run
Spend
~610 cr including the instructive failures

The hardest replication and the one that taught the most. One locked overhead camera, eleven exercise vignettes joined by jump cuts, a kinetic type ladder, and a phone raised to the lens. Three beat types, three completely different pipelines.

The segment recipe — exercise beats

  1. Find the hidden cuts. What looks like a continuous take is eleven ~1s pose vignettes under identical framing. They only appear at a low scene-detection threshold, and every boundary was verified frame-by-frame against a cut-verification grid. Round 2's "flying legs" were the model physically bridging poses that were never connected.
  2. Chop at those boundaries so every reference segment is genuinely continuous.
  3. Slow each segment to ~6.4s — H3 refuses durations under 6, and duration 4 fails silently (four wasted jobs before that was pinned down).
  4. Run H3 with the slowed segment as the motion reference, at a matching duration, with the charsheet plus a per-beat pose still as identity references.
  5. Retime back to real speed with motion-blended resampling (minterpolate mi_mode=blend), not plain frame decimation.

The phone bridge, and the end card

Four rules, each paid for

Reference length must match output length. Give H3 a 3s reference for a 6s output and it invents choreography to fill the gap.

A single shared anchor still over-anchors. Later beats collapse toward it. Per-beat pose stills fixed it.

Never drop image references to dodge moderation. With weak identity conditioning, H3 clones the reference video's subject — one beat came back as the Sweat athlete on a hot-pink mat, and shipped that way until it was caught. The fix is to swap in any modest approved still of the same character, plus an explicit "she is NOT the person in the reference video".

Retime gently, and only where it's earned. Compressing a 6.5s master into a 1.5s beat decimates 80% of the frames and reads as jitter. Pick a window of the slow master and compress ~2.5×. Acted and idle beats play at 1:1 with a window trim; only exercise-form beats get compressed at all. Every one of those pacing fixes cost zero credits — it's a re-render, not a regeneration.

Also settled here: Kling Motion Control cannot read bird's-eye footage. The pose tracker misreads the foreshortening and returns ghost duplicate arms and a standing profile glued to the mat.

Remake 07 · 28 Aug

Sweat rank01 — the $4 talking head

Source
Sweat's rank-1 creative · single-take UGC talking head pitching a price
Ours
ads/four-dollars-talkinghead/final-v4.mp4 · 25.3s
Status
Shippable — pending two claim checks
Spend
~1,400 cr

A format remake rather than a shot remake: same four-beat skeleton — price hook, absurdity anchor, finger-counted value stack, CTA — with our offer inside it. The interesting work was almost entirely in performance and audio.

Steps

  1. Measure the source voiceover word by word — pauses, stretched words, energy dynamics — and turn that measurement into explicit delivery direction.
  2. Write the script line and the delivery direction into the same text prompt, one MiniMax H3 generation per beat from an approved Nano Banana Pro still.
  3. Transcribe the outputs back and check the read landed verbatim, measured the same way as the source. All four beats passed first take.
  4. Audition the voice for real — 10+ takes — and choose H3's own voice over our Roxie TTS preset, which reads as announcer TTS the moment it's on camera.
  5. Finish the audio with a zero-credit chain: per-clip EQ match to the source's spectrum, near-dry convolution reverb at 80ms, the source ad's own room tone harvested from its single silent gap and spectrally resynthesised, mixed ~6 dB under its noise floor, then loudnorm to −26 LUFS. Music was A/B'd and dropped.
  6. Build captions from the output audio's word timestamps — auto-phrased into 2–5 word pills at real pause boundaries, swapped on phrase boundaries, in the source's own caption system.
  7. Cut 12 shots on exact word timestamps, alternating native crops, with one tighter punch-in reserved for the price line. Dead air trimmed at every join.
Pre-ship checks live in the README

"$4 per week" is $49 ÷ 12 weeks — it moves if the price ladder moves. "Free to try for a week" is only true while the 7-day trial is live on the deployed paywall. Both are written into the file as blocking checks rather than remembered.

Cost note worth carrying: H3 pricing swung from 5 to 58 credits per second within a single day. Probe the rate before regenerating anything.

Remake 08 · 29 Aug

Sweat rank08 — at-home workouts for women over 50

Source
Sweat rank-8 running Meta ad, "At-home Pilates" · 14.2s · five continuous movement shots
Ours
ads/50queen-rank08-athome-workouts/50queen-rank08-remake-final.mp4 · 13.7s
Status
Final — 29 Aug
Spend
~160 cr (failed jobs auto-refunded)

The clearest demonstration that there is no single technique. Five movement beats needed five different generation strategies, and each was approved before the next was attempted.

BeatMovementTechnique
S1Wide squat, overhead reachCarbon copy — Kling Motion Control driven by the source at its native 360×640, palindrome-padded past the ~4s minimum
S2Side plank → one-arm bridgeSlow driver — source retimed to 2× slow-mo, output retimed back, halving per-frame displacement so pose tracking survives floor work. Driver starts late to dodge moderation on a hip close-up
S3Same squat, other sideMirrored S1 — driver h-flipped with a mirrored start frame; the room stays true because the background comes from the start frame
S4Down-dog → upward dogPinned endpoints — Kling image-to-video with start frame, generated end frame, and a prompt naming the vinyasa. Motion control mangled the inversion; prompt-only i2v just held the pose
S5Kneeling gate side bendSeedance omni-reference — pose named in the prompt, source segment as motion reference, start frame for appearance
S6End cardComposited — pure ffmpeg and PIL, phone playing a real recording of the workout player

What we changed

Structure preserved exactly — five shots, one caption each, brand card with the offer. The caption ladder rewritten onto our positioning, and a Roxie voiceover added as a sound-on layer over a mute-first cut where the captions already carry the whole pitch.

Provenance, written down before it becomes a problem

Four of the six segments are motion-driven by, or mirrored from, the source footage. Structure, timing and caption rhythm aren't protectable — but the README records that before scaling spend on this creative, those segments should be re-driven from our own filmed reference of the same movements so the motion provenance is fully ours.

Remake 09 · 28 Aug

LazyFit rank04 — the 12-week at-home challenge

Source
LazyFit "28-Day Bed Pilates Challenge", rank 4 of the page's top 15
Ours
ads/50queen-12week-at-home-challenge/final/…-1080.mp4 · 30.0s
Status
Final — v9, 28 Aug
Spend
~35 stills, 12 clips, 3 TTS reads, 1 music track

A painterly/écorché illustrated ad — which meant the generation problem was as much about style conversion order as about motion.

The recipe, validated across every movement beat

  1. Segment the reference to movement boundaries using scene detection plus dense timestamped frame grids. Keyframes must come from a segment's actual first frame, never from arbitrary-rate samples.
  2. Build pose masters in the painterly style first, and reach peak poses by limb-only edits. Never pose-edit inside a derived style — it grows limbs.
  3. Convert the verified painterly poses to écorché by style conversion, mat swap included.
  4. Animate with H3 image-to-video using the rest pose as both start and end frame, with the motion described in the prompt — clips then open at the true starting position and loop seamlessly.
  5. Composite in ffmpeg: paper ground, strip pan, PIL-built captions and UI and cards, keyed hand cutout — alpha-scan the cutout's edges before animating it, because this one exits bottom-right.
  6. Audio last: per-line VO placed from STT word timestamps, ducked music, loudnorm. Captions inherit the same anchor table, so VO and captions cannot drift.

What we changed

Nine versions, each one a named review note: audible VO → ducked music → the first hand → voice swap → natural pace → real workout cards → uncropped hand → tap-collapse-tap flow → quieter music and the Week-2 unlock. The version history is the log.

Remake 10 · 28 Aug

LazyFit rank05 — the printable day cards

Source
LazyFit's #5 creative — "Printable Bed Pilates", the printer-prints-day-cards format · 37.5s
Ours
ads/50queen-beginner-friendly-at-home/final/…-1080.mp4 · 36.6s
Status
Final — v4.2
Spend
~225 cr + ~1.2k ElevenLabs characters

Steps

  1. Frame-by-frame teardown: shot table, caption and VO transcript, motion grids.
  2. Probe exactly one card and publish it as its own artifact — that probe was the go/no-go gate for the whole build, and it settled the recipe before production started.
  3. Per card: 3fps frame-grid the source segment to write the beat description → composition-matched Roxie still (Nano Banana Pro + charsheet, pose copied from the source frame) → Seedance 2.0 motion copy with the source segment as video reference, the still as image reference, and a beat prompt at source-matched duration → same-cadence grid verification → composite into programmatic card chrome.
  4. Draw everything else in code — hook sheet, captions, app screens, chrome in PIL; printer, tumble, tap and collapse animations in ffmpeg.
  5. Wire the app demo to the real product: Week-1 movement clips with their real names and durations from the content model.
  6. Voice: ElevenLabs Jessica reading the caption ladder verbatim, over a licensed bed.
The model rule this build produced

Seedance copies motion. MiniMax H3 interprets it. Kling Motion Control refuses stylized references. That one line decides the model per beat on every build since.

Two operational gotchas: the account caps at 6 concurrent jobs, so parallel sessions collide — run sequentially with backoff. And the ripped source's "PROTECTED" watermark ghosts contaminated any paper texture sampled from it, so all paper is now synthesized and the early chrome had to be rebuilt.

Adaptation worth noting: the source's "lose 30 lbs in 28 days" was carried verbatim through a deliberate carbon-copy pass — to isolate structure from copy — and then dropped for "feel healthier in 28 days" in the shipping cut. The claim class was flagged as the one Meta's health policies target, and the decision to drop it was made explicitly rather than by default.

Remake 11 · 26–28 Aug

BetterMe Men — the challenge statics

Source
A BetterMe Men static from the Ad Library, plus its sibling video from the same campaign
Ours
Four 1080×1920 statics in ads/50queen-12week-challenge-statics/
Status
Shippable — four with distinct test roles

The only non-video remake, and the only structural replication — the mechanisms were copied, none of the pixels were.

Steps

  1. Archive both source creatives locally with frames and transcript, then break them down into a seven-mechanism analysis: art-object instead of a real body, date-anchored start, challenge framing, milestone ladder, placement-native CTA, age-gate quiz landing.
  2. Write an adaptation table — mechanism by mechanism, kept or replaced, with the reason.
  3. Generate the backgrounds with Nano Banana 2 at 9:16 2K, one pass each, as expressionist oil-paint scenes.
  4. Composite the type in HTML, not in the painting, and screenshot it at 1080×1920 through the headless browser. v1's dark-on-paint failed legibility; v2 put white Avenir Next over a soft dark wash.
  5. Ship four variants with explicitly different jobs: a primary, a B-test, a clone control deliberately outside our voice rules to isolate whether the "before-body" device beats the function scenes, and a hold.
The adaptation that mattered

BetterMe's transformation promise — "unrecognisable by October" — is outside our voice rules. So the milestone ladder became function milestones: mornings feel looser → stairs without stopping → up off the floor without hands. Which happen to be the program's own path names, so the ad and the product say the same thing.

Structurally useful: the dates live in the compositor HTML, not in the paintings. Re-dating the set is a text edit and four screenshots; the backgrounds are reusable indefinitely.

Remake 12 · 29–30 Aug

Wall-of-text UGC — the $1.03 format

Sources
uskin.app's AI UGC ad (via a tweet claiming a $1.03 result) + @walkingwithnat's persona account — 759K likes on 3.6K followers
Ours
Two creatives in ads/50queen-wall-of-text-ugc/final/ · 5s and 8s
Status
Shippable — organic-first ship path
Spend
~270 cr across ~17 rounds

A faceless, phone-rough vertical video that is read, not watched. A ~60-word overlay takes 18–24 seconds against a 5–8 second loop, which forces rewatches and manufactures the watch-time signal. The copy opens with a concession a brand would never make, and the product is named exactly once, in a parenthetical.

Steps

  1. Frames with Nano Banana Pro at 2K 9:16, charsheet as the identity reference.
  2. Motion per play. Play A used Kling 3.0 with start and end frames and the "slight movement" trick — the end frame is the same tightness as the start, so the only motion is the phone tilt settling, and the end frame is regenerated from the start frame so the grip matches. Play B was a race — Kling against H3 on a jog prompt with no motion reference available — and H3 won on stride quality.
  3. Crop ~10% off the top of both, because head-bob pulls the mouth into frame.
  4. Rough the plate deliberately: centre-crop to exact 1080×1920, bounce through 480p with sensor noise, come back to 1080.
  5. Overlay plain white Arial Bold 60px, per-line centred, no stroke, no shadow, block at 46% height — one ffmpeg call, so copy variants cost nothing and can be tested freely against the same plate.
Three generation notes

At extreme close-up the image reference dominates skin texture — anchor close-ups on a frame whose skin already looks right. A reference image containing a text overlay will bleed that text into the output. And model outputs are not true 9:16 (716×1284, 716×1272) — always re-crop rather than assuming.

Ship path is different from every other build here: post organic on the page, boost small, then run winners as Reels-placement ads — and kill anything that doesn't beat the Says-Who $0.85 CPL benchmark after roughly $15 of spend.

Reference

Which model for which beat

ModelUse it forConstraints that bite
nano_banana_pro / 2 Every still — charsheets, scene frames, pose masters, end-card art Up to 14 references; holds identity off a charsheet. Two references both containing a person → identity hijack. Trained toward attractive; fights any "unflattering" brief.
minimax_h3 Interpreting motion; talking heads; image-to-video from a pinned pose duration 4 fails silently — use ≥6. Reference length must match output length. start_image can't be mixed with image_references. Drifts framing and grade from the start frame. Clones the reference video's subject if identity conditioning is weak. Outputs are not true 9:16. Moderation-rejects some wardrobe outright.
seedance_2_0 / 2_5 Copying motion faithfully; stylized and illustrated references Up to 9 image references plus video references. ~2s minimum reference — palindrome-pad shorter ones. Account caps at 6 concurrent jobs.
kling3_0 Pinned-endpoint shots — start frame and end frame both supplied The turbo variant takes exactly one reference and treats it as frame 1, so it can't take a charsheet at all.
kling3_0_motion_control Carbon-copying a real movement onto our character Cannot read bird's-eye footage. Refuses stylized references. ~4s minimum driver. Never upscale the driver — feed it at native resolution. Failed jobs auto-refund.
ElevenLabs On-camera-adjacent voiceover (Ginger, Jessica) Needs a TTS-scoped key — STT keys 401 on text_to_speech. Library voices may need a paid tier.
seed_audio / Sonilo Original music beds — no licensing exposure Source-ad audio is fine for an academic replication and must be replaced before any paid run.
Reference

Rules earned the hard way

Sequencing

Identity and moderation

Post is where the quality is

Claims and provenance

Cost