Assemble a narrated-UGC "stitch reply" ad from a config — a single spoken VO carries a verbatim testimonial while ~30 per-cut i2v clips (one creator across ~5 wardrobes in ~3 worlds, plus product B-roll) are each trimmed to their EDL window built from the VO's Whisper word boundaries and hard-concatenated via filter_complex concat (never the demuxer, which drops audio on a duration mismatch), the VO mixed over an optional sidechain-ducked instrumental bed (−20dB, 20 to 1) so the VO stays on top, karaoke-pop captions burned on every word throughout (VEED Whisper preset, re-spelled against the locked script), a landing-page scroll rendered as FFmpeg zoompan over a Playwright PNG (not i2v), and closed on the brand's real end-card PNG — never AI-rendered text. This is the FREE deterministic assembly stage (trim-to-EDL + filter_complex concat + VO and music mix + karaoke captions + landing-page zoompan + end-card append); the VO, creator, start-frames, and clips come from create-vo-elevenlabs / create-image-gpt-image-fal / create-image-fal / create-video-fal. Use for the narrated-ugc-wardrobe-stitch format.
npx gooseworks install --all # then, in Claude Code, Cursor, or Codex: /gooseworks use the render-narrated-ugc-wardrobe-stitch skill
Assemble a narrated-UGC "stitch reply" ad from a config: a fast-cut vertical testimonial where a single spoken VO carries a verbatim ~13-sentence testimonial over ONE creator across ~5 wardrobe changes in ~3 micro-worlds, interspersed with product B-roll (e.g. product macro, unboxing, a landing-page scroll), ~30 hard cuts on the VO cadence, closing on a brand end card. This capability is the FREE, deterministic assembly — trim-to-EDL, hard-concat, the VO+music mix, the karaoke-pop caption burn, the landing-page zoompan, and the end-card append.
scripts/config.example.json is one worked example (Bioma "Do NOT buy Bioma Probiotics", ~37s
1080×1920 9:16, ~30 body cuts + a ~2s end card) — its creator, voice, hook, worlds and music are
that demo's answers, not defaults; scripts/PIPELINE.md maps every config block to
its source step and scripts/README.md documents the free assembly.
The creative calls are made upstream by the user (the recipe's choices) and arrive in the config;
this assembly never picks them.
character.descriptor / character.name. The demo used a
28-year-old blonde woman.vo.voice_id / vo.settings. The demo used an excited voice
matched to its creator.vo.hook_line / vo.script_md / vo.payoff_line. The
demo used a "Do not buy <brand>" reversal.worlds.briefs. The demo used bedroom / kitchen / bathroom.audio_mix.music_brief. With "no music", skip the bed
and mix the VO alone.This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are
separate capabilities — the spoken VO (create-vo-elevenlabs) Whisper-aligned so the WORD
BOUNDARIES set the cut grid; one locked creator (create-image-gpt-image-fal anchor + ~5 wardrobe
edits chained off the anchor) + 3 world wides + per-cut start-frames (create-image-fal product
composites); and one Veo/Seedance i2v clip per cut (create-video-fal). Given the VO +
vo-final.words.json + edl.json + one clip per cut + a Playwright landing-page PNG + the brand
end-card PNG, render-narrated-ugc-wardrobe-stitch trims each clip to its EDL window, hard-concats
on the VO cadence, mixes the VO over the ducked bed, burns the karaoke-pop captions, appends the
end card → the master. Re-cuts reuse the existing VO / start-frames / clips and cost $0.
hook, feature,
reaction-insert, payoff-hold, b-roll-insert, landing-page); snap every cut window to the
word boundaries. The payoff line gets a HELD payoff-hold beat (~3× mean shot length).filter_complex concat, not the demuxer. Trim each clip to its EDL window and
hard-concat with filter_complex concat — the -f concat demuxer drops the audio when a
drawtext/scale step shaves a clip a few ms below its window. No dissolves.vo-final.words.json (VEED
Whisper preset, bold yellow), on every word; re-spell brand tokens Whisper mishears against the
locked script (Bioma demo: "synbiotic" over "symbiotic"; kept "I'ma" verbatim) — never edit the script to
match Whisper. Captions are suppressed over the end card. If VEED mis-captions a brand token,
hand-patch that sentence with local ASS karaoke.filter_complex concat, VO+music mix,
caption burn, landing-page zoompan, end-card append, loudnorm I=-14 → a 1080×1920 h264+aac
master (~37s). No paid calls, no keys.character.descriptor / character.name. The demo used avo.voice_id / vo.settings. The demo used an excited voicevo.hook_line / vo.script_md / vo.payoff_line. Theworlds.briefs. The demo used bedroom / kitchen / bathroom.audio_mix.music_brief. With "no music", skip the bedFull video production sequence with script review, actual ingredient choices, controlled generation, editing, evidence-based quality review, polish, captions and delivery. A host binding supplies project storage, authentic approvals, provider access and billing.
Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
Write the words of a short-form video ad (voiceover, dialogue, chat bubbles, on-screen lines) the way performance creative teams do instead of from a blank page. Builds the script from the buyers' own words, the beat sheet of an ad that already works and three deliberately different angles, filters them with a rule check and a second non-Claude model, and takes the strongest into the review with the other two as one-line swaps. Use it in every video ad run before any paid step, and whenever the user asks to write, rewrite or improve a video ad script or says a script sounds generic or AI-written.