Assemble a two-host fake-podcast skit ad from a config — per-line lipsync clips hard-concatenated in script order, scaled/padded to 1080×1920, WHITE bottom-center captions (up to 5 words per cue, broken on sentence punctuation, word-wrapped to stay in-frame, held at least 0.9s) built from each line's OWN ElevenLabs char-level timestamps (offset by cumulative clip start, never Whisper), and closed on a Playwright/PIL brand end card composited from the real wordmark — never AI-rendered text. This is the FREE deterministic assembly stage (concat + white captions + end card + crf28 encode); the per-line VOs, photoreal gpt-image-2 base stills, expression variants, and lipsync clips come from create-vo-elevenlabs / create-image-gpt-image-fal / create-video-fal. Use for the podcast-skit format.
npx gooseworks install --all # then, in Claude Code, Cursor, or Codex: /gooseworks use the render-podcast-skit skill
Assemble a two-host podcast skit ad from a config: two hosts at a podcast desk do a snappy back-and-forth about the product, related however the user chose (friends, interviewer + guest, a doubter won over, two fans, a friendly debate). Each line is its own lipsync clip so the edit can cut on the dialogue beat (~1.8s avg); this capability is the FREE, deterministic assembly that concatenates those clips, renders the WHITE captions, and appends the brand end card.
The creative calls are the user's, asked by the format recipe before any paid step — this assembly just renders whatever the config holds:
role. Never default to skeptic vs believer. The demo used "a doubter
won over".voices.HER / voices.HIM and
who: HER|HIM are only the host A and host B SLOTS — they fix neither gender nor role. The demo
used a young woman as host A (the doubter) and a young man as host B.scripts/config.example.json is the worked example (Ladder run-02 "Laundromat 2am", ~49s
1080×1920 9:16, ~22 lines) — copy its structure, never its creative values; scripts/PIPELINE.md maps every config block to its source step
and scripts/README.md documents the free assembly.
This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are
separate capabilities: one ElevenLabs with-timestamps VO per line (one voice per host) via
create-vo-elevenlabs; two photoreal base stills at the chosen set's desk plus ~10 expression variants
(mouths NEUTRAL/CLOSED, gpt-image-2 quality=high, not nano-banana) via
create-image-gpt-image-fal; and one lipsync clip per (still, VO) pair via create-video-fal.
Given the per-line clips + their VO timestamps + the brand wordmark SVG, render-podcast-skit
walks the scenes in script order, builds the global caption timeline, renders the WHITE captions,
hard-concats the clips, auto-appends the end card, and final-encodes crf28 → the master. Re-cuts
reuse the existing VOs / stills / clips and cost $0.
words.json by offsetting each line's char-level word timings by the cumulative clip
start, group into ≤5-word cues broken on sentence-final punctuation, and render WHITE
#FFFFFF bottom-center captions (black outline), word-wrapped to stay in-frame and held
≥0.9s — PIL PNG overlays when the host ffmpeg lacks libass (common), else ASS. (Yellow 3-word
karaoke was the old style, rejected in testing.) Whisper on the rendered clips mistimes; the VO
timestamps are ground truth.-preset slow -crf 28 + aac
96k → a 1080×1920 h264+aac master (~6MB for ~28s; the old -crf 20 produced ~16MB). No paid
calls, no keys.voices.HER / voices.HIM andFull video production sequence with script review, actual ingredient choices, controlled generation, editing, evidence-based quality review, polish, captions and delivery. A host binding supplies project storage, authentic approvals, provider access and billing.
Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
Write the words of a short-form video ad (voiceover, dialogue, chat bubbles, on-screen lines) the way performance creative teams do instead of from a blank page. Builds the script from the buyers' own words, the beat sheet of an ad that already works and three deliberately different angles, filters them with a rule check and a second non-Claude model, and takes the strongest into the review with the other two as one-line swaps. Use it in every video ad run before any paid step, and whenever the user asks to write, rewrite or improve a video ad script or says a script sounds generic or AI-written.