Assemble a narrated motion-graphic LISTICLE video ad from a config — a spoken voiceover (cloned or cast, in the chosen tone) carries a numbered listicle while N web-animated hyperframe beats (HTML plus the Web Animations API, one branded design system of alternating tiles, big hero numerals, and callouts in the chosen visual style) are rendered frame-by-frame via Playwright and anchored to the VO's word-level timestamps, optional color-graded breather windows (brand clip, AI clip, or motion-graphic) give visual breath, and captions burn ONLY inside those B-roll windows (2-word chunks, ASS Format header carrying a Name field so none drop) with the VO mixed under a low music bed. This is the FREE deterministic assembly stage (Playwright beat render plus ffmpeg concat plus window-masked caption burn plus VO-and-music mix plus final composite) — the VO, the music bed, and any AI breather clips come from create-vo-elevenlabs, create-music-elevenlabs, and create-video-fal. Use for the vo-anchored-motion-listicle format.
npx gooseworks install --all # then, in Claude Code, Cursor, or Codex: /gooseworks use the render-vo-anchored-motion-listicle skill
Assemble a narrated motion-graphic listicle ad from a config: a spoken voiceover carries a numbered listicle (hook + N points + CTA) and every visual beat is anchored to
the VO's word-level timestamps. Each beat is a web-animated hyperframe (an HTML page + the Web
Animations API driven by window.renderAt(t)) rendered to video frame-by-frame with Playwright, all
in ONE branded design system (alternating background tiles, big hero numerals, body type, decorative
accents, callouts). Optional color-graded breather windows give visual breath, and
captions burn only on the B-roll windows. The shipped master is pure motion-graphic + VO — there
is NO lipsync (the still headshot is kept only for a future lipsync variant). This capability
is the FREE, deterministic assembly — the Playwright beat render, the ffmpeg concat, the
window-masked caption burn, the VO+music mix, and the final composite.
scripts/config.example.json is the worked example (Everself "doctor-educator" listicle, ~66s
1080×1920 9:16 at 25fps) — copy its structure, never its creative values; scripts/PIPELINE.md maps every config block to its source step and
scripts/README.md documents the free assembly.
The creative calls are made upstream (the format recipe's choices, asked of the user) and arrive in
the config; this assembly just renders what it is given. The demo's values are examples only.
voice.mode / voice.voice_id.
The demo used a cloned doctor-educator.beats[].
The demo used expert tips with mistake → fix beats.design.name / callout_style /
accents / background_tiles. The demo used EDUCATOR_MOTION with glass-pill callouts.music.prompt (none = mix the VO alone). The demo used a warm lo-fi
bed at ~0.18.This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are separate
capabilities — the spoken VO (create-vo-elevenlabs, a cloned or cast voice, eleven_v3 +
atempo) whose word-level timestamps (Groq whisper-large-v3 word-level) set the timeline; the low
music bed (create-music-elevenlabs); and any AI breather clips (create-video-fal, trimmed +
color-graded; brand clips and motion-graphic breathers are free). There is no free stock-footage
source.
Given the VO + words-flat.json + the N authored hyperframe beats + the color-graded B-roll windows +
the brand wordmark SVG, render-vo-anchored-motion-listicle renders each beat frame-by-frame via
Playwright (all beats at fps 25), concats the beats + B-roll, burns the window-masked captions, mixes
the VO under the low music bed, and composites → the master. Re-cuts reuse the existing VO / beats /
B-roll and cost $0.
window.renderAt(t); Playwright screenshots it frame-by-frame and
ffmpeg encodes it. This is NOT i2v — it is deterministic web motion graphics._shared.css (palette + type + alternating
tiles + accents + callout style) so N beats read as one designed reel; alternate only the background
tile, keep numerals / body / accents / callouts consistent.Format: header MUST carry a Name field — without it the leading-comma bug eats the
first field and captions silently drop. If the host ffmpeg lacks libass, render the cues as timed
PIL PNG overlays (ffmpeg overlay=…:enable='between(t,st,en)') at the same placement.none) sits ~0.18 vol under it. No ducking needed at that level.I=-14 → a 1080×1920 25fps h264 crf18 + aac 192k master. No paid calls, no keys.voice.mode / voice.voice_id.beats[].design.name / callout_style /music.prompt (none = mix the VO alone). The demo used a warm lo-fiFull video production sequence with script review, actual ingredient choices, controlled generation, editing, evidence-based quality review, polish, captions and delivery. A host binding supplies project storage, authentic approvals, provider access and billing.
Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
Write the words of a short-form video ad (voiceover, dialogue, chat bubbles, on-screen lines) the way performance creative teams do instead of from a blank page. Builds the script from the buyers' own words, the beat sheet of an ad that already works and three deliberately different angles, filters them with a rule check and a second non-Claude model, and takes the strongest into the review with the other two as one-line swaps. Use it in every video ad run before any paid step, and whenever the user asks to write, rewrite or improve a video ad script or says a script sounds generic or AI-written.