capabilities

Render VO Anchored Motion Listicle

Assemble an expert/educator motion-graphic LISTICLE video ad from a config — a spoken authoritative voiceover carries a numbered listicle while N web-animated hyperframe beats (HTML plus the Web Animations API, one branded design system of alternating tiles, big hero numerals, and glass-pill callouts) are rendered frame-by-frame via Playwright and anchored to the VO's word-level timestamps, periodic color-graded B-roll windows give visual breath, and captions burn ONLY inside those B-roll windows (2-word chunks, ASS Format header carrying a Name field so none drop) with the VO mixed under a low music bed. This is the FREE deterministic assembly stage (Playwright beat render plus ffmpeg concat plus window-masked caption burn plus VO-and-music mix plus final composite) — the VO, the music bed, and the stock B-roll come from create-vo-elevenlabs, create-music-elevenlabs, and media-proxy. Use for the vo-anchored-motion-listicle format.

Gooseby Athina AI
Install
Terminal
npx gooseworks install --all

# then, in Claude Code, Cursor, or Codex:
/gooseworks use the render-vo-anchored-motion-listicle skill
About This Skill

render-vo-anchored-motion-listicle

Assemble an expert/educator motion-graphic listicle ad from a config: an authoritative spoken voiceover carries a numbered listicle (hook + N points + CTA) and every visual beat is anchored to the VO's word-level timestamps. Each beat is a web-animated hyperframe (an HTML page + the Web Animations API driven by window.renderAt(t)) rendered to video frame-by-frame with Playwright, all in ONE branded design system (alternating background tiles, big hero numerals, body type, decorative SVG accents, glass-pill callouts). Periodic color-graded B-roll windows give visual breath, and captions burn only on the B-roll windows. The shipped master is pure motion-graphic + VO — there is NO lipsync (the still expert headshot is kept only for a future lipsync variant). This capability is the FREE, deterministic assembly — the Playwright beat render, the ffmpeg concat, the window-masked caption burn, the VO+music mix, and the final composite.

scripts/config.example.json is the worked example (Everself "doctor-educator" listicle, ~66s 1080×1920 9:16 at 25fps); scripts/PIPELINE.md maps every config block to its source step and scripts/README.md documents the free assembly.

Run

This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are separate capabilities — the spoken VO (create-vo-elevenlabs, a cloned or cast expert voice, eleven_v3 + atempo) whose word-level timestamps (Groq whisper-large-v3 word-level) set the timeline; the low music bed (create-music-elevenlabs); and the stock B-roll (media-proxy, trimmed + color-graded). Given the VO + words-flat.json + the N authored hyperframe beats + the color-graded B-roll windows + the brand wordmark SVG, render-vo-anchored-motion-listicle renders each beat frame-by-frame via Playwright (all beats at fps 25), concats the beats + B-roll, burns the window-masked captions, mixes the VO under the low music bed, and composites → the master. Re-cuts reuse the existing VO / beats / B-roll and cost $0.

Contract (the free assembly)

  • The spoken VO carries the listicle — it sets the timeline. The cloned/cast expert VO is the spine; Whisper-transcribe it to word-level timestamps and anchor every beat reveal to those word times. There is no on-camera human and NO lipsync in the shipped master.
  • Beats are web-animated hyperframes, rendered deterministically. Each beat is an HTML page + the Web Animations API driven by window.renderAt(t); Playwright screenshots it frame-by-frame and ffmpeg encodes it. This is NOT i2v — it is deterministic web motion graphics.
  • ALL beats share the same fps (25). A mismatched-fps beat stutters at the concat seam. Render every beat and every B-roll window at fps 25.
  • ONE design system across every beat. A single _shared.css (palette + type + alternating tiles + accents + glass-pill) so N beats read as one designed reel; alternate only the background tile, keep numerals / body / accents / pills consistent.
  • Periodic color-graded B-roll windows for breath. Trim + color-grade each stock/brand clip to the palette (fps 25). These windows are the ONLY captioned windows.
  • Captions ONLY on the B-roll windows. On the motion-graphic beats the on-screen type IS the caption — burning Whisper captions there double-stacks text. Build the caption ASS from the VO word timings, kept only inside the B-roll windows, 2-word chunks, closing a cue on any >0.4s word gap. The ASS Format: header MUST carry a Name field — without it the leading-comma bug eats the first field and captions silently drop. If the host ffmpeg lacks libass, render the cues as timed PIL PNG overlays (ffmpeg overlay=…:enable='between(t,st,en)') at the same placement.
  • Low music bed under the VO. The VO is the spine; the ElevenLabs Music bed sits ~0.18 vol under it. No ducking needed at that level.
  • Never AI-render the brand lockup. The CTA / end beat composites the brand's real wordmark, never text-in-diffusion.
  • FFmpeg composite, deterministic, FREE. Render each beat via Playwright, concat the beats + B-roll (ffmpeg demuxer), burn the window-masked caption ASS, mix the VO under the music bed, loudnorm I=-14 → a 1080×1920 25fps h264 crf18 + aac 192k master. No paid calls, no keys.

What's included

·
The spoken VO carries the listicle — it sets the timeline.* The cloned/cast expert VO is the
·
Beats are web-animated hyperframes, rendered deterministically.* Each beat is an HTML page + the
·
ALL beats share the same fps (25).* A mismatched-fps beat stutters at the concat seam. Render
·
ONE design system across every beat.* A single _shared.css (palette + type + alternating
·
Periodic color-graded B-roll windows for breath.* Trim + color-grade each stock/brand clip to
You Might Also Like

Render Stopmotion Hand Swatch Cycle

Assemble a stop-motion hand-swatch-cycle product-demo ad from a config — a sequence of still PLATES (one hand swiping a single-barrel cosmetic across a cream skin-patch, the barrel + swatch changing per plate while the hand, background, crop, and lighting stay locked) is PNG→mp4 loop-encoded at each plate's own stop-motion hold (fast motion frames 150–250ms, per-shade ~380ms, hero beats 1100–1800ms), concat-demuxed with HARD cuts into a silent master, closed on a Playwright HTML-rendered branded end card (serif tagline + sans subtitle + real logo SVG over a hero BG, never AI-rendered text), and muxed with a pre-sourced music track playing under the end card with a fade tail (no VO). This is the FREE deterministic assembly stage (loop-encode + concat-demux + end-card render + music mux); the master-anchor plate, shade plates, and end-card BG come from create-image-gpt-image-fal and the track from create-music-elevenlabs. Use for the stopmotion-hand-swatch-cycle format.

Render Split Screen Creator

Assemble a split-screen creator ad from a config — a two-zone vertical composite where a supplied AI-creator lip-sync take fills the BOTTOM ~48% while real 16:9 product/demo clips run uncropped in the TOP ~52%, each top clip contain-fit with a darkened blurred cover-scale fill of the same clip (never black bars), a 3px brand-color divider between the zones, the creator slice cover-fit per the per-scene VO timing, scenes hard-concatenated with the body audio being the concatenated creator VO slices, an end card held on the last sharp frame ~3s, then the ASSEMBLED cut transcribed with local Whisper (not the raw VO — concat drops inter-scene silence) and word-level captions burned in the chosen style. This is the FREE deterministic assembly + caption stage (two-zone composite + blurred fill + divider + hard-concat + end card + captions); the VO comes from create-vo-elevenlabs, the anchor from create-image-gpt-image-fal, and the whole-VO lip-sync from a paid VEED Fabric 1.0 take (a no-atom upstream input). Use for the split-screen-creator format.

Render Narrated UGC Wardrobe Stitch

Assemble a narrated-UGC "stitch reply" ad from a config — a single spoken VO carries a verbatim testimonial while ~30 per-cut i2v clips (one creator across ~5 wardrobes in ~3 worlds, plus product B-roll) are each trimmed to their EDL window built from the VO's Whisper word boundaries and hard-concatenated via filter_complex concat (never the demuxer, which drops audio on a duration mismatch), the VO mixed over an optional sidechain-ducked instrumental bed (−20dB, 20 to 1) so the VO stays on top, karaoke-pop captions burned on every word throughout (VEED Whisper preset, re-spelled against the locked script), a landing-page scroll rendered as FFmpeg zoompan over a Playwright PNG (not i2v), and closed on the brand's real end-card PNG — never AI-rendered text. This is the FREE deterministic assembly stage (trim-to-EDL + filter_complex concat + VO and music mix + karaoke captions + landing-page zoompan + end-card append); the VO, creator, start-frames, and clips come from create-vo-elevenlabs / create-image-gpt-image-fal / create-image-fal / create-video-fal. Use for the narrated-ugc-wardrobe-stitch format.

Newsletter

Learn to build Growth systems with AI

2-3 compounding systems per week using Claude Code, OpenClaw, and more.