Assemble a cosmic-mythology-voiceover reel from a config — one spoken voiceover carries the whole narrative while N curated stills in one chosen look are weighted beat-synced across the delivered VO duration (cut_dur = VO_dur times weight over the weight sum, so emotional beats hold longer), Ken-Burns-zoomed per still (scale 2x, center crop, zoompan, fade-in first and fade-out last), ffmpeg-concatenated, the VO composited under the picture (libx264 crf18 plus aac), the ONE on-screen hook line faded on over the open with a drawtext alpha window, and Whisper/VEED captions burned along the bottom — never in-world text on a still. This is the FREE deterministic assembly stage (weighted sequence plus Ken-Burns plus concat plus VO composite plus hook overlay plus caption burn); the VO and the stills come from create-vo-elevenlabs and create-image-fal. Use for the cosmic-mythology-voiceover format.
npx gooseworks install --all # then, in Claude Code, Cursor, or Codex: /gooseworks use the render-cosmic-mythology-voiceover skill
Assemble a cosmic-mythology-voiceover reel from a config: a faceless, cinematic storytelling video where one spoken voiceover carries the whole narrative over a slow, weighted Ken-Burns zoom across curated stills in ONE locked look, with ONE on-screen hook line and burned captions. The narrator, voice, tone, story shape, visual world and art style are the recipe's user choices — this capability only assembles what it is given (the demo: a warm contemplative "myth as teacher" reframe over deep-indigo + gold cosmic stills). This capability is the FREE, deterministic assembly — the weighted beat-sync sequencing, the Ken-Burns render, the concat, the VO composite, the hook overlay, and the caption burn.
scripts/config.example.json is the worked example — reference only, never the default
(WishAstro "Saturn isn't your villain", ~31s
1080×1920 9:16, 12 weighted Ken-Burns cuts); scripts/PIPELINE.md maps every config block to its
source step and scripts/README.md documents the free assembly.
There is a single runnable script — scripts/render.py (config-driven, ffmpeg + Pillow only, NO
API keys and NO drawtext/libass required):
python3 scripts/render.py --config config.json --vo working/vo2/vo_atempo.mp3 \
--stills-dir working/stills --out working/final.mp4 \
[--words working/vo2/words.json] [--endcard working/endcard.png]This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are
separate capabilities — the spoken VO (create-vo-elevenlabs, ElevenLabs eleven_v3 from a
tone-tagged script, atempo time-stretched so the delivered duration sets the timeline) and the 4–6
hero stills in one look pack (create-image-fal, Flux Pro 1.1, reused as repeats to reach the
~10–12 cuts). Given the VO + the stills + the per-cut weight array + the hook line, render.py
distributes the cuts across the VO duration by the weighted formula, Ken-Burns-renders each still,
concats, composites the VO, fades the hook line on over the open, burns the captions, and (if
--endcard is passed) appends a brand end card → the master + a poster. Re-cuts reuse the existing
VO / stills and cost $0. See scripts/README.md §0 for the full arg contract.
cut_dur = VO_dur × weight / Σweights —
heavier weights hold longer on the emotional beats (the open, the turn, the close); the setup
cuts run shorter. Every cut stays proportional to the whole VO.zoompan to the
configured zoom_end (~1.10); apply zoom_out on the flagged cuts; fade_in on the FIRST cut
and fade_out on the LAST. Stills are reusable — the sequence repeats a few across the cuts.render.py
does this with a PIL PNG + ffmpeg fade=…:alpha=1 (no drawtext dependency, since stock ffmpeg
often lacks it); an ffmpeg drawtext alpha window is an equivalent alternative where available.#FFFFFF. If the host ffmpeg lacks libass (no subtitles/ass filter),
render the cues as timed PIL PNG overlays (ffmpeg overlay=…:enable='between(t,st,en)') at the
same bottom placement — a free local Whisper + ffmpeg burn is the fallback to the VEED tier.crf 18 + aac 192k), burn the hook alpha-fade + the caption track →
a 1080×1920 h264+aac master (~31s). No paid calls, no keys (the local caption fallback is free).cut_dur = VO_dur × weight / Σweights —zoompan to theFull video production sequence with script review, actual ingredient choices, controlled generation, editing, evidence-based quality review, polish, captions and delivery. A host binding supplies project storage, authentic approvals, provider access and billing.
Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
Write the words of a short-form video ad (voiceover, dialogue, chat bubbles, on-screen lines) the way performance creative teams do instead of from a blank page. Builds the script from the buyers' own words, the beat sheet of an ad that already works and three deliberately different angles, filters them with a rule check and a second non-Claude model, and takes the strongest into the review with the other two as one-line swaps. Use it in every video ad run before any paid step, and whenever the user asks to write, rewrite or improve a video ad script or says a script sounds generic or AI-written.