Assemble a song-driven music-video ad from a config — a generated sung track carries the whole narration across N tableaux (one keyframe -> one i2v clip per lyric beat) with NO separate voiceover, captions synced to the song's OWN word timings (script-window, never Whisper) and the hook word landing on the chorus drop, closed on a PIL brand end card. This is the FREE deterministic assembly stage (clip cut-to-timeline + captions + end card + FFmpeg composite); the song, keyframes, and clips come from create-music-elevenlabs / create-image-fal / create-video-fal. Use for the song-driven-music-video format.
npx gooseworks install --all # then, in Claude Code, Cursor, or Codex: /gooseworks use the render-song-mv skill
Assemble a song-driven music-video ad from a config: a purpose-written, sung song is
the entire narration (no separate voiceover), and every visual beat is timed to the lyrics.
The delivered song sets the timeline; N tableaux (one keyframe → one image-to-video clip per
lyric beat, all in a single look pack) are cut to their lyric windows and hard-concatenated
on the beat, captions are built from the song's OWN word timings with the hook line landing
on the chorus drop, and the spot closes on a PIL brand end card. It reads like a tiny animated
music video, not a demo. scripts/config.example.json is one worked example (Loóna "Fall In
Love With Sleep Again", 28s paper-craft 9:16) — its style, song, vocalist, mood, protagonist and
setting are that demo's answers, not defaults; scripts/PIPELINE.md maps every config block
to its step and scripts/README.md documents the free assembly.
The creative calls are made upstream by the user (the recipe's choices) and arrive in the config;
this assembly never picks them.
look_pack, clip_engine.motion_opener. The demo used a paper-craft diorama
at night.song.prompt / song.bpm. The demo used a dreamy synth lullaby at ~80 BPM.song.prompt. The demo used a soft breathy female lead.song.prompt, the lyrics, tableaux. The demo went calm/dreamy.look_pack.style_opener. The demo used a young-woman paper-doll.tableaux[].keyframe_prompt. The demo went bedroom at night → moonlit paper village.This is the FREE, deterministic assembly stage — it spends nothing. The three paid
inputs are separate capabilities: the sung song (create-music-elevenlabs, music_v1,
force_instrumental FALSE — the lyrics ARE the script, returns mp3 + words.json), one
keyframe per tableau (create-image-fal), and one Kling 3.0 i2v clip per tableau
(create-video-fal). Given the delivered song + words.json + one clip per beat,
render-song-mv cuts each clip to its lyric window, hard-concats on the beat, builds the
lyric-synced captions, composites the PIL end card, and muxes → the master. Re-cuts reuse
the existing song / keyframes / clips and cost $0.
force_instrumental false); do not add a spoken voiceover or a
second music bed.timeline.json) — never trim the song to a pre-planned grid.audio/words.json (~3 words at lyric boundaries); accent words get captions.accent_color.
Whisper on sung audio returns "🎵 Music Playing 🎵", so it can't caption lyrics.is_hook) is timed so the
payoff word (song.hook_word) sits on the chorus drop; accent that word in the captions.style_opener + negative_tail + palette drives
every keyframe so N beats read as one film; no morph within a clip.audio_mix.climax_beat_id — set it to
this run's hook tableau; the demo's T08 is example-only), loudnorm to −14 LUFS → 1080×1920
h264+aac. No paid calls, no keys.look_pack, clip_engine.motion_opener. The demo used a paper-craft dioramasong.prompt / song.bpm. The demo used a dreamy synth lullaby at ~80 BPM.song.prompt. The demo used a soft breathy female lead.song.prompt, the lyrics, tableaux. The demo went calm/dreamy.look_pack.style_opener. The demo used a young-woman paper-doll.Full video production sequence with script review, actual ingredient choices, controlled generation, editing, evidence-based quality review, polish, captions and delivery. A host binding supplies project storage, authentic approvals, provider access and billing.
Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
Write the words of a short-form video ad (voiceover, dialogue, chat bubbles, on-screen lines) the way performance creative teams do instead of from a blank page. Builds the script from the buyers' own words, the beat sheet of an ad that already works and three deliberately different angles, filters them with a rule check and a second non-Claude model, and takes the strongest into the review with the other two as one-line swaps. Use it in every video ad run before any paid step, and whenever the user asks to write, rewrite or improve a video ad script or says a script sounds generic or AI-written.