Assemble a cinematic live-action-style music-video ad from a config — an original sung anthem carries the whole narrative while N 35mm-film-look i2v clips are each cut to their lyric window and hard-concatenated on the beat as a 3-act arc, the anthem muxed at loudnorm I=-14, cinematic lower-third serif captions built from the song's OWN word timings (never Whisper) with the hook line landing on the chorus drop, and closed on a brand end card composited from the real asset — never AI-rendered text. This is the FREE deterministic assembly stage (cut-to-window + hard concat + anthem mux + captions + end card); the anthem, keyframes, and clips come from create-music-elevenlabs / create-image-gpt-image-fal / create-video-fal. Use for the cinematic-music-video format.
npx gooseworks install --all # then, in Claude Code, Cursor, or Codex: /gooseworks use the render-cinematic-music-video skill
Assemble a cinematic music-video ad from a config: a live-action-STYLE short film where an
original sung anthem is the score and every visual beat is a shot-on-film tableau (Kodak Portra
grain, light leaks, natural light, handheld imperfection) timed to the lyrics, arranged as a
3-act arc (its shape is the user's story_arc choice) with the hook line on the chorus drop. This capability is
the FREE, deterministic assembly — cut-to-window, hard-concat, anthem mux, caption burn,
and the brand end card.
scripts/config.example.json is the worked example (Hype and Vice "Game Day Girls", ~28s
1080×1920 9:16, 14 tableaux) — copy its structure, never its creative values; scripts/PIPELINE.md maps every config block to its source step
and scripts/README.md documents the free assembly.
This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are
separate capabilities: the sung anthem (create-music-elevenlabs, force_instrumental false —
the lyrics ARE the script, returns mp3 + words_timestamps); one 35mm-film keyframe per beat
in one look pack (create-image-gpt-image-fal); and one Kling 3.0 i2v clip per beat (create-video-fal).
Given the delivered anthem + words.json + one clip per beat + the brand end-card asset,
render-cinematic-music-video cuts each clip to its lyric window, hard-concats on the beat,
muxes the anthem, burns the cinematic lower-third captions, and overlays the end card → the
master. Re-cuts reuse the existing anthem / keyframes / clips and cost $0.
The creative content this stage assembles is decided upstream by the format's choices, asked
of the user before any paid step. The worked example's values are examples, never defaults:
This assembly reads them only through the config (tableaux[] windows + captions, the anthem,
captions.accent_words, end_card.*); it hardcodes none of them.
force_instrumental false); do not add a spoken voiceover or a
second bed.words.json from the
music model's words_timestamps, chunk ~4 words at lyric boundaries, and burn cinematic
lower-third serif captions (--placement low) with the hook line accent-treated (bold-italic).
Whisper on sung audio returns "🎵 Music Playing 🎵".is_hook) is timed so the
load-bearing line sits on the chorus drop; accent that line in the captions.afade in/out + loudnorm I=-14), burn the caption ASS, overlay
the end card → a 1080×1920 h264+aac master. No paid calls, no keys.Full video production sequence with script review, actual ingredient choices, controlled generation, editing, evidence-based quality review, polish, captions and delivery. A host binding supplies project storage, authentic approvals, provider access and billing.
Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
Write the words of a short-form video ad (voiceover, dialogue, chat bubbles, on-screen lines) the way performance creative teams do instead of from a blank page. Builds the script from the buyers' own words, the beat sheet of an ad that already works and three deliberately different angles, filters them with a rule check and a second non-Claude model, and takes the strongest into the review with the other two as one-line swaps. Use it in every video ad run before any paid step, and whenever the user asks to write, rewrite or improve a video ad script or says a script sounds generic or AI-written.
Teach your coding agent to build fast, cheap value prop video ads from a brand's logo and product images, using HTML instead of a video model.
An iMessage video ad shows a text conversation on an iPhone screen. Here is how a single skill teaches Claude to build one, with sound and an end card.
A seven-step workflow for turning a concept brief into a finished animated explainer ad with Claude Code, image models and ffmpeg.