Assemble an absurdist animated-explainer video ad (~38s, 9:16) from per-scene i2v clips + their measured VO windows — retime each clip to its VO, re-encode every segment to identical 30fps/libx264/yuv420p so the concat demuxer never drops frames, concat, build a REAL-product PIL end card (never AI) with a slow Ken-Burns, mix VO (loudnorm I=-14) under music (loudnorm I=-26, volume 0.62, amix normalize=0), and burn libass captions last. FREE deterministic assembly (bash-free, Python + ffmpeg + PIL); the recipe supplies the clips, VO, music, product photo, palette, and caption table and gates the paid keyframe/clip/VO/music calls to their own capabilities. Use for the absurdist-explainer format.
npx gooseworks install --all # then, in Claude Code, Cursor, or Codex: /gooseworks use the render-absurdist-explainer skill
The free, deterministic renderer for the absurdist-explainer video ad format — an animated spot that explains one problem and how the product fixes it. The format is the mechanics: show the problem in a funny, exaggerated cartoon way, teach the product's ownable mechanism visually, show what the problem costs, bring the product in, land the fix as the climax, pay it off, end on a real-product card. One narrator voice carries the whole spot. How the story is told (problem → fix, how it works, a day in the life, the product as hero, or a villain arc) is the user's choice, made upstream. This capability is the FREE assembly stage only. All generative work (nano-banana keyframes, Seedance i2v clips, ElevenLabs VO + music) happens upstream in the recipe and is handed to this capability as files.
It ports the validated compose recipe from two reference runs (a cortisol/stress supplement and a baby-eczema cream; both happened to use the villain arc). The recipe is deterministic — iterate the cut for free, re-roll only the offending paid beat.
The creative calls are made upstream by the recipe's choices and arrive here only as
files and config values. This renderer is story- and style-agnostic: it never assumes a
story shape, a cast, a look, a narrator or a music style.
story_shape — how the story is told: problem → fix (no villain) / how it works /
a day in the life / the product as hero / villain arc (the problem as a cartoon villain
who schemes and loses). Reaches this capability as the per-scene caption text and the
VO files. Asked of the user; the demo used the villain arc.cast — who the cartoon characters are (the customer, the product as a character,
a mascot; a personified problem only on the villain arc). Reaches this capability only
inside the clips. Asked of the user.narrator — who speaks the single VO track. Arrives as the scenes[].vo files.
Asked of the user; the demo used the villain (villain arc only).visual_style — the art style of the i2v clips (scenes[].clip). Asked of the
user; the demo used Pixar-style 3D. Nothing here depends on it.narrator_voice — the voice cast for the VO. Asked of the user; the demo used a
characterful male villain voice.music — the bed at music_bed. Asked of the user; the demo used whimsical
pizzicato + woodwinds + xylophone. The mix constants below apply to any style.scale=1080:1920:force_original_aspect_ratio=increase,crop=1080:1920,fps=30,setsar=1,
then tpad=stop_mode=clone if the VO is longer than the clip, else -t trim).libx264 -crf 18 -pix_fmt yuv420p -r 30 even if already correct — a framerate mismatch makes the concat demuxer silently
drop frames.-c copy).build_endcard.py composites the REAL retail product photo
over the brand palette (flat, or sampled from the photo's own edge pixel) with a typeset
wordmark + claim rows + CTA pill in PIL ImageDraw.text — never an AI cartoon bottle,
never AI-rendered brand text. compose.py Ken-Burnses it 1.00 → 1.04 over the dwell.loudnorm I=-14 TP=-1.5, music bus loudnorm I=-26 TP=-3 then
volume=0.62, amix inputs=2 duration=first normalize=0 → master lands at
-14.5..-13.5 LUFS with the music ducked under the VO.make_captions.py emits a libass .ass (one cue per scene, Arial 64
white / 6px outline / MarginV=330, start = scene_start + 0.08s, suppressed on the end
card). compose.py burns it as the final filter so captions sit on top.scripts/build_endcard.py — PIL composite of the real product photo + typeset brand
layer (wordmark / product line / claim rows / accent CTA pill). Reads the same
config.json. Run this FIRST so end_card.image exists before compose.py.scripts/make_captions.py — emits the per-scene libass .ass from the SAME scene table
compose reads, so caption windows stay in lockstep with the cut. Run before compose.py
(or point config.captions_ass at nothing to skip captions).scripts/compose.py — the assembler: per-scene retime + identical 30fps re-encode →
concat → Ken-Burns end card → VO/music loudnorm mix → burn captions → master mp4.scripts/config.example.json — the shape of the config the recipe binds. Its values
are a labelled worked example (the demo build: a villain-arc eczema story, placeholder
brand).
Captions, end-card copy and palette come from the user's brand and choices — never copy
them as defaults.--config + a runtime work dir — NO hardcoded paths)config.json carries: scenes[] (each {id, clip, target_sec, vo, caption, atempo?}
where target_sec is the measured VO window), end_card{product_image, image, dwell_sec, zoom_to, wordmark, product_line, claims[], cta, background?}, brand_palette {primary, primary_lite, accent, grey}, music_bed, music_volume (default 0.62),
atempo (compose-stage VO speed-up, default off; the reference runs used 1.3 when the VO
read slow), captions_ass, and caption_style. See config.example.json.
ImageDraw.text. AI draws the world + characters only.target_sec is the MEASURED VO duration (ffprobe each VO mp3), never a planned word
count — VO drives the per-scene timing.volume≈0.62 to 0.70
(the two reference runs), amix normalize=0. Master target -14.5..-13.5 LUFS,
true-peak ≤ -1.5 dBFS.start = scene_start + 0.08s, suppressed on the end card (its typeset copy
carries the message — two text layers at one spot are both unreadable).watch (QC the final master — confirm every character's look holds, the single
narrator voice carries the whole spot, the motif lands ≥3×, no AI brand text leaked into a cartoon
background, the end card is the real product, and duration is within ±0.1s of the summed
windows). The recipe gates the paid create-image-fal (keyframes), create-video-fal
(Seedance i2v), create-vo-elevenlabs, and create-music-elevenlabs calls to their own
capabilities — this capability itself makes NO paid calls.
story_shape* — how the story is told: problem → fix (no villain) / how it works /cast* — who the cartoon characters are (the customer, the product as a character,narrator* — who speaks the single VO track. Arrives as the scenes[].vo files.visual_style* — the art style of the i2v clips (scenes[].clip). Asked of thenarrator_voice* — the voice cast for the VO. Asked of the user; the demo used aFull video production sequence with script review, actual ingredient choices, controlled generation, editing, evidence-based quality review, polish, captions and delivery. A host binding supplies project storage, authentic approvals, provider access and billing.
Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
Write the words of a short-form video ad (voiceover, dialogue, chat bubbles, on-screen lines) the way performance creative teams do instead of from a blank page. Builds the script from the buyers' own words, the beat sheet of an ad that already works and three deliberately different angles, filters them with a rule check and a second non-Claude model, and takes the strongest into the review with the other two as one-line swaps. Use it in every video ad run before any paid step, and whenever the user asks to write, rewrite or improve a video ad script or says a script sounds generic or AI-written.
A seven-step workflow for turning a concept brief into a finished animated explainer ad with Claude Code, image models and ffmpeg.