Assemble an editorial-motion podcast-clip ad from a config — a real clipped podcast MP3 carries the narrative while N flat limited-palette editorial-illustration keyframes (one look pack) are animated NOT by generative i2v but by DETERMINISTIC ffmpeg ken-burns (zoompan) + hard cuts (no crossfades, which expose geometric drift), each beat snapped to its spoken line, the real audio muxed, Whisper-driven captions burned only mid-sentence, and closed on a PIL brand end card — never AI-rendered text. This is the FREE deterministic assembly stage (ffmpeg ken-burns + hard concat + audio mux + captions + end card); the real audio is clipped from source and the keyframes come from create-image-fal. Use for the editorial-motion-podcast format.
npx gooseworks install --all # then, in Claude Code, Cursor, or Codex: /gooseworks use the render-editorial-motion-podcast skill
Assemble an editorial-motion podcast-clip ad from a config: a real clipped podcast audio line carries the whole narrative and every visual beat is timed to the sentence it describes, in ONE flat, strictly limited-palette editorial-illustration look pack ("a magazine spot-illustration that moves" — the style and palette are the caller's choice). The motion is not generative video but deterministic ffmpeg ken-burns on static keyframes, so it reads as a printed page that moves. This capability is that FREE, deterministic assembly — the ffmpeg motion, hard-concat, audio mux, caption burn, and PIL end card.
scripts/config.example.json is the worked example (Klarify "Rat Park", ~40.8s 1080×1920
9:16, 6 beats — its 2-tone Niemann look, cream/charcoal/sage palette and Rat Park metaphor are
that demo's picks, never defaults); scripts/PIPELINE.md maps every config block to its source step and
scripts/README.md documents the free assembly.
The creative calls are the caller's (the video-format recipe asks the user); this assembly never picks them. The demo's value is an example only:
This is the FREE, deterministic assembly stage — it spends nothing on the motion layer.
The paid inputs are separate: the real podcast MP3 is clipped from source (free ffmpeg) with
its Whisper word timings, and one editorial-illustration keyframe per beat (chained ref images
so cage/character geometry holds) comes from create-image-fal (Nano Banana). Given the
clipped audio + words.json + the per-beat keyframes + the real brand wordmark PNG,
render-editorial-motion-podcast renders each keyframe as a ken-burns segment, hard-concats
on the beat, muxes the real audio, burns the mid-sentence captions, and composites the PIL end
card → the master. Re-cuts reuse the existing audio / keyframes and cost $0.
-map 0:v:0 -map 1:a:0) — a real clipped podcast line (preferred) OR an
approved generated VO (create-vo-elevenlabs). Never a sung/generated track. (Clip-vs-generate
is the recipe's STEP-0 intake decision — if no source episode is supplied, ASK the user.)zoompan (push-in / pull-back, 1.0→~1.06×, 24fps); Seedance/Kling are photoreal-trained
and invent naturalistic middle states that collapse the flat limited-palette look look. Never -loop 1 with
zoompan d=N (it balloons the duration); feed a single image and clamp with -t + trim.frosted-subtle
captions while the speaker talks; leave silent/reflective beats and the end card uncaptioned.
THREE mandatory rules (each bit us in prod — bake them in):
end = min(last_word_end + ~0.15, next_start - 0.03)). Two boxes must never stack at the
same spot; an end-tail bleeding into the next window is the #1 caption bug.look_pack.caption_safe_area). If a finished keyframe's subject
intrudes into the caption band, deterministically shift the subject UP into the empty top
space (PIL: paste up ~0.24H onto a canvas pre-filled with the exact paper color from a clean
corner) — never let the box sit on the subject.ass/subtitles filter), but check ffmpeg -filters
first: many builds (Homebrew) lack libass/drawtext. If absent, use the deterministic
overlay fallback — render each line as a transparent PNG (frosted rounded box + white
text, PIL) and composite via the ffmpeg overlay filter with timed
enable='between(t,st,en)' windows. Same look, no libass.Full video production sequence with script review, actual ingredient choices, controlled generation, editing, evidence-based quality review, polish, captions and delivery. A host binding supplies project storage, authentic approvals, provider access and billing.
Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
Write the words of a short-form video ad (voiceover, dialogue, chat bubbles, on-screen lines) the way performance creative teams do instead of from a blank page. Builds the script from the buyers' own words, the beat sheet of an ad that already works and three deliberately different angles, filters them with a rule check and a second non-Claude model, and takes the strongest into the review with the other two as one-line swaps. Use it in every video ad run before any paid step, and whenever the user asks to write, rewrite or improve a video ad script or says a script sounds generic or AI-written.