Assemble a split-screen creator ad from a config — a two-zone vertical composite where a supplied AI-creator lip-sync take fills the BOTTOM ~48% while real 16:9 product/demo clips run uncropped in the TOP ~52%, each top clip contain-fit with a darkened blurred cover-scale fill of the same clip (never black bars), a 3px brand-color divider between the zones, the creator slice cover-fit per the per-scene VO timing, scenes hard-concatenated with the body audio being the concatenated creator VO slices, an end card held on the last sharp frame ~3s, then the ASSEMBLED cut transcribed with local Whisper (not the raw VO — concat drops inter-scene silence) and word-level captions burned in the chosen style. This is the FREE deterministic assembly + caption stage (two-zone composite + blurred fill + divider + hard-concat + end card + captions); the VO comes from create-vo-elevenlabs, the anchor from create-image-gpt-image-fal, and the whole-VO lip-sync from a paid VEED Fabric 1.0 take (a no-atom upstream input). Use for the split-screen-creator format.
npx gooseworks install --all # then, in Claude Code, Cursor, or Codex: /gooseworks use the render-split-screen-creator skill
Assemble a split-screen creator ad from a config: a two-zone vertical (1080×1920, 9:16, ~40s) format where an AI creator talking-head anchors the BOTTOM ~48% of the frame and real 16:9 product/demo clips run uncropped in the TOP ~52%, a 3px brand-color divider between the zones. The creator delivers the whole VO cold-to-camera and each top clip proves the claim its VO line makes. This capability is the FREE, deterministic assembly + captions — the two-zone composite (contain-fit + blurred-cover fill + divider + creator slice), the hard-concat, the end card, and the word-level caption burn from the assembled cut.
scripts/config.example.json is one worked example (Perplexity concept-10
"Bloomberg terminal", ~40s 1080×1920 9:16, 6 scenes + an end card) — copy its
structure, never its creative values;
scripts/PIPELINE.md maps every config block to its source step and
scripts/README.md documents the free assembly.
The creative calls are made upstream by the user (the calling format's recipe
choices) and arrive here as config. This capability never picks them.
creator_gender / creator_age / creator_look — who the creator is
(creator.brief); shapes the anchor and the voice upstream. Here it only
affects framing checks. The worked example used a man in his 20s.setting — where the creator films from (creator.brief.vibe_setting).
The worked example used a lived-in room with window daylight.tone — the script and VO delivery; upstream only. The worked example
was a confident, fast explainer.caption_style — captions.style (serif-accent, kinetic-pop,
neon-glow, clean-bubble). Burned here. The worked example shipped
serif-accent.If a creative field is empty in the config, ask for it — don't fall back to the worked example's value.
This is the FREE, deterministic assembly + captions stage — it spends
nothing. The paid inputs are separate steps — the VO (create-vo-elevenlabs,
ElevenLabs eleven_v3 with-timestamps, sliced into per-scene windows), the
photoreal MEDIUM chest-up AI-creator anchor (create-image-gpt-image-fal,
gpt-image-2) — shot at a natural webcam distance (headroom + shoulders, real
room), not a plain-background close-up headshot (see the anchor note below), and
the whole-VO lip-sync (a paid VEED Fabric 1.0 @ 720p take — a no-atom step,
image_url = the anchor, audio_url = the vo mp3, ~$0.15/sec, ~$5.90 for a 39s
VO; run its calls sequentially, veed/fabric-1.0 storage-auths 403 under
parallel load). Given the creator lip-sync take + the per-scene VO timing + one
16:9 top clip per scene + the scene-1 hook graphic + the end-card clip,
render-split-screen-creator composites the two zones, hard-concats the scenes,
appends the end card, transcribes the assembled cut, and burns the captions → the
master. Re-cuts reuse the existing VO / lip-sync / clips and cost $0.
top_height ~998. Keep every stacked height EVEN
(998 + 4 divider + 918 = 1920) — libx264 rejects odd dimensions.ugc-walk-and-talk). VEED Fabric handles photoreal fine (unlike Seedance).top_start/top_end) to the on-message segment that proves its VO line.
Never loop a short clip — set the window and the assembler speed-fits it to
the scene (looping replays into a sparse/black tail).timing.json; the
lip-sync drives the mouth.endcard.clip_end),
not the black tail.captions.style (serif-accent, kinetic-pop, …). Keep the -precaption cut + the
.ass sidecar so captions restyle without re-rendering the composite. If the
host ffmpeg lacks libass, render the cues as timed PIL PNG overlays (ffmpeg
overlay=…:enable='between(t,st,en)') at the same placement.loudnorm I=-14 → a
1080×1920 h264+aac master. No paid calls, no keys — the VEED Fabric lip-sync is
a supplied input, produced upstream.creator_gender / creator_age / creator_look* — who the creator issetting* — where the creator films from (creator.brief.vibe_setting).tone* — the script and VO delivery; upstream only. The worked examplecaption_style* — captions.style (serif-accent, kinetic-pop,Full video production sequence with script review, actual ingredient choices, controlled generation, editing, evidence-based quality review, polish, captions and delivery. A host binding supplies project storage, authentic approvals, provider access and billing.
Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
Write the words of a short-form video ad (voiceover, dialogue, chat bubbles, on-screen lines) the way performance creative teams do instead of from a blank page. Builds the script from the buyers' own words, the beat sheet of an ad that already works and three deliberately different angles, filters them with a rule check and a second non-Claude model, and takes the strongest into the review with the other two as one-line swaps. Use it in every video ad run before any paid step, and whenever the user asks to write, rewrite or improve a video ad script or says a script sounds generic or AI-written.