Assemble a multi-scene GRWM beauty-demo ad from a config — a locked-identity creator applies ~5 products step by step while a SEPARATE ElevenLabs voiceover narrates and every scene cut is snapped to the VO's product-name word-starts (Whisper word-level timestamps), then ~5 Playwright product overlay cards (real PDP-verified taglines) are composited onto the master each on its product-NAME word-start, the SEPARATE VO is mixed on top of a ducked music bed at loudnorm I=-14, clean-white 3-words/cue captions are burned, and the video closes on a flat-lay end card. This is the FREE deterministic assembly stage (re-cut to the VO word-starts, hard-concat, Playwright card render + card composite, VO plus music mix, caption burn, flat-lay end card); the VO, scene clips, product cutouts, and music come from create-music-elevenlabs / create-image-gpt-image-fal / create-video-fal. Use for the glassy-matte-grwm format.
npx gooseworks install --all # then, in Claude Code, Cursor, or Codex: /gooseworks use the render-glassy-matte-grwm skill
Assemble a multi-scene GRWM beauty-demo ad from a config — a locked-identity creator applies ~5 makeup/skincare products step by step in a chosen setting, a separate ElevenLabs voiceover narrates the routine, and every scene cut is snapped to the VO's product-name word-starts, with ~5 Playwright product overlay cards on the product-name beats, a ducked music bed, burned captions, and a flat-lay end card. This capability is the FREE, deterministic assembly — the Whisper-driven re-cut + hard-concat, the Playwright card render + card composite, the VO + music mix, the caption burn, and the flat-lay end card.
This is the multi-scene beauty demo, distinct from the single-take apparel outfit-reveal
(ugc-grwm, one Seedance reference-to-video call with native lip-sync and minimal post). Here the
timeline is driven by a SEPARATE VO and the scenes are re-cut to its word-starts.
scripts/config.example.json is the worked example (DIBS Beauty "5-Step Glassy Matte Routine",
~32s 1080×1920 9:16, 12 VO-snapped cuts + 5 product cards); scripts/PIPELINE.md maps every
config block to its source step and scripts/README.md documents the free assembly.
The creative calls come from the recipe's choices, asked of the user before any paid step. The
demo's picks are examples, never defaults. This capability only assembles what those choices
produced; it hardcodes none of them.
assets.creator_anchor, the look in the scene beats). The
demo used a young woman.vo.voice_id / vo.voice_name). The demo used ElevenLabs Karri.assets.vanity_world). The demo used a sage-green vanity.vo.script, vo.style). The demo was upbeat and friendly.music.prompt; "none" = VO only, skip the music mix). The demo
used a warm acoustic bed.Card colours come from palette (the brand kit), not the demo's cream + pink.
This is the FREE, deterministic assembly stage — it spends nothing. The paid inputs are
separate capabilities — the SEPARATE narration VO (create-music-elevenlabs, or a user-supplied
mp3; word-level Whisper timestamps set the timeline), ~7 Seedance scene clips one per product step
(create-video-fal), the ~5 white-bg product cutouts + the flat-lay end-card still
(create-image-gpt-image-fal), and the ducked music bed. Given the VO + .words.json + one clip
per step + the ~5 product cutouts + the music bed, render-glassy-matte-grwm re-cuts each clip to
its VO word-start window, hard-concats on the cut, renders + composites the product cards on the
product-name beats, mixes the VO over the ducked music, burns the captions, and appends the
flat-lay end card → the master. Re-cuts reuse the existing VO / clips / cutouts and cost $0.
-c:v libx264 -crf 20 — -c copy corrupts the duration when zoompan/PNG clips are in the
chain.-loop 1 -t <dur> — without it the PNG emits one frame at t=0 and the fade/enable filters
silently no-op (cards go invisible).loudnorm I=-14. When choices.music is "none", loudnorm the VO alone. If the host ffmpeg lacks a filter, apad/atrim to length before the mix.overlay=…:enable='between(t,st,en)' at the same
placement.assets.creator_anchor, the look in the scene beats). Thevo.voice_id / vo.voice_name). The demo used ElevenLabs Karri.assets.vanity_world). The demo used a sage-green vanity.vo.script, vo.style). The demo was upbeat and friendly.music.prompt; "none" = VO only, skip the music mix). The demoFull video production sequence with script review, actual ingredient choices, controlled generation, editing, evidence-based quality review, polish, captions and delivery. A host binding supplies project storage, authentic approvals, provider access and billing.
Build a vox-pop street interview video ad. An interviewer with a handheld mic asks passers-by one question about the brand's product, they give blunt wrong guesses, one gives the real answer, and the cut lands on a branded end card. Generates the takes through the GooseWorks fal proxy (Seedance 2.0 with native voice), then grades, re-cuts, captions and gates them locally. Use for the street-interview format.
Write the words of a short-form video ad (voiceover, dialogue, chat bubbles, on-screen lines) the way performance creative teams do instead of from a blank page. Builds the script from the buyers' own words, the beat sheet of an ad that already works and three deliberately different angles, filters them with a rule check and a second non-Claude model, and takes the strongest into the review with the other two as one-line swaps. Use it in every video ad run before any paid step, and whenever the user asks to write, rewrite or improve a video ad script or says a script sounds generic or AI-written.