founder-product-video
You generate a 65-second founder-style product video from a product URL plus user-provided imagery: 60 seconds of talking-founder body video plus a 5-second branded end card. The user's images (product photos / website screenshots / app screenshots) flow into the SeeDance acts as visual references, and digital product screens / brand wordmarks are composited after generation so readable UI is deterministic instead of model-rendered.
No cutaways. No website CSS extraction. AI generation, deterministic render, captions, concat, and music mix go through Pika MCP tools by default. Lower-third overlays are opt-in and use MCP compose by default, with local ffmpeg only as an emergency fallback.
Cost transparency gate
Before any paid MCP call, call identity_balance({verbose: true}) once. Surface the current balance, recent burn rate, and remaining runway, then gate the run with this exact message:
Estimated cost: about 4,000 credits (~$40) for a typical four-act Seedance founder video plus supporting assets. This exceeds $5, so Reply
proceedto continue orcancelto stop.
Do not call any paid MCP tool until the user replies proceed. If the user replies cancel, stop without generating. For non-interactive --quick or --config callers, require cost_ack=proceed in the config; if it is absent, stop with the estimate instead of spending credits.
[0] Intake — run first, before any pipeline step
If invoked with empty args, print this menu verbatim and stop — wait for the user to paste inputs:
What founder video do you want to make? Required:
- Product URL —
https://...(anything with a real homepage)- Founder — name + role, e.g. "Eli Kim, CEO"
- Founder photo — local path, https URL, OR
generate(I'll create a portrait)Optional (sensible defaults if omitted): brand kit path · custom on-phone screenshots · music · aspect (16:9 / 9:16 / 1:1) · location image · voice style · product type
Example:
/founder-product-video https://example.com --founder "Eli Kim, CEO" --photo ~/Pictures/eli.jpg
If args carry partial input in interactive mode, skip the menu and gather the missing required fields by asking one at a time — ask, wait, ask the next. Don't bundle questions into one block. If the user supplies a field unprompted (e.g. they pasted a URL in the trigger message), skip that question and confirm the value back to them once at the end. Don't start the pipeline until all required fields are answered. If the non-interactive fast lane applies, use step [0.5] instead.
[0.5] Non-interactive fast lane
Use this path when the caller passes --quick or --config <path>, or when the
caller states they are running from CI, a subagent, a batch job, or any other
non-interactive harness.
This section has precedence over the interactive ask/wait instructions below.
When it applies, use this fast lane and do not fall through to the multi-turn
intake unless url or founder name/role is truly missing.
--config <path>points to a JSON file with pre-baked values for the canonical input contract:url,brand_kit_pathorbuild_brand,founder_name,founder_role,founder_photo,assets,music_url,aspect_ratio,location_image_url,voice_style,product_type, andlower_third.--quickmeans use defaults for optional extras, auto-build the brand kit withbuild-a-brand --quickifbrand-kitis omitted, and usefounder_photo = "generate"when no photo is supplied.- For
--quickor--config, do not stop for confirmation at the brand-kit branch, founder-photo generation prompt, optional-extras prompt, script choices, or end-card/caption defaults. Record assumptions inline and continue. - If
urlor founder name/role cannot be found in args or config, stop once with a single compact missing-fields list instead of starting a multi-turn Q&A loop.
1. Product URL (required) — https://.... Used to (a) derive the brief in step [1] and (b) feed the brand-kit branch below.
2. Brand kit (required) — interactive mode: ask "Do you already have a brand kit folder, or should I build one first?"
- If a path -> use it (
state.brand_kit_path = <path>). Accept eitherbrand.jsonor an exportedbuild-a-brandkit containingbrand.md,tokens/tokens.json, and logo assets. - If "build" -> invoke the
build-a-brandskill on the URL/brief and wait for the exported brand kit. This is a full identity workflow and may pause for user choices; surface those prompts in interactive mode. - Fast lane: if config provides
brand_kit_path, use it. If config setsbuild_brandor--quickomitsbrand-kit, invokebuild-a-brand --quickon the URL/brief and wait for the exported brand kit; do not surface build-a-brand prompts or stop for identity choices. After either branch, setstate.brand_kit_path. Only stop with a single compact missing-fields list if there is no path and the brand kit cannot be built.
3. Founder identity (required) — interactive mode: ask all three together:
founder_name— e.g. "Avery"founder_role— e.g. "CEO, ExampleCo"founder_photo— local path / https URL / OR the literal stringgenerateto auto-create a portrait. Ifgenerate, prompt the user for a 1-line vibe ("warm, casual smart attire" / "Pixar-style 3D animation" / etc.) — this becomes the seed prompt forgenerate_imagein step [4].- Fast lane: use founder values from args/config. If
founder_photois omitted, setfounder_photo = "generate"and use a neutral founder-portrait vibe derived from the product tone; do not stop for a separate photo-vibe prompt.
Default to no lower-third so the happy path stays lean. If the user explicitly asks for a lower-third, record state.lower_third = true; in interactive mode confirm that edit_video_compose will add the transparent overlay after rendering.
4. Optional extras — interactive mode: offer these once as a single message, then proceed without waiting if no answer comes back in the same turn. Fast lane: use the defaults below without asking.
- Custom imagery — list of
assets(product photos / app screenshots) shown on the founder's phone. Default if omitted: use screenshots frombrand.json.screenshotswhen present, otherwise look for obvious screenshots or product images inside the brand kit, otherwise capture the product URL withcapture_website(mode:"screenshot")before step [2]. Do not proceed to script or SeeDance without real product UI / product imagery unlessproduct_typeis explicitlyserviceand the user accepts an environment-only video. In the fast lane, if no supplied/brand-kit/captured asset exists, stop once with a compact missing-assets error instead of silently shipping a generic talking-head video. - Music — local path / https URL / OR
generate(instrumental, ~60s). Default:generatevia Kling background mode. - Lower-third — optional. Default: off. If enabled, render the transparent
.movthrough MCP and overlay it onto the body withedit_video_compose. - Aspect ratio —
16:9(default),9:16,1:1. - Location — defaults to a flat seamless backdrop in
state.brand.colors.accent(clean studio-shoot look, character against a single brand color, whatever the brand's accent is). Override with a path / URL / text description if the user wants office, outdoor, etc. - Voice style — VO direction string for SeeDance, e.g. "warm authentic founder energy, conversational". Default: derived from
brief.tone. - Product type —
digital | physical_apparel | physical_object | consumable | service. Default: auto-derived in step [2] from asset analyses.
After Stage 0 completes, store all gathered values in state.inputs. If you already created a local work directory for this run, optionally persist the same object as <workdir>/inputs.json; do not require a predefined work-directory environment variable. Then enter the pipeline at step [1].
[0.6] Avatar-type probe for founder photos
Before any paid generate_reference_video call, run this Avatar-type probe on the resolved founder photo/avatar URL after local upload or user-supplied URL normalization. This applies to any founder photo used as the character reference — whether supplied via --photo or generated.
Call analyze_media once:
Route from the result:
- recognized IP / copyright risk -> STOP only when
avatar_typeis"recognized_ip", orrecognized_characternames a specific character (for example"Batman"), or when bothmoderation_riskis"high"andrecommendationis"reject". Treatrecognized_character: null, empty string,"none","unknown","n/a", and low/mediummoderation_riskas not enough to stop by themselves. Run this check before the real/stylized routes. A chibi Batman is still Batman even whenavatar_typeis stylized / illustrated. - real human / AI-generated realistic -> proceed normally.
- stylized / illustrated -> proceed with a visible warning that stylized avatars may be less reliable for Seedance likeness and moderation, then continue only if the user supplied or accepted that avatar.
- trademarked / copyrighted -> STOP before generation. Surface this message:
Your founder photo appears to be a trademarked character ([X]). Most video providers will moderate this and refuse to generate. Pass --photo <real-looking-photo-url> to override.For this skill,--photo <real-looking-photo-url>is the accepted concrete flag; you may also mention the cross-skill--avatar <real-looking-photo-url>wording because users may know that convention.
Required inputs (canonical contract)
After Stage 0, these are the fields downstream steps consume:
url— the product website (https://). Drives step [1] brief.brand_kit_path— brand kit folder. Required. End card AND lower-third consumebrand.jsonwhen present, otherwisebrand.md,tokens/tokens.json, and logo assets from abuild-a-brandexport. See step [4.5].founder_name+founder_role+founder_photo— required from intake. Step [4] normalizesfounder_photointofounder_photo_urlandcharacter_urlbefore any SeeDance call.assets— optional array of{ url, role?, caption? }. Defaults to the screenshots captured by the brand kit; when the brand kit has no screenshots, capture the product URL before step [2] and store the returnedimage_urlas a real product UI asset.roleis a hint string mapping the asset to a script beat (hero,feature_a,cta, etc.).location_image_url— optional. Defaults to a generated solid-color backdrop instate.brand.colors.accent.music_url— optional. Defaults togenerate(Kling 60s background bed in step [7]).aspect_ratio— default16:9.voice_style— optional, defaults tobrief.tone.product_type— optional, auto-derived in step [2].
State
Keep a simple state object as you work and save every CDN URL there so a partial run can be resumed. Treat task_status value completed as the successful terminal state (failed and cancelled are the failure terminals), then unwrap result.structuredContent when present. The final video lives on Pika's CDN; no local workspace is required unless MCP compose is unavailable and you explicitly trigger the local lower-third fallback in step [8b].
Long-running task_status polling
When any long-running generation or edit call returns a task_id with or without an initial status, including {task_id}, {task_id, status: "queued"}, or an initial queued, running, or processing status, record the task id and start time immediately in state.
- Call
task_status({task_id})in a tight loop until terminal (completed | failed | cancelled). No manual sleep and no Bash polling; the worker holds each status call open. - Emit ONE visible progress line every 60s while status is
queued,running, orprocessing:Seedance i2v queued for {N}m {S}s... still processing. Replace the provider/stage label when polling music, captions, render, concat, mix, or edit tasks. - On
completed, unwrap the returned result URL and save it intostate. - On
failedorcancelled, surface failure to the user withtask_id, status, and the last status message. - After 15 min total from the original submit, call
task_cancel({task_id})if the task is still non-terminal, then surface failure to the user. If cancel reports the task is already terminal, call status once more and report that terminal result. - Do not submit a duplicate request while the original task is still
queued,running, orprocessing.
Pipeline overview
Operational notes
Keep the main workflow focused on sequencing. Historical server validation details live in references/ops-notes.md; only the active constraints stay here:
- Use a unique
seedper SeeDance act (101, 202, 303, 404). Identical generation params can replay cached failures. - Kling music bed generation uses
provider: "kling-audio",mode: "text_to_audio",background: true, andduration_seconds: 60; the MCP worker generates one 10s Kling seed and extends it locally. - If SeeDance rejects a real-person founder photo, re-roll the founder ref with stronger stylization rather than retrying the same rejected reference.
- For local brand-kit logos, upload only logo-appropriate raster assets (
image/png,image/jpeg, orimage/webp). Do not send SVGs toupload_asset; choose the PNG export frombuild-a-brandor rasterize first. - Use CSS
background-image: url(...)for CDN-hosted logo/photo assets in end-card HTML;<img crossorigin>is blocked by CDN CORS. - Use server-side deterministic tools for captions, lower-third compose, concat, and mix. Local ffmpeg is only the fallback if
edit_video_composeis unavailable whilestate.lower_third = true. - Decompose every 15s act into 3 time-coded sub-shots. Single-shot acts look static.
- Open each act with the style-match location framing and repeat the same
WARDROBE LOCK:sentence across all 4 act prompts.
[1] Analyze brief
Save the result as brief. You'll reference brief.product_name, brief.tagline, brief.key_features, brief.tone, brief.call_to_action throughout.
[2] Resolve product UI assets, then analyze each asset + derive product_type
Before analyzing assets, normalize assets so product reveal shots have a real visual reference:
- Use any caller-provided
assetsfirst. - If none were supplied, read screenshots from
brand.json.screenshotswhen present. - If
brand.jsonhas no screenshots, look for obvious raster screenshots or product images inside the brand kit (screenshots/,assets/,product/, or image files named likehero,screen,app,dashboard,product). - When
assetsis empty after those checks, callcapture_websiteon the product URL:
If the product is likely mobile-first, also run a second capture with mobile: true and keep both URLs when available. Save these as real product UI assets before continuing.
Do not proceed to script writing, generate_reference_video, or any paid SeeDance call without at least one real product UI / product imagery asset, unless the caller explicitly set product_type: "service" and accepted an environment-only video. If capture fails or returns no image_url, surface: Could not capture <url>. Please provide screenshots or hosted product assets; founder-product-video will not silently ship without real product UI.
For each entry in assets, run analyze_media to extract content + visual style + asset type. Run all in parallel in one tool batch:
asset_type decoder:
digital_screen— app UI / website screenshot / SaaS dashboard / mobile app capturephysical_apparel— t-shirts, hoodies, hats, anything wearable (model + garment)physical_object— gadgets, accessories, packaged goods, anything held in handconsumable— food, beverages, supplements (something used/eaten/drunk)infographic— chart, diagram, data viz, illustrationother— anything else; describe and pick best fit
Save as asset_analyses[i].
After analysis, build:
If usable_product_assets is empty and you have not already tried a URL capture, call capture_website(mode:"screenshot") on the product URL, append the returned image_url to assets, run analyze_media on that capture, save the analysis at the same new index, and rebuild usable_product_assets. The captured screenshot must therefore have its own asset_index; do not reuse a logo/hero/infographic index for a product reveal.
Do not auto-derive product_type = "service" from logo-only, hero-only, infographic, abstract brand, or other assets. Those are not real product reveal anchors. If usable_product_assets is empty after the capture attempt, stop unless the caller explicitly set product_type: "service" and accepted an environment-only video. Surface the missing-assets error instead of falling through to the service reveal pattern.
Derive product_type
Look at the dominant asset_type across usable_product_assets:
If user passed product_type explicitly, use that and skip auto-derivation; service is valid only when explicitly chosen/accepted. The product_type value drives which shots are picked in step [3] and how the founder reveals the product in step [5]. Get this right or the video shows the wrong thing on screen.
Product type → reveal pattern (the most important table in this skill)
product_type (set in step [2]) controls which shots to pick AND how the asset is revealed in each shot. The reveal beat in the SeeDance prompt is product-specific; using the wrong one makes the founder hold a phone for a t-shirt brand.
For digital products, the SeeDance prompt is only responsible for the founder, camera move, phone gesture, and a blank / neutral screen placeholder. Do not ask Seedance or the video model to render readable brand wordmark or product UI text. The real screenshot, product UI, and exact brand spelling are composited in step [5.5].
For physical products, every shot where the product appears in frame should pass that asset as a reference image; otherwise SeeDance tends to invent a generic-looking product. For apparel, if the founder is wearing a t-shirt and the script says "we make t-shirts", the founder's t-shirt needs to reference one of the assets even in shots that are not reveal moments. Pass the asset URLs in reference_images for those acts and write prompt language like "the founder is wearing the t-shirt from @Image3 — print matches exactly".
[3] Write script + character voice + per-shot asset + per-line beats (you do this — no model call)
Three sub-products, all written by you (Claude) in one inline JSON:
character_voice_profile— 3-4 lines describing the character's DEFAULT delivery (carries through every act for consistency)- Per-shot
asset_index+reveal_beat— what asset is visible in this shot and how it's revealed - Per-shot
beats[]— line-by-line acting direction withemotion+physical+ silence beats between sentences
This is what separates a generic AI-talking-head from a character that actually feels intentional. Read all four sub-sections below ([3.0] founder voice, [3a] character voice profile, [3b] beats, [3c] transitions, [3c.1] acting energy, [3d] full JSON) before writing.
[3.0] Founder voice — write a PITCH, not a feature list
The single most common failure mode in this skill is dialogue that reads like a marketing-page bullet list ("It can reason. Code. Even write your emails. No proxies. No selectors. No maintenance. Plug it into LangChain. LlamaIndex. MCP. Twenty-four thousand stars on GitHub. MIT licensed. Production-grade.") — clean copy, but it's not how a founder talks. Field feedback: "the script sounds like a list of features, not like a founder would sell their product on camera."
Real founders pitching their own product on camera use:
- First-person ownership — "I built", "we shipped", "we use it ourselves", "honestly we just want this everywhere"
- A personal stake or origin moment — Act 1 should reference a frustration the founder lived through, NOT the product abstractly. "Every time I tried building X, I hit the same wall" beats "X is hard."
- Conversational connectives — "look", "honestly", "the thing is", "so", "actually", trailing "..." for thinking. These are throwaway words in writing but the breath of natural speech.
- A "bet" framing for the product — "what if X just worked?", "we asked ourselves", "the whole idea was". Founders frame their product as an answer to a question they asked themselves, not as a list of capabilities.
- One concrete anchor — a specific number, a specific time, a specific scenario. "24 thousand devs starred it last year" beats "it's popular." "At 3 AM the layout breaks" beats "scrapers are unreliable."
- Invitation-energy CTA — "come try us", "go play with it", "we just want it everywhere". NOT "stop scraping. start extracting." (that's a Don Draper tagline, not a founder).
Banned patterns (each was empirically called out by the user, do not repeat):
Allowed patterns (use these instead):
Structure (4 acts, ~30-40 words per act = 120-160 words total, ~50-60s spoken):
- Act 1: Personal stake / pain. First person. Reference a specific frustration the founder lived through. Land on the problem named cleanly.
- Act 2: The bet. "So we built X. The idea was — what if [pain] just worked?" One sentence on what it actually does (URL → data, prompt → image, etc).
- Act 3: Proof + community. ONE specific number (stars, customers, ARR). One casual mention of integrations or where it's used. Tone: quiet confidence, not bragging.
- Act 4: Invitation. "If [reader's situation] — come try us. [URL]. [One inviting line]." End on warmth, not a tagline.
Self-test before approving the script. Read each act's dialogue out loud. If you'd be embarrassed to say it on camera as the founder, rewrite it. If it sounds like a 30-second commercial voiceover, rewrite it. If a paragraph has more than two punctuation periods in a row of short fragments, rewrite it.
[3a] Derive character_voice_profile + wardrobe_lock
Two separate fields, both required:
character_voice_profile (3-4 lines) — how the character delivers EVERYTHING: cadence, default expression, signature gestures, hand habits, pause behavior, when smiles arrive. Look at brief.tone + the character reference image (character_image_url or your generated founder ref) + product_type. This is an actor's "circumstance" — not what they're saying, but who they are. It carries through all 4 acts so consistency feels intentional, not accidental.
wardrobe_lock (1 sentence) — what the character is wearing in every act. SeeDance reads @Image1 fresh for each 15s generation and may interpret different clothing between acts. The wardrobe_lock sentence is repeated verbatim in every act's prompt to keep clothing consistent. Read what the founder is wearing in the reference photo and describe it explicitly. Example: "wearing the same charcoal hoodie over a dark band tee throughout all 4 acts, black-framed glasses on". Without an explicit wardrobe lock, later acts can invent different clothing even when the first act matches @Image1.
Tonal-template starters (orchestrator picks/customizes from the brief tone):
Worked example — developer-tool founder (casual tone, 3D Pixar 20s woman):
"Casual confidence, like explaining the product to a friend at coffee. Slight smirk default. Eyebrow flicks on key reveals. Hand goes to chin when thinking, opens out flat on the big 'meet the A P I' claim. Pauses with held eye contact rather than filling silence. Lands punchlines deadpan and lets a small smile arrive a beat after."
Worked example — streetwear founder (playful/edgy tone):
"Sharp dry wit. Talks fast in clipped sentences with sudden pauses. Default slight smirk with one raised eyebrow. Eye-rolls on the pain points ('boring', 'generic'). Mischievous grin breaks through on punchlines but disappears immediately. Hands stay mostly still — the FACE does the work."
[3b] Per-line beats[] — line-level direction, not act-level
Each shot's dialogue is broken into beats. Each beat is one short sentence (or a deliberate silence) with its own emotion + physical direction. Silence between beats is part of the performance — fill it with held looks, micro-expressions, gesture transitions.
A beat with text: "(beat)" is silent (no spoken text) — it just describes what happens visually during the natural pause between sentences. Use these between dialogue beats that need a held moment for emphasis.
When the SeeDance prompt is built in step [5], beats become the per-shot acting direction (the dialogue text without (beat) markers becomes the <<<voice_1>>> payload).
[3c] transition_from_prev — choreograph continuous camera motion between shots in the same clip
The fundamental SeeDance limitation: each 15s SeeDance generation renders ONE virtual environment with ONE virtual camera. When a multi-shot prompt declares "Shot C, then Shot A" without specifying a continuous camera move between them, SeeDance defaults to re-framing the same camera position (zoom or crop). The result reads as a jump zoom, not a real cut — same background, character at different sizes.
The fix: every shot beyond the first in an act must declare a transition_from_prev field — a one-line description of the continuous camera motion that takes us from the previous shot's framing to this one. SeeDance then has to render an actual move-through-space, which means different parts of the room appear behind the character across the clip.
Pattern: name the camera's start position, name where it ends up, name the move that connects them. Movement verbs that work: dolly, pull back, push in, orbit, arc, glide, crane up, crane down, tilt up, tilt down, drift left/right.
Examples:
The first shot in an act has no transition_from_prev — it establishes the framing. Every subsequent shot in that act gets one.
SeeDance can do hard cuts within a single 15s clip when prompted explicitly. Write Hard cut: between time-coded sub-shots (instead of Transition:) for distinct framing changes — SeeDance honors this and renders a real cut, not a re-frame. Reserve Transition: for continuous-motion handoffs where you want the camera to glide between framings. Pattern: hard cuts feel like a real edited piece (different framings, different camera angles, different acting energy); transitions feel like a single moving long take.
[3c.1] Acting energy floor — every beat needs explicit body movement
A frequent failure mode: beats are written with only facial micro-expressions ("slight nod", "eyebrow flick", "eyes hold camera"). SeeDance renders this as a near-frozen founder — eyes barely move, no presence. Result reads as "static, frozen, no excitement."
Rule: every beat's physical field needs at least one of:
- A hand or arm gesture (open palm, count on fingers, dismissive flick, point at self/camera, hand to chest, wider arm sweep, hand-to-temple thinking)
- A torso shift (lean forward, lean back, slight body turn, shoulder shift)
- A head action LARGER than a micro-flick (turn left/right and back, tilt 8°+, slow head shake, head bob on rhythm)
- A directional eye flick combined with eyebrow movement (look down then snap up to camera, etc.)
Facial-only beats are acceptable only for:
- Silent
(beat)markers between spoken sentences (those are meant to be still — the held look is the point) - Final landing beat at end of an act when the camera is already moving (camera does the work)
When you write the SeeDance prompt, make sure the assembled "Acting beats" block reads physically dense — if you scan it and see five beats in a row that all say "slight nod" or "small smirk" with no other movement, the founder will look frozen. Rewrite with bigger movement.
[3d] Script JSON
Per-shot asset assignment rules:
- For digital products: only shots
CandEget anasset_index(phone-reveal moments). These are overlay pointers for step [5.5] and are excluded from the Seedance reference image array. Other shots show the founder without a specific UI reference. - For physical products: any shot where the product is visible in frame gets an
asset_index. The asset is the source of truth for what that product looks like. Acts can reuse the same asset across multiple shots, OR show a different asset per shot to demo product variety. - For service products: no
asset_indexanywhere — the script relies on dialogue + environment.
Every non-null asset_index must come from usable_product_assets[*].asset_index. Never assign a logo-only, hero-only, infographic, abstract brand, or other asset index to a product reveal shot. When using product details in dialogue, prefer usable_product_assets[*].analysis.key_elements and usable_product_assets[*].analysis.visible_copy so the spoken pitch tracks the same real product artifact shown on screen.
Reference image array per act = the union of Seedance-safe asset URLs across that act's shots. Digital digital_screen assets are not Seedance-safe; they stay overlay-only for step [5.5]. Within the SeeDance prompt, refer to non-digital assets by their position in the array as @Image3, @Image4 (positions 3+ — positions 1 and 2 are always character + location refs). The orchestrator computes this mapping when building the prompt in step [5].
Dialogue rules — TOTAL across all 4 acts must read aloud in 55–60s (~150 wpm = ~150 words total, ~37 per act). Short, punchy, speakable. Avoid em-dashes (founders don't speak them). Use natural contractions. Reference the user's actual product features (drawn from usable_product_assets[*].analysis.key_elements and usable_product_assets[*].analysis.visible_copy), not invented ones.
TTS pronunciation rewrites
SeeDance's native lip-sync TTS reads <<<voice_1>>> text literally — it has no semantic awareness that "Ari" is a name, "Vercel" is pronounced "ver-SELL", "UGC" is an acronym to spell, or "example.com" is a URL. Rewrite the dialogue text the way you want it pronounced, then submit. Apply these substitutions:
When in doubt about how a brand pronounces an acronym (NASA vs N.A.S.A.) or non-obvious product name, check the brand's own website / videos. Default to spelled-out letters for unclear acronyms and a simple phonetic alias for non-phonetic names. For domains that include the brand, combine both rewrites: Vercel.com becomes ver-SELL dot com in the <<<voice_1>>> payload.
Worked example — original script vs TTS-safe rewrite:
Keep the rewrites in the dialogue text only — your script JSON's dialogue field is what flows into the SeeDance prompt verbatim and then into each <<<voice_1>>> block. Visual surfaces keep the canonical original spelling: state.brand.name, brief.product_name, product UI overlays, wordmarks, end card text, URLs, filenames, and QA expectations must stay unchanged. Never replace deterministic visual text with the phonetic alias.
Shot type reference
Each shot has variants depending on product_type. Pick the variant that matches.
⚠️ Avoid film-industry shot terminology that SeeDance interprets literally. SeeDance reads named shot types as a literal recipe — including any implied subjects the term carries. Specifically:
- ❌ "Two-shot" → adds a SECOND PERSON to the frame (term means "shot with two subjects" in film, but SeeDance just sees "two" + "person").
- ❌ "Three-shot" — same trap.
- ❌ "Over-shoulder" / "OTS" → adds a phantom shoulder/back-of-head in the foreground for the character to interact with. The character then performs toward that phantom person, not toward the camera.
- ❌ "Master shot" — can be misread as "the master / their boss".
- ✅ "Medium shot", "Close-up", "Wide" are safe; they're commonly used colloquially.
The fix in every case is plain language about what the camera sees, not film vocabulary that implies extra subjects. Examples:
- "Over-shoulder reveal" → "the character holds her phone up toward camera, screen facing the viewer at chest height"
- "Two-shot" → "wide framing of the character with [product/logo/environment] in frame"
- "POV" → "low camera angle from the character's eyeline"
[4] Founder + custom location reference images
Prepare character_url before any SeeDance call:
- If
founder_photois an HTTPS URL, setfounder_photo_url = character_url = founder_photo. - If
founder_photois a local path, upload it withupload_asset, then setfounder_photo_url = character_url = public_url. - If
founder_photoisgenerate, callgenerate_imageand use the returned URL:
Handle location only when the user supplied a custom location:
- If
location_image_urlis an HTTPS URL, setlocation_url = location_image_url. - If it is a local path, upload it with
upload_assetand setlocation_url = public_url. - If it is a text description, generate a custom location reference with
generate_image. - If no custom location was supplied, do nothing here. Step [4.5] generates the default brand-accent backdrop after
state.brandexists.
Save the resulting URLs into state. If SeeDance later rejects the founder ref on content policy, see "Known infra quirks" — re-roll with stronger stylization.
[4.5] Brand-kit ingestion (always — Stage 0 guarantees brand_kit_path)
brand_kit_path is required by Stage 0 — either user-supplied or built first with build-a-brand. Parse it once and reuse across the end card and the lower-third. If the folder is missing in interactive mode, ask the user to provide or rebuild the brand kit before continuing. In the non-interactive fast lane, try the build-a-brand --quick branch first; if that cannot produce a kit, stop once with a single compact missing-fields list.
Preferred source is brand.json when present. Otherwise extract from a build-a-brand export:
brand.mdfor name, tagline, voice, typography names, and logo descriptions.tokens/tokens.jsonfor colors and font tokens.logo/assets for wordmark, symbol/icon, and lockups.
Extract into state.brand:
Upload step is required when the brand-kit assets are local files. Without public wordmark/icon URLs, the HTML rendered by render_html_animation can't reach them. Use the MCP upload_asset flow with raster logo files only and save the returned public_url values on state.brand. upload_asset rejects image/svg+xml; if the best logo is an SVG, pick the sibling PNG export from the brand kit or rasterize the SVG to PNG via html_to_png by inlining the SVG inside an HTML <svg> block and using the returned PNG public_url.
Brand-accent backdrop default location — when location_url was not set by Step [4], render a solid-color PNG via html_to_png using state.brand.colors.accent. Match the requested video aspect so the reference is not cropped later:
Save the returned file_url as location_url. This produces the clean studio-shoot aesthetic — character against a flat seamless wall in the brand's own accent color, no furniture or background detail. This is more reliable than AI-generated office sets and renders SeeDance reliably.
[5] Generate 4 SeeDance acts in parallel
CRITICAL: Step [5] requires generate_reference_video (multi-ref). The @Image1 / @Image2 / @Image3 tokens in this skill's prompt template are only resolved by generate_reference_video. If you call generate_video (single-image i2v) by mistake, @ImageN tokens are silently ignored and Seedance will hallucinate UI from prompt text alone, typically misspelling the product name. Stop and fix the tool call before firing any paid video generation.
For product_type: "digital", this is now a style/gesture generation step, not the final readable UI step. Do not ask Seedance to spell the product name, render a readable brand wordmark, or reproduce product-screen copy. Phone reveals must ask for a blank neutral screen placeholder, because real product UI is composited in step [5.5]. Do not pass digital_screen assets in reference_images; any digital_screen asset index is an overlay pointer for step [5.5] only, even if a mixed or explicitly non-digital run also has physical assets.
For each act, collect the union of asset_index values across all shots in that act to build the reference_images array:
The asset entry's position in reference_images determines its @ImageN token (positions 1, 2 are character + location; usable product assets start at position 3). When writing the prompt, map each non-digital shot's asset_index to its matching seedance_asset_entries position. For any digital_screen asset, do not reference an @ImageN token in the Seedance prompt; use the blank phone-screen placeholder language and let step [5.5] consume the original act_asset_entries entry.
Prompt template
Refer to the character through the @Image1 reference, not descriptive prompt prose. When the prompt says both "young creative streetwear founder" and @Image1 is the character ref, the two descriptions can fight — SeeDance may try to satisfy both by inventing a second figure. Same rule applies to the location and @Image2. Let the reference images carry the visual identity.
Each act's prompt has three layers, top to bottom:
- Character + location identity (always the same opening line).
- Character voice — the act prompt repeats
script.character_voice_profileverbatim. This carries the actor's personality through every shot. - Per-shot blocks — each shot gets its composition (or
reveal_beatif defined), then a per-line "Acting beats" list built from the shot'sbeats[], then the dialogue inside<<<voice_1>>>.
⚠️ The opener line, WARDROBE LOCK:, Background context:, <<<voice_1>>>, Transition: / Hard cut:, and the closing Native lip-synced dialogue audio, no music overlay. line are all load-bearing — each is documented in ## Load-bearing phrases near the bottom with the specific failure mode it prevents. Paraphrasing any of them silently breaks the recipe (literal backdrop, sung lyrics, jump zooms, wardrobe drift, etc.). Leave them verbatim; the connective prose around them is yours to compose.
Template:
Notes on the template:
- Open with "The character (matching @Image1) in a setting whose visual style ... match @Image2" — never "inside the location matching @Image2". SeeDance reads the LITERAL backdrop from @Image2 if you say "inside the location" — see Known infra quirks.
- Don't describe the character in prose — SeeDance reads visual identity from
@Image1. EXCEPTION: lock wardrobe explicitly via theWARDROBE LOCKline (e.g. "wearing the same charcoal hoodie over a dark band tee throughout"). @Image1 alone doesn't lock wardrobe across separate 15s generations. - Don't add aesthetic adjectives that contradict the references. If
@Image1is a 3D Pixar-style character, don't add "photorealistic" anywhere in the prompt. - The CHARACTER VOICE + WARDROBE LOCK lines appear once per prompt, before the shots. They prime SeeDance for the actor's whole vibe + clothing continuity.
- Each shot block includes a
Background context:line describing a distinct physical position in the space — different wall, different window, different background feature than the other shots in this act and ideally distinct from the other acts. This forces SeeDance to render scene variety while keeping the brand aesthetic anchored to @Image2. (beat)markers stay OUT of the<<<voice_1>>>payload. They're acting direction only — the natural pause between sentences in the dialogue text is where they happen.- Every shot beyond the first in an act gets a
Transition: …line built fromshot.transition_from_prev. This narrates the camera move between framings, forcing SeeDance to render real spatial motion (different parts of the room behind the character) instead of an unmotivated re-frame that reads as a jump zoom.
Worked example — physical_apparel, Act 2 (with voice + beats)
Act 2 has shots [C, A]. Shot C has asset_index: 0 and a reveal_beat about the search-history t-shirt. Shot A has no asset_index. Asset 0 = @Image3. Character voice profile is the streetwear-founder example from step [3a].
Physical product consistency across shots
For physical_apparel, if the character is wearing the brand's product across multiple acts (e.g. acts where no specific product reveal happens but the character is still in a brand t-shirt), pick one hero shirt asset and pass it to those acts too with prompt language like "the character is wearing the t-shirt from @Image3 — print matches exactly". Otherwise SeeDance invents a generic-looking shirt, defeating the asset coverage goal.
Fire all 4 acts in parallel — 3 sub-shots per act
Fire all 4 acts in parallel in a single tool batch. Every act prompt should contain 3 time-coded sub-shots; single-shot 15s clips render as static, frozen-looking founder regardless of how detailed the prompt is. If a run looks "very static, no body language, camera work boring", the fix is decomposing each act into 3 time-coded sub-shots within the prompt.
Default decomposition: 3 sub-shots per act, sized by dialogue density (e.g. (0-4s), (4-9s), (9-15s)).
Each sub-shot needs:
- A different camera framing — never two consecutive sub-shots with the same shot type. Mix medium / close-up / wide / lower-angle. The user reads variety as production value.
- A different camera motion — push-in, pull-back, lateral arc, orbit, handheld, static-held, rack-focus. Not all "slow push-in".
- A different physical position in the space — see Location-reference rule in "Known infra quirks". Each shot must describe a different wall / window / feature visible behind the character so SeeDance moves through the space instead of reproducing one literal backdrop.
- An ENERGETIC physical action per beat — see [3c.1] above. Hand gestures, leans, head turns, shoulder shifts. No facial-only beats.
- The dialogue sub-portion for that window, wrapped in
<<<voice_1>>>...<<<voice_1>>>tokens inside the shot's block.
SeeDance firing pattern:
Notes:
- Each act runs ~3-8 min on SeeDance. If generation completes asynchronously, follow the MCP tool's returned status handle until the act reaches a terminal state.
- The full prompt (opening line + CHARACTER VOICE + 3 sub-shots with Acting beats +
<<<voice_1>>>per shot + Transition lines) goes in the singlepromptparameter. SeeDance has noshots:[]array — the multi-shot structure is encoded in prose. - Use unique seeds (101, 202, 303, 404) so identical-looking calls don't hash to the same cached task ID. The
seedparameter is seedance-only.
Save the 4 returned URLs in submission order as act_urls = [act1, act2, act3, act4].
[5.5] Deterministic digital UI overlays
This step runs after all four Seedance acts are available and before step [6]. For efficiency, run the cross-act identity QA below first and use the final accepted act_urls; if identity QA later retries an act, discard the old overlay outputs and rerun this step. Its job is to make digital product reveals pixel-grounded: the phone / product-screen visual seen by the viewer comes from the real captured screenshot or rendered wordmark, not from Seedance.
If product_type !== "digital", set ui_grounded_act_urls = act_urls and continue.
If product_type === "digital":
-
Build
digital_reveal_shotsfrom every scripted shot whose type isCorE. Every scriptedCorEdigital reveal must have a non-nullasset_indexand that index must resolve to adigital_screenentry inusable_product_assets. If any scripted C/E digital reveal hasasset_index: null, stop and surface the script shot; do not proceed to Seedance concat or copy rawact_urls. Builddigital_reveal_planfrom the validated reveal shots. Each planned overlay stores:act,shot,asset_index,asset_urlvisible_copyfromusable_product_assets[*].analysis.visible_copybrand_name = state.brand.name- the expected text list: exact
brand_nameplus any short, readablevisible_copyphrases that matter for the pitch
If
digital_reveal_planis empty, stop and surface the script shots plususable_product_assets; do not silently ship a digital product video without real UI. If any scripted C/E digital reveal points at a missing or non-digital_screenasset, stop before composing. Do not copy rawact_urlsthrough when a digital reveal is missing its overlay; only acts with no digital reveal shot may pass through unchanged. -
For each unique
asset_url, render a 15-second overlay clip withrender_html_animationusingformat: "mov"when alpha is needed. Save the returned URL asui_overlay_clip_url.- The HTML should place the real screenshot / captured UI inside a rounded phone-screen frame with the correct aspect ratio.
- If the real screenshot's brand wordmark is too small to read after scaling, include an exact deterministic wordmark row using
state.brand.name; do not invent a shorter alias. - Keep the overlay background transparent or visually isolated so it works as a deterministic phone/UI overlay over the generated act.
-
Composite the overlay onto each planned act with
edit_video_compose:
Use the phone-screen placement when the generated phone target is stable. If the phone target is not stable enough to cover cleanly, use a stable overlay-panel placement that does not cover the founder's face or the caption area. Save each edited result into ui_grounded_act_urls[act - 1]; untouched acts copy through from act_urls.
-
When a short exact brand label is still needed after the compose pass, call
edit_text_overlayon that act withtext: state.brand.name. This is only for deterministic brand spelling, not captions. Save the updated URL back intoui_grounded_act_urls[act - 1]. -
Run OCR QA on each overlaid reveal act before concat with
extract_frameandanalyze_media:
Only brand_name_visible: "yes", brand_name_exact: "yes", visible_copy_ok: "yes", and verdict: "clean" can continue. If OCR says the brand is missing, misspelled, garbled, expected visible copy is missing/replaced, or verdict is degraded / catastrophic, stop and surface the overlay_qa_frame_url, ui_overlay_clip_url, expected brand_name, expected visible_copy, and the QA JSON. Do not proceed to step [6] with raw or AI-imagined phone text.
Cross-act identity QA before step [6]
Before edit_concat, run a cross-act identity check on a single comparable visual artifact. The goal is to catch a founder swap while the bad act can still be regenerated, not after the 60s body is already stitched.
- Extract one representative frame from each of the three sub-shot windows in every act. One frame at the middle of the act is not enough; a founder swap can happen in the first or final sub-shot and still pass.
- Render a single identity contact sheet with
html_to_png. Use CSSbackground-image: url(...)for each remote image. The sheet must include thirteen labeled panels in one image:founder reference(founder_photo_url/character_url), thenact 1 shot 1,act 1 shot 2,act 1 shot 3, throughact 4 shot 3. Save the returned file asidentity_contact_sheet_url.
- Run one
analyze_mediacall onidentity_contact_sheet_url:
If the contact sheet cannot be rendered, stop and surface the contact-sheet failure; do not proceed with unverified identity. Only a clean yes / yes / yes identity QA can proceed. If the QA returns same_founder_as_reference: "no" or same_founder_as_reference: "unclear", same_founder_across_acts: "no" or same_founder_across_acts: "unclear", wardrobe_consistent_with_lock: "no" or wardrobe_consistent_with_lock: "unclear", bad_act_numbers non-empty, bad_frame_labels non-empty, or verdict: "degraded" or verdict: "catastrophic", do not proceed to concat. Retry only the bad act(s) once with the same prompt, same reference_images, and a new seed (original_seed + 1000) plus one added prompt sentence: Identity continuity is mandatory: the on-camera founder remains the same person as @Image1 for the entire act. After retry, extract all three frame times for the retried act(s), rebuild the full contact sheet, and rerun the same QA. If any act_urls entry changes after a bad-act retry, discard ui_grounded_act_urls and any previous overlay/OCR evidence, rerun step [5.5] against the final act_urls, then rerun its OCR QA before step [6]. If the retry still fails cross-act identity QA, stop and surface identity_contact_sheet_url, the bad act URL(s), bad_frame_labels, and the QA JSON; do not ship a stitched video with a different founder.
Duration floor and partial-act recovery
All 4 act_urls are required before step [6]. Do not concat a partial act list. Three completed acts plus the end card produce a ~50s asset, which misses the 55s duration floor and must not be reported as a successful founder video.
If one SeeDance act reaches a failure terminal, reaches cancelled, or is
terminal or successfully cancelled by the long-running polling contract while
other acts completed:
- Retry the missing act once with the same prompt,
reference_images,duration,sound,resolution, andaspect_ratio, but a new seed (original_seed + 1000). Do not rerun successful acts. - If the retry completes, insert that URL into the original act slot and
continue with
act_urls = [act1, act2, act3, act4]. - If the retry cannot complete, stop and surface the upstream SeeDance timeout.
You may return completed act URLs as a diagnostic preview, but do not deliver
a partial concat as
final_url, do not call it production-ready, and do not proceed to step [6].
Do not retry a stalled act while its original task is still queued, running,
or processing. Continue polling it with visible progress, then call
task_cancel({task_id}) at the 15 min total ceiling before
using the one missing-act retry.
[6] Stitch acts into 60s base
Save as base_url (60 seconds, 16:9, native dialogue audio). ui_grounded_act_urls equals act_urls for non-digital products; for digital products it contains the step [5.5] overlaid reveal acts. Do not feed the raw act_urls into concat when a digital UI overlay plan exists.
[7] Generate background music — INSTRUMENTAL, target ~60s
The pattern is fixed. The sound is a per-brand creative decision. Use Kling audio because it supports background: true and the MCP worker now handles the long-bed gap: it generates one 10s Kling seed, then extends it locally with ffmpeg looping/crossfade to the requested 60s. WHAT genre / instrumentation / mood goes in the prompt is your job — pick something that fits the brand's tone, the founder's voice, and the product. Do NOT copy the piano example below verbatim; that's one possible sound, not a template.
Fixed pattern (don't change):
provider: "kling-audio"mode: "text_to_audio"background: trueduration_seconds: 60- No
lyricsfield. Kling uses the prompt only. - No second provider call or manual tiling. The MCP worker owns the one-seed extension path.
promptmust be<= 200characters. Kling rejects longer prompts with validation code1201, so keep the style snapshot short.
Creative decision (per video — pick from brief.tone + script vibe + product context):
Examples of what the music register might be for different brands. Don't use these literally — match the vibe of YOUR brand:
When in doubt: read brief.tone (technical / casual / playful / professional / disruptive) and pick a register that wouldn't feel weird next to the founder's voice profile + the script's emotional arc. Match the energy, don't fight it.
Canonical call (substitute YOUR creative direction into prompt, staying <= 200 chars):
How it works (load-bearing):
- Kling generates one 10s seed: the paid provider call stays short and background-capable.
- The worker extends locally: MCP loops/crossfades the 10s seed with local ffmpeg to reach
duration_seconds: 60, then returns the extended CDNaudio_url. promptfield carries the style snapshot. Include "soft instrumental background bed", "no vocals", and "leave room for narration" so the mix does not fight the founder dialogue. Keep it <= 200 characters; if Kling returns1201, shorten the same style idea instead of retrying the long prompt.
Save as music_url. Read result.duration_seconds:
- If
>= 55s→ mix it. Expected path with the Kling worker extension. - If
< 55s→ do not silently accept a short bed. Retry once with the samekling-audiocall; if it still returns short, stop and surface the tool result because the worker extension path did not satisfy the contract.
Banned anti-patterns (each empirically caused a failure):
- ❌
provider: "minimax-music"for the default generated bed — MiniMax does not supportbackground: trueand can foreground the melody under narration. - ❌ Manual 10s tiling in the skill — the MCP worker already extends one 10s Kling seed locally.
- ❌ Putting duration only in
prompttext ("60 second instrumental") — useduration_seconds: 60. - ❌ Adding a
lyricsfield — this path is Kling prompt-only. - ❌ Copying the same piano-and-pads example for every brand — the recipe is the pattern, not the sound. Pick a register per brand.
[8] Composition layer — lower-third + subtitles
Pipeline ordering note — music mix happens in step [10], AFTER end-card concat. Mixing music into the body before the end card is concatenated leaves the end card silent (the music track ends at the cut). Always: overlays on body → end card → concat → THEN mix music over the full assembled clip.
Use MCP tools first. add_captions handles subtitle timing and burn-in server-side; render_html_animation handles authored HTML motion; edit_video_compose handles transparent lower-third overlays.
Default paths
If state.lower_third is false or unset, skip [8a] and [8b]. This keeps the default path fully MCP-native.
Do not call local Whisper/caption scripts or chained edit_text_overlay for captions. If exact original-script spelling matters, pass manual subtitles[] only when you already have exact timed segments from a trusted source; otherwise prefer the add_captions auto waterfall.
[8a] Render the lower-third (only if state.lower_third = true)
Skip this sub-step unless state.lower_third = true. Render via render_html_animation with format: "mov" (ProRes 4444 with yuva420p — preserves alpha). Do NOT use format: "webm" — HyperFrames currently emits webm as VP9 pix_fmt=yuv420p with no alpha channel, so "transparent" areas come out as literal black pixels and the composited LT shows a black box outside the pill. The .mov ProRes path is the only reliable alpha path right now.
- Native dimensions: 800×220 (matches the placement size on a 1280×720 frame, so no scaling artifacts)
- Pill:
state.brand.colors.primarybg (default#0d0d0d),state.brand.colors.highlightborder (default#fefbcf),state.brand.colors.accentdrop shadow (default#cfc3ff), 18px border-radius - Two-line text:
founder_name(Space Grotesk 800, 80px, white) +founder_role(Space Grotesk 500, 28px, butter) - NO logo inside the pill — brand mark lives in the end card; lower-third is about the person
- CSS
@keyframesonly: slide in 0–0.6s fromtranslateX(-900px), subtle box-shadow pulse around 3s, slide out 4–5s. Do not use GSAP for lower-third animation; the same per-frame seek concerns as the end card apply. - Save URL as
lower_third_url(file extension.mov)
If you need to verify the .mov, inspect the video stream pixel format and confirm it contains alpha (yuva...). If it is plain yuv..., the alpha was dropped — switch render format or re-render.
[8b] Lower-third overlay
Use this only when the lower-third is enabled. Default to the MCP compose path:
Save the returned url as body_with_lower_third_url.
Local fallback contract, only if MCP compose is unavailable:
- Download
base_urlandlower_third_urlonly for this local composition step. - Overlay the 800×220 lower-third at
x=50,y=video_height - 220 - 100, enabled fort=0..5s. - Preserve the original body audio without re-encoding so lip-sync stays exact.
- Use visually lossless H.264 settings for the local checkpoint.
- Upload the checkpoint with
upload_assetand save the returnedpublic_urlasbody_with_lower_third_url.
If subtitles are requested too, call add_captions(video_url: body_with_lower_third_url, ...) and save its returned url as body_with_overlays_url. If not, body_with_overlays_url = body_with_lower_third_url.
[8c] Captions via MCP
Default call:
Save returned url as body_with_overlays_url. The returned transcript is useful for QA, but the video URL is the pipeline artifact.
[9] Animated end card (5s) — author inline HTML, render via HyperFrames
We do NOT use generate_slide_animation here. That tool delegates HTML authoring to a slide-card LLM, which routinely adds corner clutter (top-left wordmarks, bottom-right URLs), picks wrong aspects, and produces animations that don't reliably play through HyperFrames' per-frame seek. Instead, the orchestrator authors the end-card HTML directly and renders it via render_html_animation. Same engine the lower-third uses.
Hard rules — empirically verified, do not deviate
These are NOT stylistic preferences. Each was discovered by rendering, extracting frames, comparing to expectations, and observing a specific failure. Reverting any of them will reproduce a known bug.
-
Aspect ratio matches the body video. Read
aspect_ratio(default16:9). Computedata-width×data-heightfor the#stage:16:9 → 1920×1080,9:16 → 1080×1920,1:1 → 1080×1080. Hardcoding the wrong orientation produces a side-bar concat where the body and end card play side-by-side instead of in sequence. -
No corner clutter. No top-left wordmark. No bottom URL. No icon squircle next to the title. The end card is a single centered message + CTA. The brand mark is implied by the typography and palette; an explicit logo competes with the title and reads as cluttered. (If the user explicitly asks for a logo, place it integrated into the centered stack — never in a corner.)
-
Use CSS
@keyframesfor the entrance animation, not GSAPtl.from(). HyperFrames Chrome seeks per-frame via BeginFrame; GSAP'stl.from()records its initial state lazily on first play and never fires under seek-only playback — every frame renders the static FINAL state with no entrance motion. CSS@keyframesare tied to Chrome's compositor clock and animate deterministically per frame. Declare duration through the#stage/#carddata-durationattributes; do not add a GSAP script just to establish timing. Repeated frame captures with GSAP showed identical frames at t=0 and t=2s; switching to CSS@keyframesfixes it. -
All entrance animations must complete by
t = duration - 0.5s. Half-second hold so the final state sits readable before the cut. Withduration: 5sthat's allanimation-delay + animation-duration <= 4.5s. -
Use absolute positioning for animated elements, not flex. Flex layout doesn't fully settle by frame 0 in HyperFrames Chrome, which compounds the GSAP-from() bug above — elements pop in mid-animation as flex finishes its second pass. Pure absolute positioning gives stable, deterministic layout from frame 0.
-
Give each
@font-facea uniquefont-familyname; don't rely on weight matching. Declaring two@font-face { font-family: "telka"; ... font-weight: 700/500; }blocks is supposed to let CSSfont-weight: 500pick the 500 face — but in HyperFrames Chrome the matching is unreliable for some weights. Use distinct families:"telka-ext-900","telka-700","telka-500". When weight matching fails, the tagline can render in a serif fallback; switching the tagline to an explicitly loaded family avoids that. -
telkaextended-900-normal.woff2andtelka-700-normal.woff2are KNOWN-WORKING in HyperFrames Chrome.telka-500-normal.woff2is KNOWN-BROKEN — it parses successfully via fontTools (correct OS/2.usWeightClass=500, valid cmap, correct family name) but Chrome silently rejects the @font-face declaration and falls back to system serif. Workaround: usetelka-700for the tagline too (same family, slightly heavier — visually still on-brand). If a future end card needs medium weight, test the candidate woff2 by rendering+extracting frame 30 BEFORE shipping. Don't trust that "Telka 400" or "Telka 300" will work just because Telka 700 does. -
Inline brand-kit fonts as base64. Pika CDN doesn't accept font uploads (mime allowlist) and doesn't send CORS headers, so
@font-faceURL references fail in HyperFrames Chrome and the font silently falls back to system serif. Subset each woff2 withpyftsubsetto just the glyphs in title+tagline+CTA, base64-encode, embed asdata:font/woff2;base64,.... Subsetted files are typically 5–10 KB each. -
Don't write CSS
font-familyfallback chains for brand designs. If the brand font fails to load, a fallback chain hides the failure — you ship Helvetica thinking it's Telka. Usefont-family: "telka-700"alone (no fallback). Then a font load failure renders Chrome's default serif, which is visually obvious and triggers a fix. -
Composition contract — the HyperFrames contract:
<div id="stage" data-composition-id="main" data-start="0" data-duration="5" data-width="W" data-height="H">wraps a SINGLE direct child<div id="card" class="clip" data-start="0" data-duration="5" data-track-index="0">which contains everything else. Visible timed elements must includeclass="clip"because HyperFrames uses it for visibility control, and the clip must be nested inside the composition root, not a sibling. Multi-tracked direct children of#stageinteract poorly with frame seeking. (Fixed by mirroring the working lower-third structure.) -
Runtime readiness hook — include a small compatibility hook before
</body>:window.__hf = { duration: 5, seek: (t) => { document.documentElement.style.setProperty("--hf-time", String(t)); } };. CSS@keyframesstill drive the visual animation, but the hook makes the prod frame-capture path ready when it probes forwindow.__hf. If the worker reportswindow.__hf not ready after 45000ms, treat the HTML as invalid forrender_html_animation; fix the composition contract or hook and rerender. Do not fall back to a static PNG. -
Always extract frames at t=0, t=1s, t=2s after rendering and visually compare. If frames 0 and 2 look identical, the entrance animation isn't running. If the tagline looks like a serif, the brand font didn't load. Don't trust the URL alone. Don't ship without this check. (User caught these failures three renders in a row before frame extraction was added.)
Build steps
Layout (centered stack — no corners)
Animation timing reference (5s end card, CSS @keyframes)
All implemented as CSS animation: name duration easing delay forwards on the corresponding element. The element's pre-animation CSS state IS the "from" — no JS needed.
Reference implementation pattern: use a centered-stack HTML recipe. Author the filled HTML inline in the orchestrator and pass it directly to render_html_animation; do not call a bundled helper script or rely on a separate presets/ directory.
If the brand-kit lacks fonts (state.brand.fonts is null) — fall back to system -apple-system, sans-serif for tagline/CTA but DROP the title down to a system-display weight. Don't render brand-typography end cards with fallback fonts; they always look wrong. Flag this in the deliver step so the user knows the brand-kit is incomplete.
Save the returned MP4 URL as end_card_url. end_card_url must be an MP4 video segment, not a static PNG, because step [10] concatenates it with the body video. Download it only if you need local visual QA frames or a local fallback assembly.
[10] Assemble body + end card + music via MCP
Use the server-side deterministic edit tools for final assembly. The current MCP server edit_concat normalizes mismatched inputs before concat, and edit_audio_mix preserves the original video audio while mixing the music track.
Mix music after concat, never before, so the score continues through the end card. If edit_audio_mix fails because the music file is too short or malformed, set final_url = assembled_url, surface the music issue, and still run step [10.5] before delivery; do not rerun expensive SeeDance acts.
[10.5] Final duration floor
Before reporting final_url to the user, probe the assembled result:
Save the result as final_duration_seconds. It must be >= 55 and <= 75
before reporting final_url as the completed deliverable.
If final_duration_seconds is under 55s, treat the run as a failed partial
assembly: do not deliver the URL as final, do not mark the skill complete, and
return to the missing-act recovery above. If all 4 acts were present but the
probe is still under 55s, stop and surface the concat/provider truncation for
investigation instead of padding with unrelated footage.
[11] Final URL
Save final_url into state. If a local fallback assembly produced the final MP4, upload that checkpoint through upload_asset and replace final_url with the returned public_url.
Post-flight quality gate
Before declaring success, call analyze_media on final_url and ask for a structured verdict:
- If
verdictisclean, deliver the final URL normally. - If
verdictisdegraded, deliver the final URL plus thequality_warningso the user can review before publishing. - If
verdictiscatastrophic, do not call the video complete; surface the verdict andre_roll_suggestioninstead of declaring success.
[12] Asset bundle (optional)
Return the final URL first. If the user asks for editable source assets, provide a bundle containing:
final_urlact_urlscharacter_url,location_url, and any product/screenshot asset URLsmusic_urlend_card_url- optional
lower_third_url/body_with_overlays_url - the script, brand quick-reference, and act-to-asset map
Why this matters:
- The user can swap a layer (founder photo, sub timing, music) and re-render WITHOUT re-fetching everything
- The brand kit is the single source of truth for any future video for the same brand — re-using it is one folder copy away
- The act-level
.mp4s are valuable on their own (e.g. the user might want to clip just one act for a specific channel) - Users often ask to save the generated refs next to the video so they can reuse those assets. Codified as default.
[13] Deliver
Report final_url to the user. Include:
- Asset bundle URL/path only if the user asked for one
- Total duration from
final_duration_seconds(~65s = 60s body + 5s end card) - Intermediate URLs or local files useful for reruns:
base_url, optionalbody_with_overlays_url,music_url,end_card_url,final_url, and eachact_urls[i] - The brief (
brief.product_name/brief.tagline) so the user can confirm the model picked up the right product - A 1-line summary of which asset went into which act, so the user can confirm placement
Verification gates
Failure modes
Except for the one missing-act retry documented in "Duration floor and partial-act recovery", stop and surface on first verification failure. Don't auto-retry expensive calls (SeeDance acts run 3–8 min each — repeated failures burn credits).
Recovering from upstream 5xx on analyze_brief / generate_image / generate_reference_video / generate_music / upload_asset
If any paid generation, brief analysis, music, render, edit, or asset MCP call returns:
code: "provider_5xx"ANDretry_class: "retry_after_backoff"- Or HTTP 502 / 503 / 504 from any upstream provider (Seedance, Kling, OpenAI, Gemini, storage)
Do this:
- Wait 5 seconds.
- Re-call the exact same MCP tool with the exact same arguments. Do not rewrite the brief, product name, founder photo, act prompt, seed, reference-image order, music prompt, or upload payload.
- If the retry also fails with 5xx, abort and surface to the user: "Provider returned a transient upstream error twice. Try again in 1-2 minutes; this usually clears on its own."
Do not retry more than once. This path is for transient provider outages; changing creative inputs after a 5xx creates a different artifact and can duplicate spend.
Recovering from upstream 4xx / moderation_blocked
If generate_image returns an upstream 4xx or moderation_blocked while creating a founder/location/product reference:
- Do NOT retry the same prompt; moderation and most 4xx validation failures are deterministic.
- If the asset is not user-critical, try a fallback provider once only when the substitute still matches the brief. For founder identity photos, no fallback provider should alter the user's identity; ask for a safer/user-supplied reference instead.
- If the fallback provider also fails or would change the product/founder identity, surface to the user: "Image provider declined this reference prompt. Provide a different reference image or choose a less recognizable/stylized direction."
For Seedance act moderation, follow the duration-floor / missing-act recovery only where it applies. Do not run repeated blind act retries.
Recovering from upstream 429 (rate limit)
If any upstream returns HTTP 429 with a backoff hint:
- Wait the hinted backoff, or 30 seconds if no hint is provided.
- Re-call the exact same MCP tool with the exact same arguments.
- Do not retry more than once. If it still returns 429, abort and surface the rate-limit message.
capture_website returning empty / page-not-loaded
This skill usually extracts product facts with analyze_brief, but URL brief helpers may call capture_website. If capture_website returns 200 but action_bboxes is empty or recording_viewport is 0x0:
- Do NOT retry; the page failed to render in the capture environment.
- Surface: "Could not capture <url>. The page may be blocked / paywalled / require auth. Please provide a product brief, screenshots, or hosted assets instead."
upload_asset network / auth failure
If upload_asset fails while converting founder photos, product assets, brand logos, lower-thirds, or fallback local assemblies to hosted URLs, do not continue with local filesystem paths in reference_images or HTML. Retry once only for the 5xx or 429 classes above. For auth_error, unsupported MIME, network failure, or repeated upload failure, stop and ask for a hosted URL or supported raster export.
Long-running task_status exceeding ceiling
Each async MCP call returns either an inline result or {task_id, status} for polling. Use these ceilings before deciding a task is stuck:
- Seedance i2v: 10 min per call
- Kling audio: 5 min per call
- gpt-image-2 high quality: 3 min per call
- render/edit/upload helpers: 5 min per call
Use whichever is earlier: the provider's ceiling x 1.5 or any skill-specific hard polling cap, including the 15 min total cap in the Long-running task_status polling contract above. If task_status returns status: "processing" or status: "queued" past that earlier limit, call task_cancel({task_id}) and surface: "Provider taking unusually long; aborting. Try again."
Load-bearing phrases
These strings go into the SeeDance prompt (or HTML render) verbatim. Each was empirically validated — paraphrasing breaks the recipe silently. When editing prompts, search for these anchors and leave them intact.
What NOT to do
- Don't open the SeeDance prompt with "inside the location matching @Image2" — reproduces the literal backdrop across all acts. Use "in a setting whose visual style ... match @Image2".
- Don't describe the character in prompt prose — @Image1 carries identity. Prose conflicts produce phantom figures or wrong outfits. Exception: the
WARDROBE LOCK:line. - Don't use film-industry shot terms — "Two-shot" / "Three-shot" / "Over-shoulder" / "OTS" / "Master shot" trigger SeeDance phantom-subject artifacts. Describe what the camera sees in plain language.
- Don't render the lower-third as webm — alpha not preserved (HyperFrames emits yuv420p). Use
format: "mov"(ProRes 4444 yuva). Verify withffprobe \| grep pix_fmt. - Don't use
generate_slide_animationfor the end card — that tool's slide-card LLM adds corner clutter and produces animations that don't seek deterministically. Author inline HTML and render viarender_html_animation. - Don't chain pika MCP
edit_text_overlay/ overlay calls for arbitrary text/caption composition — that cascades quality loss and can introduce lip-sync drift. Useadd_captionsfor captions and a single compose/local pass for bounded lower-thirds. Exception: step [5.5] deterministic digital UI overlays deliberately uses one boundededit_video_composepass plus optional exact-brandedit_text_overlay, followed by OCR QA, to prevent Seedance-rendered brand/UI text. - Don't use local
ffmpeg concat -c copyfor final assembly unless MCP is unavailable — the old audio-drop bug was in local concat behavior. Default to MCPedit_concat+edit_audio_mix. - Don't fire SeeDance with identical params across acts — the MCP idempotency cache hashes to the same task ID and replays old results (sometimes failures). Pass unique
seedper act. - Don't use MiniMax for the default generated bed — it does not support
background: trueand can foreground the melody under narration. - Don't manually tile the 10s Kling clip in the skill — the MCP worker owns the local ffmpeg extension path for
duration_seconds: 60. - Don't copy the example music sound for every brand — the recipe is the Kling background-bed call, not the specific instrumentation. Pick a register that matches
brief.tone(see step [7] table).
Engine choice: seedance-only (with caveats)
SeeDance (fal-seedance-2-i2v via generate_reference_video provider: "seedance") is the sole video engine. Picked over alternatives after testing:
- vs Kling v3-omni: Kling has a true
shots[]hard-cut array (cleaner multi-shot) but rejects theseedparameter (cache-busting harder), and 4 × pro 1080p outputs sum >50MB and exceed theedit_concatupload cap (forces local concat). Kling does have more permissive content policy for real-person photos — it's a worth keeping in mind as a fallback if SeeDance's intermittent 422 becomes a hard block. - vs Happy Horse
happyhorse-1.0-r2v(Alibaba DashScope): produced clean 1080p with native lip-sync but the multi-shot prompt direction was weaker and visibly less cinematic than SeeDance. - SeeDance wins because: native
<<<voice_1>>>lip-sync, acceptsseed(cache-busting), permissive enough on real-person photos that 95%+ runs pass content filter, single 15s prompt with time-coded sub-shots gives enough variation for a talking-head register.
If SeeDance is down or its content filter starts rejecting your founder ref repeatedly, the documented fallback is to re-roll the founder portrait with stronger stylization (Pixar / Disney 3D aesthetic). Switching engines mid-pipeline changes too many assumptions in the prompt, concat, and asset-size flow.
Runtime expectations
Wall-clock budget per step. Total run is ~12–18 minutes, dominated by the parallel SeeDance batch.
Defaults
- 4 × 15s SeeDance acts, parallel, unique seeds (101, 202, 303, 404)
- Character identity comes from the
@Image1reference — avoid describing the character in prompt prose (no "Founder Avery", no "young creative streetwear founder"). Open every prompt with "The character (matching @Image1) in a setting whose visual style, palette, lighting and materials match @Image2." Using "inside the location matching @Image2" makes SeeDance reproduce the literal backdrop across all acts. The one exception is aWARDROBE LOCK:line in every act prompt to keep clothing consistent across the 4 separate 15s generations. - Avoid film-industry shot terminology that SeeDance reads literally — never write "Two-shot", "Three-shot", or "Master shot". Shot F is "Brand context shot".
- Every script has a
character_voice_profile(3-4 lines describing default delivery — cadence, signature gestures, pause behavior). Repeated verbatim asCHARACTER VOICE: …in every act's SeeDance prompt. - Every shot has
beats[], not act-levelacting. Each beat hastext+emotion+physical. Silent(beat)entries direct what happens BETWEEN spoken sentences (held looks, micro-expressions, gesture transitions). Beats are emitted as a per-shot "Acting beats" block in the SeeDance prompt. - Every shot beyond the first in an act has
transition_from_prev— a continuous-camera-move description that takes us from the previous framing to this one (dolly, arc, push past, pull back, orbit). Without this, multi-shot acts read as jump zooms because SeeDance reframes the same virtual camera position instead of moving through space. - Apply the TTS pronunciation rewrites to dialogue text before joining into the
<<<voice_1>>>payload (Ari → Airy, API → A P I, example.com → example dot com, etc.). See "TTS pronunciation rewrites" in step [3]. - 16:9, 1080p (SeeDance
resolution: "1080p") - User-provided assets are revealed product-type-appropriately:
digital→ blank on-phone placeholders in shots C and E, then real UI / wordmark composited in step [5.5]physical_apparel→ founder wears/holds the actual t-shirts; assets passed as refs to EVERY shot where the shirt appears (not just one reveal beat)physical_object→ founder holds the product up; assets passed to all shots where product is visibleconsumable→ founder uses/eats/drinks; same patternservice→ no asset reveal; environment + dialogue only
- Music: target ~60s instrumental — call
generate_musicwithprovider: "kling-audio",mode: "text_to_audio",background: true, andduration_seconds: 60. Keepprompt<= 200 chars and include "soft instrumental background bed", "no vocals", and "leave room for narration". Retry once if under 55s, then stop and surface instead of accepting a short bed. - 5s end card via
render_html_animation— author inline HTML per step [9], inline brand-kit fonts as base64. Sources brand fromstate.brand(set in step [4.5]) → real logo, real palette, real fonts. - Captions via
add_captions. Use server-side word-level caption burn-in by default. Font choices are the tool-supported set (inter,bebas-neue,noto-cjk); use brand accent colors for highlight/outline instead of local custom font drawtext. - Lower-third fallback. Off by default. If
state.lower_third = true, render a 5s branded pill bottom-left viarender_html_animation(format:"mov"); the final overlay onto the body usesedit_video_compose, or one local ffmpeg pass only if MCP compose is unavailable. - Final assembly via MCP. Use
edit_concatfor body + end card, thenedit_audio_mixfor music. Local concat/mix is a fallback, not the canonical path. - Provider:
seedanceonly. Reference tokens are@Image1/@Image2/@Image3. Native lip-sync via<<<voice_1>>>...<<<voice_1>>>tokens per sub-shot. Real-person founder photos pass the content filter the vast majority of the time; intermittent 422 → re-roll with stronger stylization.

