🎬 Seedance 2.0 Cinema Expert
The definitive skill for "Director-Level" AI video orchestration. Seedance 2.0 is not a descriptive model; it is an instructional model. It responds best to technical cinematography, physics directives, and precise camera grammar.
Core Competencies
- Text-to-Video (t2v): Generate cinematic video from a Director Brief — Chinese, Global, or VIP tier.
- Image-to-Video (i2v): Animate 1–9 reference images — Chinese, Global (smart mode), or VIP tier.
- Video Extension (extend): Seamlessly continue an existing Seedance 2.0 video (Chinese tier).
- First & Last Frame (first-last): Interpolate a fluid video between a start image and end image (Global/VIP).
- Omni Reference (omni): Full multimodal reference with images + audio + character refs (all tiers).
- Omni Reference Training (omni-train): Train a custom persistent character for identity-consistent generation.
- Character Sheet (character): Build a reusable character from 1–3 images (Chinese tier).
- Video Edit (video-edit): Edit an existing video with a prompt + optional reference images (Chinese tier).
- Watermark Removal (watermark-remove): Strip Seedance 2.0 watermarks (basic or Pro).
🏷️ Tiers
Add --fast to any Global or VIP call to use the fast-queue variant (lower latency, same quality).
📥 Input Limits
Output: 4–15 seconds, auto-generated sound, 480p–720p.
⚠️ Restrictions
- No realistic human faces in uploaded images/videos (except character/omni-train modes).
--mode extendrequires arequest_idfrom a priorseedance-v2.0-t2vorseedance-v2.0-i2vjob.--mode first-lastrequires--tier globalor--tier vip.- Global/VIP omni does not support video references (images + audio only).
--qualityapplies to Chinese tier only.
🔗 Core Syntax: The @ Reference System
Assign explicit roles to each uploaded asset. Tags differ by mode.
Chinese Tier (i2v, omni)
Global/VIP Omni (omni-reference-no-video / vip-omni-reference)
Character References (all tiers)
Role Assignment Table
Multi-Reference Combination
🏗️ Technical Specification: The Director Brief
Structure prompts using this six-component hierarchy. Order matters — composition first, texture and micro-motion last:
Seedance 2.0 generates audio natively. Always include an Audio directive — even one sentence. Without it the model generates random ambient sound that may not match your scene.
Time-Segmented Prompts (Recommended for 10s+ videos)
Break prompts into timed segments for precise control:
Single-beat rule: Each segment should contain one action. 4–7s = one beat. 10–15s = 3–4 beats maximum. Overloading a segment with multiple narrative changes degrades output quality.
Negative Prompting
Seedance 2.0 supports appending negative guidance directly in the prompt. Use plain language at the end:
Common negative additions:
Avoid: abrupt cuts, scene changes, multiple locations.(for single-take shots)Avoid: human faces, realistic people.(for product-only content)Avoid: fast motion, blur, unstable framing.(for smooth product reveals)
🎥 Camera Language Reference
Basic Movements
Advanced Techniques
Shot Sizes
🧠 Prompt Optimization Protocol
The Agent MUST transform user intent into a technical "Director Brief" before execution.
- Technical Grammar: Use camera terms: Dolly In/Out, Crane Shot, Whip Pan, Tracking Shot, Anamorphic Lens, Shallow Depth of Field, High-Speed Dive, Orbital Arc.
- Physics Directives: Use "caustic patterns," "volumetric rays," or "subsurface scattering" instead of "good lighting."
- Timecode Notation: For multi-beat scenes, use
[00:00-00:05s]format to specify timing. - Tag References: If files provided, use: "Replicate the camera movement of @video1 while maintaining the visual style of @image1." (lowercase, 1-based index)
- ORDER MATTERS: Tokens at the start define composition; tokens at the end define texture and micro-motion.
- Multi-Image i2v: Provide up to 9 reference images. The model blends aspects (style, identity, environment) across all inputs.
- Audio is mandatory: Seedance 2.0 generates audio natively. Always include an Audio line — music genre/tone, key SFX, ambient texture. Silent direction = random audio.
- Single-beat discipline: Each timed segment = one action. Cramming two narrative beats into 4s degrades physics and motion consistency.
🎭 Capability-Specific Patterns
1. Character Consistency
2. Camera Movement Replication
3. Video Extension (Forward)
4. Video Extension (Reverse / Prepend)
5. Video Editing (Modify Existing)
6. Music Beat-Matching
7. Dialogue / Voice Acting
8. One-Take / Long Take
9. E-commerce / Product Showcase
10. Science / Educational Visualization
11. FPV First-Person Shot
12. Cinematic Drone Flythrough
🎨 Prompt Templates
Cinematic Film
Product Ad (15s)
Short Drama (15s)
Dance / Beat-Sync (13s)
Scenery Montage (15s)
Advertising / Product Motion
Action / Physics
Character Consistency (Martial Arts)
🎚️ Style & Quality Modifiers
Visual Style
Cinematic quality, film grain, shallow depth of field2.35:1 widescreen, 24fpsInk wash painting style/Anime style/PhotorealisticHigh saturation neon colors, cool-warm contrast4K medical CGI, semi-transparent visualization
Mood / Atmosphere
Tense and suspenseful/Warm and healing/Epic and grandComedy with exaggerated expressionsDocumentary tone, restrained narration
Audio Direction
Background music: grand and majesticSound effects: footsteps, crowd noise, car soundsVoice tone reference @Video1Beat-synced transitions matching music rhythm
❌ Common Mistakes to Avoid
- Vague references: Don't say "reference @Video1" — specify WHAT to reference (camera? action? effects? rhythm?)
- Conflicting instructions: Don't ask for "static camera" and "orbit shot" in the same segment.
- Overloading: Don't pack too many scenes into 4–5 seconds — keep it physically plausible.
- Missing @ assignments: If you upload 5 images, make sure each one is referenced with a clear purpose.
- Ignoring audio: Sound design dramatically improves output — always include audio direction.
- Forgetting duration: Match prompt complexity to the selected generation length.
- Real faces: Don't upload real human photos — the system will block them.
- Keyword soup: DO NOT use "8k, masterpiece, trending." Use technical descriptions instead.
- Discontinuous action: Avoid "The man runs and then he stops." Use fluid transitional language.
- Missing audio direction: Seedance 2.0 generates audio natively — always specify music tone, SFX, or ambience. Skipping it produces random sound.
- Narrative overload per segment: Each timed segment should contain one action beat. Multiple scene changes in 4s produce degraded physics and motion artifacts.
- FPV without continuous motion: FPV requires a motion-rich environment to work — a static room with FPV intent will not trigger the immersive effect. Pair FPV with corridors, streets, natural terrain, or product flyovers.
- Drone without a destination: Drone shots need a resolve point — specify what the camera descends toward or arrives at. "Drone shot" alone produces aimless floating.
🚀 Protocol: All Modes
Mode 1: Text-to-Video (t2v)
Mode 2: Image-to-Video (i2v)
Mode 3: Extend Video (Chinese tier)
Mode 4: First & Last Frame (Global/VIP)
Mode 5: Omni Reference (omni)
Mode 6: Train Omni Reference Character (omni-train)
Mode 7: Character Sheet (character, Chinese tier)
Mode 8: Video Edit (video-edit, Chinese tier)
Mode 9: Watermark Removal (watermark-remove)
Async Pattern
⚙️ Implementation Details
Endpoint Reference
Parameter Differences by Tier
This skill acts as a Cinematographic Wrapper that translates creative intent into high-fidelity technical instructions for the muapi core.

