HappyHorse Prompt Studio
A 4-phase guided skill that turns "I want to make a video" into a production-ready HappyHorse prompt — starting from inspiration, not from a blank page.
Overview
This skill guides the Agent through a structured conversation:
Phase 1 · Inspiration Menu (灵感菜单)
Start every conversation here. Before asking any questions, show the user what HappyHorse can do. Present these as four "flavors" — each one a door into a different creative world.
Use the language the user is using (JP/CN/EN). The descriptions below are in English for the Agent's reference — translate them to match the user's language.
Flavor A · "让你的角色开口说话"
Voiced Manga Drama (漫画配音剧 / ボイスコミック)
You have a manga, a webtoon, or an original story. You've drawn the characters, written the dialogue — now you want them to speak.
Upload 2-3 character reference images + a short script. HappyHorse generates a 15-30 second voiced drama where characters talk, emote, and stay visually consistent across cuts. Lip-sync included.
Vibe: Movie dub meets manga animation. Your characters, their voices.
Flavor B · "一张立绘,开口自我介绍"
Character Voice PV (角色语音 PV / キャラボイス PV)
You have a game character, a VTuber, or an original OC. You want a short 8-10 second PV where they introduce themselves — or let out a battle cry.
Upload 1-3 character art images + a line or two. HappyHorse generates a voiced, lip-synced character PV.
Vibe: Character reveal trailer. One illustration, one voice, one PV.
Flavor C · "让一格漫画活过来"
Manga Panel Motion (漫画分格动态化 / コマ動画化)
You have manga panels, comic pages, or illustrated scenes. You want to turn them into 5-10 second motion clips — perfect for social media.
Upload one panel as the first frame. HappyHorse animates it while preserving your art style.
Vibe: Your drawing, but it breathes. Hair moves, eyes blink, wind blows.
Flavor D · "你的虚拟偶像,30 秒成 MV"
Virtual Idol MV (虚拟偶像 MV / バーチャルアイドル MV)
You have a virtual idol, a VTuber group, or an original idol project. You want an MV — with stage lighting, lip-sync singing, and multi-shot choreography.
Upload 3-5 multi-angle character images + a licensed song segment. HappyHorse generates a 30-second MV clip.
Vibe: Your idol, center stage. No Live2D. No MMD. Just one prompt.
⚠️ Note: This scenario requires the strongest compliance guardrails. We'll check together.
How to present the menu
Present the four flavors conversationally, not as a dry list. Something like:
"HappyHorse can bring your characters to life in a few different ways. Think of it as four flavors:
A · Voiced Drama — your manga characters talk to each other, with voice and lip-sync B · Character PV — your game character or OC introduces itself out loud C · Panel Motion — a single manga panel starts moving, hair blowing, eyes blinking D · Idol MV — your virtual idol performs a 30-second MV on stage
Which one sounds closest to what you're imagining? Or tell me about your project and I'll suggest."
If the user already knows what they want, skip to Phase 2.
Phase 2 · Discovery (需求发现)
Once a flavor is chosen (or the user describes their own scenario), ask these questions. Ask them conversationally, not as a form. Group related questions together.
2.1 Character & World (角色与世界)
- What's your character's name and role? (protagonist / antagonist / side character)
- What do they look like? (hair, eyes, outfit, accessories, any signature items)
- What's their personality vibe? (cool / energetic / shy / mysterious / cheerful)
- Where does the scene take place? (school rooftop / fantasy castle / neon city / café / etc.)
2.2 Scene Intent (场景意图)
- What's happening in this scene? (a confession / a battle / a quiet moment / a group dance)
- What emotion should the viewer feel? (heart-fluttering / adrenaline / nostalgic / hype / calm)
- How long should the output be? (5s / 10s / 15s / 30s)
2.3 Voice & Sound (声音与音频)
- Does your character speak? If yes:
- What language? (Japanese / Chinese / English)
- Voice type? (young woman / young man / child / mature / elderly)
- Voice color? (bright / low / soft / powerful / cool)
- What do they say? (provide the exact line, or ask me to suggest)
- Background audio? (silence / ambient sounds / BGM style)
2.4 Visual Style (视觉风格)
- Art style reference? (anime / photorealistic / Pixar / watercolor / pixel art / etc.)
- Color palette? (warm / cool / neon / pastel / high-contrast)
- Camera preference? (close-up / medium / wide / rotating / slow push / static)
2.5 Compliance Quick-Check (合规快检)
Before proceeding, verify:
- ☐ Is the character your own original creation or properly licensed?
- ☐ Is the character depicted as 18 or older (especially for idol scenarios)?
- ☐ Is the outfit SFW (no suggestive or revealing clothing)?
- ☐ Is the scene SFW (no sensitive locations like bedrooms/pools)?
- ☐ If there's music, is it licensed or original (not a commercial song)?
If any answer is NO, pause and suggest an alternative — don't proceed with a non-compliant prompt.
Phase 3 · Prompt Assembly (Prompt 组装)
Now build the prompt using the HappyHorse Formula:
3.1 The Formula (公式)
3.2 R2V Character Consistency Syntax
When the user provides multiple reference images, use this syntax:
Or when referencing a specific character in a multi-character scene:
Key rules:
- Always use
@「Image n」to lock character identity across shots - Describe what each reference image shows (正面 / 側面 / 表情差分)
- End with:
キャラの顔・髪・衣装が変わらない(character's face/hair/outfit stays unchanged)
3.3 Video-Edit Style Unification
When the user wants to unify style across multiple shots:
Key rule: always add 100% 保持 (100% preserved) constraints for things that must not change.
3.4 Language Rules
Japanese-specific tips:
- Use
ネイティブな日本語to ensure natural Japanese (not translation-style) - Specify voice color with JP adjectives:
明るく元気な少女声,低めの落ち着いた青年声,柔らかい囁くような声 - Keep dialogue in
「」brackets - Avoid mixing languages in dialogue unless intentionally bilingual
3.5 Prompt Templates by Flavor
Flavor A · Voiced Manga Drama
Flavor B · Character Voice PV
Flavor C · Manga Panel Motion
Flavor D · Virtual Idol MV
3.6 Assembling the Output
Present the final prompt in a code block so the user can copy it directly. Include:
- The prompt itself (in the user's language)
- A brief breakdown of what each part does
- Suggested model variant (t2v / i2v / r2v / video-edit)
- Estimated cost (720P: ¥0.9/sec, 1080P: ¥1.6/sec)
Example output format:
[PROMPT HERE]
Phase 4 · Quality Check (质量检查)
Before finalizing, run through this checklist silently. If anything fails, fix before presenting.
4.1 Prompt Quality
- ☐ Does the prompt follow the Scene + Subject + Motion + Audio + Quality structure?
- ☐ Is the camera movement explicitly stated? (not left to chance)
- ☐ Is the voice type described with specific adjectives? (not vague)
- ☐ Is the dialogue in the correct brackets for the language? (「」 for JP, "" for EN)
- ☐ Is the "stays unchanged" constraint included at the end?
- ☐ Is the prompt length between 150-300 characters? (too short = under-specified; too long = hard to control)
4.2 Compliance Check
- ☐ No existing anime/manga/game IP referenced?
- ☐ No real person likeness?
- ☐ Character depicted as adult?
- ☐ Outfit is SFW?
- ☐ Scene location is SFW?
- ☐ If music is involved, it's licensed/original?
4.3 Optimization Tips
If the prompt looks good, offer these pro-tips:
- "Try 3 variants" — HappyHorse results vary; generating 3-5 and picking the best is standard practice
- "Start 720P, finish 1080P" — do test runs at 720P (cheaper), then re-generate the winner at 1080P
- "Shorter lines = better lip-sync" — if the voice line is over 15 characters, consider splitting into two shots
- "Specific beats vague" — "camera slowly pushes from full-body to chest close-up" beats "camera moves"
Free-Form Mode (自由模式)
If the user's scenario doesn't fit Flavors A-D, use the formula directly:
- Ask: "What's your scene? Describe it like you're telling a friend about a movie you just watched."
- Extract: scene, subject, motion, audio, quality from their description
- Assemble using the formula
- Apply the quality check
This mode is especially useful for:
- Product advertisements
- Educational explainers
- Abstract / artistic videos
- Non-character-driven content
Common Pitfalls (常见问题)
CLI Quick-Start (for users who want to run it immediately)
If the user has bailian-cli installed, they can run the prompt directly:
Example Interactions
Example 1 · First-time user (flavor A)
昼休みの学校の屋上、青空と白い雲、風が心地よい。 桜色のロングヘアの少女がフェンスに寄りかかり、こちらを見て笑っている。
少女が手を振り、カメラがゆっくり寄る。 [少女、ネイティブな日本語、明るく元気な若い女性声、嬉しそう] 言う: 「ねえ!来てくれたんだ!」
背景に風の音、遠くで校庭のざわめき、明るいピアノの BGM。 映画級質感、キャラの顔・髪・制服が変わらない。
Example 2 · Experienced user (free-form)
Final Notes for the Agent
- Always start with Phase 1 unless the user is clearly experienced and already knows what they want
- Be creative with descriptions — don't just ask "what's the scene?", say "paint the picture for me — where are we, what time of day, what's the vibe?"
- Suggest, don't just ask — if the user seems unsure, offer defaults: "How about a sunset rooftop scene with a gentle breeze?"
- Show the prompt in a code block so it's easy to copy
- Always offer to iterate — "Want me to adjust the voice tone? Change the camera angle? Add a second character?"
- Keep compliance friendly, not scary — "Just to make sure everything's smooth, is this your original character?" not "COMPLIANCE CHECK: CONFIRM IP STATUS"


