When to Use
- User wants to convert text to spoken audio
- User asks for "read aloud", "TTS", "text to speech", "voice narration"
- User says "朗读", "配音", "语音合成"
- User wants multi-speaker scripted audio or dialogue
When NOT to Use
- User wants a podcast-style discussion with topic exploration (use
/podcast) - User wants an explainer video with visuals (use
/explainer) - User wants to generate an image (use
/image-gen)
Purpose
Convert text into natural-sounding speech audio. Two paths:
- Quick mode (
--mode direct): Single voice, low-latency, sync. For casual chat, reading snippets, instant audio. - Script mode (
--mode smart): Multi-speaker, per-segment voice assignment. For dialogue, audiobooks, scripted content.
Hard Constraints
- Always check CLI auth following
shared/cli-authentication.md - Follow
shared/cli-patterns.mdfor CLI execution, errors, and interaction patterns - Never hardcode speaker IDs in CLI calls — use built-in defaults from
shared/speaker-selection.mdas fallback only; fetch from the speakers CLI when the user wants to change voice - Always read config following
shared/config-pattern.mdbefore any interaction - Always follow
shared/speaker-selection.mdfor speaker selection (text table + free-text input) - Never save files to
~/Downloads/or/tmp/as primary output — save artifacts to the current working directory with friendly topic-based names (seeshared/config-pattern.md§ Artifact Naming)
<HARD-GATE>
Use the AskUserQuestion tool for every multiple-choice step — do NOT print options as plain text. Ask one question at a time. Wait for the user's answer before proceeding to the next step. After all parameters are collected, summarize the choices and ask the user to confirm. Do NOT call any generation CLI command until the user has explicitly confirmed.
</HARD-GATE>
Mode Detection
Determine the mode from the user's input automatically before asking any questions:
Interaction Flow
Step -1: CLI Auth Check
Follow shared/cli-authentication.md. If the CLI is not installed or the user is not logged in, auto-install and auto-login — never ask the user to run commands manually.
Then follow shared/cli-authentication.md § Auth Mode Detection to determine AUTH_MODE and set:
All subsequent CLI calls use $CMD_PREFIX instead of hardcoded listenhub tts.
Step 0: Config Setup
Follow shared/config-pattern.md Step 0 (Zero-Question Boot).
If file doesn't exist — silently create with defaults and proceed:
Do NOT ask any setup questions. Proceed directly to the Interaction Flow.
If file exists — read config silently and proceed:
Setup Flow (user-initiated reconfigure only)
Only run when the user explicitly asks to reconfigure. Display current settings:
Then ask:
-
outputMode: Follow
shared/output-mode.md§ Setup Flow Question. -
Language (optional): "默认语言?"
- "中文 (zh)"
- "English (en)"
- "每次手动选择" → keep
null
After collecting answers, save immediately:
Generation Speed
Default: 1.0x (original speed). Never ask about speed.
Only pass --speed when the user explicitly asks for a faster or slower reading —
"慢一点"、"快一点"、"1.25 倍速"、"read it faster". Otherwise omit the flag.
- Range: any value from
0.5to2.0, at most two decimals — a continuous range, not fixed steps. - Common values:
0.5,0.75,1(default),1.25,1.5,2; in-between values such as0.85or1.35work too. - Meaning: the speaking rate of the generated audio, not a player playback rate.
- Map vague wording conservatively: "慢一点" →
0.85, "快一点" →1.25, "慢很多" →0.5, "快很多" →1.75. A number the user names is passed through unchanged.
Show the speed in the confirmation summary only when it is not 1.
Quick Mode — $CMD_PREFIX create --mode direct
Step 1: Extract text
Get the text to convert. If the user hasn't provided it, ask:
"What text would you like me to read aloud?"
Step 2: Determine voice
- If
config.defaultSpeakers.{language}[0]is set → use it silently (skip to Step 4) - If not set → use the built-in default from
shared/speaker-selection.mdfor the detected language (skip to Step 4) - Only show speaker selection if the user explicitly asks to change voice
Step 3: Save preference
After the user explicitly selects a new voice (not when using defaults):
Step 4: Confirm
Step 5: Generate
For short text, pass inline:
For long text, write to a temp file first (see shared/cli-patterns.md § Long Text Input):
Step 6: Present result
Read OUTPUT_MODE from config. Follow shared/output-mode.md for behavior.
inline or both: Display the audioUrl as a clickable link.
Present:
download or both: Also download the file. Generate a topic slug from the text content following shared/config-pattern.md § Artifact Naming.
Present:
Script Mode — $CMD_PREFIX create --mode smart
Step 1: Get scripts
Determine whether the user already has a scripts array:
-
Already provided (JSON or clear segments): parse and display for confirmation
-
Not yet provided: help the user structure segments. Ask:
"Please provide the script with speaker assignments. Format: each line as
SpeakerName: text content. I'll convert it."Once the user provides the script, parse it into speaker-annotated text.
Step 2: Assign voices per character
For each unique character in the script:
- If
config.defaultSpeakers.{language}has saved voices → auto-assign silently (one per character in order) - If not set → use built-in defaults from
shared/speaker-selection.md(Primary for first character, Secondary for second) - Only show speaker selection if the user explicitly asks to change voices
Step 3: Save preferences
After all voices are assigned (if any were new):
Step 4: Confirm
Step 5: Generate
Format the script text with speaker markers and submit. For multi-speaker scripts, include speaker names inline in the text. Run with run_in_background: true since script mode may take longer.
Submit (foreground) with --no-wait:
For long scripts, write to a temp file first:
Poll (background) with run_in_background: true and timeout: 600000:
Step 6: Present result
When the background task completes, parse the result:
Read OUTPUT_MODE from config. Follow shared/output-mode.md for behavior.
inline or both: Display the audioUrl and subtitlesUrl as clickable links.
Present:
download or both: Also download the file. Generate a topic slug following shared/config-pattern.md § Artifact Naming.
Present:
Updating Config
When saving preferences, merge into .listenhub/tts/config.json — do not overwrite unchanged keys.
- Quick voice: set
defaultSpeakers.{language}[0]to the selectedspeakerId - Script voices: set
defaultSpeakers.{language}to the full array assigned this session - Language: set
languageif the user explicitly specifies it
API Reference
- CLI execution patterns:
shared/cli-patterns.md - CLI authentication:
shared/cli-authentication.md - Speaker list:
shared/cli-speakers.md - Speaker selection guide:
shared/speaker-selection.md - Config pattern:
shared/config-pattern.md - Output mode:
shared/output-mode.md
Composability
- Invokes: speakers CLI (for speaker selection)
- Invoked by: explainer (for voiceover)
Examples
Quick mode:
"TTS this: The server will be down for maintenance at midnight."
- Detect: Quick mode (plain text, "TTS this")
- Read config:
defaultSpeakers.enis empty - Use built-in default: Mars (
cozy-man-english) - Confirm → user approves
- Generate:
- Present: display
audioUrlas link (inline mode)
Script mode:
"帮我做一段双人对话配音,A说:欢迎大家,B说:谢谢邀请"
- Detect: Script mode ("双人对话")
- Parse segments: A -> "欢迎大家", B -> "谢谢邀请"
- Read config:
defaultSpeakers.zhempty - Use built-in defaults: 原野 (Primary) + 高晴 (Secondary)
- Confirm → user approves
- Generate:
- Poll in background until complete
- Present:
audioUrl,subtitlesUrl, duration

