Tts

marswaveai/skills/tts

作者 marswaveai0dd6cecda529无许可证78 个星标收录于 2026年10月9日更新于 2026年10月9日仓库2周前更新

Text-to-speech and voice narration. Triggers on: "朗读这段", "配音", "TTS", "语音合成", "text to speech", "read this aloud", "convert to speech", "voice narration", "read aloud".

仅含说明Design & Creative
AI 生成的概览

通过 listenhub CLI 将文本转换为语音,支持单音色快速模式与多角色脚本模式。

功能
借助 listenhub 命令行工具把文本转换为自然流畅的语音音频。它提供用于短文本的单音色快速模式,以及为多个角色分配不同音色的脚本模式,适用于对话或有声书。它可以返回音频链接,按需将 MP3 下载到当前工作目录,脚本模式下还会给出字幕与时长。它会把音色与输出偏好保存在本地配置文件中。
适用场景
当用户希望朗读文本、把文字转成语音或制作语音旁白时使用。它也适合多角色脚本音频或对话。它不适用于播客式讨论、讲解视频或图像生成。
运行要求
需要 PATH 中提供 listenhub CLI 可执行文件、已登录的 listenhub 账号以及访问该服务的网络。它使用 jq、curl 等 shell 工具。该技能不附带脚本,仅为说明文档,并包含共享参考文档。

When to Use

  • User wants to convert text to spoken audio
  • User asks for "read aloud", "TTS", "text to speech", "voice narration"
  • User says "朗读", "配音", "语音合成"
  • User wants multi-speaker scripted audio or dialogue

When NOT to Use

  • User wants a podcast-style discussion with topic exploration (use /podcast)
  • User wants an explainer video with visuals (use /explainer)
  • User wants to generate an image (use /image-gen)

Purpose

Convert text into natural-sounding speech audio. Two paths:

  1. Quick mode (--mode direct): Single voice, low-latency, sync. For casual chat, reading snippets, instant audio.
  2. Script mode (--mode smart): Multi-speaker, per-segment voice assignment. For dialogue, audiobooks, scripted content.

Hard Constraints

  • Always check CLI auth following shared/cli-authentication.md
  • Follow shared/cli-patterns.md for CLI execution, errors, and interaction patterns
  • Never hardcode speaker IDs in CLI calls — use built-in defaults from shared/speaker-selection.md as fallback only; fetch from the speakers CLI when the user wants to change voice
  • Always read config following shared/config-pattern.md before any interaction
  • Always follow shared/speaker-selection.md for speaker selection (text table + free-text input)
  • Never save files to ~/Downloads/ or /tmp/ as primary output — save artifacts to the current working directory with friendly topic-based names (see shared/config-pattern.md § Artifact Naming)

<HARD-GATE>

Use the AskUserQuestion tool for every multiple-choice step — do NOT print options as plain text. Ask one question at a time. Wait for the user's answer before proceeding to the next step. After all parameters are collected, summarize the choices and ask the user to confirm. Do NOT call any generation CLI command until the user has explicitly confirmed.

</HARD-GATE>

Mode Detection

Determine the mode from the user's input automatically before asking any questions:

SignalMode
"多角色", "脚本", "对话", "script", "dialogue", "multi-speaker"Script
Multiple characters mentioned by name or roleScript
Input contains structured segments (A: ..., B: ...)Script
Single paragraph of text, no character markersQuick
"读一下", "read this", "TTS", "朗读" with plain textQuick
AmbiguousQuick (default)

Interaction Flow

Step -1: CLI Auth Check

Follow shared/cli-authentication.md. If the CLI is not installed or the user is not logged in, auto-install and auto-login — never ask the user to run commands manually.

Then follow shared/cli-authentication.md § Auth Mode Detection to determine AUTH_MODE and set:

bash
if [ "$AUTH_MODE" = "openapi" ]; then  CMD_PREFIX="listenhub openapi tts"else  CMD_PREFIX="listenhub tts"fi

All subsequent CLI calls use $CMD_PREFIX instead of hardcoded listenhub tts.

Step 0: Config Setup

Follow shared/config-pattern.md Step 0 (Zero-Question Boot).

If file doesn't exist — silently create with defaults and proceed:

bash
mkdir -p ".listenhub/tts"echo '{"outputMode":"inline","language":null,"defaultSpeakers":{}}' > ".listenhub/tts/config.json"CONFIG_PATH=".listenhub/tts/config.json"CONFIG=$(cat "$CONFIG_PATH")

Do NOT ask any setup questions. Proceed directly to the Interaction Flow.

If file exists — read config silently and proceed:

bash
CONFIG_PATH=".listenhub/tts/config.json"[ ! -f "$CONFIG_PATH" ] && CONFIG_PATH="$HOME/.listenhub/tts/config.json"CONFIG=$(cat "$CONFIG_PATH")

Setup Flow (user-initiated reconfigure only)

Only run when the user explicitly asks to reconfigure. Display current settings:

当前配置 (tts):  输出方式:{inline / download / both}  语言偏好:{zh / en / 未设置}  默认主播:{speakerName / 使用内置默认}

Then ask:

  1. outputMode: Follow shared/output-mode.md § Setup Flow Question.

  2. Language (optional): "默认语言?"

    • "中文 (zh)"
    • "English (en)"
    • "每次手动选择" → keep null

After collecting answers, save immediately:

bash
NEW_CONFIG=$(echo "$CONFIG" | jq --arg m "$OUTPUT_MODE" '. + {"outputMode": $m}')# Save language if user chose one (not "每次手动选择")if [ "$LANGUAGE" != "null" ]; then  NEW_CONFIG=$(echo "$NEW_CONFIG" | jq --arg lang "$LANGUAGE" '. + {"language": $lang}')fiecho "$NEW_CONFIG" > "$CONFIG_PATH"CONFIG=$(cat "$CONFIG_PATH")

Generation Speed

Default: 1.0x (original speed). Never ask about speed.

Only pass --speed when the user explicitly asks for a faster or slower reading — "慢一点"、"快一点"、"1.25 倍速"、"read it faster". Otherwise omit the flag.

  • Range: any value from 0.5 to 2.0, at most two decimals — a continuous range, not fixed steps.
  • Common values: 0.5, 0.75, 1 (default), 1.25, 1.5, 2; in-between values such as 0.85 or 1.35 work too.
  • Meaning: the speaking rate of the generated audio, not a player playback rate.
  • Map vague wording conservatively: "慢一点" → 0.85, "快一点" → 1.25, "慢很多" → 0.5, "快很多" → 1.75. A number the user names is passed through unchanged.

Show the speed in the confirmation summary only when it is not 1.

Quick Mode — $CMD_PREFIX create --mode direct

Step 1: Extract text

Get the text to convert. If the user hasn't provided it, ask:

"What text would you like me to read aloud?"

Step 2: Determine voice

  • If config.defaultSpeakers.{language}[0] is set → use it silently (skip to Step 4)
  • If not set → use the built-in default from shared/speaker-selection.md for the detected language (skip to Step 4)
  • Only show speaker selection if the user explicitly asks to change voice

Step 3: Save preference

After the user explicitly selects a new voice (not when using defaults):

Question: "Save {voice name} as your default voice for {language}?"Options:  - "Yes" — update .listenhub/tts/config.json  - "No" — use for this session only

Step 4: Confirm

Ready to generate:
  Text: "{first 80 chars}..."  Voice: {voice name}  Speed: {speed}x        # omit this line when speed is 1
Proceed?

Step 5: Generate

For short text, pass inline:

bash
RESULT=$($CMD_PREFIX create --text "{text}" --mode direct --speaker "{name}" --lang {lang} [--speed {0.5-2.0}] --json 2>/tmp/lh-err)EXIT_CODE=$?
if [ $EXIT_CODE -ne 0 ]; then  ERROR=$(cat /tmp/lh-err)  case $EXIT_CODE in    2) echo "Auth error: run 'listenhub auth login'" ;;    3) echo "Timeout: try --no-wait" ;;    *) echo "Error: $ERROR" ;;  esac  rm -f /tmp/lh-errfirm -f /tmp/lh-err
AUDIO_URL=$(echo "$RESULT" | jq -r '.audioUrl')

For long text, write to a temp file first (see shared/cli-patterns.md § Long Text Input):

bash
cat > /tmp/lh-content.txt << 'ENDCONTENT'Long text content goes here...ENDCONTENT
RESULT=$($CMD_PREFIX create --text "$(cat /tmp/lh-content.txt)" --mode direct --speaker "{name}" --lang {lang} [--speed {0.5-2.0}] --json)AUDIO_URL=$(echo "$RESULT" | jq -r '.audioUrl')
rm -f /tmp/lh-content.txt

Step 6: Present result

Read OUTPUT_MODE from config. Follow shared/output-mode.md for behavior.

inline or both: Display the audioUrl as a clickable link.

Present:

Audio generated!
在线收听:{audioUrl}

download or both: Also download the file. Generate a topic slug from the text content following shared/config-pattern.md § Artifact Naming.

bash
SLUG="{topic-slug}"  # e.g. "server-maintenance-notice"NAME="${SLUG}.mp3"# Dedup: if file exists, append -2, -3, etc.BASE="${NAME%.*}"; EXT="${NAME##*.}"; i=2while [ -e "$NAME" ]; do NAME="${BASE}-${i}.${EXT}"; i=$((i+1)); donecurl -sS -o "$NAME" "$AUDIO_URL"

Present:

Audio generated!
已保存到当前目录:  {NAME}

Script Mode — $CMD_PREFIX create --mode smart

Step 1: Get scripts

Determine whether the user already has a scripts array:

  • Already provided (JSON or clear segments): parse and display for confirmation

  • Not yet provided: help the user structure segments. Ask:

    "Please provide the script with speaker assignments. Format: each line as SpeakerName: text content. I'll convert it."

    Once the user provides the script, parse it into speaker-annotated text.

Step 2: Assign voices per character

For each unique character in the script:

  • If config.defaultSpeakers.{language} has saved voices → auto-assign silently (one per character in order)
  • If not set → use built-in defaults from shared/speaker-selection.md (Primary for first character, Secondary for second)
  • Only show speaker selection if the user explicitly asks to change voices

Step 3: Save preferences

After all voices are assigned (if any were new):

Question: "Save these voice assignments for future sessions?"Options:  - "Yes" — update defaultSpeakers in .listenhub/tts/config.json  - "No" — use for this session only

Step 4: Confirm

Ready to generate:
  Characters:    {name}: {voice}    {name}: {voice}  Segments: {count}  Speed: {speed}x        # omit this line when speed is 1  Title: (auto-generated)
Proceed?

Step 5: Generate

Format the script text with speaker markers and submit. For multi-speaker scripts, include speaker names inline in the text. Run with run_in_background: true since script mode may take longer.

Submit (foreground) with --no-wait:

bash
RESULT=$($CMD_PREFIX create --text "{formatted script with speaker markers}" --mode smart --speaker "{name1}" --speaker "{name2}" --lang {lang} [--speed {0.5-2.0}] --no-wait --json)ID=$(echo "$RESULT" | jq -r '.id')echo "Submitted: $ID"

For long scripts, write to a temp file first:

bash
cat > /tmp/lh-content.txt << 'ENDCONTENT'SpeakerA: First line of dialogueSpeakerB: Second line of dialogue...ENDCONTENT
RESULT=$($CMD_PREFIX create --text "$(cat /tmp/lh-content.txt)" --mode smart --speaker "{name1}" --speaker "{name2}" --lang {lang} [--speed {0.5-2.0}] --no-wait --json)ID=$(echo "$RESULT" | jq -r '.id')
rm -f /tmp/lh-content.txt

Poll (background) with run_in_background: true and timeout: 600000:

bash
ID="<id-from-above>"for i in $(seq 1 60); do  RESULT=$(listenhub creation get "$ID" --json 2>/dev/null)  STATUS=$(echo "$RESULT" | jq -r '.status // "processing"')
  case "$STATUS" in    completed) echo "$RESULT"; exit 0 ;;    failed) echo "FAILED: $RESULT" >&2; exit 1 ;;    *) sleep 10 ;;  esacdoneecho "TIMEOUT" >&2; exit 2

Step 6: Present result

When the background task completes, parse the result:

bash
AUDIO_URL=$(echo "$RESULT" | jq -r '.audioUrl')SUBTITLES_URL=$(echo "$RESULT" | jq -r '.subtitlesUrl // empty')DURATION=$(echo "$RESULT" | jq -r '.audioDuration // empty')CREDITS=$(echo "$RESULT" | jq -r '.credits // empty')

Read OUTPUT_MODE from config. Follow shared/output-mode.md for behavior.

inline or both: Display the audioUrl and subtitlesUrl as clickable links.

Present:

Audio generated!
在线收听:{audioUrl}字幕:{subtitlesUrl}时长:{audioDuration / 1000}s消耗积分:{credits}

download or both: Also download the file. Generate a topic slug following shared/config-pattern.md § Artifact Naming.

bash
SLUG="{topic-slug}"  # e.g. "welcome-dialogue"NAME="${SLUG}.mp3"# Dedup: if file exists, append -2, -3, etc.BASE="${NAME%.*}"; EXT="${NAME##*.}"; i=2while [ -e "$NAME" ]; do NAME="${BASE}-${i}.${EXT}"; i=$((i+1)); donecurl -sS -o "$NAME" "$AUDIO_URL"

Present:

已保存到当前目录:  {NAME}

Updating Config

When saving preferences, merge into .listenhub/tts/config.json — do not overwrite unchanged keys.

  • Quick voice: set defaultSpeakers.{language}[0] to the selected speakerId
  • Script voices: set defaultSpeakers.{language} to the full array assigned this session
  • Language: set language if the user explicitly specifies it

API Reference

  • CLI execution patterns: shared/cli-patterns.md
  • CLI authentication: shared/cli-authentication.md
  • Speaker list: shared/cli-speakers.md
  • Speaker selection guide: shared/speaker-selection.md
  • Config pattern: shared/config-pattern.md
  • Output mode: shared/output-mode.md

Composability

  • Invokes: speakers CLI (for speaker selection)
  • Invoked by: explainer (for voiceover)

Examples

Quick mode:

"TTS this: The server will be down for maintenance at midnight."

  1. Detect: Quick mode (plain text, "TTS this")
  2. Read config: defaultSpeakers.en is empty
  3. Use built-in default: Mars (cozy-man-english)
  4. Confirm → user approves
  5. Generate:
    bash
    RESULT=$($CMD_PREFIX create --text "The server will be down for maintenance at midnight." --mode direct --speaker "Mars" --lang en --json)AUDIO_URL=$(echo "$RESULT" | jq -r '.audioUrl')
  6. Present: display audioUrl as link (inline mode)

Script mode:

"帮我做一段双人对话配音,A说:欢迎大家,B说:谢谢邀请"

  1. Detect: Script mode ("双人对话")
  2. Parse segments: A -> "欢迎大家", B -> "谢谢邀请"
  3. Read config: defaultSpeakers.zh empty
  4. Use built-in defaults: 原野 (Primary) + 高晴 (Secondary)
  5. Confirm → user approves
  6. Generate:
    bash
    RESULT=$($CMD_PREFIX create --text "A: 欢迎大家B: 谢谢邀请" --mode smart --speaker "原野" --speaker "高晴" --lang zh --no-wait --json)ID=$(echo "$RESULT" | jq -r '.id')
  7. Poll in background until complete
  8. Present: audioUrl, subtitlesUrl, duration

来源与署名

来源:marswaveai/skills位于tts提交0dd6cec

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架