Audio Generation

PostPlusAI/postplus-skills/skills/10-routing/audio-generation

作者 PostPlusAI7f28d1494958f69136a7d6fe85fdea3942a300f3无许可证收录于 2026年10月9日更新于 2026年10月9日

Plan TTS, voice cloning, voice change, translated dub, or lip-sync audio. Resolve voice and reference policy before handing a ready request to voice-batch-runner.

仅含说明AI & Agents
AI 生成的概览

规划语音合成、声音克隆、配音与唇形同步音频请求,并交给执行运行器处理。

功能
该控制类技能将音频请求归类为 tts、change_voice、translate_dub、voice_clone_take、podcast_audio 或 lip_sync_handoff 等任务类别。它会确定脚本、声音与参考策略,包括哪些参考音频具有约束力、哪些仅作灵感参考。随后生成包含 taskClass、scriptPolicy、voicePolicy、referencePolicy、runnerHandoff、nextVideoHandoff 和 mustNotDo 等字段的交接产物。它本身不提交任务。
适用场景
当最终产物是生成的音频,或为视频渲染准备的音频时使用,包括语音合成、声音设计、声音克隆、变声、翻译配音、播客音频或唇形同步交接。不适用于转写已有音频,或已规范化可直接执行的请求。
运行要求
不附带脚本或工具,完全由指令驱动。需要用户意图、来源证据和自有输入产物,并会交接给 voice-batch-runner、video-batch-runner 或 audio-transcription 等配套技能。

Audio Generation

Use When

  • The desired final asset is generated audio or audio prepared for a video render.
  • The request includes TTS, voice design, voice cloning, voice change, translated dub, podcast audio, or lip-sync handoff.
  • The next decision is audio task class, reference policy, and runner handoff.

Do Not Use When

  • The user needs speech-to-text from existing audio. Use audio-transcription; use its local subtitle reference when an existing timed transcript needs subtitle files.
  • The voice request is already normalized for execution. Use voice-batch-runner.
  • The final work is a full video production pipeline. Use video-batch-runner after the audio handoff is clear.

Core Boundary

This is the audio generation controller. It does not submit jobs.

It must classify the task and hand off execution. It must not let a runner invent voice strategy, translation policy, or lip-sync intent.

Task Classes

Task classUse whenHandoff
ttsnew spoken audio from scriptvoice-batch-runner with voice design rules
change_voicepreserve script, alter voice identity or deliveryreference contract, then voice-batch-runner
translate_dubtranslate and dub source audiorequire language, meaning-preservation, and timing policy
voice_clone_takeapproved reference voice should preserve timbrebind reference audio, then voice-batch-runner
podcast_audiospeaker-led or conversational audiocreate voice/script handoff before video assembly
lip_sync_handoffaudio drives talking-head or UGC rendervoice-batch-runner, then video-batch-runner

Reference Rules

  • Approved voice reference audio is binding.
  • Accent, energy, cadence, or genre examples are inspiration-only unless the user explicitly binds them.
  • Source audio used only for translation meaning is not a voice identity binding unless stated.
  • Excluded voices, music, or effects must not enter the runner request.

Routing Table

If not audio-generationSend to
Transcribe existing audioaudio-transcription
Need generated image/video around audiovideo-batch-runner
Need normalized hosted voice executionvoice-batch-runner
Need lip-sync video after audiovideo-batch-runner

Output Shape

Return:

  • taskClass
  • scriptPolicy
  • voicePolicy
  • referencePolicy
  • runnerHandoff
  • nextVideoHandoff when lip-sync or video assembly follows
  • mustNotDo

Stop Conditions

  • Stop when required user intent, source evidence, or owned input artifacts are missing and guessing would change the result.
  • Do not ask voice-batch-runner to decide the creative role of the voice.

Public Command Boundary

  • Choose the smallest matching command or workflow from the user input and run it directly.

  • This public skill is instruction-driven. Produce the controller handoff artifact directly from the available evidence.

  • Do not call private provider/runtime paths or unpublished local tools.

  • If the CLI returns a quote-confirmation challenge, obtain user approval for its scope and cost before running postplus quote confirm --json --challenge-file <challenge.json> and retry with the returned token.

来源与署名

来源:PostPlusAI/postplus-skills位于skills/10-routing/audio-generation提交7f28d14

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架