Omnivoice Tts

作者 reason-machines2384a003145a無授權條款83 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Expert skill for OmniVoice, a massively multilingual zero-shot TTS model supporting 600+ languages with voice cloning and voice design capabilities.

僅含說明AI & Agents
AI 產生的概覽

指導使用 OmniVoice 多語言零樣本 TTS 模型進行語音複製、音色設計與語音生成。

功能
此技能提供執行 OmniVoice 的說明與程式碼範例,這是一個支援 600 多種語言的零樣本文字轉語音模型。內容涵蓋安裝、Python API、命令列工具、以參考音訊進行語音複製、透過屬性字串進行音色設計、自動音色生成、批次推論以及輸出處理。產出為合成的語音音訊檔案,通常為 24 kHz 的 WAV 輸出。
適用情境
當你需要用 OmniVoice 生成語音、從參考音訊複製音色,或依描述屬性設計合成音色時使用。也適合批次或多語言的語音生成任務。
執行需求
需要 Python 3.9+ 與 PyTorch 2.8+,支援 CUDA、Apple Silicon MPS 或 CPU;需要 omnivoice 套件(或原始碼安裝)與 torchaudio;需要來自 HuggingFace(k2-fsa/OmniVoice)的模型權重,可選擇使用鏡像位址;自動轉寫參考音訊需 openai-whisper。此技能未附帶指令碼,僅為說明文件。

OmniVoice TTS Skill

Skill by ara.so — Daily 2026 Skills collection.

OmniVoice is a state-of-the-art zero-shot TTS model supporting 600+ languages, built on a diffusion language model-style architecture. It supports voice cloning (from reference audio), voice design (via text attributes), and auto voice generation with RTF as low as 0.025.


Installation

Requirements

  • Python 3.9+
  • PyTorch 2.8+
  • CUDA (recommended) or Apple Silicon (MPS) or CPU

pip (recommended)

bash
# Step 1: Install PyTorch for your platform
# NVIDIA GPU (CUDA 12.8)pip install torch==2.8.0+cu128 torchaudio==2.8.0+cu128 --extra-index-url https://download.pytorch.org/whl/cu128
# Apple Siliconpip install torch==2.8.0 torchaudio==2.8.0
# Step 2: Install OmniVoicepip install omnivoice
# Or from source (latest)pip install git+https://github.com/k2-fsa/OmniVoice.git
# Or editable dev installgit clone https://github.com/k2-fsa/OmniVoice.gitcd OmniVoicepip install -e .

uv

bash
git clone https://github.com/k2-fsa/OmniVoice.gitcd OmniVoiceuv sync# With mirror: uv sync --default-index "https://mirrors.aliyun.com/pypi/simple"

HuggingFace Mirror (if blocked)

bash
export HF_ENDPOINT="https://hf-mirror.com"

Core Concepts

ModeWhat you provideUse case
Voice Cloningref_audio + ref_textClone a speaker from a short audio clip
Voice Designinstruct stringDescribe speaker attributes (no audio needed)
Auto Voicenothing extraModel picks a random voice

Python API

Load the Model

python
from omnivoice import OmniVoiceimport torchimport torchaudio
# NVIDIA GPUmodel = OmniVoice.from_pretrained(    "k2-fsa/OmniVoice",    device_map="cuda:0",    dtype=torch.float16)
# Apple Siliconmodel = OmniVoice.from_pretrained(    "k2-fsa/OmniVoice",    device_map="mps",    dtype=torch.float16)
# CPU (slower)model = OmniVoice.from_pretrained(    "k2-fsa/OmniVoice",    device_map="cpu",    dtype=torch.float32)

Voice Cloning

python
# With manual reference transcription (faster, more accurate)audio = model.generate(    text="Hello, this is a test of zero-shot voice cloning.",    ref_audio="ref.wav",    ref_text="Transcription of the reference audio.",)
# Without ref_text — Whisper auto-transcribes ref_audioaudio = model.generate(    text="Hello, this is a test of zero-shot voice cloning.",    ref_audio="ref.wav",)
# audio is a list of torch.Tensor, shape (1, T) at 24kHztorchaudio.save("out.wav", audio[0], 24000)

Voice Design

python
# Describe speaker via comma-separated attributesaudio = model.generate(    text="Hello, this is a test of zero-shot voice design.",    instruct="female, low pitch, british accent",)torchaudio.save("out.wav", audio[0], 24000)

Supported attributes:

  • Gender: male, female
  • Age: child, young, middle-aged, elderly
  • Pitch: very low pitch, low pitch, high pitch, very high pitch
  • Style: whisper
  • English accents: american accent, british accent, australian accent, etc.
  • Chinese dialects: 四川话, 陕西话, etc.

Auto Voice

python
audio = model.generate(text="This is a sentence without any voice prompt.")torchaudio.save("out.wav", audio[0], 24000)

Generation Parameters

python
audio = model.generate(    text="Hello world.",    ref_audio="ref.wav",    ref_text="Reference text.",    num_step=32,      # diffusion steps; use 16 for faster (slightly lower quality)    speed=1.2,        # speaking rate multiplier (>1 faster, <1 slower)    duration=8.0,     # fix output duration in seconds (overrides speed))

Non-Verbal Symbols

python
# Insert expressive non-verbal sounds inlineaudio = model.generate(    text="[laughter] You really got me. I didn't see that coming at all.")

Supported tags: [laughter], [sigh], [confirmation-en], [question-en], [question-ah], [question-oh], [question-ei], [question-yi], [surprise-ah], [surprise-oh], [surprise-wa], [surprise-yo], [dissatisfaction-hnn]

Pronunciation Control

python
# Chinese: pinyin with tone numbers (inline, uppercase)audio = model.generate(    text="这批货物打ZHE2出售后他严重SHE2本了,再也经不起ZHE1腾了。")
# English: CMU dict pronunciation in brackets (uppercase)audio = model.generate(    text="You could probably still make [IH1 T] look good.")

CLI Tools

Web Demo

bash
omnivoice-demo --ip 0.0.0.0 --port 8001omnivoice-demo --help  # all options

Single Inference

bash
# Voice Cloning (ref_text optional; omit for Whisper auto-transcription)omnivoice-infer \    --model k2-fsa/OmniVoice \    --text "This is a test for text to speech." \    --ref_audio ref.wav \    --ref_text "Transcription of the reference audio." \    --output hello.wav
# Voice Designomnivoice-infer \    --model k2-fsa/OmniVoice \    --text "This is a test for text to speech." \    --instruct "male, British accent" \    --output hello.wav
# Auto Voiceomnivoice-infer \    --model k2-fsa/OmniVoice \    --text "This is a test for text to speech." \    --output hello.wav

Batch Inference (Multi-GPU)

bash
omnivoice-infer-batch \    --model k2-fsa/OmniVoice \    --test_list test.jsonl \    --res_dir results/

JSONL format (test.jsonl):

jsonl
{"id": "sample_001", "text": "Hello world", "ref_audio": "/path/to/ref.wav", "ref_text": "Reference transcript"}{"id": "sample_002", "text": "Voice design example", "instruct": "female, british accent"}{"id": "sample_003", "text": "Auto voice example"}{"id": "sample_004", "text": "Speed controlled", "ref_audio": "/path/to/ref.wav", "speed": 1.2}{"id": "sample_005", "text": "Duration fixed", "ref_audio": "/path/to/ref.wav", "duration": 10.0}{"id": "sample_006", "text": "With language hint", "ref_audio": "/path/to/ref.wav", "language_id": "en", "language_name": "English"}

JSONL field reference:

FieldRequiredDescription
id✅Unique identifier
text✅Text to synthesize
ref_audio❌Path to reference audio (voice cloning)
ref_text❌Transcript of ref audio
instruct❌Speaker attributes (voice design)
language_id❌Language code, e.g. "en"
language_name❌Language name, e.g. "English"
duration❌Fixed output duration in seconds
speed❌Speaking rate multiplier (ignored if duration set)

Common Patterns

Full Voice Cloning Pipeline

python
from omnivoice import OmniVoiceimport torchimport torchaudiofrom pathlib import Path
def clone_voice(ref_audio_path: str, texts: list[str], output_dir: str):    model = OmniVoice.from_pretrained(        "k2-fsa/OmniVoice",        device_map="cuda:0",        dtype=torch.float16    )    Path(output_dir).mkdir(parents=True, exist_ok=True)
    for i, text in enumerate(texts):        audio = model.generate(            text=text,            ref_audio=ref_audio_path,            # ref_text omitted: Whisper auto-transcribes            num_step=32,            speed=1.0,        )        out_path = f"{output_dir}/output_{i:04d}.wav"        torchaudio.save(out_path, audio[0], 24000)        print(f"Saved: {out_path}")
clone_voice(    ref_audio_path="speaker.wav",    texts=["Hello world.", "Second sentence.", "Third sentence."],    output_dir="outputs/")

Batch Processing from a List

python
import jsonfrom omnivoice import OmniVoiceimport torchimport torchaudio
model = OmniVoice.from_pretrained("k2-fsa/OmniVoice", device_map="cuda:0", dtype=torch.float16)
items = [    {"id": "s1", "text": "English sentence.", "instruct": "female, american accent"},    {"id": "s2", "text": "Another sentence.", "ref_audio": "ref.wav"},    {"id": "s3", "text": "Auto voice.", },]
for item in items:    kwargs = {"text": item["text"]}    if "ref_audio" in item:        kwargs["ref_audio"] = item["ref_audio"]    if "ref_text" in item:        kwargs["ref_text"] = item["ref_text"]    if "instruct" in item:        kwargs["instruct"] = item["instruct"]
    audio = model.generate(**kwargs)    torchaudio.save(f"{item['id']}.wav", audio[0], 24000)

Voice Design Combinations

python
designs = [    "male, elderly, low pitch",    "female, child, high pitch",    "male, whisper",    "female, british accent, high pitch",    "male, american accent, middle-aged",]
for design in designs:    audio = model.generate(        text="The quick brown fox jumps over the lazy dog.",        instruct=design,    )    safe_name = design.replace(", ", "_").replace(" ", "-")    torchaudio.save(f"design_{safe_name}.wav", audio[0], 24000)

Fast Inference (Lower Diffusion Steps)

python
# Default: num_step=32 (high quality)# Fast: num_step=16 (slightly lower quality, ~2x faster)audio = model.generate(    text="Fast inference example.",    ref_audio="ref.wav",    num_step=16,)

Output Format

  • Sample rate: 24,000 Hz
  • Type: list[torch.Tensor], each tensor shape (1, T)
  • Save: use torchaudio.save(path, audio[0], 24000)

Troubleshooting

HuggingFace download fails

bash
export HF_ENDPOINT="https://hf-mirror.com"

CUDA out of memory

python
# Use float16 (not float32)model = OmniVoice.from_pretrained("k2-fsa/OmniVoice", device_map="cuda:0", dtype=torch.float16)# Or reduce batch size / text length in batch inference

Whisper ASR not available for ref_text auto-transcription

bash
pip install openai-whisper

Wrong pronunciation in Chinese

Use inline pinyin with tone numbers directly in the text string:

python
# Format: PINYINTONE_NUMBER within the sentencetext = "这批货物打ZHE2出售"

Audio quality issues

  • Increase num_step to 32 or 64
  • Provide ref_text manually instead of relying on auto-transcription
  • Use a clean, noise-free reference audio clip (3–15 seconds recommended)

Apple Silicon (MPS) issues

python
# Use mps device explicitlymodel = OmniVoice.from_pretrained("k2-fsa/OmniVoice", device_map="mps", dtype=torch.float16)

Model & Resources

ResourceLink
HuggingFace Modelk2-fsa/OmniVoice
HuggingFace Spacehttps://huggingface.co/spaces/k2-fsa/OmniVoice
Paper (arXiv)https://arxiv.org/abs/2604.00688
Demo Pagehttps://zhu-han.github.io/omnivoice
Supported Languagesdocs/languages.md in repo
Voice Design Attributesdocs/voice-design.md in repo
Generation Parametersdocs/generation-parameters.md in repo
Training/Eval Examplesexamples/ in repo

來源與署名

來源:reason-machines/trending-skills位於skills/omnivoice-tts提交2384a00

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架