OmniVoice TTS Skill
Skill by ara.so — Daily 2026 Skills collection.
OmniVoice is a state-of-the-art zero-shot TTS model supporting 600+ languages, built on a diffusion language model-style architecture. It supports voice cloning (from reference audio), voice design (via text attributes), and auto voice generation with RTF as low as 0.025.
Installation
Requirements
- Python 3.9+
- PyTorch 2.8+
- CUDA (recommended) or Apple Silicon (MPS) or CPU
pip (recommended)
uv
HuggingFace Mirror (if blocked)
Core Concepts
Python API
Load the Model
Voice Cloning
Voice Design
Supported attributes:
- Gender:
male,female - Age:
child,young,middle-aged,elderly - Pitch:
very low pitch,low pitch,high pitch,very high pitch - Style:
whisper - English accents:
american accent,british accent,australian accent, etc. - Chinese dialects:
四川话,陕西话, etc.
Auto Voice
Generation Parameters
Non-Verbal Symbols
Supported tags:
[laughter], [sigh], [confirmation-en], [question-en], [question-ah],
[question-oh], [question-ei], [question-yi], [surprise-ah], [surprise-oh],
[surprise-wa], [surprise-yo], [dissatisfaction-hnn]
Pronunciation Control
CLI Tools
Web Demo
Single Inference
Batch Inference (Multi-GPU)
JSONL format (test.jsonl):
JSONL field reference:
Common Patterns
Full Voice Cloning Pipeline
Batch Processing from a List
Voice Design Combinations
Fast Inference (Lower Diffusion Steps)
Output Format
- Sample rate: 24,000 Hz
- Type:
list[torch.Tensor], each tensor shape(1, T) - Save: use
torchaudio.save(path, audio[0], 24000)
Troubleshooting
HuggingFace download fails
CUDA out of memory
Whisper ASR not available for ref_text auto-transcription
Wrong pronunciation in Chinese
Use inline pinyin with tone numbers directly in the text string:
Audio quality issues
- Increase
num_stepto 32 or 64 - Provide
ref_textmanually instead of relying on auto-transcription - Use a clean, noise-free reference audio clip (3–15 seconds recommended)



