Sentencepiece

orchestra-research/ai-research-skills/02-tokenization/sentencepiece

作者 orchestra-research773a52944ba4MIT13K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.

僅含說明AI & Agents
AI 產生的概覽

指導訓練與使用 SentencePiece 分詞器,實現與語言無關的多語言文字分詞。

功能
此技能說明如何安裝與使用 SentencePiece,在原始文字上訓練 BPE 或 Unigram 模型,並將文字編碼或解碼為子詞片段或 ID。內容涵蓋詞彙量、字元覆蓋率、使用者自訂符號與子詞正則化等設定,並提供效能基準與支援的模型。它產出訓練好的分詞器模型與分詞後的文字,而非獨立應用程式。
適用情境
適用於建置多語言或中日韓分詞器、需要可重現的確定性分詞,或希望在無需語言特定前置處理的原始文字上訓練的情境。也適用於使用依賴 SentencePiece 的 T5、ALBERT、XLNet 或 mBART 等模型時。
執行需求
需要 sentencepiece 套件,可選用 transformers;C++ 建置方式需要 CMake 與編譯器。訓練與分詞需要本機語料資料。此技能僅附參考文件,不含指令碼。

SentencePiece - Language-Independent Tokenization

Unsupervised tokenizer that works on raw text without language-specific preprocessing.

When to use SentencePiece

Use SentencePiece when:

  • Building multilingual models (no language-specific rules)
  • Working with CJK languages (Chinese, Japanese, Korean)
  • Need reproducible tokenization (deterministic vocabulary)
  • Want to train on raw text (no pre-tokenization needed)
  • Require lightweight deployment (6MB memory, 50k sentences/sec)

Performance:

  • Speed: 50,000 sentences/sec
  • Memory: ~6MB for loaded model
  • Languages: All (language-independent)

Use alternatives instead:

  • HuggingFace Tokenizers: Faster training, more flexibility
  • tiktoken: OpenAI models (GPT-3.5/4)
  • BERT WordPiece: English-centric tasks

Quick start

Installation

bash
# Pythonpip install sentencepiece
# C++ (requires CMake)git clone https://github.com/google/sentencepiece.gitcd sentencepiecemkdir build && cd buildcmake .. && make -j $(nproc)sudo make install

Train model

bash
# Command-line (BPE with 8000 vocab)spm_train --input=data.txt --model_prefix=m --vocab_size=8000 --model_type=bpe
# Python APIimport sentencepiece as spm
spm.SentencePieceTrainer.train(    input='data.txt',    model_prefix='m',    vocab_size=8000,    model_type='bpe')

Training time: ~1-2 minutes for 100MB corpus

Encode and decode

python
import sentencepiece as spm
# Load modelsp = spm.SentencePieceProcessor(model_file='m.model')
# Encode to piecespieces = sp.encode('This is a test', out_type=str)print(pieces)  # ['▁This', '▁is', '▁a', '▁test']
# Encode to IDsids = sp.encode('This is a test', out_type=int)print(ids)  # [284, 47, 11, 1243]
# Decodetext = sp.decode(ids)print(text)  # "This is a test"

Language-independent design

Whitespace as symbol (▁)

python
text = "Hello world"pieces = sp.encode(text, out_type=str)print(pieces)  # ['▁Hello', '▁world']
# Decode preserves spacesdecoded = sp.decode_pieces(pieces)print(decoded)  # "Hello world"

Key principle: Treat text as raw Unicode, whitespace = ▁ (meta symbol)

Tokenization algorithms

BPE (Byte-Pair Encoding)

python
spm.SentencePieceTrainer.train(    input='data.txt',    model_prefix='bpe_model',    vocab_size=16000,    model_type='bpe')

Used by: mBART

Unigram (default)

python
spm.SentencePieceTrainer.train(    input='data.txt',    model_prefix='unigram_model',    vocab_size=8000,    model_type='unigram')

Used by: T5, ALBERT, XLNet

Training configuration

Essential parameters

python
spm.SentencePieceTrainer.train(    input='corpus.txt',    model_prefix='m',    vocab_size=32000,    model_type='unigram',    character_coverage=0.9995,  # 1.0 for CJK    user_defined_symbols=['[SEP]', '[CLS]'],    unk_piece='<unk>',    num_threads=16)

Character coverage

Language TypeCoverageRationale
English0.9995Most common chars
CJK (Chinese)1.0All characters needed
Multilingual0.9995Balance

Encoding options

Subword regularization

python
# Sample different tokenizationsfor _ in range(3):    pieces = sp.encode('tokenization', out_type=str, enable_sampling=True, alpha=0.1)    print(pieces)
# Output (different each time):# ['▁token', 'ization']# ['▁tok', 'en', 'ization']

Use case: Data augmentation for robustness.

Common patterns

T5-style training

python
spm.SentencePieceTrainer.train(    input='c4_corpus.txt',    model_prefix='t5',    vocab_size=32000,    model_type='unigram',    user_defined_symbols=[f'<extra_id_{i}>' for i in range(100)],    unk_id=2,    eos_id=1,    pad_id=0)

Integration with transformers

python
from transformers import T5Tokenizer
# T5 uses SentencePiece internallytokenizer = T5Tokenizer.from_pretrained('t5-base')inputs = tokenizer('translate English to French: Hello', return_tensors='pt')

Performance benchmarks

Training speed

CorpusBPE (16k)Unigram (8k)
100 MB1-2 min3-4 min
1 GB10-15 min30-40 min

Tokenization speed

  • SentencePiece: 50,000 sentences/sec
  • HF Tokenizers: 200,000 sentences/sec (4× faster)

Supported models

T5 family: t5-base, t5-large (32k vocab, Unigram) ALBERT: albert-base-v2 (30k vocab, Unigram) XLNet: xlnet-base-cased (32k vocab, Unigram) mBART: facebook/mbart-large-50 (250k vocab, BPE)

References

  • Training Guide [blocked] - Detailed options, corpus preparation
  • Algorithms [blocked] - BPE vs Unigram, subword regularization

Resources

來源與署名

來源:orchestra-research/ai-research-skills位於02-tokenization/sentencepiece提交773a529

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架