Sentencepiece

orchestra-research/ai-research-skills/02-tokenization/sentencepiece

by orchestra-research773a52944ba4MIT13K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 3 months ago

Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.

Instructions onlyAI & Agents

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
references/algorithms.md4.1 KBtext/markdown
references/training.md6.1 KBtext/markdown
SKILL.md5.5 KBtext/markdown

Source and attribution

Source:orchestra-research/ai-research-skillsin02-tokenization/sentencepieceat commit773a529

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal