Sentencepiece

orchestra-research/ai-research-skills/02-tokenization/sentencepiece

作者 orchestra-research773a52944ba4MIT13K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.

僅含說明AI & Agents
  1. 773a52944ba4目前提交 773a529發布於 2026年10月8日

來源與署名

來源:orchestra-research/ai-research-skills位於02-tokenization/sentencepiece提交773a529

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架