Huggingface Tokenizers

orchestra-research/ai-research-skills/02-tokenization/huggingface-tokenizers

作者 orchestra-research773a52944ba4MIT13K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫3 個月前更新

Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.

僅公開檔案列表。將技能安裝到工作區後即可檢視檔案內容。

路徑大小類型
references/algorithms.md14.8 KBtext/markdown
references/integration.md15 KBtext/markdown
references/pipeline.md16.4 KBtext/markdown
references/training.md14.2 KBtext/markdown
SKILL.md13.4 KBtext/markdown

來源與署名

來源:orchestra-research/ai-research-skills位於02-tokenization/huggingface-tokenizers提交773a529

授權條款: MIT

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架