Huggingface Tokenizers

orchestra-research/ai-research-skills/02-tokenization/huggingface-tokenizers

by orchestra-research773a52944ba4MIT13K starsListed Oct 8, 2026Updated Oct 8, 2026Repository updated 3 months ago

Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.

Instructions onlySoftware Development

Only the file list is public. File contents are available once the skill is installed in a workspace.

PathSizeType
references/algorithms.md14.8 KBtext/markdown
references/integration.md15 KBtext/markdown
references/pipeline.md16.4 KBtext/markdown
references/training.md14.2 KBtext/markdown
SKILL.md13.4 KBtext/markdown

Source and attribution

Source:orchestra-research/ai-research-skillsin02-tokenization/huggingface-tokenizersat commit773a529

License: MIT

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal