Long Context: Extending Transformer Context Windows
When to Use This Skill
Use Long Context techniques when you need to:
- Process long documents (32k, 64k, 128k+ tokens) with transformer models
- Extend context windows of pre-trained models (LLaMA, Mistral, etc.)
- Implement efficient positional encodings (RoPE, ALiBi)
- Train models with length extrapolation capabilities
- Deploy models that handle variable-length inputs efficiently
- Fine-tune existing models for longer contexts with minimal compute
Key Techniques: RoPE (Rotary Position Embeddings), YaRN, ALiBi (Attention with Linear Biases), Position Interpolation
Papers: RoFormer (arXiv 2104.09864), YaRN (arXiv 2309.00071), ALiBi (arXiv 2108.12409), Position Interpolation (arXiv 2306.15595)
Installation
Quick Start
RoPE (Rotary Position Embeddings)
ALiBi (Attention with Linear Biases)
Position Interpolation for LLaMA
Core Concepts
1. RoPE (Rotary Position Embeddings)
How it works:
- Encodes absolute position via rotation matrix
- Provides relative position dependency in attention
- Enables length extrapolation
Mathematical formulation:
Advantages:
- Decaying inter-token dependency with distance
- Compatible with linear attention
- Better extrapolation than absolute position encodings
2. YaRN (Yet another RoPE extensioN)
Key innovation:
- NTK-aware interpolation (Neural Tangent Kernel)
- Attention temperature scaling
- Efficient context extension (10× less tokens vs baselines)
Parameters:
Performance:
- Extends LLaMA to 128k tokens
- 2.5× less training steps than baselines
- State-of-the-art context window extension
3. ALiBi (Attention with Linear Biases)
Core idea:
- No positional embeddings added to tokens
- Apply distance penalty directly to attention scores
- Bias proportional to key-query distance
Formula:
Advantages:
- 11% faster training vs sinusoidal embeddings
- 11% less memory usage
- Strong length extrapolation (train 1k, test 2k+)
- Inductive bias towards recency
4. Position Interpolation
Technique:
- Linearly down-scale position indices
- Interpolate within trained range (vs extrapolate beyond)
- Minimal fine-tuning required
Formula:
Results:
- LLaMA 7B-65B extended to 32k tokens
- 1000 fine-tuning steps sufficient
- 600× better stability than extrapolation
Method Comparison
Implementation Patterns
HuggingFace Transformers Integration
Custom RoPE Implementation
Fine-tuning for Long Context
Minimal Fine-tuning (Position Interpolation)
YaRN Fine-tuning
Best Practices
1. Choose the Right Method
2. Scaling Factor Selection
3. Fine-tuning Data
4. Avoid Common Pitfalls
Production Deployment
Inference with Long Context
Memory Optimization
Resources
- RoPE Paper: https://arxiv.org/abs/2104.09864 (RoFormer)
- YaRN Paper: https://arxiv.org/abs/2309.00071
- ALiBi Paper: https://arxiv.org/abs/2108.12409 (Train Short, Test Long)
- Position Interpolation: https://arxiv.org/abs/2306.15595
- HuggingFace RoPE Utils: https://github.com/huggingface/transformers/blob/main/src/transformers/modeling_rope_utils.py
- YaRN Implementation: https://github.com/jquesnelle/yarn
- Together AI Blog: https://www.together.ai/blog/llama-2-7b-32k
See Also
references/rope.md- Detailed RoPE implementation and theoryreferences/extension_methods.md- YaRN, ALiBi, Position Interpolation comparisonsreferences/fine_tuning.md- Complete fine-tuning guide for context extension


