Model Merging: Combining Pre-trained Models
When to Use This Skill
Use Model Merging when you need to:
- Combine capabilities from multiple fine-tuned models without retraining
- Create specialized models by blending domain-specific expertise (math + coding + chat)
- Improve performance beyond single models (often +5-10% on benchmarks)
- Reduce training costs - no GPUs needed, merges run on CPU
- Experiment rapidly - create new model variants in minutes, not days
- Preserve multiple skills - merge without catastrophic forgetting
Success Stories: Marcoro14-7B-slerp (best on Open LLM Leaderboard 02/2024), many top HuggingFace models use merging
Tools: mergekit (Arcee AI), LazyMergekit, Model Soup
Installation
Quick Start
Simple Linear Merge
SLERP Merge (Best for 2 Models)
Core Concepts
1. Merge Methods
Linear (Model Soup)
- Simple weighted average of parameters
- Fast, works well for similar models
- Can merge 2+ models (
w1 + w2 + ... = 1)
SLERP (Spherical Linear Interpolation)
- Interpolates along sphere in weight space
- Preserves magnitude of weight vectors
- Best for merging 2 models
- Smoother than linear
Task Arithmetic
- Extract "task vectors" (fine-tuned - base)
- Combine task vectors, add to base
- Good for merging multiple specialized models (
merged = base + α₁·tv₁ + α₂·tv₂)
TIES-Merging
- Task arithmetic + sparsification
- Resolves sign conflicts in parameters
- Best for merging many task-specific models
DARE (Drop And REscale)
- Randomly drops fine-tuned parameters
- Rescales remaining parameters
- Reduces redundancy, maintains performance
2. Configuration Structure
Merge Methods Guide
Linear Merge
Best for: Simple model combinations, equal weighting
SLERP Merge
Best for: Two models, smooth interpolation
Layer-specific SLERP:
Task Arithmetic
Best for: Combining specialized skills
TIES-Merging
Best for: Many models, resolving conflicts
DARE Merge
Best for: Reducing redundancy
Advanced Patterns
Layer-wise Merging
MoE from Merged Models
Tokenizer Merging
Best Practices
1. Model Compatibility
2. Weight Selection
Unsupervised Coefficient Tuning (no labeled data needed)
Instead of manual search, use generation consistency: merge with several candidate coefficients, generate responses on a small unlabeled subset, and pick the coefficient whose outputs are most similar to those of its neighbors. Consistent outputs signal a stable, well-performing merge region (AdaMMS, arXiv:2503.23733).
See references/coefficient-tuning.md [blocked] for the full algorithm, similarity metrics, multi-coefficient search, and end-to-end pipeline.
3. Method Selection
4. Density Tuning (TIES/DARE)
5. Layer-specific Merging
Preserve the base model's first/last layers (often best left untouched) and merge only the middle via merge_method: passthrough with slices — see the Layer-wise Merging pattern above.
Evaluation & Testing
Benchmark Merged Models
Common Benchmarks
- Open LLM Leaderboard: General capabilities
- MT-Bench: Multi-turn conversation
- MMLU: Multitask accuracy
- HumanEval: Code generation
- GSM8K: Math reasoning
Production Deployment
Save and Upload
Quantize Merged Model
Common Pitfalls
- Mismatched architectures — only merge models that share the same architecture (e.g., don't mix Llama and Mistral).
- Over-weighting one model (e.g.,
0.95 / 0.05) — keep weights balanced, typically in the 0.3–0.7 range. - Skipping evaluation — always benchmark a merged model before deploying (see the Evaluation & Testing section above).
Resources
- mergekit GitHub: https://github.com/arcee-ai/mergekit
- HuggingFace Tutorial: https://huggingface.co/blog/mlabonne/merge-models
- LazyMergekit: Automated merging notebook
- TIES Paper: https://arxiv.org/abs/2306.01708
- DARE Paper: https://arxiv.org/abs/2311.03099
See Also
references/methods.md- Deep dive into merge algorithmsreferences/examples.md- Real-world merge configurationsreferences/evaluation.md- Benchmarking and testing strategiesreferences/coefficient-tuning.md- Unsupervised coefficient search via generation consistency (AdaMMS, arXiv:2503.23733)


