dflash-mlx Speculative Decoding
Skill by ara.so — Daily 2026 Skills collection.
DFlash implements lossless speculative decoding for MLX on Apple Silicon. A small draft model (~1B params) generates 16 tokens in parallel using block diffusion; the target model verifies all 16 in a single forward pass. Tokens are only emitted after target verification — output is lossless (every token is the target model's greedy argmax).
Typical speedups: 1.7x–4.1x over baseline mlx_lm depending on model size and context length. Acceptance rates hover around 87–90% for Qwen3.5 models.
Installation
Requires Python 3.10+, MLX 0.31.1+, Apple Silicon Mac.
Key CLI Commands
Generate text
OpenAI-compatible server
Benchmark
Outputs per-run JSON reports with tok/s, acceptance rate, and speedup vs baseline.
Supported Model Pairs
Draft models are auto-resolved from a registry — no --draft flag needed for listed pairs. Models without a matching draft are rejected at startup.
Python API Usage
Streaming generation
Full generation with stats
Custom draft block size and context
OpenAI client against dflash-serve
Tool calling (via dflash-serve)
Common Patterns
Side-by-side demo (baseline vs DFlash)
Integrating with Open WebUI
- Start
dflash-serve --model Qwen/Qwen3.5-9B --port 8000 - In Open WebUI settings → Connections → add OpenAI API with URL
http://localhost:8000/v1 - Select model
Qwen/Qwen3.5-9Bin the chat UI
Works the same for Continue, aider, OpenCode, and any OpenAI-compatible client.
Override draft for unsupported models
Disable thinking tokens for Qwen3.5
Architecture Notes
- Tape-replay rollback: For hybrid GatedDeltaNet + attention models (Qwen3.5), dflash records an innovation tape during verify and replays only accepted steps via a custom Metal kernel — avoids full state snapshots.
- JIT SDPA 2-pass: For contexts ≥ 1024 tokens, a custom Metal attention kernel maintains numerical alignment with stock MLX attention.
- Greedy acceptance: Keeps the longest correct prefix from the 16 drafted tokens, rejects the rest. No temperature/sampling on verification — strictly lossless.
- Qwen3 (pure attention) models work but don't benefit from tape-replay rollback (that's GatedDeltaNet-specific).
Troubleshooting
Model rejected at startup
→ Pass --draft org/ModelName-DFlash explicitly, or use a model from the supported pairs table.
Low acceptance rate (< 80%)
- Usually caused by very long context (4096+). Try
--dflash-max-ctx 8192to extend the fallback threshold. - Qwen3 (non-3.5) models have lower acceptance than Qwen3.5 hybrid models.
Numerical divergence / output differs from pure AR
- Expected behavior: "Output can still differ from pure AR because of MLX dispatch divergence, but no unverified token is ever emitted."
- If outputs seem wrong (not just different), ensure MLX 0.31.1+ is installed:
python -c "import mlx; print(mlx.__version__)"
Server not accepting connections
Out of memory with large models
- Use 4-bit quantized variants:
mlx-community/Qwen3.5-27B-4bitinstead of the full model. - The draft model loads alongside the target — budget ~1–2GB extra for the draft.
Benchmark results JSON location


