Interpreters and Bytecode VMs
Purpose
Guide agents through implementing efficient bytecode interpreters and simple JITs in C/C++: dispatch strategies, VM architecture choices, and performance patterns.
Triggers
- "How do I implement a fast bytecode dispatch loop?"
- "What is the difference between switch dispatch and computed goto?"
- "How do I implement a register-based vs stack-based VM?"
- "How do I add basic JIT compilation to my interpreter?"
- "Why is my interpreter slow?"
Workflow
1. VM architecture choice
Stack-based: easier to implement, compile to; code generation is simpler. More instructions per expression. Register-based: fewer dispatch iterations; needs register allocation in the compiler; better cache behaviour for complex expressions.
2. Dispatch loop strategies
Switch dispatch (simplest, baseline)
Problem: switch compiles to a single indirect branch from a jump table. Modern CPUs can mispredict it heavily because the same indirect branch is used for all opcodes.
Computed goto (GCC/Clang extension — fastest portable approach)
Each opcode ends with its own indirect branch. The CPU can train the branch predictor per-opcode, dramatically improving prediction rates.
Note: &&label is a GCC/Clang extension, not standard C. Use #ifdef __GNUC__ to guard and fall back to switch for other compilers.
Direct threaded code (most aggressive)
Each bytecode word is a function pointer or label address; the VM is the fetch-decode-execute loop itself.
3. Value representation
Tagged pointer: Store type tag in low bits of pointer (pointer alignment guarantees ≥ 2 bits free):
NaN boxing (64-bit): Store non-double values in NaN bit patterns:
Used by V8 (formerly), LuaJIT, JavaScriptCore.
4. Stack management
5. Inline caching (IC)
Inline caching speeds up property lookups and method dispatch by caching the last observed type at each call site.
Polymorphic IC (PIC): cache up to N (typ. 4) type-method pairs.
6. Simple JIT (mmap + machine code)
For x86-64: allocate executable memory, write machine code bytes, call it.
On macOS Apple Silicon (M-series): use pthread_jit_write_protect_np() or MAP_JIT flag.
7. Performance tips
- Dispatch: Use computed goto over switch on GCC/Clang
- Values: Use NaN boxing or tagged pointers; avoid boxing/unboxing in hot paths
- Stack: Keep stack pointer in a callee-saved register (
register Value *sp asm("r15")— GCC global register variable) - Locals access: Keep frequently accessed locals in VM registers (struct fields), not stack
- Profiling: Use
perfor sampling to find dispatch overhead vs actual work - Specialisation: Generate specialised handler variants for common type combinations (int+int add vs generic add)
- Trace recording: Trace JITs (LuaJIT approach) compile hot traces instead of full functions
For a benchmark of dispatch strategies, see references/benchmarks.md [blocked].
Related skills
- Use
skills/profilers/linux-perfto profile the interpreter dispatch loop - Use
skills/low-level-programming/assembly-x86to understand JIT output - Use
skills/runtimes/fuzzingto fuzz the bytecode parser/loader - Use
skills/compilers/llvmfor LLVM IR-based JIT (MCJIT / ORC JIT)


