CPU Cache Optimization
Purpose
Guide agents through cache-aware programming: diagnosing cache misses with perf, data layout transformations (AoS→SoA), false sharing detection and fixes, prefetching, and cache-friendly algorithm design.
Triggers
- "My program has high cache miss rates — how do I fix it?"
- "What is false sharing and how do I detect it?"
- "Should I use AoS or SoA data layout?"
- "How do I measure cache performance with perf?"
- "How do I use __builtin_prefetch?"
- "My multithreaded program is slower than single-threaded due to cache"
Workflow
1. Measure cache performance
2. Cache line basics
- Cache line size: 64 bytes on x86-64, ARM (most platforms)
- L1 cache: 32–64 KB, ~4 cycles latency
- L2 cache: 256 KB–1 MB, ~12 cycles latency
- L3 cache: 6–64 MB, ~40 cycles latency
- Main memory: ~200–300 cycles latency
3. AoS vs SoA data layout
4. Common cache-unfriendly patterns
5. False sharing
False sharing occurs when two threads write to different variables that share a cache line, causing constant cache-line invalidations.
6. Prefetching
Manual prefetch hints to hide memory latency:
Prefetching rules:
- Prefetch too early = cache evicted before use
- Prefetch too late = no benefit
- Prefetch distance = memory latency / time per iteration (typically 8–32 elements)
7. Cache-friendly algorithm design
For perf cache event reference and false sharing detection patterns, see references/cache-counters.md [blocked].
Related skills
- Use
skills/profilers/linux-perfforperf statandperf recordcache measurements - Use
skills/profilers/valgrind— cachegrind simulates cache behaviour - Use
skills/low-level-programming/simd-intrinsics— SoA layout pairs with SIMD vectorization - Use
skills/low-level-programming/memory-modelfor false sharing in concurrent contexts


