Linux perf
Purpose
Guide agents through perf for CPU profiling: sampling, hardware counter measurement, hotspot identification, and integration with flamegraph generation.
Triggers
- "Which function is consuming the most CPU?"
- "How do I measure cache misses / IPC?"
- "How do I use
perfto find hotspots?" - "How do I generate a flamegraph from perf data?"
- "perf shows
[unknown]or[kernel]frames"
Workflow
1. Prerequisites
Compile the target with debug symbols for useful frame data:
2. perf stat — quick counters
Interpret perf stat output:
- IPC (instructions per cycle) < 1.0: memory-bound or stalled pipeline
- cache-miss rate > 5%: significant cache pressure
- branch-miss rate > 5%: branch predictor struggling
3. perf record — sampling
4. perf report — interactive analysis
Navigation in TUI:
Enter— expand a symbola— annotate (show assembly with hit counts)s— show source (needs debug info)d— filter by DSO (library)t— filter by thread?— help
5. perf annotate — hot instructions
High hit count on a mov or vmovdqa suggests a cache miss at that load.
6. perf top — live profiling
7. Feed into flamegraphs
See skills/profilers/flamegraphs for reading flamegraphs and interpreting results.
8. Common issues
9. Useful events
For a counter reference and interpretation guide, see references/events.md [blocked].
Related skills
- Use
skills/profilers/flamegraphsfor SVG flamegraph generation and reading - Use
skills/profilers/valgrindfor cache simulation and memory profiling - Use
skills/compilers/gccorskills/compilers/clangfor PGO from perf data (AutoFDO)


