Hardware Performance Counters
Purpose
Guide agents through hardware performance counter analysis: collecting PMU events with perf stat -e, using the PAPI library for portable counter access, interpreting cache miss rates and branch misprediction ratios, computing IPC, and correlating events to source lines with perf annotate.
Triggers
- "How do I measure cache miss rate with perf?"
- "How do I count branch mispredictions?"
- "How do I compute IPC (instructions per clock) with perf?"
- "How do I use the PAPI library for hardware counters?"
- "How do I see which source lines cause the most cache misses?"
- "How do I measure memory bandwidth with performance counters?"
Workflow
1. perf stat — basic counter collection
2. Specifying PMU events with -e
3. Key metrics and thresholds
4. Raw PMU events (CPU-specific)
For events not in the generic aliases, use raw event codes:
5. Source-level annotation with perf record/annotate
6. PAPI — Portable API for hardware counters
PAPI provides a portable C API across different CPU architectures:
Common PAPI presets:
7. Intel PCM (Performance Counter Monitor)
Related skills
- Use
skills/profilers/intel-vtune-amd-uproffor guided microarchitecture analysis - Use
skills/profilers/linux-perffor perf record/report and flamegraph generation - Use
skills/low-level-programming/cpu-cache-optfor applying cache optimization patterns - Use
skills/low-level-programming/simd-intrinsicsfor improving FLOPS/cycle metrics


