Hardware Counters

mohitmishra786/low-level-dev-skills/skills/profilers/hardware-counters

作者 mohitmishra786bdc58472fa9f無授權條款253 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫3 個月前更新

Hardware performance counter skill for low-level CPU analysis. Use when collecting PMU events with perf stat, using the PAPI library, measuring cache miss rates and branch misprediction ratios, computing IPC, or correlating PMU events to source lines. Activates on queries about hardware counters, PMU events, perf stat -e, PAPI, cache miss rate, branch misprediction, IPC measurement, or CPU performance events.

AI 產生的概覽

指導使用 perf stat、PAPI 與 Intel PCM 進行硬體效能計數器分析,量測 IPC、快取未命中等 CPU 指標。

功能
此技能提供在 CPU 上蒐集與解讀硬體效能計數器(PMU)資料的指引。內容涵蓋 perf stat 事件蒐集、原始 PMU 事件代碼、使用 perf record 與 annotate 進行原始碼層級標註、PAPI C 函式庫以及 Intel PCM。它說明 IPC、快取未命中率、分支預測錯誤率與 MPKI 等衍生指標,並提供健康與需留意數值的參考門檻。
適用情境
適用於使用硬體計數器量測 CPU 效能的情境,例如蒐集 PMU 事件、計算 IPC、檢查快取未命中率或分支預測錯誤率,以及將 PMU 事件關聯到原始碼行。也適用於使用 perf stat、PAPI 或 Intel PCM 的情況。
執行需求
需要 Linux perf 工具(perf stat、perf record、perf annotate);PAPI 範例還需要 PAPI 函式庫與標頭檔以及 C 編譯器;Intel PCM 部分需要從原始碼建置。依核心設定而定,部分計數器蒐集可能需要提高權限。此技能不附帶指令碼,僅為說明文件。

Hardware Performance Counters

Purpose

Guide agents through hardware performance counter analysis: collecting PMU events with perf stat -e, using the PAPI library for portable counter access, interpreting cache miss rates and branch misprediction ratios, computing IPC, and correlating events to source lines with perf annotate.

Triggers

  • "How do I measure cache miss rate with perf?"
  • "How do I count branch mispredictions?"
  • "How do I compute IPC (instructions per clock) with perf?"
  • "How do I use the PAPI library for hardware counters?"
  • "How do I see which source lines cause the most cache misses?"
  • "How do I measure memory bandwidth with performance counters?"

Workflow

1. perf stat — basic counter collection

bash
# Basic hardware event summaryperf stat ./prog
# Output:#  Performance counter stats for './prog':##      1,234,567,890      instructions#        456,789,012      cycles#         12,345,678      cache-misses         #    1.23 % of all cache refs#         23,456,789      branch-misses        #    2.34 % of all branches##       0.456789012 seconds time elapsed
# Derived metrics (computed from the output)# IPC = instructions / cycles = 1,234,567,890 / 456,789,012 ≈ 2.70# CPI = cycles / instructions ≈ 0.37

2. Specifying PMU events with -e

bash
# Specific hardware eventsperf stat -e instructions,cycles,cache-misses,branch-misses ./prog
# L1/L2/L3 cache eventsperf stat -e \  L1-dcache-loads,L1-dcache-load-misses,\  L2-loads,L2-load-misses,\  LLC-loads,LLC-load-misses \  ./prog
# Memory bandwidth (Intel)perf stat -e \  uncore_imc/cas_count_read/,\  uncore_imc/cas_count_write/ \  ./prog
# TLB missesperf stat -e dTLB-loads,dTLB-load-misses,iTLB-loads,iTLB-load-misses ./prog
# Branch misprediction rateperf stat -e branches,branch-misses ./prog# Rate = branch-misses / branches × 100%
# Available events (varies by CPU)perf list hardware          # generic hardware eventsperf list cache             # cache eventsperf list pmu               # raw PMU events for your CPU

3. Key metrics and thresholds

MetricFormulaHealthyConcerning
IPCinstructions / cycles> 2.0 (modern x86)< 1.0
L1 miss rateL1-misses / L1-accesses< 1%> 5%
LLC miss rateLLC-misses / LLC-accesses< 1%> 10%
Branch miss ratebranch-misses / branches< 1%> 5%
MPKImisses per 1K instructions—L3 MPKI > 10 = memory bound
bash
# Compute MPKI (Misses Per Kilo-Instructions)perf stat -e instructions,LLC-load-misses ./prog# MPKI = LLC-load-misses / (instructions / 1000)

4. Raw PMU events (CPU-specific)

For events not in the generic aliases, use raw event codes:

bash
# Intel: use perf list or look up in Intel SDM# Format: rXXYY where XX=umask, YY=event codeperf stat -e r0124 ./prog    # example Intel raw event
# List Intel events with ocperf (OpenCL Perf Events)pip install ocperfocperf.py list | grep "mem_load"
# Use libpfm4 for event namespfm_ls | grep "MEM_LOAD"perf stat -e $(pfm_ls | grep "MEM_LOAD_RETIRED.L3_MISS") ./prog
# AMD: similar approachperf stat -e r04041 ./prog   # AMD raw event

5. Source-level annotation with perf record/annotate

bash
# Record with hardware eventsperf record -e LLC-load-misses -g ./prog
# Annotate: show source lines sorted by cache miss countperf annotate --stdio
# Interactive (requires debug symbols)perf report# Press 'a' on a function to annotate it
# Combined: record hotspot + annotateperf record -e cycles:u -g ./progperf annotate --symbol=my_function --stdio 2>/dev/null | head -40
# Example annotate output:# Percent | Source code#   45.23 |     for (int i = 0; i < N; i++)#    3.12 |         sum += data[i];   ← cache miss here (strided access)

6. PAPI — Portable API for hardware counters

PAPI provides a portable C API across different CPU architectures:

c
#include <papi.h>#include <stdio.h>
int main(void) {    int Events[] = {PAPI_TOT_INS, PAPI_TOT_CYC,                    PAPI_L2_TCM,  PAPI_BR_MSP};    long long values[4];
    if (PAPI_library_init(PAPI_VER_CURRENT) != PAPI_VER_CURRENT) {        fprintf(stderr, "PAPI init failed\n");        return 1;    }
    PAPI_start_counters(Events, 4);
    // --- Code to measure ---    do_work();    // -----------------------
    PAPI_stop_counters(values, 4);
    printf("Instructions:      %lld\n", values[0]);    printf("Cycles:            %lld\n", values[1]);    printf("IPC:               %.2f\n", (double)values[0]/values[1]);    printf("L2 cache misses:   %lld\n", values[2]);    printf("Branch mispred:    %lld\n", values[3]);
    return 0;}
bash
# Build with PAPIgcc -O2 -g -o prog prog.c -lpapi
# Available PAPI events on your systempapi_avail -a | head -30papi_native_avail | grep "L3"    # native events with "L3"

Common PAPI presets:

PresetEvent
PAPI_TOT_INSTotal instructions
PAPI_TOT_CYCTotal cycles
PAPI_L1_DCML1 data cache misses
PAPI_L2_TCML2 total cache misses
PAPI_L3_TCML3 total cache misses
PAPI_BR_MSPBranch mispredictions
PAPI_TLB_DMData TLB misses
PAPI_FP_INSFloating point instructions
PAPI_VEC_INSVector/SIMD instructions

7. Intel PCM (Performance Counter Monitor)

bash
# Intel PCM — system-wide counters, no root required on modern kernelsgit clone https://github.com/intel/pcmcd pcm && cmake -S . -B build && cmake --build build
# Measure memory bandwidth./build/bin/pcm-memory 1    # sample every 1 second
# Core utilization + IPC./build/bin/pcm 1
# Cache miss breakdown per socket./build/bin/pcm 1 -csv | head -20

Related skills

  • Use skills/profilers/intel-vtune-amd-uprof for guided microarchitecture analysis
  • Use skills/profilers/linux-perf for perf record/report and flamegraph generation
  • Use skills/low-level-programming/cpu-cache-opt for applying cache optimization patterns
  • Use skills/low-level-programming/simd-intrinsics for improving FLOPS/cycle metrics

來源與署名

來源:mohitmishra786/low-level-dev-skills位於skills/profilers/hardware-counters提交bdc5847

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架