Linux Perf

mohitmishra786/low-level-dev-skills/skills/profilers/linux-perf

作者 mohitmishra786bdc58472fa9f無授權條款253 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫3 個月前更新

Linux perf profiler skill for CPU performance analysis. Use when collecting sampling profiles with perf record, generating perf report, measuring hardware counters (cache misses, branch mispredicts, IPC), identifying hot functions, or feeding perf data into flamegraph tools. Activates on queries about perf, Linux performance counters, PMU events, off-CPU profiling, perf stat, perf annotate, or sampling-based profiling on Linux.

AI 產生的概覽

指導代理使用 Linux perf 進行 CPU 效能分析、硬體計數器量測、熱點定位與火焰圖資料準備。

功能
此技能提供使用 Linux perf 工具進行 CPU 效能分析的說明。內容涵蓋安裝 perf、調整 perf_event_paranoid 權限、使用除錯符號編譯,以及執行 perf stat、perf record、perf report、perf annotate 和 perf top。它也說明如何匯出 perf script 輸出以產生火焰圖,並排解未知框架或權限錯誤等常見問題。
適用情境
適用於調查 CPU 熱點、量測快取未命中、分支預測失敗或 IPC 等硬體計數器,或對 Linux 行程進行取樣分析。也適用於 perf 輸出顯示未知框架或核心框架,或準備將 perf 資料用於火焰圖工具的情境。
執行需求
需要系統安裝 Linux perf 工具,通常透過 linux-perf 或 perf 套件,並具備足夠權限,例如 root 或降低 perf_event_paranoid 設定。為獲得有用的分析輸出,建議使用除錯符號和框架指標。產生火焰圖需要外部 FlameGraph 指令碼以及複製它們的網路存取。此技能不附帶指令碼。

Linux perf

Purpose

Guide agents through perf for CPU profiling: sampling, hardware counter measurement, hotspot identification, and integration with flamegraph generation.

Triggers

  • "Which function is consuming the most CPU?"
  • "How do I measure cache misses / IPC?"
  • "How do I use perf to find hotspots?"
  • "How do I generate a flamegraph from perf data?"
  • "perf shows [unknown] or [kernel] frames"

Workflow

1. Prerequisites

bash
# Installsudo apt install linux-perf    # Debian/Ubuntu (version-matched)sudo dnf install perf          # Fedora/RHEL
# Check permissions# By default perf requires root or paranoid level ≤ 1cat /proc/sys/kernel/perf_event_paranoid# 2 = only CPU stats (not kernel), 1 = user+kernel, 0 = all, -1 = no restrictions
# Temporarily lower (session only)sudo sysctl -w kernel.perf_event_paranoid=1
# Persistentecho 'kernel.perf_event_paranoid=1' | sudo tee /etc/sysctl.d/99-perf.confsudo sysctl -p /etc/sysctl.d/99-perf.conf

Compile the target with debug symbols for useful frame data:

bash
gcc -g -O2 -fno-omit-frame-pointer -o prog main.c# -fno-omit-frame-pointer: essential for frame-pointer-based unwinding# Alternative: compile with DWARF CFI and use --call-graph=dwarf

2. perf stat — quick counters

bash
# Basic hardware countersperf stat ./prog
# With specific eventsperf stat -e cache-misses,cache-references,instructions,cycles,branch-misses ./prog
# Wall-clock comparison: N runsperf stat -r 5 ./prog
# Attach to existing processperf stat -p 12345 sleep 10

Interpret perf stat output:

  • IPC (instructions per cycle) < 1.0: memory-bound or stalled pipeline
  • cache-miss rate > 5%: significant cache pressure
  • branch-miss rate > 5%: branch predictor struggling

3. perf record — sampling

bash
# Default: sample at 1000 Hz (cycles event)perf record -g ./prog
# Specify frequencyperf record -F 999 -g ./prog
# Specific eventperf record -e cache-misses -g ./prog
# Attach to running processperf record -F 999 -g -p 12345 sleep 30
# Off-CPU profiling (time spent waiting)perf record -e sched:sched_switch -ag sleep 10
# DWARF call graphs (better for binaries without frame pointers)perf record -F 999 --call-graph=dwarf ./prog
# Save to named fileperf record -o myapp.perf.data -g ./prog

4. perf report — interactive analysis

bash
perf report                          # reads perf.dataperf report -i myapp.perf.dataperf report --no-children            # self time only (not cumulative)perf report --sort comm,dso,sym      # sort by fieldsperf report --stdio                  # non-interactive text output

Navigation in TUI:

  • Enter — expand a symbol
  • a — annotate (show assembly with hit counts)
  • s — show source (needs debug info)
  • d — filter by DSO (library)
  • t — filter by thread
  • ? — help

5. perf annotate — hot instructions

bash
# Show assembly with hit percentagesperf annotate sym_name
# From report: press 'a' on a symbol# Or directly:perf annotate -i perf.data --symbol=hot_function --stdio

High hit count on a mov or vmovdqa suggests a cache miss at that load.

6. perf top — live profiling

bash
# Live top, like 'top' but for functionssudo perf top -g
# Filter by processsudo perf top -p 12345

7. Feed into flamegraphs

bash
# Generate perf script outputperf script > out.perf
# Use Brendan Gregg's FlameGraph toolsgit clone https://github.com/brendangregg/FlameGraph./FlameGraph/stackcollapse-perf.pl out.perf > out.folded./FlameGraph/flamegraph.pl out.folded > flamegraph.svg
# Open flamegraph.svg in browser

See skills/profilers/flamegraphs for reading flamegraphs and interpreting results.

8. Common issues

ProblemCauseFix
Permission deniedperf_event_paranoid too highLower paranoid level or run with sudo
[unknown] framesMissing frame pointers or debug infoRecompile with -fno-omit-frame-pointer or use --call-graph=dwarf
[kernel] everywhereKernel symbols not visibleUse sudo perf record; install linux-image-$(uname -r)-dbgsym
No kallsymsKernel symbols unavailable`echo 0
Empty report for short programProgram exits too fastUse -F 9999 or instrument longer workload
DWARF unwinding slowLarge DWARF stackLimit with --call-graph dwarf,512

9. Useful events

bash
# List all available eventsperf list
# Common hardware eventscyclesinstructionscache-referencescache-missesbranch-instructionsbranch-missesstalled-cycles-frontendstalled-cycles-backend
# Software eventscontext-switchescpu-migrationspage-faults
# Tracepoints (requires root)sched:sched_switchsyscalls:sys_enter_read

For a counter reference and interpretation guide, see references/events.md [blocked].

Related skills

  • Use skills/profilers/flamegraphs for SVG flamegraph generation and reading
  • Use skills/profilers/valgrind for cache simulation and memory profiling
  • Use skills/compilers/gcc or skills/compilers/clang for PGO from perf data (AutoFDO)

來源與署名

來源:mohitmishra786/low-level-dev-skills位於skills/profilers/linux-perf提交bdc5847

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架