Linux Perf

mohitmishra786/low-level-dev-skills/skills/profilers/linux-perf

作者 mohitmishra786bdc58472fa9f无许可证253 个星标收录于 2026年10月9日更新于 2026年10月9日仓库3个月前更新

Linux perf profiler skill for CPU performance analysis. Use when collecting sampling profiles with perf record, generating perf report, measuring hardware counters (cache misses, branch mispredicts, IPC), identifying hot functions, or feeding perf data into flamegraph tools. Activates on queries about perf, Linux performance counters, PMU events, off-CPU profiling, perf stat, perf annotate, or sampling-based profiling on Linux.

AI 生成的概览

指导代理使用 Linux perf 进行 CPU 性能分析、硬件计数器测量、热点定位和火焰图数据准备。

功能
该技能提供使用 Linux perf 工具进行 CPU 性能分析的说明。内容涵盖安装 perf、调整 perf_event_paranoid 权限、使用调试符号编译,以及运行 perf stat、perf record、perf report、perf annotate 和 perf top。它还说明如何导出 perf script 输出以生成火焰图,并排查未知帧或权限错误等常见问题。
适用场景
适用于调查 CPU 热点、测量缓存未命中、分支预测失败或 IPC 等硬件计数器,或对 Linux 进程进行采样分析。也适用于 perf 输出显示未知帧或内核帧,或准备将 perf 数据用于火焰图工具的场景。
运行要求
需要系统安装 Linux perf 工具,通常通过 linux-perf 或 perf 软件包,并具备足够权限,例如 root 或降低 perf_event_paranoid 设置。为获得有用的分析输出,建议使用调试符号和帧指针。生成火焰图需要外部 FlameGraph 脚本以及克隆它们的网络访问。该技能不附带脚本。

Linux perf

Purpose

Guide agents through perf for CPU profiling: sampling, hardware counter measurement, hotspot identification, and integration with flamegraph generation.

Triggers

  • "Which function is consuming the most CPU?"
  • "How do I measure cache misses / IPC?"
  • "How do I use perf to find hotspots?"
  • "How do I generate a flamegraph from perf data?"
  • "perf shows [unknown] or [kernel] frames"

Workflow

1. Prerequisites

bash
# Installsudo apt install linux-perf    # Debian/Ubuntu (version-matched)sudo dnf install perf          # Fedora/RHEL
# Check permissions# By default perf requires root or paranoid level ≤ 1cat /proc/sys/kernel/perf_event_paranoid# 2 = only CPU stats (not kernel), 1 = user+kernel, 0 = all, -1 = no restrictions
# Temporarily lower (session only)sudo sysctl -w kernel.perf_event_paranoid=1
# Persistentecho 'kernel.perf_event_paranoid=1' | sudo tee /etc/sysctl.d/99-perf.confsudo sysctl -p /etc/sysctl.d/99-perf.conf

Compile the target with debug symbols for useful frame data:

bash
gcc -g -O2 -fno-omit-frame-pointer -o prog main.c# -fno-omit-frame-pointer: essential for frame-pointer-based unwinding# Alternative: compile with DWARF CFI and use --call-graph=dwarf

2. perf stat — quick counters

bash
# Basic hardware countersperf stat ./prog
# With specific eventsperf stat -e cache-misses,cache-references,instructions,cycles,branch-misses ./prog
# Wall-clock comparison: N runsperf stat -r 5 ./prog
# Attach to existing processperf stat -p 12345 sleep 10

Interpret perf stat output:

  • IPC (instructions per cycle) < 1.0: memory-bound or stalled pipeline
  • cache-miss rate > 5%: significant cache pressure
  • branch-miss rate > 5%: branch predictor struggling

3. perf record — sampling

bash
# Default: sample at 1000 Hz (cycles event)perf record -g ./prog
# Specify frequencyperf record -F 999 -g ./prog
# Specific eventperf record -e cache-misses -g ./prog
# Attach to running processperf record -F 999 -g -p 12345 sleep 30
# Off-CPU profiling (time spent waiting)perf record -e sched:sched_switch -ag sleep 10
# DWARF call graphs (better for binaries without frame pointers)perf record -F 999 --call-graph=dwarf ./prog
# Save to named fileperf record -o myapp.perf.data -g ./prog

4. perf report — interactive analysis

bash
perf report                          # reads perf.dataperf report -i myapp.perf.dataperf report --no-children            # self time only (not cumulative)perf report --sort comm,dso,sym      # sort by fieldsperf report --stdio                  # non-interactive text output

Navigation in TUI:

  • Enter — expand a symbol
  • a — annotate (show assembly with hit counts)
  • s — show source (needs debug info)
  • d — filter by DSO (library)
  • t — filter by thread
  • ? — help

5. perf annotate — hot instructions

bash
# Show assembly with hit percentagesperf annotate sym_name
# From report: press 'a' on a symbol# Or directly:perf annotate -i perf.data --symbol=hot_function --stdio

High hit count on a mov or vmovdqa suggests a cache miss at that load.

6. perf top — live profiling

bash
# Live top, like 'top' but for functionssudo perf top -g
# Filter by processsudo perf top -p 12345

7. Feed into flamegraphs

bash
# Generate perf script outputperf script > out.perf
# Use Brendan Gregg's FlameGraph toolsgit clone https://github.com/brendangregg/FlameGraph./FlameGraph/stackcollapse-perf.pl out.perf > out.folded./FlameGraph/flamegraph.pl out.folded > flamegraph.svg
# Open flamegraph.svg in browser

See skills/profilers/flamegraphs for reading flamegraphs and interpreting results.

8. Common issues

ProblemCauseFix
Permission deniedperf_event_paranoid too highLower paranoid level or run with sudo
[unknown] framesMissing frame pointers or debug infoRecompile with -fno-omit-frame-pointer or use --call-graph=dwarf
[kernel] everywhereKernel symbols not visibleUse sudo perf record; install linux-image-$(uname -r)-dbgsym
No kallsymsKernel symbols unavailable`echo 0
Empty report for short programProgram exits too fastUse -F 9999 or instrument longer workload
DWARF unwinding slowLarge DWARF stackLimit with --call-graph dwarf,512

9. Useful events

bash
# List all available eventsperf list
# Common hardware eventscyclesinstructionscache-referencescache-missesbranch-instructionsbranch-missesstalled-cycles-frontendstalled-cycles-backend
# Software eventscontext-switchescpu-migrationspage-faults
# Tracepoints (requires root)sched:sched_switchsyscalls:sys_enter_read

For a counter reference and interpretation guide, see references/events.md [blocked].

Related skills

  • Use skills/profilers/flamegraphs for SVG flamegraph generation and reading
  • Use skills/profilers/valgrind for cache simulation and memory profiling
  • Use skills/compilers/gcc or skills/compilers/clang for PGO from perf data (AutoFDO)

来源与署名

来源:mohitmishra786/low-level-dev-skills位于skills/profilers/linux-perf提交bdc5847

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架