Linux Perf

mohitmishra786/low-level-dev-skills/skills/profilers/linux-perf

by mohitmishra786bdc58472fa9fNo license253 starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated 3 months ago

Linux perf profiler skill for CPU performance analysis. Use when collecting sampling profiles with perf record, generating perf report, measuring hardware counters (cache misses, branch mispredicts, IPC), identifying hot functions, or feeding perf data into flamegraph tools. Activates on queries about perf, Linux performance counters, PMU events, off-CPU profiling, perf stat, perf annotate, or sampling-based profiling on Linux.

AI-generated overview

Guides agents through Linux perf for CPU profiling, hardware counters, hotspot analysis and flamegraph data.

What it does
This skill provides instructions for using the Linux perf tool to profile CPU performance. It covers installing perf, adjusting perf_event_paranoid permissions, compiling with debug symbols, and running perf stat, perf record, perf report, perf annotate and perf top. It also explains how to export perf script output for flamegraph generation and how to troubleshoot common issues such as unknown frames or permission errors.
When to use it
Use it when investigating CPU hotspots, measuring hardware counters such as cache misses, branch mispredicts or IPC, or sampling Linux processes. It is also relevant when perf output shows unknown or kernel frames, or when preparing perf data for flamegraph tools.
Requirements
Requires the Linux perf tool installed on the system, typically via linux-perf or perf packages, and sufficient permissions such as root or a lowered perf_event_paranoid setting. Debug symbols and frame pointers are recommended for useful profiling output. Flamegraph generation requires external FlameGraph scripts and network access to clone them. No scripts ship with this skill.

Linux perf

Purpose

Guide agents through perf for CPU profiling: sampling, hardware counter measurement, hotspot identification, and integration with flamegraph generation.

Triggers

  • "Which function is consuming the most CPU?"
  • "How do I measure cache misses / IPC?"
  • "How do I use perf to find hotspots?"
  • "How do I generate a flamegraph from perf data?"
  • "perf shows [unknown] or [kernel] frames"

Workflow

1. Prerequisites

bash
# Installsudo apt install linux-perf    # Debian/Ubuntu (version-matched)sudo dnf install perf          # Fedora/RHEL
# Check permissions# By default perf requires root or paranoid level ≤ 1cat /proc/sys/kernel/perf_event_paranoid# 2 = only CPU stats (not kernel), 1 = user+kernel, 0 = all, -1 = no restrictions
# Temporarily lower (session only)sudo sysctl -w kernel.perf_event_paranoid=1
# Persistentecho 'kernel.perf_event_paranoid=1' | sudo tee /etc/sysctl.d/99-perf.confsudo sysctl -p /etc/sysctl.d/99-perf.conf

Compile the target with debug symbols for useful frame data:

bash
gcc -g -O2 -fno-omit-frame-pointer -o prog main.c# -fno-omit-frame-pointer: essential for frame-pointer-based unwinding# Alternative: compile with DWARF CFI and use --call-graph=dwarf

2. perf stat — quick counters

bash
# Basic hardware countersperf stat ./prog
# With specific eventsperf stat -e cache-misses,cache-references,instructions,cycles,branch-misses ./prog
# Wall-clock comparison: N runsperf stat -r 5 ./prog
# Attach to existing processperf stat -p 12345 sleep 10

Interpret perf stat output:

  • IPC (instructions per cycle) < 1.0: memory-bound or stalled pipeline
  • cache-miss rate > 5%: significant cache pressure
  • branch-miss rate > 5%: branch predictor struggling

3. perf record — sampling

bash
# Default: sample at 1000 Hz (cycles event)perf record -g ./prog
# Specify frequencyperf record -F 999 -g ./prog
# Specific eventperf record -e cache-misses -g ./prog
# Attach to running processperf record -F 999 -g -p 12345 sleep 30
# Off-CPU profiling (time spent waiting)perf record -e sched:sched_switch -ag sleep 10
# DWARF call graphs (better for binaries without frame pointers)perf record -F 999 --call-graph=dwarf ./prog
# Save to named fileperf record -o myapp.perf.data -g ./prog

4. perf report — interactive analysis

bash
perf report                          # reads perf.dataperf report -i myapp.perf.dataperf report --no-children            # self time only (not cumulative)perf report --sort comm,dso,sym      # sort by fieldsperf report --stdio                  # non-interactive text output

Navigation in TUI:

  • Enter — expand a symbol
  • a — annotate (show assembly with hit counts)
  • s — show source (needs debug info)
  • d — filter by DSO (library)
  • t — filter by thread
  • ? — help

5. perf annotate — hot instructions

bash
# Show assembly with hit percentagesperf annotate sym_name
# From report: press 'a' on a symbol# Or directly:perf annotate -i perf.data --symbol=hot_function --stdio

High hit count on a mov or vmovdqa suggests a cache miss at that load.

6. perf top — live profiling

bash
# Live top, like 'top' but for functionssudo perf top -g
# Filter by processsudo perf top -p 12345

7. Feed into flamegraphs

bash
# Generate perf script outputperf script > out.perf
# Use Brendan Gregg's FlameGraph toolsgit clone https://github.com/brendangregg/FlameGraph./FlameGraph/stackcollapse-perf.pl out.perf > out.folded./FlameGraph/flamegraph.pl out.folded > flamegraph.svg
# Open flamegraph.svg in browser

See skills/profilers/flamegraphs for reading flamegraphs and interpreting results.

8. Common issues

ProblemCauseFix
Permission deniedperf_event_paranoid too highLower paranoid level or run with sudo
[unknown] framesMissing frame pointers or debug infoRecompile with -fno-omit-frame-pointer or use --call-graph=dwarf
[kernel] everywhereKernel symbols not visibleUse sudo perf record; install linux-image-$(uname -r)-dbgsym
No kallsymsKernel symbols unavailable`echo 0
Empty report for short programProgram exits too fastUse -F 9999 or instrument longer workload
DWARF unwinding slowLarge DWARF stackLimit with --call-graph dwarf,512

9. Useful events

bash
# List all available eventsperf list
# Common hardware eventscyclesinstructionscache-referencescache-missesbranch-instructionsbranch-missesstalled-cycles-frontendstalled-cycles-backend
# Software eventscontext-switchescpu-migrationspage-faults
# Tracepoints (requires root)sched:sched_switchsyscalls:sys_enter_read

For a counter reference and interpretation guide, see references/events.md [blocked].

Related skills

  • Use skills/profilers/flamegraphs for SVG flamegraph generation and reading
  • Use skills/profilers/valgrind for cache simulation and memory profiling
  • Use skills/compilers/gcc or skills/compilers/clang for PGO from perf data (AutoFDO)

Source and attribution

Source:mohitmishra786/low-level-dev-skillsinskills/profilers/linux-perfat commitbdc5847

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal