Intel Vtune Amd Uprof

mohitmishra786/low-level-dev-skills/skills/profilers/intel-vtune-amd-uprof

by mohitmishra786bdc58472fa9fNo license253 starsListed Oct 9, 2026Updated Oct 9, 2026Repository updated 3 months ago

Intel VTune and AMD uProf profiling skill for microarchitecture analysis. Use when analyzing hotspots, microarchitecture bottlenecks, memory access patterns, pipeline stalls, or using the roofline model. Covers VTune Community Edition (free) and AMD uProf as a free alternative. Activates on queries about VTune, uProf, microarchitecture analysis, pipeline stalls, memory bandwidth, roofline model, or hardware performance analysis.

Instructions onlySoftware Development
AI-generated overview

Guides CPU microarchitecture profiling with Intel VTune and AMD uProf, covering hotspots, stalls, memory access and roofline analysis.

What it does
This skill provides instructions for profiling CPU code with Intel VTune Profiler (free Community Edition) and AMD uProf. It covers setup and CLI commands, analysis types such as hotspots, microarchitecture exploration, memory access and threading, and how to read metrics like IPC, pipeline stalls and DRAM bandwidth. It also explains the roofline model and how to compare achieved performance against hardware limits.
When to use it
Use it when investigating CPU hotspots, microarchitecture bottlenecks, pipeline stalls, cache or memory bandwidth issues, or when applying the roofline model. It is also relevant when choosing AMD uProf as a free alternative to VTune on AMD CPUs.
Requirements
Requires Intel VTune Profiler (Community Edition) or AMD uProf installed, plus optional tools such as likwid-perfctr for manual roofline measurements. Programs should be built with debug symbols for meaningful results. It ships no scripts; it is instructions only.

Intel VTune & AMD uProf

Purpose

Guide agents through CPU microarchitecture profiling with Intel VTune Profiler (free Community Edition) and AMD uProf: hotspot identification, microarchitecture analysis, memory access pattern optimization, pipeline stall diagnosis, and roofline model analysis.

Triggers

  • "How do I use Intel VTune to profile my code?"
  • "What are pipeline stalls and how do I reduce them?"
  • "How do I analyze memory bandwidth with VTune?"
  • "What is the roofline model and how do I use it?"
  • "How do I use AMD uProf as a free alternative to VTune?"
  • "My code has good cache hit rates but is still slow"

Workflow

1. VTune setup (free Community Edition)

bash
# Download Intel VTune Profiler (Community Edition — free)# https://www.intel.com/content/www/us/en/developer/tools/oneapi/vtune-profiler.html
# Install on Linuxsource /opt/intel/oneapi/vtune/latest/env/vars.sh
# CLI usagevtune -collect hotspots ./progvtune -collect microarchitecture-exploration ./progvtune -collect memory-access ./prog
# View results in GUIvtune-gui &# File → Open Result → select .vtune directory
# Or use amplxe-cl (legacy CLI)amplxe-cl -collect hotspots ./progamplxe-cl -report hotspots -r result/

2. Analysis types

AnalysisWhat it findsWhen to use
HotspotsCPU-bound functionsFirst step — find where time is spent
Microarchitecture ExplorationIPC, pipeline stalls, retired instructionsAfter hotspot — why is the hotspot slow?
Memory AccessCache misses, DRAM bandwidth, NUMAMemory-bound code
ThreadingLock contention, parallel efficiencyMultithreaded code
HPC PerformanceVectorization, memory, rooflineHPC / scientific code
I/ODisk and network bottlenecksI/O-bound code

3. Hotspot analysis

bash
# Collect and report hotspotsvtune -collect hotspots -result-dir hotspots_result ./prog
# Report top functions by CPU timevtune -report hotspots -r hotspots_result -format csv | head -20
# CLI output example:# Function       CPU Time  Module# compute_fft    4.532s    libfft.so# matrix_mult    2.108s    prog# parse_input    0.234s    prog

Build with debug info for meaningful symbols:

bash
gcc -O2 -g ./prog.c -o prog     # symbols visible in VTunegcc -O2 -g -gsplit-dwarf -fno-omit-frame-pointer ./prog.c -o prog  # better stacks

4. Microarchitecture exploration — pipeline stalls

bash
vtune -collect microarchitecture-exploration -r micro_result ./progvtune -report summary -r micro_result

Key metrics to examine:

MetricMeaningGood value
IPC (Instructions Per Clock)How many instructions retire per cyclex86: aim for > 2.0
CPI (Clocks Per Instruction)Inverse of IPCLower is better
Bad SpeculationBranch mispredictions< 5%
Front-End BoundInstruction decode bottleneck< 15%
Back-End BoundExecution unit or memory stall< 30%
RetiringUseful work fraction> 70% ideal
Memory Bound% cycles waiting for memory< 20%
Pipeline Analysis (Top-Down Methodology):├── Retiring (good, useful work)├── Bad Speculation (branch mispredictions)├── Front-End Bound│   ├── Fetch Latency (I-cache misses, branch mispredicts)│   └── Fetch Bandwidth└── Back-End Bound    ├── Memory Bound    │   ├── L1 Bound → L1 cache misses    │   ├── L2 Bound → L2 cache misses    │   ├── L3 Bound → L3 cache misses    │   └── DRAM Bound → main memory bandwidth limited    └── Core Bound → ALU/compute bound

5. Memory access analysis

bash
# Collect memory access profilevtune -collect memory-access -r mem_result ./prog
# Key output sections:# - Memory Bound: % time waiting for memory# - LLC (Last Level Cache) Miss Rate# - DRAM Bandwidth: GB/s achieved vs theoretical peak# - NUMA: cross-socket accesses (for multi-socket systems)

Reading DRAM bandwidth:

DRAM Bandwidth: 18.4 GB/sPeak Theoretical: 51.2 GB/sUtilization: 36% — likely not DRAM-bound

If DRAM-bound: optimize data layout (AoS → SoA), reduce working set, improve spatial locality.

6. AMD uProf — free alternative for AMD CPUs

bash
# Download AMD uProf# https://www.amd.com/en/developer/uprof.html
# CLI profilingAMDuProfCLI collect --config tbp ./prog          # time-based profilingAMDuProfCLI collect --config assess ./prog       # microarchitecture assessmentAMDuProfCLI collect --config memory ./prog       # memory access
# Generate reportAMDuProfCLI report -i /tmp/uprof_result/ -o report.html
# Open GUIAMDuProf &

AMD uProf metrics map to VTune equivalents:

  • Retired Instructions → IPC analysis
  • Branch Mispredictions → Bad Speculation
  • L1/L2/L3 Cache Misses → Memory Bound levels
  • Data Cache Accesses → Cache efficiency

7. Roofline model

The roofline model shows whether code is compute-bound or memory-bound by comparing achieved performance against hardware limits:

Performance (GFLOPS/s)     |                    _______________Peak |                 /Perf |              /  compute bound     |           /     |        /     |     /  memory bandwidth bound     |  /     +------------------------------→        Arithmetic Intensity (FLOPS/Byte)
bash
# VTune roofline collectionvtune -collect hpc-performance -r roofline_result ./prog# Then: VTune GUI → Roofline view
# For manual calculation:# Arithmetic Intensity = FLOPS / memory_bytes_accessed# Peak FLOPS = CPUs × cores × freq × FLOPS_per_cycle_per_core# Peak BW = from hardware spec (e.g., 51.2 GB/s for DDR4-3200 dual channel)
# likwid-perfctr for manual roofline data (Linux)likwid-perfctr -C 0 -g FLOPS_DP ./prog          # double-precision FLOPSlikwid-perfctr -C 0 -g MEM ./prog               # memory bandwidth

Related skills

  • Use skills/profilers/hardware-counters for raw PMU event collection with perf stat
  • Use skills/profilers/linux-perf for perf-based profiling on Linux
  • Use skills/low-level-programming/cpu-cache-opt for memory access pattern optimization
  • Use skills/low-level-programming/simd-intrinsics for vectorization to increase FLOPS

Source and attribution

Source:mohitmishra786/low-level-dev-skillsinskills/profilers/intel-vtune-amd-uprofat commitbdc5847

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal