Intel Vtune Amd Uprof

mohitmishra786/low-level-dev-skills/skills/profilers/intel-vtune-amd-uprof

作者 mohitmishra786bdc58472fa9f無授權條款253 個星標收錄於 2026年10月9日更新於 2026年10月9日儲存庫3 個月前更新

Intel VTune and AMD uProf profiling skill for microarchitecture analysis. Use when analyzing hotspots, microarchitecture bottlenecks, memory access patterns, pipeline stalls, or using the roofline model. Covers VTune Community Edition (free) and AMD uProf as a free alternative. Activates on queries about VTune, uProf, microarchitecture analysis, pipeline stalls, memory bandwidth, roofline model, or hardware performance analysis.

AI 產生的概覽

指導使用 Intel VTune 與 AMD uProf 進行 CPU 微架構效能分析,涵蓋熱點、管線停滯、記憶體存取與 roofline 分析。

功能
此技能提供使用 Intel VTune Profiler(免費社群版)與 AMD uProf 對 CPU 程式碼進行效能分析的指引。內容包含安裝與命令列用法、熱點分析、微架構探索、記憶體存取與執行緒分析等分析類型,以及如何解讀 IPC、管線停滯與 DRAM 頻寬等指標。它也說明 roofline 模型,以及如何將實際效能與硬體上限進行比較。
適用情境
適用於排查 CPU 熱點、微架構瓶頸、管線停滯、快取或記憶體頻寬問題,或套用 roofline 模型時。在 AMD CPU 上選擇 AMD uProf 作為 VTune 的免費替代方案時也適用。
執行需求
需要安裝 Intel VTune Profiler(社群版)或 AMD uProf,可選工具如 likwid-perfctr 用於手動 roofline 量測。程式應以除錯符號編譯以獲得有意義的結果。此技能不附帶指令碼,僅為說明文件。

Intel VTune & AMD uProf

Purpose

Guide agents through CPU microarchitecture profiling with Intel VTune Profiler (free Community Edition) and AMD uProf: hotspot identification, microarchitecture analysis, memory access pattern optimization, pipeline stall diagnosis, and roofline model analysis.

Triggers

  • "How do I use Intel VTune to profile my code?"
  • "What are pipeline stalls and how do I reduce them?"
  • "How do I analyze memory bandwidth with VTune?"
  • "What is the roofline model and how do I use it?"
  • "How do I use AMD uProf as a free alternative to VTune?"
  • "My code has good cache hit rates but is still slow"

Workflow

1. VTune setup (free Community Edition)

bash
# Download Intel VTune Profiler (Community Edition — free)# https://www.intel.com/content/www/us/en/developer/tools/oneapi/vtune-profiler.html
# Install on Linuxsource /opt/intel/oneapi/vtune/latest/env/vars.sh
# CLI usagevtune -collect hotspots ./progvtune -collect microarchitecture-exploration ./progvtune -collect memory-access ./prog
# View results in GUIvtune-gui &# File → Open Result → select .vtune directory
# Or use amplxe-cl (legacy CLI)amplxe-cl -collect hotspots ./progamplxe-cl -report hotspots -r result/

2. Analysis types

AnalysisWhat it findsWhen to use
HotspotsCPU-bound functionsFirst step — find where time is spent
Microarchitecture ExplorationIPC, pipeline stalls, retired instructionsAfter hotspot — why is the hotspot slow?
Memory AccessCache misses, DRAM bandwidth, NUMAMemory-bound code
ThreadingLock contention, parallel efficiencyMultithreaded code
HPC PerformanceVectorization, memory, rooflineHPC / scientific code
I/ODisk and network bottlenecksI/O-bound code

3. Hotspot analysis

bash
# Collect and report hotspotsvtune -collect hotspots -result-dir hotspots_result ./prog
# Report top functions by CPU timevtune -report hotspots -r hotspots_result -format csv | head -20
# CLI output example:# Function       CPU Time  Module# compute_fft    4.532s    libfft.so# matrix_mult    2.108s    prog# parse_input    0.234s    prog

Build with debug info for meaningful symbols:

bash
gcc -O2 -g ./prog.c -o prog     # symbols visible in VTunegcc -O2 -g -gsplit-dwarf -fno-omit-frame-pointer ./prog.c -o prog  # better stacks

4. Microarchitecture exploration — pipeline stalls

bash
vtune -collect microarchitecture-exploration -r micro_result ./progvtune -report summary -r micro_result

Key metrics to examine:

MetricMeaningGood value
IPC (Instructions Per Clock)How many instructions retire per cyclex86: aim for > 2.0
CPI (Clocks Per Instruction)Inverse of IPCLower is better
Bad SpeculationBranch mispredictions< 5%
Front-End BoundInstruction decode bottleneck< 15%
Back-End BoundExecution unit or memory stall< 30%
RetiringUseful work fraction> 70% ideal
Memory Bound% cycles waiting for memory< 20%
Pipeline Analysis (Top-Down Methodology):├── Retiring (good, useful work)├── Bad Speculation (branch mispredictions)├── Front-End Bound│   ├── Fetch Latency (I-cache misses, branch mispredicts)│   └── Fetch Bandwidth└── Back-End Bound    ├── Memory Bound    │   ├── L1 Bound → L1 cache misses    │   ├── L2 Bound → L2 cache misses    │   ├── L3 Bound → L3 cache misses    │   └── DRAM Bound → main memory bandwidth limited    └── Core Bound → ALU/compute bound

5. Memory access analysis

bash
# Collect memory access profilevtune -collect memory-access -r mem_result ./prog
# Key output sections:# - Memory Bound: % time waiting for memory# - LLC (Last Level Cache) Miss Rate# - DRAM Bandwidth: GB/s achieved vs theoretical peak# - NUMA: cross-socket accesses (for multi-socket systems)

Reading DRAM bandwidth:

DRAM Bandwidth: 18.4 GB/sPeak Theoretical: 51.2 GB/sUtilization: 36% — likely not DRAM-bound

If DRAM-bound: optimize data layout (AoS → SoA), reduce working set, improve spatial locality.

6. AMD uProf — free alternative for AMD CPUs

bash
# Download AMD uProf# https://www.amd.com/en/developer/uprof.html
# CLI profilingAMDuProfCLI collect --config tbp ./prog          # time-based profilingAMDuProfCLI collect --config assess ./prog       # microarchitecture assessmentAMDuProfCLI collect --config memory ./prog       # memory access
# Generate reportAMDuProfCLI report -i /tmp/uprof_result/ -o report.html
# Open GUIAMDuProf &

AMD uProf metrics map to VTune equivalents:

  • Retired Instructions → IPC analysis
  • Branch Mispredictions → Bad Speculation
  • L1/L2/L3 Cache Misses → Memory Bound levels
  • Data Cache Accesses → Cache efficiency

7. Roofline model

The roofline model shows whether code is compute-bound or memory-bound by comparing achieved performance against hardware limits:

Performance (GFLOPS/s)     |                    _______________Peak |                 /Perf |              /  compute bound     |           /     |        /     |     /  memory bandwidth bound     |  /     +------------------------------→        Arithmetic Intensity (FLOPS/Byte)
bash
# VTune roofline collectionvtune -collect hpc-performance -r roofline_result ./prog# Then: VTune GUI → Roofline view
# For manual calculation:# Arithmetic Intensity = FLOPS / memory_bytes_accessed# Peak FLOPS = CPUs × cores × freq × FLOPS_per_cycle_per_core# Peak BW = from hardware spec (e.g., 51.2 GB/s for DDR4-3200 dual channel)
# likwid-perfctr for manual roofline data (Linux)likwid-perfctr -C 0 -g FLOPS_DP ./prog          # double-precision FLOPSlikwid-perfctr -C 0 -g MEM ./prog               # memory bandwidth

Related skills

  • Use skills/profilers/hardware-counters for raw PMU event collection with perf stat
  • Use skills/profilers/linux-perf for perf-based profiling on Linux
  • Use skills/low-level-programming/cpu-cache-opt for memory access pattern optimization
  • Use skills/low-level-programming/simd-intrinsics for vectorization to increase FLOPS

來源與署名

來源:mohitmishra786/low-level-dev-skills位於skills/profilers/intel-vtune-amd-uprof提交bdc5847

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架