Intel Vtune Amd Uprof

mohitmishra786/low-level-dev-skills/skills/profilers/intel-vtune-amd-uprof

作者 mohitmishra786bdc58472fa9f无许可证253 个星标收录于 2026年10月9日更新于 2026年10月9日仓库3个月前更新

Intel VTune and AMD uProf profiling skill for microarchitecture analysis. Use when analyzing hotspots, microarchitecture bottlenecks, memory access patterns, pipeline stalls, or using the roofline model. Covers VTune Community Edition (free) and AMD uProf as a free alternative. Activates on queries about VTune, uProf, microarchitecture analysis, pipeline stalls, memory bandwidth, roofline model, or hardware performance analysis.

AI 生成的概览

指导使用 Intel VTune 和 AMD uProf 进行 CPU 微架构性能分析,涵盖热点、流水线停顿、内存访问与 roofline 分析。

功能
该技能提供使用 Intel VTune Profiler(免费社区版)和 AMD uProf 对 CPU 代码进行性能分析的指导。内容包括安装与命令行用法、热点分析、微架构探索、内存访问和线程分析等分析类型,以及如何解读 IPC、流水线停顿和 DRAM 带宽等指标。它还讲解 roofline 模型,以及如何将实际性能与硬件上限进行对比。
适用场景
适用于排查 CPU 热点、微架构瓶颈、流水线停顿、缓存或内存带宽问题,或应用 roofline 模型时。在 AMD CPU 上选择 AMD uProf 作为 VTune 的免费替代方案时也适用。
运行要求
需要安装 Intel VTune Profiler(社区版)或 AMD uProf,可选工具如 likwid-perfctr 用于手动 roofline 测量。程序应带调试符号编译以获得有意义的结果。该技能不附带脚本,仅为说明文档。

Intel VTune & AMD uProf

Purpose

Guide agents through CPU microarchitecture profiling with Intel VTune Profiler (free Community Edition) and AMD uProf: hotspot identification, microarchitecture analysis, memory access pattern optimization, pipeline stall diagnosis, and roofline model analysis.

Triggers

  • "How do I use Intel VTune to profile my code?"
  • "What are pipeline stalls and how do I reduce them?"
  • "How do I analyze memory bandwidth with VTune?"
  • "What is the roofline model and how do I use it?"
  • "How do I use AMD uProf as a free alternative to VTune?"
  • "My code has good cache hit rates but is still slow"

Workflow

1. VTune setup (free Community Edition)

bash
# Download Intel VTune Profiler (Community Edition — free)# https://www.intel.com/content/www/us/en/developer/tools/oneapi/vtune-profiler.html
# Install on Linuxsource /opt/intel/oneapi/vtune/latest/env/vars.sh
# CLI usagevtune -collect hotspots ./progvtune -collect microarchitecture-exploration ./progvtune -collect memory-access ./prog
# View results in GUIvtune-gui &# File → Open Result → select .vtune directory
# Or use amplxe-cl (legacy CLI)amplxe-cl -collect hotspots ./progamplxe-cl -report hotspots -r result/

2. Analysis types

AnalysisWhat it findsWhen to use
HotspotsCPU-bound functionsFirst step — find where time is spent
Microarchitecture ExplorationIPC, pipeline stalls, retired instructionsAfter hotspot — why is the hotspot slow?
Memory AccessCache misses, DRAM bandwidth, NUMAMemory-bound code
ThreadingLock contention, parallel efficiencyMultithreaded code
HPC PerformanceVectorization, memory, rooflineHPC / scientific code
I/ODisk and network bottlenecksI/O-bound code

3. Hotspot analysis

bash
# Collect and report hotspotsvtune -collect hotspots -result-dir hotspots_result ./prog
# Report top functions by CPU timevtune -report hotspots -r hotspots_result -format csv | head -20
# CLI output example:# Function       CPU Time  Module# compute_fft    4.532s    libfft.so# matrix_mult    2.108s    prog# parse_input    0.234s    prog

Build with debug info for meaningful symbols:

bash
gcc -O2 -g ./prog.c -o prog     # symbols visible in VTunegcc -O2 -g -gsplit-dwarf -fno-omit-frame-pointer ./prog.c -o prog  # better stacks

4. Microarchitecture exploration — pipeline stalls

bash
vtune -collect microarchitecture-exploration -r micro_result ./progvtune -report summary -r micro_result

Key metrics to examine:

MetricMeaningGood value
IPC (Instructions Per Clock)How many instructions retire per cyclex86: aim for > 2.0
CPI (Clocks Per Instruction)Inverse of IPCLower is better
Bad SpeculationBranch mispredictions< 5%
Front-End BoundInstruction decode bottleneck< 15%
Back-End BoundExecution unit or memory stall< 30%
RetiringUseful work fraction> 70% ideal
Memory Bound% cycles waiting for memory< 20%
Pipeline Analysis (Top-Down Methodology):├── Retiring (good, useful work)├── Bad Speculation (branch mispredictions)├── Front-End Bound│   ├── Fetch Latency (I-cache misses, branch mispredicts)│   └── Fetch Bandwidth└── Back-End Bound    ├── Memory Bound    │   ├── L1 Bound → L1 cache misses    │   ├── L2 Bound → L2 cache misses    │   ├── L3 Bound → L3 cache misses    │   └── DRAM Bound → main memory bandwidth limited    └── Core Bound → ALU/compute bound

5. Memory access analysis

bash
# Collect memory access profilevtune -collect memory-access -r mem_result ./prog
# Key output sections:# - Memory Bound: % time waiting for memory# - LLC (Last Level Cache) Miss Rate# - DRAM Bandwidth: GB/s achieved vs theoretical peak# - NUMA: cross-socket accesses (for multi-socket systems)

Reading DRAM bandwidth:

DRAM Bandwidth: 18.4 GB/sPeak Theoretical: 51.2 GB/sUtilization: 36% — likely not DRAM-bound

If DRAM-bound: optimize data layout (AoS → SoA), reduce working set, improve spatial locality.

6. AMD uProf — free alternative for AMD CPUs

bash
# Download AMD uProf# https://www.amd.com/en/developer/uprof.html
# CLI profilingAMDuProfCLI collect --config tbp ./prog          # time-based profilingAMDuProfCLI collect --config assess ./prog       # microarchitecture assessmentAMDuProfCLI collect --config memory ./prog       # memory access
# Generate reportAMDuProfCLI report -i /tmp/uprof_result/ -o report.html
# Open GUIAMDuProf &

AMD uProf metrics map to VTune equivalents:

  • Retired Instructions → IPC analysis
  • Branch Mispredictions → Bad Speculation
  • L1/L2/L3 Cache Misses → Memory Bound levels
  • Data Cache Accesses → Cache efficiency

7. Roofline model

The roofline model shows whether code is compute-bound or memory-bound by comparing achieved performance against hardware limits:

Performance (GFLOPS/s)     |                    _______________Peak |                 /Perf |              /  compute bound     |           /     |        /     |     /  memory bandwidth bound     |  /     +------------------------------→        Arithmetic Intensity (FLOPS/Byte)
bash
# VTune roofline collectionvtune -collect hpc-performance -r roofline_result ./prog# Then: VTune GUI → Roofline view
# For manual calculation:# Arithmetic Intensity = FLOPS / memory_bytes_accessed# Peak FLOPS = CPUs × cores × freq × FLOPS_per_cycle_per_core# Peak BW = from hardware spec (e.g., 51.2 GB/s for DDR4-3200 dual channel)
# likwid-perfctr for manual roofline data (Linux)likwid-perfctr -C 0 -g FLOPS_DP ./prog          # double-precision FLOPSlikwid-perfctr -C 0 -g MEM ./prog               # memory bandwidth

Related skills

  • Use skills/profilers/hardware-counters for raw PMU event collection with perf stat
  • Use skills/profilers/linux-perf for perf-based profiling on Linux
  • Use skills/low-level-programming/cpu-cache-opt for memory access pattern optimization
  • Use skills/low-level-programming/simd-intrinsics for vectorization to increase FLOPS

来源与署名

来源:mohitmishra786/low-level-dev-skills位于skills/profilers/intel-vtune-amd-uprof提交bdc5847

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架