Intel VTune & AMD uProf
Purpose
Guide agents through CPU microarchitecture profiling with Intel VTune Profiler (free Community Edition) and AMD uProf: hotspot identification, microarchitecture analysis, memory access pattern optimization, pipeline stall diagnosis, and roofline model analysis.
Triggers
- "How do I use Intel VTune to profile my code?"
- "What are pipeline stalls and how do I reduce them?"
- "How do I analyze memory bandwidth with VTune?"
- "What is the roofline model and how do I use it?"
- "How do I use AMD uProf as a free alternative to VTune?"
- "My code has good cache hit rates but is still slow"
Workflow
1. VTune setup (free Community Edition)
2. Analysis types
3. Hotspot analysis
Build with debug info for meaningful symbols:
4. Microarchitecture exploration — pipeline stalls
Key metrics to examine:
5. Memory access analysis
Reading DRAM bandwidth:
If DRAM-bound: optimize data layout (AoS → SoA), reduce working set, improve spatial locality.
6. AMD uProf — free alternative for AMD CPUs
AMD uProf metrics map to VTune equivalents:
Retired Instructions→ IPC analysisBranch Mispredictions→ Bad SpeculationL1/L2/L3 Cache Misses→ Memory Bound levelsData Cache Accesses→ Cache efficiency
7. Roofline model
The roofline model shows whether code is compute-bound or memory-bound by comparing achieved performance against hardware limits:
Related skills
- Use
skills/profilers/hardware-countersfor raw PMU event collection with perf stat - Use
skills/profilers/linux-perffor perf-based profiling on Linux - Use
skills/low-level-programming/cpu-cache-optfor memory access pattern optimization - Use
skills/low-level-programming/simd-intrinsicsfor vectorization to increase FLOPS


