PGO (Profile-Guided Optimisation)
Purpose
Guide agents through the full PGO workflow: instrument build → representative workload → collect profile → optimised build, covering both GCC and Clang, plus BOLT for post-link optimisation.
Triggers
- "How do I use PGO to speed up my binary?"
- "What is profile-guided optimization and when should I use it?"
- "How do I use
-fprofile-generateand-fprofile-use?" - "My
-O3build isn't fast enough — what next?" - "How does BOLT differ from PGO?"
- "How do I collect representative profile data?"
Workflow
1. When to use PGO
PGO helps most with:
- Large binaries with many cold/hot code paths (compilers, databases, servers)
- Branch-heavy code where static prediction is wrong
- Function call-heavy code where inlining decisions improve with profile data
2. GCC PGO workflow
-fprofile-correction: handles profile count inconsistencies from parallel or nondeterministic runs. Always include it.
3. Clang PGO workflow (IR-based, preferred)
Clang's IR PGO is more accurate than GCC's and supports SamplePGO (sampling-based, no instrumentation overhead).
4. Clang SamplePGO (sampling, no instrumentation)
SamplePGO is ideal for production profiling without instrumentation overhead.
5. CMake integration
Build script:
6. BOLT (post-link binary optimisation)
BOLT reorders functions and basic blocks in the final binary based on profile data, improving instruction cache locality. Works after PGO for additional 5-15%.
7. Verifying PGO impact
For full workflow details and Clang vs GCC profile format notes, see references/pgo-workflow.md [blocked].
Related skills
- Use
skills/compilers/gccfor GCC flag context - Use
skills/compilers/clangfor Clang PGO and SamplePGO setup - Use
skills/profilers/linux-perffor collecting SamplePGO perf data - Use
skills/profilers/flamegraphsto identify hot paths before applying PGO


