SIMD Intrinsics
Purpose
Guide agents through SIMD: reading auto-vectorization output, writing SSE2/AVX2/NEON intrinsics, runtime CPU feature detection, and choosing between compiler auto-vectorization and manual intrinsics.
Triggers
- "How do I check if my loop is being auto-vectorized?"
- "How do I write SSE2/AVX2 intrinsics?"
- "Auto-vectorization failed — how do I fix it?"
- "How do I check for CPU features at runtime?"
- "Should I use intrinsics or let the compiler vectorize?"
- "How do I write NEON intrinsics for ARM?"
Workflow
1. Check auto-vectorization
Common auto-vectorization blockers:
2. Runtime CPU feature detection
3. SSE2 / SSE4.2 intrinsics (x86)
4. AVX2 intrinsics (x86)
Compile with: gcc -O2 -mavx2 -mfma src/simd.c
5. NEON intrinsics (ARM/AArch64)
Compile with: gcc -O2 -march=armv8-a+simd src/simd.c
6. Choose auto-vectorization vs intrinsics
7. Alignment and performance
For Intel Intrinsics Guide reference and NEON lookup tables, see references/intel-intrinsics-guide.md [blocked].
Related skills
- Use
skills/compilers/gccfor-march,-msse4.2,-mavx2flags - Use
skills/compilers/clangfor vectorization remarks and auto-vectorization control - Use
skills/profilers/linux-perfto measure SIMD impact with perf stat counters - Use
skills/low-level-programming/assembly-x86for reading SIMD assembly output


