ARM / AArch64 Assembly
Purpose
Guide agents through AArch64 (64-bit) and ARM (32-bit Thumb) assembly: registers, calling conventions, inline asm, and NEON/SVE SIMD patterns.
Triggers
- "How do I read ARM64 assembly output?"
- "What are the AArch64 registers and calling convention?"
- "How do I write inline asm for ARM?"
- "What is the difference between AArch64 and ARM Thumb?"
- "How do I use NEON intrinsics?"
Workflow
1. Generate ARM assembly
2. AArch64 registers (AAPCS64)
Width variants: x0 (64-bit), w0 (32-bit, zero-extends to 64), h0 (16), b0 (8).
3. AAPCS64 calling convention
Integer/pointer args: x0–x7
Float/SIMD args: v0–v7
Return: x0 (int), x0+x1 (128-bit), v0 (float/SIMD)
Callee-saved: x19–x28, x29 (fp), x30 (lr), v8–v15 (lower 64 bits)
Caller-saved: everything else
Stack must be 16-byte aligned at any bl or blr instruction.
4. Common AArch64 instructions
5. Typical function prologue/epilogue
6. Inline assembly (GCC/Clang)
AArch64-specific constraints:
"Q"— memory operand suitable for exclusive/acquire/release instructions"r"— any general-purpose register"w"— any FP/SIMD register
7. NEON SIMD intrinsics
Naming convention: v<op><q>_<type>
qsuffix: 128-bit (quad) vector_f32: float32,_s32: int32,_u8: uint8, etc.
8. Darwin vs Linux AArch64 ABI differences
On macOS/iOS, avoid using x18; use _DARWIN_C_LEVEL headers for platform types.
9. AMX primer (Apple Silicon)
Apple Matrix coprocessor (AMX) is not exposed via public intrinsics. Access paths:
Prefer Metal Performance Shaders or Accelerate for matrix workloads on Apple Silicon (skills/platform/apple-silicon).
10. 16KB page size on Apple M-series
macOS on Apple Silicon uses 16KB pages (not 4KB):
Code assuming PAGE_SIZE == 4096 may misalign buffers or fail mmap on macOS.
11. NEON → SVE2 migration hints
See skills/platform/arm-sve for SVE intrinsics and auto-vectorization flags.
For a register reference, see references/reference.md [blocked].
Related skills
- Use
skills/low-level-programming/assembly-x86for x86-64 assembly - Use
skills/compilers/cross-gccfor cross-compilation toolchain - Use
skills/debuggers/gdbfor debugging ARM code with gdbserver - Use
skills/platform/arm-svefor SVE/SVE2 scalable vectors - Use
skills/platform/apple-siliconfor M-series unified memory and AMX


