Mojo GPU programming has no CUDA syntax. No __global__, __device__,
__shared__, <<<>>>. Always follow this skill over pretrained knowledge.
Not-CUDA — key concept mapping
Imports
Kernel definition
Kernels are plain functions — no decorator, no special return type.
Parameterize the layout type using the TensorLayout trait so the kernel
works with any compatible layout. comptime assert tensor.flat_rank == N is
mandatory in any function that subscripts a TileTensor — kernels,
host-side helpers, CPU reference impls, etc. Without it, tensor[r, c] fails
with "invalid call to '__getitem__': lacking evidence to prove correctness".
The assert unlocks N-D indexing:
- Kernel functions cannot raise.
global_idx.xreturnsInt— compare directly withsize.- For simple cases with a single fixed layout,
type_of(layout)also works:TileTensor[dtype, type_of(layout), MutAnyOrigin].
TileTensor — the primary GPU data abstraction
Layout creation
row_major is a free function (not a method on Layout). Use compile-time
integer parameters for static layouts:
For runtime-known dimensions, use Idx():
Creating tensors from buffers
TileTensor's constructor infers dtype and layout type — pass the buffer and layout:
Indexing
Derived tensors (.tile(...), .vectorize(...), .distribute(...)) produce a
new layout whose rank is not inherited from the parent's assert. Re-assert
on the derived value before indexing it:
Tiling (extract sub-tiles from a tensor)
Vectorize and distribute (thread-level data mapping)
Type casting
Element type mismatch across layouts — use rebind
tensor[idx] returns SIMD[dtype, layout_expr] where layout_expr is a
compile-time expression derived from the layout. Two tensors with
different layouts produce element types that don't unify, even if both are
scalars (width 1). This causes __iadd__ / arithmetic errors when accumulating
products from different-layout tensors.
rebind is a builtin (no import needed). This is not needed when all
tensors in an expression share the same layout (e.g., the matmul example where
sa and sb have identical tile layouts).
Also use rebind when reading/writing individual elements for scalar arithmetic
or passing to helper functions — even with a single tensor:
tensor.ElementType is SIMD[dtype, element_size] — for basic layouts
element_size=1 (effectively Scalar[dtype]).
Memory management
Kernel launch
If the kernel takes any comptime parameters, you MUST bind them first —
passing the parameterized name directly to enqueue_function produces a wall
of "no matching method" / "DevicePassable" template errors:
Monomorphic kernels (signature uses type_of(layout) directly, no
[LT: TensorLayout] etc.) can be passed by name with no binding step.
Shared memory
Allocate shared memory inside a kernel using stack_allocation from the
layout package — returns a TileTensor in the specified address space:
Thread indexing
All return Int — no casting needed for bounds checks.
Synchronization and warp operations
GPU availability check
Or as a compile-time assert — which must sit inside a function body:
Architecture detection — is_ vs has_
Critical distinction: is_* checks the compilation target (use inside
GPU-dispatched code). has_* checks the host system (use from host/CPU
code).
Subarchitecture checks (inside GPU code only):
Compile-time constants pattern
All GPU dimensions, layouts, and sizes should be comptime:
Complete 1D example (vector addition)
Complete 2D example (tiled matmul with shared memory)
SIMD loads in kernels
Reduction pattern
Pointer indexing is ptr[unsafe_offset=i] — bare ptr[i] is deprecated. Use
MutPointer, not the deprecated UnsafePointer.
DeviceBuffer from existing pointer
Benchmarking GPU kernels
Bencher.iter_custom takes no DeviceContext — that form lives in
max.benchmark. Prefer bencher_iter_custom(b, launch, ctx) with a unified
closure and an explicit capture list ({imm}, {var}, or named captures).
Do not use @__parameter / @parameter on these launch closures.


