Performance optimization
Performance work is a measurement discipline, not a bag of tricks. The method is always the
same: profile → find the one bottleneck → fix that → measure again. This skill teaches that
loop and the highest-leverage fixes (pooling, batching, allocation control, asset budgets), and
points you at each engine's profiler. It pairs with physics-tuning for simulation cost.
When to use
- Use when the frame rate is low or uneven, the game stutters/hitches, or it must hit a target (60 FPS desktop, 30/60 mobile) and currently doesn't.
- Use to decide what to optimize: profile, read the frame budget, and identify whether the CPU or GPU is the bottleneck before changing any code.
- Use to apply specific fixes: object pooling, draw-call/batch reduction, removing per-frame allocations and GC spikes, and setting asset budgets.
When not to use: for physics jitter/tunneling/timestep specifically, use physics-tuning.
For the engine's concrete profiler UI and rendering settings, use that
engine skill (godot-export covers some build settings; engine cores cover the rest). This skill
is the cross-engine method and the shared fixes.
The golden rule: measure first, never guess
Most performance "fixes" applied without profiling target the wrong thing and add complexity for no gain. Do not optimize code you have not measured. Open the profiler, find the single biggest cost in a representative scene on representative hardware, and fix that. Re-measure to confirm the fix helped before moving on. Profile a release/optimized build where it matters — editor and debug builds lie (editor overhead, no compiler optimization).
Core workflow
- Define the target and reproduce. State the goal (e.g. 60 FPS = 16.67 ms/frame) and find a repeatable worst-case scene. "Sometimes slow" is unfixable; a reproducible spike is fixable.
- Profile before touching code. Run the engine profiler and read the frame: total frame time, and the split between CPU (game logic, physics, scripts) and GPU (rendering).
- Find the bottleneck — CPU or GPU. If GPU time ≫ CPU, attack draw calls/overdraw/shaders/ resolution. If CPU time dominates, attack scripts/physics/allocations. Fixing the wrong side does nothing.
- Fix the single biggest cost. Prefer an algorithmic win (do less work, cache, spatial partition, run less often) over micro-optimizing a hot line. Apply the matching shared fix (pooling, batching, allocation removal).
- Re-measure on the same scene/hardware. Confirm the number moved. Keep or revert based on data, not intuition.
- Set budgets so it stays fixed. Per-frame ms budgets per subsystem, plus asset budgets (texture sizes, triangle counts, draw-call ceilings); add a perf check to verification.
- Report measured numbers. State before/after frame time, the bottleneck found, and the fix — never "should be faster". If you could only measure in-editor, say so.
Patterns
1. Frame budget math (turn "feels slow" into a number)
2. Measure with the engine profiler (do this before any fix)
3. Object pooling (stop allocating/freeing in hot loops)
4. Cut draw calls (the most common GPU-side win)
5. Kill per-frame allocations (GC spikes = stutter)
Pitfalls
- Optimizing without profiling. The intuitive culprit is usually wrong. Measure first, every time.
- Profiling the editor / a debug build. Editor overhead and unoptimized code mislead. Profile a release build on target hardware for real numbers.
- Fixing the wrong side. Micro-optimizing CPU code when the GPU is the bottleneck (or vice versa) changes nothing. Check the CPU-vs-GPU split first.
- Micro-optimizing over algorithm. Shaving a function when an O(n²) loop or a per-frame full-scene query is the real cost. Reduce the work, don't polish it.
- Instantiate/free in hot loops. Spawning and destroying bullets/particles every frame causes fragmentation and GC spikes. Pool them.
- Per-frame allocations / LINQ / boxing in
Update(C#) feed the GC → periodic hitches. Cache and reuse. - Draw-call explosion from unique materials and unbatched sprites/meshes. Atlas, share materials, instance, batch.
- Overdraw from stacked transparents/particles/full-screen effects re-shading pixels.
- No budgets. Without per-subsystem ms and asset ceilings, performance silently regresses; enforce them in your build/CI checks.
- Optimizing too early. Don't contort a prototype for performance before it's fun or measured.
References
- For per-engine profiler walkthroughs, the CPU-vs-GPU triage flowchart, a complete pooling
manager, batching/instancing rules per engine, allocation/GC guidance, LOD/culling, and asset
budgets (texture sizes, triangle counts, audio, mobile thermals), read
references/profiling-and-budgets.md.
Related skills
physics-tuning— simulation cost, fixed-step budget, sleeping bodies, broadphase layers.godot-export— release/build settings that affect measured performance.procedural-gen,game-ai— common CPU hotspots (generation, pathfinding) to budget and defer.roguelike,tower-defense,survival-crafting— entity-heavy genres that need pooling/budgets.


