M10 Performance

作者 actionbook5c40d3ad7851無授權條款1.5K 個星標收錄於 2026年10月8日更新於 2026年10月8日儲存庫6 週前更新

CRITICAL: Use for performance optimization. Triggers: performance, optimization, benchmark, profiling, flamegraph, criterion, slow, fast, allocation, cache, SIMD, make it faster, 性能优化, 基准测试

AI 產生的概覽

指導 Rust 效能最佳化:先量測瓶頸,再選擇資料結構、記憶體配置與快取策略。

功能
此技能提供一套效能最佳化決策框架,核心是先做效能分析再改程式碼。它把減少配置、提升快取區域性、平行化、避免複製與減少間接尋址等目標,對應到具體的實作選擇。內容還包括效能分析與基準測試工具、最佳化優先順序、常用技巧與反模式,並指向所有權、並行與資源管理等相關技能。
適用情境
當程式碼執行過慢,需要判斷該最佳化什麼以及最佳化是否值得時使用。適用於效能分析、基準測試以及記憶體配置或快取調校,尤其是 Rust 專案。
執行需求
不附帶任何指令碼或工具,僅為說明性內容。文中提到的工具(cargo bench、criterion、perf、flamegraph、heaptrack、valgrind)皆為選用,使用時需環境中已安裝。

Performance Optimization

Layer 2: Design Choices

Core Question

What's the bottleneck, and is optimization worth it?

Before optimizing:

  • Have you measured? (Don't guess)
  • What's the acceptable performance?
  • Will optimization add complexity?

Performance Decision → Implementation

GoalDesign ChoiceImplementation
Reduce allocationsPre-allocate, reusewith_capacity, object pools
Improve cacheContiguous dataVec, SmallVec
ParallelizeData parallelismrayon, threads
Avoid copiesZero-copyReferences, Cow<T>
Reduce indirectionInline datasmallvec, arrays

Thinking Prompt

Before optimizing:

  1. Have you measured?

    • Profile first → flamegraph, perf
    • Benchmark → criterion, cargo bench
    • Identify actual hotspots
  2. What's the priority?

    • Algorithm (10x-1000x improvement)
    • Data structure (2x-10x)
    • Allocation (2x-5x)
    • Cache (1.5x-3x)
  3. What's the trade-off?

    • Complexity vs speed
    • Memory vs CPU
    • Latency vs throughput

Trace Up ↑

To domain constraints (Layer 3):

"How fast does this need to be?"    ↑ Ask: What's the performance SLA?    ↑ Check: domain-* (latency requirements)    ↑ Check: Business requirements (acceptable response time)
QuestionTrace ToAsk
Latency requirementsdomain-*What's acceptable response time?
Throughput needsdomain-*How many requests per second?
Memory constraintsdomain-*What's the memory budget?

Trace Down ↓

To implementation (Layer 1):

"Need to reduce allocations"    ↓ m01-ownership: Use references, avoid clone    ↓ m02-resource: Pre-allocate with_capacity
"Need to parallelize"    ↓ m07-concurrency: Choose rayon or threads    ↓ m07-concurrency: Consider async for I/O-bound
"Need cache efficiency"    ↓ Data layout: Prefer Vec over HashMap when possible    ↓ Access patterns: Sequential over random access

Quick Reference

ToolPurpose
cargo benchMicro-benchmarks
criterionStatistical benchmarks
perf / flamegraphCPU profiling
heaptrackAllocation tracking
valgrind / cachegrindCache analysis

Optimization Priority

1. Algorithm choice     (10x - 1000x)2. Data structure       (2x - 10x)3. Allocation reduction (2x - 5x)4. Cache optimization   (1.5x - 3x)5. SIMD/Parallelism     (2x - 8x)

Common Techniques

TechniqueWhenHow
Pre-allocationKnown sizeVec::with_capacity(n)
Avoid cloningHot pathsUse references or Cow<T>
Batch operationsMany small opsCollect then process
SmallVecUsually smallsmallvec::SmallVec<[T; N]>
Inline buffersFixed-size dataArrays over Vec

Common Mistakes

MistakeWhy WrongBetter
Optimize without profilingWrong targetProfile first
Benchmark in debug modeMeaninglessAlways --release
Use LinkedListCache unfriendlyVec or VecDeque
Hidden .clone()Unnecessary allocsUse references
Premature optimizationWasted effortMake it work first

Anti-Patterns

Anti-PatternWhy BadBetter
Clone to avoid lifetimesPerformance costProper ownership
Box everythingIndirection costStack when possible
HashMap for small setsOverheadVec with linear search
String concat in loopO(n^2)String::with_capacity or format!

Related Skills

WhenSee
Reducing clonesm01-ownership
Concurrency optionsm07-concurrency
Smart pointer choicem02-resource
Domain requirementsdomain-*

來源與署名

來源:actionbook/rust-skills位於skills/m10-performance提交5c40d3a

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架