M10 Performance

actionbook/rust-skills/skills/m10-performance

作者 actionbook5c40d3ad7851无许可证1.5K 个星标收录于 2026年10月8日更新于 2026年10月8日仓库6周前更新

CRITICAL: Use for performance optimization. Triggers: performance, optimization, benchmark, profiling, flamegraph, criterion, slow, fast, allocation, cache, SIMD, make it faster, 性能优化, 基准测试

AI 生成的概览

指导 Rust 性能优化:先测量瓶颈,再选择数据结构、内存分配与缓存策略。

功能
该技能提供一套性能优化决策框架,核心是先做性能分析再改代码。它把减少分配、提升缓存局部性、并行化、避免拷贝和减少间接寻址等目标对应到具体实现选择。内容还包括性能分析与基准测试工具、优化优先级顺序、常用技巧与反模式,并指向所有权、并发和资源管理等相关技能。
适用场景
当代码运行过慢,需要判断优化什么以及优化是否值得时使用。适用于性能分析、基准测试以及内存分配或缓存调优,尤其是 Rust 项目。
运行要求
不附带任何脚本或工具,仅为说明性内容。文中提到的工具(cargo bench、criterion、perf、flamegraph、heaptrack、valgrind)均为可选,使用时需环境中已安装。

Performance Optimization

Layer 2: Design Choices

Core Question

What's the bottleneck, and is optimization worth it?

Before optimizing:

  • Have you measured? (Don't guess)
  • What's the acceptable performance?
  • Will optimization add complexity?

Performance Decision → Implementation

GoalDesign ChoiceImplementation
Reduce allocationsPre-allocate, reusewith_capacity, object pools
Improve cacheContiguous dataVec, SmallVec
ParallelizeData parallelismrayon, threads
Avoid copiesZero-copyReferences, Cow<T>
Reduce indirectionInline datasmallvec, arrays

Thinking Prompt

Before optimizing:

  1. Have you measured?

    • Profile first → flamegraph, perf
    • Benchmark → criterion, cargo bench
    • Identify actual hotspots
  2. What's the priority?

    • Algorithm (10x-1000x improvement)
    • Data structure (2x-10x)
    • Allocation (2x-5x)
    • Cache (1.5x-3x)
  3. What's the trade-off?

    • Complexity vs speed
    • Memory vs CPU
    • Latency vs throughput

Trace Up ↑

To domain constraints (Layer 3):

"How fast does this need to be?"    ↑ Ask: What's the performance SLA?    ↑ Check: domain-* (latency requirements)    ↑ Check: Business requirements (acceptable response time)
QuestionTrace ToAsk
Latency requirementsdomain-*What's acceptable response time?
Throughput needsdomain-*How many requests per second?
Memory constraintsdomain-*What's the memory budget?

Trace Down ↓

To implementation (Layer 1):

"Need to reduce allocations"    ↓ m01-ownership: Use references, avoid clone    ↓ m02-resource: Pre-allocate with_capacity
"Need to parallelize"    ↓ m07-concurrency: Choose rayon or threads    ↓ m07-concurrency: Consider async for I/O-bound
"Need cache efficiency"    ↓ Data layout: Prefer Vec over HashMap when possible    ↓ Access patterns: Sequential over random access

Quick Reference

ToolPurpose
cargo benchMicro-benchmarks
criterionStatistical benchmarks
perf / flamegraphCPU profiling
heaptrackAllocation tracking
valgrind / cachegrindCache analysis

Optimization Priority

1. Algorithm choice     (10x - 1000x)2. Data structure       (2x - 10x)3. Allocation reduction (2x - 5x)4. Cache optimization   (1.5x - 3x)5. SIMD/Parallelism     (2x - 8x)

Common Techniques

TechniqueWhenHow
Pre-allocationKnown sizeVec::with_capacity(n)
Avoid cloningHot pathsUse references or Cow<T>
Batch operationsMany small opsCollect then process
SmallVecUsually smallsmallvec::SmallVec<[T; N]>
Inline buffersFixed-size dataArrays over Vec

Common Mistakes

MistakeWhy WrongBetter
Optimize without profilingWrong targetProfile first
Benchmark in debug modeMeaninglessAlways --release
Use LinkedListCache unfriendlyVec or VecDeque
Hidden .clone()Unnecessary allocsUse references
Premature optimizationWasted effortMake it work first

Anti-Patterns

Anti-PatternWhy BadBetter
Clone to avoid lifetimesPerformance costProper ownership
Box everythingIndirection costStack when possible
HashMap for small setsOverheadVec with linear search
String concat in loopO(n^2)String::with_capacity or format!

Related Skills

WhenSee
Reducing clonesm01-ownership
Concurrency optionsm07-concurrency
Smart pointer choicem02-resource
Domain requirementsdomain-*

来源与署名

来源:actionbook/rust-skills位于skills/m10-performance提交5c40d3a

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架