Vector Index Tuning

作者 wshobson46891e7e60da无许可证收录于 2026年10月8日更新于 2026年10月8日

Optimize vector index performance for latency, recall, and memory. Use when tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure.

仅含说明AI & Agents
AI 生成的概览

指导向量索引调优,以优化延迟、召回率和内存占用。

功能
该技能提供面向生产环境的向量索引性能优化指导。内容涵盖按数据规模选择索引类型、HNSW 参数(M、efConstruction、efSearch)调优、量化策略,以及基准测试、监控和维护方面的最佳实践。它指向一个参考文件以获取模板和详细示例。
适用场景
适用于调优 HNSW 参数、选择量化策略、降低搜索延迟或内存占用、在召回率与速度之间权衡,或将向量搜索扩展到超大规模数据集的场景。
运行要求
无需脚本,仅为说明性内容。它会引用一个随附的详细说明文件以获取模板和示例。

Vector Index Tuning

Guide to optimizing vector indexes for production performance.

When to Use This Skill

  • Tuning HNSW parameters
  • Implementing quantization
  • Optimizing memory usage
  • Reducing search latency
  • Balancing recall vs speed
  • Scaling to billions of vectors

Core Concepts

1. Index Type Selection

Data Size           Recommended Index────────────────────────────────────────< 10K vectors  →    Flat (exact search)10K - 1M       →    HNSW1M - 100M      →    HNSW + Quantization> 100M         →    IVF + PQ or DiskANN

2. HNSW Parameters

ParameterDefaultEffect
M16Connections per node, ↑ = better recall, more memory
efConstruction100Build quality, ↑ = better index, slower build
efSearch50Search quality, ↑ = better recall, slower search

3. Quantization Types

Full Precision (FP32): 4 bytes × dimensionsHalf Precision (FP16): 2 bytes × dimensionsINT8 Scalar:           1 byte × dimensionsProduct Quantization:  ~32-64 bytes totalBinary:                dimensions/8 bytes

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's

  • Benchmark with real queries - Synthetic may not represent production
  • Monitor recall continuously - Can degrade with data drift
  • Start with defaults - Tune only when needed
  • Use quantization - Significant memory savings
  • Consider tiered storage - Hot/cold data separation

Don'ts

  • Don't over-optimize early - Profile first
  • Don't ignore build time - Index updates have cost
  • Don't forget reindexing - Plan for maintenance
  • Don't skip warming - Cold indexes are slow

来源与署名

来源:wshobson/agents位于plugins/llm-application-dev/skills/vector-index-tuning提交46891e7

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架