Vector Index Tuning

作者 wshobson46891e7e60da無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Optimize vector index performance for latency, recall, and memory. Use when tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure.

僅含說明AI & Agents
AI 產生的概覽

指導向量索引調校,以最佳化延遲、召回率與記憶體用量。

功能
此技能提供針對正式環境的向量索引效能最佳化指引。內容涵蓋依資料規模選擇索引類型、HNSW 參數(M、efConstruction、efSearch)調校、量化策略,以及基準測試、監控與維護的最佳實務。它會指向一份參考檔案以取得範本與詳細範例。
適用情境
適用於調校 HNSW 參數、選擇量化策略、降低搜尋延遲或記憶體用量、在召回率與速度之間取捨,或將向量搜尋擴展到超大規模資料集的情境。
執行需求
不需要指令碼,僅為說明性內容。它會引用一份隨附的詳細說明檔案以取得範本與範例。

Vector Index Tuning

Guide to optimizing vector indexes for production performance.

When to Use This Skill

  • Tuning HNSW parameters
  • Implementing quantization
  • Optimizing memory usage
  • Reducing search latency
  • Balancing recall vs speed
  • Scaling to billions of vectors

Core Concepts

1. Index Type Selection

Data Size           Recommended Index────────────────────────────────────────< 10K vectors  →    Flat (exact search)10K - 1M       →    HNSW1M - 100M      →    HNSW + Quantization> 100M         →    IVF + PQ or DiskANN

2. HNSW Parameters

ParameterDefaultEffect
M16Connections per node, ↑ = better recall, more memory
efConstruction100Build quality, ↑ = better index, slower build
efSearch50Search quality, ↑ = better recall, slower search

3. Quantization Types

Full Precision (FP32): 4 bytes × dimensionsHalf Precision (FP16): 2 bytes × dimensionsINT8 Scalar:           1 byte × dimensionsProduct Quantization:  ~32-64 bytes totalBinary:                dimensions/8 bytes

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's

  • Benchmark with real queries - Synthetic may not represent production
  • Monitor recall continuously - Can degrade with data drift
  • Start with defaults - Tune only when needed
  • Use quantization - Significant memory savings
  • Consider tiered storage - Hot/cold data separation

Don'ts

  • Don't over-optimize early - Profile first
  • Don't ignore build time - Index updates have cost
  • Don't forget reindexing - Plan for maintenance
  • Don't skip warming - Cold indexes are slow

來源與署名

來源:wshobson/agents位於plugins/llm-application-dev/skills/vector-index-tuning提交46891e7

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架