Vector Index Tuning

by wshobson46891e7e60daNo licenseListed Oct 8, 2026Updated Oct 8, 2026

Optimize vector index performance for latency, recall, and memory. Use when tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure.

Instructions onlyAI & Agents
AI-generated overview

Guides tuning of vector indexes for latency, recall, and memory in production search systems.

What it does
This skill provides guidance on optimizing vector indexes for production performance. It covers index type selection by data size, HNSW parameter tuning (M, efConstruction, efSearch), quantization strategies, and best practices for benchmarking, monitoring, and maintenance. It points to a reference file for templates and worked examples.
When to use it
Use it when tuning HNSW parameters, choosing quantization strategies, reducing search latency or memory use, balancing recall against speed, or scaling vector search to very large collections.
Requirements
No scripts; instructions only. It references a bundled details file for templates and worked examples.

Vector Index Tuning

Guide to optimizing vector indexes for production performance.

When to Use This Skill

  • Tuning HNSW parameters
  • Implementing quantization
  • Optimizing memory usage
  • Reducing search latency
  • Balancing recall vs speed
  • Scaling to billions of vectors

Core Concepts

1. Index Type Selection

Data Size           Recommended Index────────────────────────────────────────< 10K vectors  →    Flat (exact search)10K - 1M       →    HNSW1M - 100M      →    HNSW + Quantization> 100M         →    IVF + PQ or DiskANN

2. HNSW Parameters

ParameterDefaultEffect
M16Connections per node, ↑ = better recall, more memory
efConstruction100Build quality, ↑ = better index, slower build
efSearch50Search quality, ↑ = better recall, slower search

3. Quantization Types

Full Precision (FP32): 4 bytes × dimensionsHalf Precision (FP16): 2 bytes × dimensionsINT8 Scalar:           1 byte × dimensionsProduct Quantization:  ~32-64 bytes totalBinary:                dimensions/8 bytes

Templates and detailed worked examples

Full template library and detailed worked examples live in references/details.md. Read that file when you need the concrete templates.

Best Practices

Do's

  • Benchmark with real queries - Synthetic may not represent production
  • Monitor recall continuously - Can degrade with data drift
  • Start with defaults - Tune only when needed
  • Use quantization - Significant memory savings
  • Consider tiered storage - Hot/cold data separation

Don'ts

  • Don't over-optimize early - Profile first
  • Don't ignore build time - Index updates have cost
  • Don't forget reindexing - Plan for maintenance
  • Don't skip warming - Cold indexes are slow

Source and attribution

Source:wshobson/agentsinplugins/llm-application-dev/skills/vector-index-tuningat commit46891e7

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal