Qdrant Search Speed Optimization

作者 qdrant6a03d0ce8f55无许可证收录于 2026年10月8日更新于 2026年10月8日

Diagnoses and fixes slow Qdrant search. Use when someone reports 'search is slow', 'high latency', 'queries take too long', 'low QPS', 'throughput too low', 'filtered search is slow', or 'search was fast but now it's slow'. Also use when search performance degrades after config changes or data growth.

仅含说明DevOps & Cloud
AI 生成的概览

诊断 Qdrant 向量搜索变慢的问题,并针对延迟、吞吐量和过滤查询给出配置修复建议。

功能
该技能通过区分延迟、吞吐量、过滤搜索和优化器相关原因,逐步诊断 Qdrant 搜索变慢的问题。它列出诊断步骤,例如重复运行同一查询、关闭 payload 和向量返回、逐个移除过滤条件,以及测试 indexed_only。随后给出配置调整建议,如 HNSW 调优、量化、payload 索引、段数量、副本和优化器预算,并列出应避免的做法。
适用场景
当有人反馈搜索慢、延迟高、查询耗时过长、QPS 低、吞吐量不足、过滤搜索慢,或搜索原本很快后来变慢时使用。配置变更或数据增长后性能下降的情况同样适用。
运行要求
不附带脚本,仅为说明文档。要落实其中的建议,需要一个可供代理检查或重新配置的 Qdrant 部署,文档中还包含指向外部 Qdrant 文档页面的链接。

Diagnose a problem

There the multiple possible reasons for search performance degradation. The most common ones are:

  • Memory pressure: if the working set exceeds available RAM
  • Complex requests (e.g. high hnsw_ef, complex filters without payload index)
  • Competing background processes (e.g. optimizer still running after bulk upload)
  • Problem with the cluster (e.g. network issues, hardware degradation)

Single Query Too Slow (Latency)

Use when: individual queries take too long regardless of load.

Diagnostic steps:

  • Check if second run of the same request is significantly faster (indicates memory pressure)
  • Try the same query with with_payload: false and with_vectors: false to see if payload retrieval is the bottleneck
  • If request uses filters, try to remove them one by one to identify if a specific filter condition is the bottleneck

Common fixes:

Can't Handle Enough QPS (Throughput)

Use when: system can't serve enough queries per second under load.

Filtered Search Is Slow

Use when: filtered search is significantly slower than unfiltered. Most common SA complaint after memory.

  • Create payload index on the filtered field Payload index
  • Use is_tenant=true for primary filtering condition: Tenant index
  • Try ACORN algorithm for complex filters: ACORN
  • Avoid using nested filtering conditions as a primary filter. It might force qdrant to read raw payload values instead of using index.
  • If payload index was added after HNSW build, trigger re-index to create filterable subgraph links

Optimize search performance with parallel updates

Diagnostic steps

  • Try to run the same query with indexed_only=true parameter, if the query is significantly faster, it means that the optimizer is still running and has not yet indexed all segments.
  • If CPU or IO usage is high even with no queries, it also indicates that the optimizer is still running.

Recommended configuration changes

  • reduce optimizer_cpu_budget to reserve more CPU for queries
  • Use prevent_unoptimized=true to prevent creating segments with a large amount of unindexed data for searches. Instead, once a segment reaches the so called indexing_threshold, all additional points will be added in ‘deferred state’.

Learn more here

What NOT to Do

  • Set quantization to not stay in RAM (disk thrashing on every search): avoid memory: cold/cached on Qdrant 1.19 or newer, always_ram: false on 1.18 or older
  • Put HNSW on disk for latency-sensitive production (only for cold storage)
  • Increase segment count for throughput (opposite: fewer = better)
  • Create payload indexes on every field (wastes memory)
  • Blame Qdrant before checking optimizer status

来源与署名

来源:qdrant/skills位于skills/qdrant-performance-optimization/search-speed-optimization提交6a03d0c

许可证: 无许可证

内容归原作者所有。SourceWeft 从公开仓库中收录这些内容。

举报或申请下架