Qdrant Scaling Query Volume

作者 qdrant6a03d0ce8f55無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Guides Qdrant query volume scaling. Use when someone asks 'query returns too many results', 'scroll performance', 'large limit values', 'paginating search results', 'fetching many vectors', or 'high cardinality results'.

AI 產生的概覽

說明 Qdrant 如何透過基於 Poisson 分布的次取樣來擴展跨分片的大 limit 查詢。

功能
此技能針對 Qdrant 在查詢使用較大 limit 且涉及多個分片時的查詢量擴展提供指引。它說明如何讓每個分片回傳以 Poisson 分布統計計算出的較小 limit,再合併結果,而不是向每個分片要求完整 limit。它也說明此策略的啟用條件,以及在結果可能略微不完整與減少分片間資料傳輸之間的取捨。
適用情境
當有人詢問查詢回傳過多結果、scroll 效能、較大的 limit 值、搜尋結果分頁、取得大量向量或高基數結果時使用。它適用於多分片、啟用自動分片且非精確的查詢,且 limit 加 offset 達到次取樣門檻的情況。
執行需求
不需要指令碼或工具,僅為說明性指示。

Scaling for Query Volume

Problem: When a query has a large limit (e.g. 1000) and there are multiple shards (e.g. 10), naively each shard must return the full 1000 results — totaling 10,000 scored points transferred and merged. This is wasteful since data is randomly distributed across auto-shards.

Core idea

Instead of asking every shard for the full limit, ask each shard for a smaller limit computed via Poisson distribution statistics, then merge. This is safe because auto-sharding guarantees random, independent data distribution.

When it activates

  • More than 1 shard
  • Auto-sharding is in use (all queried shards share the same shard key)
  • The request's limit + offset >= SHARD_QUERY_SUBSAMPLING_LIMIT (128)
  • The query is not exact

Key tradeoff

The strategy trades a small probability of slightly incomplete results for a large reduction in inter-shard data transfer, especially for high-limit queries across many shards. The 1.2x safety factor and the 99.9% Poisson threshold keep the error rate very low — comparable to inaccuracies already introduced by approximate vector indices like HNSW.

來源與署名

來源:qdrant/skills位於skills/qdrant-scaling/scaling-query-volume提交6a03d0c

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架