Qdrant Scaling Query Volume

by qdrant6a03d0ce8f55No licenseListed Oct 8, 2026Updated Oct 8, 2026

Guides Qdrant query volume scaling. Use when someone asks 'query returns too many results', 'scroll performance', 'large limit values', 'paginating search results', 'fetching many vectors', or 'high cardinality results'.

AI-generated overview

Explains how Qdrant scales large-limit queries across shards using Poisson-based subsampling.

What it does
This skill provides guidance on scaling Qdrant query volume when a query uses a large limit across multiple shards. It describes asking each shard for a smaller limit computed with Poisson distribution statistics and then merging the results, instead of requesting the full limit from every shard. It also covers when the strategy activates and the tradeoff between slightly incomplete results and reduced inter-shard data transfer.
When to use it
Use it when someone asks about queries returning too many results, scroll performance, large limit values, paginating search results, fetching many vectors, or high cardinality results. It applies to multi-shard, auto-sharded, non-exact queries where limit plus offset reaches the subsampling threshold.
Requirements
No scripts or tools are required; it is instructions only.

Scaling for Query Volume

Problem: When a query has a large limit (e.g. 1000) and there are multiple shards (e.g. 10), naively each shard must return the full 1000 results — totaling 10,000 scored points transferred and merged. This is wasteful since data is randomly distributed across auto-shards.

Core idea

Instead of asking every shard for the full limit, ask each shard for a smaller limit computed via Poisson distribution statistics, then merge. This is safe because auto-sharding guarantees random, independent data distribution.

When it activates

  • More than 1 shard
  • Auto-sharding is in use (all queried shards share the same shard key)
  • The request's limit + offset >= SHARD_QUERY_SUBSAMPLING_LIMIT (128)
  • The query is not exact

Key tradeoff

The strategy trades a small probability of slightly incomplete results for a large reduction in inter-shard data transfer, especially for high-limit queries across many shards. The 1.2x safety factor and the 99.9% Poisson threshold keep the error rate very low — comparable to inaccuracies already introduced by approximate vector indices like HNSW.

Source and attribution

Source:qdrant/skillsinskills/qdrant-scaling/scaling-query-volumeat commit6a03d0c

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal