Qdrant Scaling

by qdrant6a03d0ce8f55No licenseListed Oct 8, 2026Updated Oct 8, 2026

Guides Qdrant scaling decisions. Use when someone asks 'how many nodes do I need', 'data doesn't fit on one node', 'need more throughput or QPS', 'CPU is pegged / can't keep up with the request rate', 'one query is slow / p99 or tail latency too high', 'cluster is slow', 'too many tenants', 'vertical or horizontal', 'how to shard', 'need to add capacity', 'large limit / pagination / scroll is slow', or 'only recent data matters / expiring old vectors / retention window'.

Instructions onlyDevOps & Cloud
AI-generated overview

Routes Qdrant scaling questions to guidance on data volume, QPS, latency, query volume and sharding.

What it does
This skill acts as a routing index for Qdrant scaling guidance. It matches a user's reported symptom, such as data not fitting on one node, low throughput, high tail latency, large result sets, many tenants or retention needs, to a specific guidance file and instructs the agent to read that file and answer from it. It also notes that latency and throughput favor opposite segment-count settings.
When to use it
Use it when someone asks how many Qdrant nodes they need, whether to scale vertically or horizontally, how to shard, how to raise QPS, how to reduce p99 latency, how to handle large limits or pagination, or how to expire old vectors.
Requirements
Requires the referenced guidance files to be present and readable; the skill declares Read, Grep and Glob tools and ships no scripts of its own.

Qdrant Scaling

Route first, then answer. Match the user's symptom in the table, Read that file, and answer from it. Do not answer from this page alone: it contains routing only, not the guidance. If two rows match, read both.

The user saysRead
Data does not fit on a single node, running out of disk or memory as the dataset growsscaling-data-volume/SKILL.md
Need to shard the collection across more nodes, data outgrew one nodescaling-data-volume/SKILL.md
Cannot handle enough parallel queries, need higher QPS or throughputscaling-qps/SKILL.md
Can't hold the request rate, CPU is peggedscaling-qps/SKILL.md
A single query is too slow, need to cut the tail latency of individual requestsminimize-latency/SKILL.md
p99 or tail latency too high, but traffic/QPS is fineminimize-latency/SKILL.md
Queries return very large result sets and slow downscaling-query-volume/SKILL.md
Large limit, top-1000 queries, pagination, scroll across shardsscaling-query-volume/SKILL.md
Many tenants or customers, one collection each, tenant isolationscaling-data-volume/tenant-scaling/SKILL.md
Only recent data matters, retention, expiring old vectors, time-based rotationscaling-data-volume/sliding-time-window/SKILL.md
Single node no longer fits the workload, before deciding to shardscaling-data-volume/vertical-scaling/SKILL.md
Already vertically maxed out, need more nodes, reshardingscaling-data-volume/horizontal-scaling/SKILL.md

Latency and throughput pull opposite ways on segment count. For latency, increase segments toward the CPU core count (default_segment_number: 16). For throughput, use fewer and larger segments (default_segment_number: 2). Applying the wrong direction makes the reported problem worse.

Source and attribution

Source:qdrant/skillsinskills/qdrant-scalingat commit6a03d0c

License: No license

Content belongs to its original authors. SourceWeft indexes it from a public repository.

Report or request removal