Qdrant Scaling

作者 qdrant6a03d0ce8f55無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Guides Qdrant scaling decisions. Use when someone asks 'how many nodes do I need', 'data doesn't fit on one node', 'need more throughput or QPS', 'CPU is pegged / can't keep up with the request rate', 'one query is slow / p99 or tail latency too high', 'cluster is slow', 'too many tenants', 'vertical or horizontal', 'how to shard', 'need to add capacity', 'large limit / pagination / scroll is slow', or 'only recent data matters / expiring old vectors / retention window'.

僅含說明DevOps & Cloud
AI 產生的概覽

將 Qdrant 擴縮容問題導向資料量、QPS、延遲、查詢量與分片相關指南。

功能
這個技能是 Qdrant 擴縮容指南的路由索引。它會把使用者描述的症狀(例如資料放不進單一節點、吞吐量不足、尾端延遲過高、結果集過大、租戶過多或需要資料保留)對應到特定的指南檔案,並要求代理讀取該檔案後據此回答。它也指出延遲與吞吐量在分段數量設定上方向相反。
適用情境
當有人詢問需要多少 Qdrant 節點、該垂直或水平擴充、如何分片、如何提高 QPS、如何降低 p99 延遲、如何處理大 limit 或分頁,或如何讓舊向量過期時使用。
執行需求
需要所引用的指南檔案存在且可讀取;這個技能宣告使用 Read、Grep 與 Glob 工具,本身不附帶指令碼。

Qdrant Scaling

Route first, then answer. Match the user's symptom in the table, Read that file, and answer from it. Do not answer from this page alone: it contains routing only, not the guidance. If two rows match, read both.

The user saysRead
Data does not fit on a single node, running out of disk or memory as the dataset growsscaling-data-volume/SKILL.md
Need to shard the collection across more nodes, data outgrew one nodescaling-data-volume/SKILL.md
Cannot handle enough parallel queries, need higher QPS or throughputscaling-qps/SKILL.md
Can't hold the request rate, CPU is peggedscaling-qps/SKILL.md
A single query is too slow, need to cut the tail latency of individual requestsminimize-latency/SKILL.md
p99 or tail latency too high, but traffic/QPS is fineminimize-latency/SKILL.md
Queries return very large result sets and slow downscaling-query-volume/SKILL.md
Large limit, top-1000 queries, pagination, scroll across shardsscaling-query-volume/SKILL.md
Many tenants or customers, one collection each, tenant isolationscaling-data-volume/tenant-scaling/SKILL.md
Only recent data matters, retention, expiring old vectors, time-based rotationscaling-data-volume/sliding-time-window/SKILL.md
Single node no longer fits the workload, before deciding to shardscaling-data-volume/vertical-scaling/SKILL.md
Already vertically maxed out, need more nodes, reshardingscaling-data-volume/horizontal-scaling/SKILL.md

Latency and throughput pull opposite ways on segment count. For latency, increase segments toward the CPU core count (default_segment_number: 16). For throughput, use fewer and larger segments (default_segment_number: 2). Applying the wrong direction makes the reported problem worse.

來源與署名

來源:qdrant/skills位於skills/qdrant-scaling提交6a03d0c

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架