Together Kueue

作者 togethercomputer644d38225bbd無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Install and use the Kueue job-queueing controller on a Together AI Kubernetes GPU cluster to gate jobs on quota. Covers installing Kueue, defining ResourceFlavor, ClusterQueue, and LocalQueue quota, submitting jobs to a queue, and watching quota admit or suspend them. Reach for it when a Together cluster's GPU pool must be shared across teams or workloads by quota, admitting jobs only when capacity is free, rather than letting every job start immediately. Pins Kueue v0.18.3 (API kueue.x-k8s.io/v1beta2).

僅含說明DevOps & Cloud
AI 產生的概覽

在 Together AI Kubernetes GPU 叢集上安裝並設定 Kueue,依 GPU 配額排隊執行工作。

功能
提供在 Together AI Kubernetes GPU 叢集上安裝 Kueue 工作排隊控制器的說明,並定義 ResourceFlavor、ClusterQueue 與 LocalQueue 配額物件。說明如何將工作提交至佇列、選擇性掛載共用儲存,以及驗證工作是被准入還是因配額而暫停。也列出常見設定陷阱,例如涵蓋資源與 queue-name 標籤。
適用情境
當 Together GPU 叢集的 GPU 集區需要依配額在團隊或工作負載之間共用,且僅在容量空閒時准入工作時使用。適合希望工作排隊等待而非全部立即啟動的情境。
執行需求
需要處於 Ready 狀態且已設定 kubectl 的 Together Kubernetes GPU 叢集、暴露 nvidia.com/gpu 的 GPU 節點,以及用於套用 Kueue 資訊清單或安裝 Helm chart 的網路存取。此技能不含指令碼,僅為說明文件。

Kueue on Together GPU clusters

Kueue is a Kubernetes-native job queueing controller. It holds jobs in a queue and admits them only when their quota is free, suspending the rest. Use it to share a fixed GPU pool across teams or workloads without overcommitting it. Unlike Volcano, Kueue does not replace the scheduler; it gates when jobs start by toggling their suspend flag.

Public cookbook: https://docs.together.ai/docs/kueue-on-gpu-clusters. Pair this skill with the together-gpu-clusters skill, which covers creating the cluster and configuring kubectl.

Preconditions

  • A Together Kubernetes GPU cluster in the Ready state, with kubectl pointed at it (tg beta clusters get-credentials <cluster_id> --set-default-context).
  • GPU nodes expose nvidia.com/gpu (NVIDIA device plugin preinstalled on Together clusters).

Install

Pin the version. The API group is kueue.x-k8s.io/v1beta2 in v0.18.x; older examples using v1beta1 will not apply.

bash
kubectl apply --server-side -f https://github.com/kubernetes-sigs/kueue/releases/download/v0.18.3/manifests.yamlkubectl -n kueue-system wait --for=condition=Available deployment/kueue-controller-manager --timeout=240s

--server-side is required; the bundled CRDs exceed the client-side apply annotation limit. Helm is an alternative: helm install kueue oci://registry.k8s.io/kueue/charts/kueue --version 0.18.3 --namespace kueue-system --create-namespace.

Define quotas

Three objects: ResourceFlavor (a node class), ClusterQueue (the quota pool), and LocalQueue (the namespaced entry point users submit to).

yaml
apiVersion: kueue.x-k8s.io/v1beta2kind: ResourceFlavormetadata:  name: gpu-flavor---apiVersion: kueue.x-k8s.io/v1beta2kind: ClusterQueuemetadata:  name: gpu-cluster-queuespec:  namespaceSelector: {}            # accept jobs from every namespace  resourceGroups:    - coveredResources: ["cpu", "memory", "nvidia.com/gpu"]  # MUST list every resource jobs request      flavors:        - name: gpu-flavor         # must match the ResourceFlavor          resources:            - name: "cpu"              nominalQuota: 64            - name: "memory"              nominalQuota: 512Gi            - name: "nvidia.com/gpu"              nominalQuota: 8       # admit at most 8 GPUs worth of jobs at once---apiVersion: kueue.x-k8s.io/v1beta2kind: LocalQueuemetadata:  namespace: default  name: gpu-queuespec:  clusterQueue: gpu-cluster-queue

Attach shared storage (optional)

Together provisions a static PersistentVolume named after the cluster's shared volume. Bind a ReadWriteMany PVC to it so queued jobs share datasets and checkpoints, then mount it in the job (below). Storage is not a quota resource, so it does not affect admission.

yaml
apiVersion: v1kind: PersistentVolumeClaimmetadata:  name: shared-pvcspec:  accessModes: ["ReadWriteMany"]  storageClassName: shared-wekafs   # Together default shared storage class  volumeName: <shared-volume-name>  # the static PV named after your shared volume  resources:    requests:      storage: 100Gi

Submit a job to a queue

Add the queue-name label and start the job suspended. Kueue flips suspend to false on admission. The volumes/volumeMounts blocks are optional; drop them if the job needs no shared storage.

yaml
apiVersion: batch/v1kind: Jobmetadata:  name: gpu-job  namespace: default  labels:    kueue.x-k8s.io/queue-name: gpu-queue   # route to the LocalQueuespec:  parallelism: 1  completions: 1  suspend: true                            # Kueue unsuspends on admission  template:    spec:      restartPolicy: Never      containers:        - name: worker          image: nvidia/cuda:12.4.0-base-ubuntu22.04          command: ["bash", "-c", "nvidia-smi -L; sleep 180"]          resources:            requests:              cpu: "2"              memory: "8Gi"            limits:              nvidia.com/gpu: 6          volumeMounts:            - name: shared              mountPath: /mnt/shared      volumes:        - name: shared          persistentVolumeClaim:            claimName: shared-pvc

Verify

  • kubectl get workloads shows ADMITTED True once the job fits.
  • An over-quota job stays suspend: true with no pod. This is correct queueing, not a failure. Read the reason:
bash
WORKLOAD=$(kubectl get workloads -o name | grep <job-name>)kubectl get "$WORKLOAD" -o jsonpath='{.status.conditions[?(@.type=="QuotaReserved")].message}'
  • Freeing quota (a running job finishes or is deleted) auto-admits the next queued job. Do not resubmit.

Rules and gotchas

  • Pin the version, and use API kueue.x-k8s.io/v1beta2 for v0.18.x.
  • The ClusterQueue's coveredResources MUST include every resource a job requests. A GPU job that also requests CPU and memory is never admitted if the ClusterQueue only covers nvidia.com/gpu. This is the most common failure.
  • flavors[].name must exactly match a ResourceFlavor name, or the ClusterQueue admits nothing.
  • The kueue.x-k8s.io/queue-name label must name a LocalQueue in the job's own namespace. A missing or wrong label means the job runs immediately, bypassing quota.
  • Volcano and Kueue can coexist: Kueue gates jobs on the default scheduler; Volcano owns pods with schedulerName: volcano. Do not point one job at both.

Reference

來源與署名

來源:togethercomputer/skills位於skills/together-kueue提交644d382

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架