Kueue on Together GPU clusters
Kueue is a Kubernetes-native job queueing controller. It holds jobs in a queue and admits them only when their quota is free, suspending the rest. Use it to share a fixed GPU pool across teams or workloads without overcommitting it. Unlike Volcano, Kueue does not replace the scheduler; it gates when jobs start by toggling their suspend flag.
Public cookbook: https://docs.together.ai/docs/kueue-on-gpu-clusters. Pair this skill with the together-gpu-clusters skill, which covers creating the cluster and configuring kubectl.
Preconditions
- A Together Kubernetes GPU cluster in the
Readystate, withkubectlpointed at it (tg beta clusters get-credentials <cluster_id> --set-default-context). - GPU nodes expose
nvidia.com/gpu(NVIDIA device plugin preinstalled on Together clusters).
Install
Pin the version. The API group is kueue.x-k8s.io/v1beta2 in v0.18.x; older examples using v1beta1 will not apply.
--server-side is required; the bundled CRDs exceed the client-side apply annotation limit. Helm is an alternative: helm install kueue oci://registry.k8s.io/kueue/charts/kueue --version 0.18.3 --namespace kueue-system --create-namespace.
Define quotas
Three objects: ResourceFlavor (a node class), ClusterQueue (the quota pool), and LocalQueue (the namespaced entry point users submit to).
Attach shared storage (optional)
Together provisions a static PersistentVolume named after the cluster's shared volume. Bind a ReadWriteMany PVC to it so queued jobs share datasets and checkpoints, then mount it in the job (below). Storage is not a quota resource, so it does not affect admission.
Submit a job to a queue
Add the queue-name label and start the job suspended. Kueue flips suspend to false on admission. The volumes/volumeMounts blocks are optional; drop them if the job needs no shared storage.
Verify
kubectl get workloadsshowsADMITTED Trueonce the job fits.- An over-quota job stays
suspend: truewith no pod. This is correct queueing, not a failure. Read the reason:
- Freeing quota (a running job finishes or is deleted) auto-admits the next queued job. Do not resubmit.
Rules and gotchas
- Pin the version, and use API
kueue.x-k8s.io/v1beta2for v0.18.x. - The ClusterQueue's
coveredResourcesMUST include every resource a job requests. A GPU job that also requests CPU and memory is never admitted if the ClusterQueue only coversnvidia.com/gpu. This is the most common failure. flavors[].namemust exactly match aResourceFlavorname, or the ClusterQueue admits nothing.- The
kueue.x-k8s.io/queue-namelabel must name aLocalQueuein the job's own namespace. A missing or wrong label means the job runs immediately, bypassing quota. - Volcano and Kueue can coexist: Kueue gates jobs on the default scheduler; Volcano owns pods with
schedulerName: volcano. Do not point one job at both.
Reference
- Kueue docs: https://kueue.sigs.k8s.io/docs/
- Installation: https://kueue.sigs.k8s.io/docs/installation/
- Concepts (ResourceFlavor, ClusterQueue, LocalQueue, Workload): https://kueue.sigs.k8s.io/docs/concepts/
- Running jobs (Job, JobSet, RayJob, MPIJob): https://kueue.sigs.k8s.io/docs/tasks/run/jobs/


