Volcano on Together GPU clusters
Volcano is a Kubernetes-native batch scheduler. Its key property is gang scheduling: a group of pods is placed all-at-once or not at all, so distributed training never starts with only some workers running. Use it when a job needs N pods to run together or not at all.
Public cookbook: https://docs.together.ai/docs/volcano-on-gpu-clusters. Pair this skill with the together-gpu-clusters skill, which covers creating the cluster and configuring kubectl.
Preconditions
- A Together Kubernetes GPU cluster in the
Readystate, withkubectlpointed at it (tg beta clusters get-credentials <cluster_id> --set-default-context). - GPU nodes expose
nvidia.com/gpu; the NVIDIA device plugin is preinstalled on Together clusters. Confirm with:
Install
Pin the version. Do not use master.
A default queue is created automatically (kubectl get queue). Helm is an alternative: helm repo add volcano-sh https://volcano-sh.github.io/helm-charts && helm install volcano volcano-sh/volcano -n volcano-system --create-namespace.
Create a queue
A queue caps and weights the resources a set of jobs can use.
Attach shared storage (optional)
Together provisions a static PersistentVolume named after the cluster's shared volume. Bind a ReadWriteMany PVC to it so every worker shares datasets and checkpoints, then mount it in the job (below).
Run a gang-scheduled job
The gang guarantee comes from minAvailable on a Volcano Job (vcjob). Set schedulerName: volcano and a queue. The volumes/volumeMounts blocks are optional; drop them if the job needs no shared storage.
Verify
kubectl get vcjob <name>should showSTATUS RunningandRUNNINGSequal tominAvailable.kubectl get podgroupshows the gang; aRunningphase means the whole group is placed.- A gang that cannot fit stays
Inqueuewith zero pods running (all-or-nothing). This is the signal Volcano is working, not a failure. Never "fix" it by loweringminAvailablebelow the job's real parallelism.
Rules and gotchas
- Always pin the Volcano version in the install URL.
schedulerName: volcanomust be on the pod template (on avcjobit goes underspec). Without it, the default scheduler grabs the pods and there is no gang guarantee.- A job whose total request exceeds the queue's
capabilityis rejected. Raise the cap or split the job. - To gang-schedule non-vcjob workloads (for example MPIJobs from the preinstalled MPI Operator), set
schedulerName: volcanoand attach aPodGroup. See https://volcano.sh/en/docs/podgroup/. - Volcano and Kueue can coexist on one cluster: Volcano owns pods with
schedulerName: volcano; Kueue admits jobs on the default scheduler. Pick one per workload; do not point a single job at both.
Reference
- Volcano docs: https://volcano.sh/en/docs/
- Installation: https://volcano.sh/en/docs/installation/
- Gang scheduling: https://volcano.sh/en/docs/gang_scheduling/
- Queue: https://volcano.sh/en/docs/queue/
- vcjob: https://volcano.sh/en/docs/vcjob/


