Together Volcano

作者 togethercomputer644d38225bbd無授權條款收錄於 2026年10月8日更新於 2026年10月8日

Install and use the Volcano batch scheduler on a Together AI Kubernetes GPU cluster for gang scheduling. Covers installing Volcano, creating queues, submitting all-or-nothing gang-scheduled jobs (vcjobs), and verifying placement. Reach for it when a job on a Together cluster needs its pods scheduled all-at-once or not at all, such as distributed multi-node training, rather than per-pod best-effort scheduling. Pins Volcano v1.15.0.

僅含說明DevOps & Cloud
AI 產生的概覽

在 Together AI Kubernetes GPU 叢集上安裝並使用 Volcano 批次排程器,實現成組排程。

功能
此技能指導在 Together AI Kubernetes GPU 叢集上安裝 Volcano v1.15.0,並說明如何建立佇列、掛載共用儲存以及提交成組排程的 vcjob。它提供佇列、持久卷宣告和全有或全無工作的 YAML 清單,以及驗證排程結果的步驟。它也列出規則與注意事項,例如固定版本號和設定 schedulerName: volcano。
適用情境
當 Together 叢集上的工作需要其 Pod 全部一起排程或全部不排程時使用,例如分散式多節點訓練。它面向已處於 Ready 狀態且已設定 kubectl 的叢集,並與 together-gpu-clusters 技能搭配使用。
執行需求
一個處於 Ready 狀態並已設定 kubectl 的 Together AI Kubernetes GPU 叢集;GPU 節點需暴露 nvidia.com/gpu,且已預裝 NVIDIA 裝置外掛。需要網路存取以取得 Volcano 安裝清單或 Helm 圖表,可選地需要一個用於 ReadWriteMany 儲存的共用 PersistentVolume。僅為說明文件,不附帶指令碼。

Volcano on Together GPU clusters

Volcano is a Kubernetes-native batch scheduler. Its key property is gang scheduling: a group of pods is placed all-at-once or not at all, so distributed training never starts with only some workers running. Use it when a job needs N pods to run together or not at all.

Public cookbook: https://docs.together.ai/docs/volcano-on-gpu-clusters. Pair this skill with the together-gpu-clusters skill, which covers creating the cluster and configuring kubectl.

Preconditions

  • A Together Kubernetes GPU cluster in the Ready state, with kubectl pointed at it (tg beta clusters get-credentials <cluster_id> --set-default-context).
  • GPU nodes expose nvidia.com/gpu; the NVIDIA device plugin is preinstalled on Together clusters. Confirm with:
bash
kubectl get nodes -o custom-columns='NODE:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu'

Install

Pin the version. Do not use master.

bash
kubectl apply -f https://raw.githubusercontent.com/volcano-sh/volcano/v1.15.0/installer/volcano-development.yamlkubectl -n volcano-system wait --for=condition=Available deployment --all --timeout=180s

A default queue is created automatically (kubectl get queue). Helm is an alternative: helm repo add volcano-sh https://volcano-sh.github.io/helm-charts && helm install volcano volcano-sh/volcano -n volcano-system --create-namespace.

Create a queue

A queue caps and weights the resources a set of jobs can use.

yaml
apiVersion: scheduling.volcano.sh/v1beta1kind: Queuemetadata:  name: researchspec:  reclaimable: true          # give idle capacity back to other queues  weight: 1                  # relative share under contention  capability:                # hard ceiling for this queue    nvidia.com/gpu: 8

Attach shared storage (optional)

Together provisions a static PersistentVolume named after the cluster's shared volume. Bind a ReadWriteMany PVC to it so every worker shares datasets and checkpoints, then mount it in the job (below).

yaml
apiVersion: v1kind: PersistentVolumeClaimmetadata:  name: shared-pvcspec:  accessModes: ["ReadWriteMany"]  storageClassName: shared-wekafs   # Together default shared storage class  volumeName: <shared-volume-name>  # the static PV named after your shared volume  resources:    requests:      storage: 100Gi

Run a gang-scheduled job

The gang guarantee comes from minAvailable on a Volcano Job (vcjob). Set schedulerName: volcano and a queue. The volumes/volumeMounts blocks are optional; drop them if the job needs no shared storage.

yaml
apiVersion: batch.volcano.sh/v1alpha1kind: Jobmetadata:  name: gpu-gangspec:  minAvailable: 4            # schedule all 4 pods together or none  schedulerName: volcano  queue: research  policies:    - event: PodEvicted      action: RestartJob     # restart the whole gang if a pod is evicted  tasks:    - replicas: 4      name: worker      template:        spec:          restartPolicy: OnFailure          containers:            - name: worker              image: nvidia/cuda:12.4.0-base-ubuntu22.04              command: ["bash", "-c", "nvidia-smi -L; sleep 300"]              resources:                limits:                  nvidia.com/gpu: 2   # 4 x 2 = 8 GPUs              volumeMounts:                - name: shared                  mountPath: /mnt/shared          volumes:            - name: shared              persistentVolumeClaim:                claimName: shared-pvc

Verify

  • kubectl get vcjob <name> should show STATUS Running and RUNNINGS equal to minAvailable.
  • kubectl get podgroup shows the gang; a Running phase means the whole group is placed.
  • A gang that cannot fit stays Inqueue with zero pods running (all-or-nothing). This is the signal Volcano is working, not a failure. Never "fix" it by lowering minAvailable below the job's real parallelism.

Rules and gotchas

  • Always pin the Volcano version in the install URL.
  • schedulerName: volcano must be on the pod template (on a vcjob it goes under spec). Without it, the default scheduler grabs the pods and there is no gang guarantee.
  • A job whose total request exceeds the queue's capability is rejected. Raise the cap or split the job.
  • To gang-schedule non-vcjob workloads (for example MPIJobs from the preinstalled MPI Operator), set schedulerName: volcano and attach a PodGroup. See https://volcano.sh/en/docs/podgroup/.
  • Volcano and Kueue can coexist on one cluster: Volcano owns pods with schedulerName: volcano; Kueue admits jobs on the default scheduler. Pick one per workload; do not point a single job at both.

Reference

來源與署名

來源:togethercomputer/skills位於skills/together-volcano提交644d382

授權條款: 無授權條款

內容歸原作者所有。SourceWeft 從公開儲存庫中收錄這些內容。

檢舉或申請下架