GPU Workload Placement
Nova can place GPU-accelerated workloads across one or more Kubernetes clusters, allowing workloads to run where suitable GPU resources are available.
When to Use This
Use this pattern when:
- GPU capacity is distributed across multiple clusters
- Different clusters provide different GPU types or configurations
- AI/ML workloads need to run where GPU resources are available
- GPU workloads may need to move as capacity changes
- Related application components should be co-located with GPU-backed services
How Nova Helps
Nova evaluates available-resource placement policies and selects a workload cluster with sufficient GPU resources.
GPU-aware placement works with standard Kubernetes resource requests, including:
nvidia.com/gpuamd.com/gpunvidia.com/mig-*(NVIDIA MIG mixed strategy)
Alternatively, GPU-aware placement works with Dynamic Resource Allocation (DRA) for NVIDIA GPUs. DRA is a K8s feature for requesting, configuring, and sharing specialized devices like GPUs via allocating ResouceClaims to matching available ResourceSlices. DRA is available starting with Kubernetes 1.34. NVIDIA's DRA Driver for GPUs works with its GPU Operator to support a number of interesting use cases, including single-node topology-aware placement via the match-attribute constraint. Examples of using DRA include the following, which are available in the elotl/try-nova repo, in the examples/dra-gpu directory.
dra-gpu-test.count1.yaml: Pod with ResourceClaimTemplate requesting an NVIDIA GPU
dra-gpu-test.count2.yaml: Pod with ResourceClaimTemplate requesting 2 NVIDIA GPUs
dra-gpu-test.count2.constraint.yaml: Pod with ResourceClaimTemplate requesting 2 NVIDIA GPUs with a match-attribute constraint
dra-gpu-test.t4.yaml: Pod with ResourceClaimTemplate requesting an NVIDIA T4 GPU
dra-gpu-test.firstavail.yaml: Pod with ResourceClaimTemplate requesting the first available of 2 alternate NVIDIA GPU requests
dra-gpu-test.cap10g.yaml: Pod with ResourceClaimTemplate requesting an NVIDIA GPU with memory capacity greater than 10Gi
dra-gpu-test.mig1.yaml: Pod with ResourceClaimTemplate requesting an NVIDIA MIG GPU
dra-gpu-test.2podsshare.yaml: Two pods sharing a ResourceClaim for an NVIDIA GPU
dra-gpu-test.semver.yaml: Pod with ResourceClaimTemplate requesting an NVIDIA GPU with a driver version greater than 550.127.8
Note that Nova also respects workload constraints such as nodeSelector, which can be used to target specific GPU characteristics.
Multi-Host Topology-Aware Placement (Multi-Node NVLink)
For rack-scale systems with Multi-Node NVLink (MNNVL) -- GH200 NVL32, GB200/GB300 NVL72, and Vera Rubin NVL72 --
NVIDIA's DRA driver exposes fast GPU interconnect across nodes via a ComputeDomain custom resource. See NVIDIA's
Enabling Multi-Node NVLink on Kubernetes for GB200 and Beyond
for the full mechanism. In short:
- A
ComputeDomainobject declares achannel.resourceClaimTemplate.name; its controller auto-creates thatResourceClaimTemplate, which the GPU workload's pods reference (viaspec.resourceClaims) to get an isolated IMEX communication channel. - The workload declares a required pod affinity term with
topologyKey: nvidia.com/gpu.clique, so all its pods land on nodes that share the same NVLink partition (thenvidia.com/gpu.cliquenode label, set by the NVIDIA GPU Operator).
To have Nova place such a workload, put the ComputeDomain and the GPU workload in the same SchedulePolicy
schedule group (i.e. give them the same GroupBy label value). Nova will:
- Create the
ComputeDomainbefore the workload on the target cluster, since the workload'sResourceClaimTemplatereference depends on it. - When doing resource-aware placement, only pick a cluster that has enough distinct nodes sharing one
nvidia.com/gpu.cliquevalue to host every replica of the workload, one pod per node, honoring each replica's GPU request (whether a staticnvidia.com/gpulimit or a DRA GPU claim) per node, along with any other resource requirements and constraints.
Current limitations
- NVIDIA's DRA driver currently supports only one pod per node per
ComputeDomain; Nova's placement check enforces this by construction (it never places two replicas of a clique-affine workload on the same node). - NVIDIA's DRA driver currently supports at most one
ComputeDomainper node. Nova does not track this across separate schedule groups or reconciles (only current per-cluster available capacity), so this constraint is best-effort from Nova's side, consistent with Nova's handling of other DRA constraints it cannot fully verify ahead of time -- if the driver rejects a claim due to this, Nova's normal reschedule-on-placement-failure handling applies.
Considerations
GPU placement depends on the workload clusters being prepared to run GPU workloads. This includes:
- GPU-enabled nodes
- Appropriate GPU drivers
- GPU operators, such as the NVIDIA GPU Operator, where applicable
- Accurate resource requests in workload manifests
Available GPU resources can be viewed through the Nova cluster inventory, for example by using:
kubectl --context=nova get clusters -o wide
This will display GPU, CPU and Memory resources:
NAME K8S-VERSION K8S-CLUSTER NOVA-CREATED PROVIDER REGION ZONE AVAIL-CPU AVAIL-MEM AVAIL-NVIDIAGPU AVAIL-AMDGPU READY IDLE STANDBY
wlc-1 1.35 worklc-12232 false azure eastus eastus-2 16019m 102957284Ki 3 0 True False False
wlc-2 1.35 worklc-30337 false azure eastus eastus-2 12516m 91274704Ki 3 0 True False False
NOTE: In some [rare] cases, workload clusters require that Kubernetes objects which use NVIDIA GPUs be configured with runtimeClassName set to nvidia.
When installing the Nova agent on such clusters, include the option "--add-nvidia-runtime-class" [new in Nova 1.5.6] to have the Nova agent add
runtimeClassName:nvidia to objects using NVIDIA GPUs that Nova schedules on those clusters. This option currently applies to native Kubernetes
objects using NVIDIA GPUs via NVIDIA plug-in specification, i.e., via nvidia.com/gpu or nvidia.com/mig* set to a value > 0.