Utilizing GPUs and EFAs with Dynamic Resource Allocation

Feature State: Alpha

Karpenter can provision nodes for pods that request NVIDIA GPUs and Elastic Fabric Adapters (EFAs) through Dynamic Resource Allocation (DRA). To enable this support, set settings.enableDRA in the Helm chart.

Overview

With device plugins, a pod asks for devices as an integer count of an extended resource, such as nvidia.com/gpu: 1. With DRA, a pod references a ResourceClaim, which describes the devices it needs. A claim can select devices by their attributes using CEL, require several devices to share an attribute (for example, a GPU and an EFA on the same PCIe root), and request a portion of a device’s capacity. A DRA driver runs on each node and publishes the node’s devices and their attributes as ResourceSlices. The scheduler allocates each claim from the devices in those ResourceSlices.

The drivers only publish ResourceSlices once a node is running, but Karpenter has to choose an instance type before the node exists. To close that gap, Karpenter includes metadata for each supported instance type that describes the devices the NVIDIA and EFA DRA drivers publish on that instance type. When a pod with ResourceClaims is pending, Karpenter evaluates its claims against this metadata, along with the ResourceSlices of existing nodes. If no existing node can satisfy the claims, Karpenter launches the cheapest instance type that can. This works from zero: no GPU node has to exist first, and NodePools and EC2NodeClasses need no DRA-specific configuration.

Capabilities

DRA makes requests possible that extended resources can’t express. For example:

Installing the Drivers

Install each driver by following its install guide, then set the Helm values below.

NVIDIA

Install the NVIDIA DRA driver by following the NVIDIA DRA driver install guide or the EKS guide. Neither value below is the chart’s default, so set both in the driver’s Helm chart:

Helm valueDefaultValueWhy
gpuResourcesEnabledOverridefalsetrueRequired for GPU allocation.
resources.computeDomains.enabledtruefalseOptional. Disables ComputeDomains, which Karpenter doesn’t support provisioning for.

To share GPUs between pods, you also set the driver’s consumable shares values. See Requesting a timesliced GPU.

The chart creates the gpu.nvidia.com DeviceClass. The kubelet plugin tolerates the nvidia.com/gpu taint by default. Its default affinity only schedules it on nodes with a GPU presence label, such as nvidia.com/gpu.present=true. The EKS-optimized AL2023 NVIDIA AMI sets that label on GPU nodes, and Karpenter selects it automatically for GPU instance types with the al2023@latest alias. If your nodes use another AMI, check that it sets the label. If it doesn’t, add the label to your DRA NodePool. Otherwise the driver won’t run on the nodes Karpenter launches, and DRA pods stay pending:

apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: gpu-dra
spec:
  template:
    metadata:
      labels:
        nvidia.com/gpu.present: "true"

EFA

Install the EFA DRA driver (DRANET) by following Install the EFA DRA driver. Set this value in the aws-dranet Helm chart:

Helm valueValueWhy
tolerations[{key: nvidia.com/gpu, operator: Exists, effect: NoSchedule}]The DaemonSet doesn’t tolerate the nvidia.com/gpu taint by default, so it won’t run on tainted GPU nodes without this.

The chart creates the efa.networking.k8s.aws DeviceClass. It selects devices whose dra.net/pciDevice attribute is Elastic Fabric Adapter (EFA). DRANET publishes every network interface on a node, but Karpenter’s metadata includes only EFA devices, so request EFAs through this DeviceClass.

Enabling Karpenter Support

Set settings.enableDRA: true in Karpenter’s Helm values to enable Karpenter’s support for every DRA driver it models, currently the NVIDIA GPU and EFA drivers. For how to change the values of an existing installation, see the upgrade guide.

This enables the DRANVIDIAGPU and DRAEFA feature gates and sets settings.ignoreDRARequests to false. The chart always grants Karpenter read access to DeviceClasses, ResourceClaims, and ResourceSlices.

settings.enableDRA assumes both drivers run on your DRA nodes. If a driver isn’t running, Karpenter launches nodes for its claims that can never serve them. If you only run one of the drivers, enable its gate individually instead.

Example: Requesting a specific GPU via DRA

This ResourceClaimTemplate requests one GPU that’s Hopper or newer (CUDA compute capability 9.0 or later) with at least 100Gi of memory:

apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: large-gpu
spec:
  spec:
    devices:
      requests:
      - name: gpu
        exactly:
          deviceClassName: gpu.nvidia.com
          count: 1
          selectors:
          - cel:
              expression: |
                device.attributes["gpu.nvidia.com"].cudaComputeCapability.compareTo(semver("9.0.0")) >= 0 &&
                device.capacity["gpu.nvidia.com"].memory.compareTo(quantity("100Gi")) >= 0                
---
apiVersion: v1
kind: Pod
metadata:
  name: inference
spec:
  tolerations:
  - key: nvidia.com/gpu
    operator: Exists
    effect: NoSchedule
  resourceClaims:
  - name: gpu
    resourceClaimTemplateName: large-gpu
  containers:
  - name: model
    image: public.ecr.aws/amazonlinux/amazonlinux:2023-minimal
    command: ["/bin/sh", "-c", "nvidia-smi && sleep infinity"]
    resources:
      claims:
      - name: gpu

Karpenter only considers instance types with a GPU that matches the selector: here, H200 (p5e, p5en) and Blackwell (p6-b200, p6-b300). It launches the cheapest one the NodePool allows. To request several GPUs, raise count. Karpenter checks that the instance type has enough matching GPUs.

You can only select on attributes that Karpenter knows before launch. See Supported Attributes by Driver.

Example: Requesting a GPU and EFA that share a PCIe Root

Enable DRA support with settings.enableDRA (or both individual gates), and install both drivers. This ResourceClaimTemplate requests one GPU and one EFA, and requires them to share a PCIe root:

apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: gpu-efa-aligned
spec:
  spec:
    devices:
      requests:
      - name: gpu
        exactly:
          deviceClassName: gpu.nvidia.com
          count: 1
      - name: efa
        exactly:
          deviceClassName: efa.networking.k8s.aws
          count: 1
      constraints:
      - requests: ["gpu", "efa"]
        matchAttribute: resource.kubernetes.io/pcieRoot

Reference the claim from a pod as in the previous example. Karpenter only launches instance types where a GPU and an EFA share a PCIe root. For example, each of the eight GPUs in a p5.48xlarge shares a PCIe root with four EFAs.

Example: Requesting a timesliced GPU

With the NVIDIA DRA driver’s consumable shares, several claims can share one GPU, and the driver time-slices between them. First, enable consumable shares by setting these values in the NVIDIA DRA driver’s Helm chart. This example splits each GPU into four shares:

Helm valueValue
featureGates.ConsumableSharestrue
consumableShares4

For the other modes and how the driver applies them, see the consumable capacity guide.

Karpenter can’t read the driver’s configuration, so set the same mode with the karpenter.k8s.aws/nvidia-consumable-capacity annotation on the EC2NodeClass:

apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
metadata:
  name: gpu-dra
  annotations:
    karpenter.k8s.aws/nvidia-consumable-capacity: "4"
spec:
  # ...

The annotation accepts the same values as the driver’s consumableShares setting:

ValueKarpenter models each GPU as
absent or disabledNot shared. Each GPU serves one claim.
positive integer NShared, with N shares. A claim takes one share by default, and its memory request defaults to 0.
memoryShared by memory. A claim with no memory request takes the GPU’s full memory.
unlimitedShared without limit. A claim’s memory request defaults to 0.

Then give each pod its own claim from a template. A plain GPU request takes one share:

apiVersion: resource.k8s.io/v1
kind: ResourceClaimTemplate
metadata:
  name: gpu-share
spec:
  spec:
    devices:
      requests:
      - name: gpu
        exactly:
          deviceClassName: gpu.nvidia.com
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: shared-inference
spec:
  replicas: 8
  selector:
    matchLabels:
      app: shared-inference
  template:
    metadata:
      labels:
        app: shared-inference
    spec:
      tolerations:
      - key: nvidia.com/gpu
        operator: Exists
        effect: NoSchedule
      resourceClaims:
      - name: gpu
        resourceClaimTemplateName: gpu-share
      containers:
      - name: model
        image: public.ecr.aws/amazonlinux/amazonlinux:2023-minimal
        command: ["/bin/sh", "-c", "nvidia-smi && sleep infinity"]
        resources:
          claims:
          - name: gpu

Karpenter packs up to four of these claims onto each GPU, so the eight replicas need only two GPUs.

In memory mode, request memory instead, for example memory: 10Gi. Shares and memory only affect scheduling. They don’t limit how much GPU memory a process actually uses.

Appendix

Reference configuration

EC2NodeClass and NodePool

apiVersion: karpenter.k8s.aws/v1
kind: EC2NodeClass
metadata:
  name: gpu-dra
  # Only for GPU sharing. Must match the driver's consumableShares setting.
  # annotations:
  #   karpenter.k8s.aws/nvidia-consumable-capacity: "4"
spec:
  role: "KarpenterNodeRole-${CLUSTER_NAME}"
  amiSelectorTerms:
  - alias: al2023@latest # Resolves to the NVIDIA variant for GPU instance types
  subnetSelectorTerms:
  - tags:
      karpenter.sh/discovery: "${CLUSTER_NAME}"
  securityGroupSelectorTerms:
  # For EFA, include a security group that allows all traffic to and from itself
  - tags:
      karpenter.sh/discovery: "${CLUSTER_NAME}"
  # Leave networkInterfaces unset. Karpenter attaches every EFA when a pod is allocated dra.net devices.
---
apiVersion: karpenter.sh/v1
kind: NodePool
metadata:
  name: gpu-dra
spec:
  template:
    spec:
      nodeClassRef:
        group: karpenter.k8s.aws
        kind: EC2NodeClass
        name: gpu-dra
      requirements:
      - key: karpenter.k8s.aws/instance-family
        operator: In
        values: ["g6", "g6e", "p5", "p5en", "p6-b200"]
      taints:
      - key: nvidia.com/gpu
        effect: NoSchedule

Pods

A pod that uses DRA devices needs:

  • A spec.resourceClaims entry that references a ResourceClaimTemplate (one claim per pod) or a ResourceClaim (shared by every pod that references it).
  • A resources.claims entry on each container that uses the devices.
  • A toleration for the nvidia.com/gpu taint.

Karpenter treats a pod as a DRA pod only if it has spec.resourceClaims or a container lists resources.claims. A pod that requests the nvidia.com/gpu extended resource isn’t treated as a DRA pod, even when a DeviceClass serves that resource through DRA.

Running DRA drivers alongside device plugins

A DRA driver and a device plugin must never manage the same device on the same node. If both run, they can each hand out the same device, oversubscribing it without any error. If your cluster also runs the NVIDIA device plugin or the EFA device plugin, keep each node on one mechanism by giving DRA and device plugin NodePools separate labels.

Karpenter’s DRA metadata applies to every NodePool that allows a supported instance type, including NodePools whose nodes run the device plugins. If a DRA pod can schedule to both, Karpenter can launch it on a device plugin node, where no DRA driver publishes devices, and the pod stays pending. Constrain DRA pods to your DRA NodePools as well.

  1. Label each NodePool with the mechanism its nodes use, for example example.com/device-manager: dra or example.com/device-manager: device-plugin:

    apiVersion: karpenter.sh/v1
    kind: NodePool
    metadata:
      name: gpu-dra
    spec:
      template:
        metadata:
          labels:
            example.com/device-manager: dra
    
  2. Set a matching nodeSelector in each chart’s Helm values:

    ChartHelm valueValue
    NVIDIA DRA driverkubeletPlugin.nodeSelectorexample.com/device-manager: dra
    EFA DRA driver (DRANET)nodeSelectorexample.com/device-manager: dra
    NVIDIA device pluginnodeSelectorexample.com/device-manager: device-plugin
    EFA device pluginnodeSelectorexample.com/device-manager: device-plugin
  3. Add a nodeSelector on example.com/device-manager: dra to pods that use DRA devices:

    spec:
      nodeSelector:
        example.com/device-manager: dra
    

Supported Instance Types

Karpenter includes DRA metadata for the instance types below. An instance type with no EFA devices listed doesn’t support EFA. Support for a new instance type requires a new Karpenter release.

Instance typeGPUGPUsEFA devices
g4dn.xlargeTesla T41-
g4dn.2xlargeTesla T41-
g4dn.4xlargeTesla T41-
g4dn.8xlargeTesla T411
g4dn.12xlargeTesla T441
g4dn.16xlargeTesla T411
g4dn.metalTesla T481
g5.xlargeNVIDIA A10G1-
g5.2xlargeNVIDIA A10G1-
g5.4xlargeNVIDIA A10G1-
g5.8xlargeNVIDIA A10G11
g5.12xlargeNVIDIA A10G41
g5.16xlargeNVIDIA A10G11
g5.24xlargeNVIDIA A10G41
g5.48xlargeNVIDIA A10G81
g5g.xlargeNVIDIA T4G1-
g5g.2xlargeNVIDIA T4G1-
g5g.4xlargeNVIDIA T4G1-
g5g.8xlargeNVIDIA T4G1-
g5g.16xlargeNVIDIA T4G2-
g5g.metalNVIDIA T4G2-
g6.xlargeNVIDIA L41-
g6.2xlargeNVIDIA L41-
g6.4xlargeNVIDIA L41-
g6.8xlargeNVIDIA L411
g6.12xlargeNVIDIA L441
g6.16xlargeNVIDIA L411
g6.24xlargeNVIDIA L441
g6.48xlargeNVIDIA L481
g6e.xlargeNVIDIA L40S1-
g6e.2xlargeNVIDIA L40S1-
g6e.4xlargeNVIDIA L40S1-
g6e.8xlargeNVIDIA L40S11
g6e.12xlargeNVIDIA L40S41
g6e.16xlargeNVIDIA L40S11
g6e.24xlargeNVIDIA L40S42
g6e.48xlargeNVIDIA L40S84
g6f.largeNVIDIA L4-3Q1-
g6f.xlargeNVIDIA L4-3Q1-
g6f.2xlargeNVIDIA L4-6Q1-
g6f.4xlargeNVIDIA L4-12Q1-
gr6.4xlargeNVIDIA L41-
gr6.8xlargeNVIDIA L411
gr6f.4xlargeNVIDIA L4-12Q1-
p3dn.24xlargeTesla V100-SXM2-32GB81
p4d.24xlargeNVIDIA A100-SXM4-40GB84
p4de.24xlargeNVIDIA A100-SXM4-80GB84
p5.4xlargeNVIDIA H100 80GB HBM311
p5.48xlargeNVIDIA H100 80GB HBM3832
p5e.48xlargeNVIDIA H200832
p5en.48xlargeNVIDIA H200816
p6-b200.48xlargeNVIDIA B20088
p6-b300.48xlargeNVIDIA B300 SXM6 AC816

Supported Attributes by Driver

Karpenter can only evaluate a claim against attributes it knows before the node launches. Some attributes can be used in two ways:

  • Selector: a CEL expression in a claim or DeviceClass that compares the attribute to a value, for example device.attributes["gpu.nvidia.com"].architecture == "Hopper".
  • matchAttribute: a constraint that requires every device allocated for a set of requests to have the same value for the attribute, without naming the value.

Other attributes can only be used with matchAttribute. They describe how devices relate to each other on a node. Their values are resolved at runtime, or are specific to a platform, so Karpenter can tell whether devices will match but not what the value will be.

NVIDIA (gpu.nvidia.com)

AttributeTypeExample valueSelectormatchAttribute
productNamestringNVIDIA H100 80GB HBM3✓✓
architecturestringHopper✓✓
brandstringNvidia✓✓
cudaComputeCapabilityversion9.0.0✓✓
typestringgpu✓✓
driverVersionversion580.82.7✓
cudaDriverVersionversion13.0.0✓
resource.kubernetes.io/pcieRootstringpci0000:10✓
CapacityExample valueSelectorCapacity request
memory81152Mi✓✓ with consumable capacity
shares4✓✓ with an integer consumable capacity mode

EFA (dra.net)

AttributeTypeExample valueSelectormatchAttribute
pciDevicestringElastic Fabric Adapter (EFA)✓✓
pciVendorstringAmazon.com, Inc.✓✓
pciSubsystemstringefa1✓✓
rdmabooltrue✓✓
numaNodeint0✓✓
resource.kubernetes.io/pcieRootstringpci0000:10✓
Last modified October 10, 2026: chore: Release v1.15.0 (#9719) (6d1426c)