<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Karpenter – Tasks</title><link>/v1.15/tasks/</link><description>Recent content in Tasks on Karpenter</description><generator>Hugo -- gohugo.io</generator><language>en</language><atom:link href="/v1.15/tasks/index.xml" rel="self" type="application/rss+xml"/><item><title>V1.15: Using Nitro Enclaves</title><link>/v1.15/tasks/nitro-enclaves/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/v1.15/tasks/nitro-enclaves/</guid><description>
&lt;p>This guide shows you how to configure Karpenter to provision nodes with AWS Nitro Enclaves enabled.&lt;/p>
&lt;p>Nitro Enclaves provide an isolated compute environment to protect and process highly sensitive data such as personally identifiable information (PII), healthcare, financial, and intellectual property data.&lt;/p>
&lt;h2 id="what-are-nitro-enclaves">What are Nitro Enclaves?&lt;/h2>
&lt;p>AWS Nitro Enclaves are isolated compute environments built on the AWS Nitro System that provide:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Isolated compute environment&lt;/strong>: Enclaves run in a separate memory space from the parent instance with no persistent storage, interactive access, or external networking&lt;/li>
&lt;li>&lt;strong>Cryptographic attestation&lt;/strong>: You can verify the enclave&amp;rsquo;s identity and integrity before trusting it with sensitive data&lt;/li>
&lt;li>&lt;strong>Reduced attack surface&lt;/strong>: No SSH access, no persistent storage, and limited communication to only the parent instance&lt;/li>
&lt;li>&lt;strong>CPU and memory isolation&lt;/strong>: Dedicated vCPUs and memory allocated from the parent instance&lt;/li>
&lt;/ul>
&lt;p>Common use cases include:&lt;/p>
&lt;ul>
&lt;li>Processing sensitive financial data and transactions&lt;/li>
&lt;li>Running confidential computing workloads&lt;/li>
&lt;li>Implementing secure key management and cryptographic operations&lt;/li>
&lt;li>Supporting compliance efforts for workloads with data-isolation requirements, such as PCI DSS or HIPAA-regulated workloads&lt;/li>
&lt;/ul>
&lt;p>For more information, see the &lt;a href="https://docs.aws.amazon.com/enclaves/">AWS Nitro Enclaves documentation&lt;/a>.&lt;/p>
&lt;p>Karpenter enables Nitro Enclaves in EC2 launch templates, excludes instance types whose EC2 &lt;code>NitroEnclavesSupport&lt;/code> value is not &lt;code>supported&lt;/code>, and models &lt;code>NodeOverlay&lt;/code> capacity. You must configure the node, install the Kubernetes device plugin, and keep the &lt;code>NodeOverlay&lt;/code> capacity consistent with the resources advertised by the device plugin and kubelet.&lt;/p>
&lt;div class="alert alert-warning" role="alert">
&lt;h4 class="alert-heading">Warning&lt;/h4>
Karpenter does not inspect your AMI, generate Nitro Enclaves bootstrap configuration, or subtract enclave CPU from the instance&amp;rsquo;s standard allocatable &lt;code>cpu&lt;/code>. Hugepage resources declared in a &lt;code>NodeOverlay&lt;/code> are subtracted from standard allocatable &lt;code>memory&lt;/code> in Karpenter&amp;rsquo;s scheduling simulation. The actual allocation happens on the node; a &lt;code>NodeOverlay&lt;/code> does not configure the node or modify the capacity reported by kubelet.
&lt;/div>
&lt;div class="alert alert-warning" role="alert">
&lt;h4 class="alert-heading">Warning&lt;/h4>
All pods and containers on the parent node can communicate with enclaves attached to that node. A taint controls scheduling; it is not a security boundary. Use a dedicated NodePool with a &lt;code>NoSchedule&lt;/code> taint to reduce which workloads can run on these nodes, and add the matching toleration only to the device-plugin DaemonSet and enclave workloads.
&lt;/div>
&lt;h2 id="prerequisites">Prerequisites&lt;/h2>
&lt;p>Before continuing:&lt;/p>
&lt;ol>
&lt;li>Enable the alpha &lt;code>NodeOverlay&lt;/code> feature gate.&lt;/li>
&lt;li>Prepare an AMI or user data that &lt;a href="https://docs.aws.amazon.com/enclaves/latest/user/kubernetes.html">configures Nitro Enclaves on the node&lt;/a>, including the allocator and the 1 GiB hugepages used by this example.&lt;/li>
&lt;li>Install the &lt;a href="https://github.com/aws/aws-nitro-enclaves-k8s-device-plugin">AWS Nitro Enclaves Kubernetes device plugin&lt;/a> on nodes labeled &lt;code>aws-nitro-enclaves-k8s-dp: enabled&lt;/code>. Configure the device-plugin DaemonSet as shown below.&lt;/li>
&lt;li>Decide how many enclave slots and CPUs and how much memory each node will advertise.&lt;/li>
&lt;/ol>
&lt;h2 id="enable-nodeoverlays">Enable NodeOverlays&lt;/h2>
&lt;p>When installing Karpenter with Helm, enable NodeOverlays through the chart settings:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">settings&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">featureGates&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">nodeOverlay&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">true&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="enable-nitro-enclaves-on-the-ec2nodeclass">Enable Nitro Enclaves on the EC2NodeClass&lt;/h2>
&lt;p>Set &lt;code>spec.enclaveOptions.enabled&lt;/code> to &lt;code>true&lt;/code>:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.k8s.aws/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">EC2NodeClass&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">enclave-enabled&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">enclaveOptions&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">enabled&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">true&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">amiFamily&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">AL2023&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">amiSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">alias&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">al2023@latest&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">subnetSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">tags&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">karpenter.sh/discovery&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;${CLUSTER_NAME}&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">securityGroupSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">tags&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">karpenter.sh/discovery&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;${CLUSTER_NAME}&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">role&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;KarpenterNodeRole-${CLUSTER_NAME}&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Switching &lt;code>enabled&lt;/code> between &lt;code>true&lt;/code> and &lt;code>false&lt;/code> changes the generated launch template and participates in EC2NodeClass drift. When &lt;code>enabled&lt;/code> is &lt;code>true&lt;/code>, Karpenter also filters the EC2 instance-type candidates to those whose &lt;code>NitroEnclavesSupport&lt;/code> value is &lt;code>supported&lt;/code>.&lt;/p>
&lt;p>The example uses an EKS-optimized AMI only to keep the EC2NodeClass concise. You must still provide the node-level Nitro Enclaves configuration, either by baking it into an AMI or supplying appropriate user data. It uses &lt;code>@latest&lt;/code> for brevity; follow the &lt;a href="/v1.15/tasks/managing-amis/#pinning-amis">AMI pinning guidance&lt;/a> for production.&lt;/p>
&lt;p>Omitting &lt;code>enclaveOptions&lt;/code> or setting &lt;code>enabled&lt;/code> to &lt;code>false&lt;/code> disables Nitro Enclaves. When &lt;code>enclaveOptions&lt;/code> is specified, &lt;code>enabled&lt;/code> is required.&lt;/p>
&lt;div class="alert alert-warning" role="alert">
&lt;h4 class="alert-heading">Warning&lt;/h4>
&lt;a href="https://docs.aws.amazon.com/enclaves/latest/user/nitro-enclave.html">Nitro Enclaves are not supported in AWS Local Zones, AWS Wavelength Zones, or AWS Outposts&lt;/a>. Configure &lt;code>subnetSelectorTerms&lt;/code> to resolve only subnets in standard Availability Zones. Karpenter filters instance types based on EC2&amp;rsquo;s &lt;code>NitroEnclavesSupport&lt;/code> value, but does not filter these unsupported locations. EC2 can reject launch attempts that target them. If only unsupported locations match, the workload remains pending.
&lt;/div>
&lt;h2 id="label-the-nodepool">Label the NodePool&lt;/h2>
&lt;p>Use the default device-plugin label on the nodes that should run the plugin:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.sh/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">NodePool&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">enclave-pool&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">template&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">labels&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">aws-nitro-enclaves-k8s-dp&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">enabled&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">taints&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">aws-nitro-enclaves-k8s-dp&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">value&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">enabled&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">effect&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">NoSchedule&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">nodeClassRef&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">group&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.k8s.aws&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">EC2NodeClass&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">enclave-enabled&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>You do not need to maintain a static list of enclave-capable instance families. The EC2NodeClass compatibility filter removes unsupported instance types. You can add normal NodePool requirements when you need tighter cost, architecture, or size constraints.&lt;/p>
&lt;h2 id="configure-the-device-plugin">Configure the Device Plugin&lt;/h2>
&lt;p>The upstream device-plugin DaemonSet selects nodes with the &lt;code>aws-nitro-enclaves-k8s-dp: enabled&lt;/code> label. Enable enclave CPU advertisement and add a toleration for the NodePool taint:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>kubectl patch daemonset aws-nitro-enclaves-k8s-daemonset &lt;span style="color:#4e9a06">\
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06">&lt;/span> --namespace kube-system &lt;span style="color:#4e9a06">\
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06">&lt;/span> --type strategic &lt;span style="color:#4e9a06">\
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06">&lt;/span> --patch &lt;span style="color:#4e9a06">&amp;#39;{
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;spec&amp;#34;: {
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;template&amp;#34;: {
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;spec&amp;#34;: {
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;tolerations&amp;#34;: [{
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;key&amp;#34;: &amp;#34;aws-nitro-enclaves-k8s-dp&amp;#34;,
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;operator&amp;#34;: &amp;#34;Equal&amp;#34;,
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;value&amp;#34;: &amp;#34;enabled&amp;#34;,
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;effect&amp;#34;: &amp;#34;NoSchedule&amp;#34;
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> }],
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;containers&amp;#34;: [{
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;name&amp;#34;: &amp;#34;aws-nitro-enclaves-k8s-dp&amp;#34;,
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;env&amp;#34;: [{
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;name&amp;#34;: &amp;#34;ENCLAVE_CPU_ADVERTISEMENT&amp;#34;,
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> &amp;#34;value&amp;#34;: &amp;#34;true&amp;#34;
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> }]
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> }]
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> }
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> }
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> }
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> }&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Without the toleration, the DaemonSet cannot run on the tainted nodes. Without CPU advertisement, the device plugin does not publish the &lt;code>aws.ec2.nitro/nitro_enclaves_cpus&lt;/code> resource used by the example.&lt;/p>
&lt;h2 id="model-enclave-resources-with-a-nodeoverlay">Model Enclave Resources with a NodeOverlay&lt;/h2>
&lt;p>Create a &lt;code>NodeOverlay&lt;/code> that matches the device-plugin label. Include numeric CPU and memory requirements so the overlay is only considered for instances larger than the enclave reservation.&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.sh/v1alpha1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">NodeOverlay&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">nitro-enclaves&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">requirements&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">aws-nitro-enclaves-k8s-dp&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">operator&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">In&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">values&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000;font-weight:bold">[&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;enabled&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">]&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.k8s.aws/instance-cpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">operator&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">Gt&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">values&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000;font-weight:bold">[&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;2&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">]&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.k8s.aws/instance-memory&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">operator&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">Gt&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">values&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000;font-weight:bold">[&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;4096&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">]&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">capacity&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">aws.ec2.nitro/nitro_enclaves&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;4&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">aws.ec2.nitro/nitro_enclaves_cpus&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;2&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">hugepages-1Gi&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;4Gi&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The &lt;code>instance-memory&lt;/code> requirement is expressed in MiB. Adjust the capacity values to match the device plugin, allocator, and hugepage configuration on your nodes.&lt;/p>
&lt;p>The overlay must describe what the device plugin and kubelet will actually advertise:&lt;/p>
&lt;ul>
&lt;li>&lt;code>aws.ec2.nitro/nitro_enclaves&lt;/code> represents the number of enclaves&lt;/li>
&lt;li>&lt;code>aws.ec2.nitro/nitro_enclaves_cpus&lt;/code> represents enclave vCPUs&lt;/li>
&lt;li>&lt;code>hugepages-1Gi&lt;/code> represents memory reserved as 1 GiB hugepages&lt;/li>
&lt;/ul>
&lt;h2 id="schedule-an-enclave-workload">Schedule an Enclave Workload&lt;/h2>
&lt;p>Request the same resources from the workload:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">Pod&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">enclave-workload&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">tolerations&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">aws-nitro-enclaves-k8s-dp&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">operator&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">Equal&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">value&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">enabled&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">effect&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">NoSchedule&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">containers&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">workload&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">image&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">public.ecr.aws/your-repository/your-image:latest&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">volumeMounts&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">hugepages&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">mountPath&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">/dev/hugepages&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">resources&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">requests&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">cpu&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">250m&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">aws.ec2.nitro/nitro_enclaves&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;1&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">aws.ec2.nitro/nitro_enclaves_cpus&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;2&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">hugepages-1Gi&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;4Gi&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">limits&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">aws.ec2.nitro/nitro_enclaves&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;1&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">aws.ec2.nitro/nitro_enclaves_cpus&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;2&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">hugepages-1Gi&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;4Gi&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">volumes&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">hugepages&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">emptyDir&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">medium&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">HugePages-1Gi&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Replace the placeholder image with an enclave application that starts an enclave. The workload does not need to select a NodePool or instance family. Karpenter matches these resource requests against the capacity declared by the &lt;code>NodeOverlay&lt;/code>.&lt;/p>
&lt;h2 id="validate-the-configuration">Validate the Configuration&lt;/h2>
&lt;p>Confirm that the &lt;code>NodeOverlay&lt;/code> is ready and passed validation:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>kubectl get nodeoverlay nitro-enclaves &lt;span style="color:#4e9a06">\
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06">&lt;/span> -o &lt;span style="color:#000">jsonpath&lt;/span>&lt;span style="color:#ce5c00;font-weight:bold">=&lt;/span>&lt;span style="color:#4e9a06">&amp;#39;{range .status.conditions[*]}{.type}={.status}{&amp;#34;\n&amp;#34;}{end}&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Verify that the output includes &lt;code>Ready=True&lt;/code> and &lt;code>ValidationSucceeded=True&lt;/code>.&lt;/p>
&lt;p>Confirm that EC2 launched the instance with enclaves enabled:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#000">INSTANCE_ID&lt;/span>&lt;span style="color:#ce5c00;font-weight:bold">=&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;i-0123456789abcdef0&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>aws ec2 describe-instances &lt;span style="color:#4e9a06">\
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06">&lt;/span> --instance-ids &lt;span style="color:#4e9a06">&amp;#34;&lt;/span>&lt;span style="color:#4e9a06">${&lt;/span>&lt;span style="color:#000">INSTANCE_ID&lt;/span>&lt;span style="color:#4e9a06">}&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;&lt;/span> &lt;span style="color:#4e9a06">\
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06">&lt;/span> --query &lt;span style="color:#4e9a06">&amp;#39;Reservations[].Instances[].EnclaveOptions&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Verify that &lt;code>Enabled&lt;/code> is &lt;code>true&lt;/code>.&lt;/p>
&lt;p>On the node, use the Nitro CLI to inspect running enclaves:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>nitro-cli describe-enclaves
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Confirm that the real node capacity matches the overlay:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#000">NODE_NAME&lt;/span>&lt;span style="color:#ce5c00;font-weight:bold">=&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;ip-10-0-1-2.us-west-2.compute.internal&amp;#34;&lt;/span>
&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>kubectl get node &lt;span style="color:#4e9a06">&amp;#34;&lt;/span>&lt;span style="color:#4e9a06">${&lt;/span>&lt;span style="color:#000">NODE_NAME&lt;/span>&lt;span style="color:#4e9a06">}&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;&lt;/span> -o json &lt;span style="color:#000;font-weight:bold">|&lt;/span> jq &lt;span style="color:#4e9a06">&amp;#39;.status.capacity | {
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> enclaves: .&amp;#34;aws.ec2.nitro/nitro_enclaves&amp;#34;,
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> enclaveCPUs: .&amp;#34;aws.ec2.nitro/nitro_enclaves_cpus&amp;#34;,
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06"> hugepages: .&amp;#34;hugepages-1Gi&amp;#34;
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#4e9a06">}&amp;#39;&lt;/span>
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>If these values are absent or differ from the &lt;code>NodeOverlay&lt;/code>, fix the node image, allocator, hugepage configuration, or device-plugin deployment. Changing only the overlay can make Karpenter provision a node for a pod that the Kubernetes scheduler cannot place.&lt;/p>
&lt;p>Confirm that the enclave workload schedules successfully:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-bash" data-lang="bash">&lt;span style="display:flex;">&lt;span>kubectl get pod enclave-workload -o wide
&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h2 id="capacity-considerations">Capacity Considerations&lt;/h2>
&lt;p>The enclave allocator takes CPU and memory from the parent instance. The &lt;code>hugepages-1Gi&lt;/code> capacity in this example reduces the allocatable memory Karpenter computes, but &lt;code>aws.ec2.nitro/nitro_enclaves_cpus&lt;/code> does not reduce the allocatable CPU Karpenter computes.&lt;/p>
&lt;p>After launch, compare standard &lt;code>cpu&lt;/code> in the NodeClaim&amp;rsquo;s &lt;code>.status.allocatable&lt;/code> with the Node&amp;rsquo;s &lt;code>.status.allocatable&lt;/code>. Keep ordinary workloads off the enclave NodePool. The sum of standard &lt;code>cpu&lt;/code> requests for every pod allowed onto the NodePool must fit the Node&amp;rsquo;s actual allocatable CPU. Separately, the sum of &lt;code>aws.ec2.nitro/nitro_enclaves_cpus&lt;/code> requests must fit the capacity declared by the &lt;code>NodeOverlay&lt;/code>. Requesting enclave CPUs does not reserve standard CPU, and a &lt;code>NoSchedule&lt;/code> taint does not change Karpenter&amp;rsquo;s CPU calculation.&lt;/p>
&lt;h2 id="additional-resources">Additional Resources&lt;/h2>
&lt;ul>
&lt;li>&lt;a href="https://docs.aws.amazon.com/enclaves/">AWS Nitro Enclaves documentation&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://github.com/aws/aws-nitro-enclaves-sdk-c">Nitro Enclaves SDK&lt;/a>&lt;/li>
&lt;li>&lt;a href="https://github.com/aws/aws-nitro-enclaves-k8s-device-plugin">AWS Nitro Enclaves Kubernetes device plugin&lt;/a>&lt;/li>
&lt;li>&lt;a href="/v1.15/concepts/nodeclasses/#specenclaveoptions">EC2NodeClass reference&lt;/a>&lt;/li>
&lt;li>&lt;a href="/v1.15/concepts/nodeoverlays/">NodeOverlay reference&lt;/a>&lt;/li>
&lt;li>&lt;a href="/v1.15/concepts/nodepools/">NodePool reference&lt;/a>&lt;/li>
&lt;li>&lt;a href="/v1.15/concepts/disruption/">Karpenter disruption documentation&lt;/a>&lt;/li>
&lt;/ul>
&lt;h2 id="follow-up">Follow-up&lt;/h2>
&lt;p>If you have questions or issues with Nitro Enclaves in Karpenter, feel free to:&lt;/p>
&lt;ul>
&lt;li>Open an issue on &lt;a href="https://github.com/aws/karpenter-provider-aws/issues/new/choose">GitHub&lt;/a>&lt;/li>
&lt;li>Ask in the &lt;a href="https://kubernetes.slack.com/archives/C02SFFZSA2K">Karpenter Slack channel&lt;/a>&lt;/li>
&lt;li>Check the &lt;a href="/v1.15/troubleshooting/">Troubleshooting Guide&lt;/a>&lt;/li>
&lt;/ul></description></item><item><title>V1.15: Managing AMIs</title><link>/v1.15/tasks/managing-amis/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/v1.15/tasks/managing-amis/</guid><description>
&lt;div class="alert alert-warning" role="alert">
&lt;h4 class="alert-heading">Important&lt;/h4>
&lt;p>Karpenter &lt;strong>heavily recommends against&lt;/strong> opting-in to use an &lt;code>amiSelectorTerm&lt;/code> with &lt;code>@latest&lt;/code> unless you are doing this in a pre-production environment or are willing to accept the risk that a faulty AMI may cause downtime in your production clusters. In general, if using a publicly released version of a well-known AMI type (like AL2, AL2023, or Bottlerocket), we recommend that you pin to a version of that AMI and deploy newer versions of that AMI type in a staged approach when newer patch versions are available.&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">amiSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">alias&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">al2023@v20240807&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>More details are described in &lt;a href="#controlling-ami-replacement">Controlling AMI Replacement&lt;/a> below.&lt;/p>
&lt;/div>
&lt;p>Understanding how Karpenter assigns AMIs to nodes can help ensure that your workloads will run successfully on those nodes and continue to run if the nodes are upgraded to newer AMIs.
Below we describe how Karpenter assigns AMIs to nodes when they are first deployed and how newer AMIs are assigned later when nodes are spun up to replace old ones.
Later, it describes the options you have to assert control over how AMIs are used by Karpenter for your clusters.&lt;/p>
&lt;p>Features for managing AMIs described here should be considered as part of the larger upgrade policies that you have for your clusters.
See &lt;a href="/v1.15/faq/#how-do-i-upgrade-an-eks-cluster-with-karpenter">How do I upgrade an EKS Cluster with Karpenter&lt;/a> for details on this process.&lt;/p>
&lt;h2 id="how-karpenter-assigns-amis-to-nodes">How Karpenter assigns AMIs to nodes&lt;/h2>
&lt;p>Here is how Karpenter assigns AMIs nodes:&lt;/p>
&lt;ul>
&lt;li>When you create an &lt;code>EC2NodeClass&lt;/code>, you are required to specify &lt;a href="/v1.15/concepts/nodeclasses/#specamiselectorterms">&lt;code>amiSelectorTerms&lt;/code>&lt;/a>. &lt;a href="/v1.15/concepts/nodeclasses/#specamiselectorterms">&lt;code>amiSelectorTerms&lt;/code>&lt;/a> allow you to select on AMIs that can be spun-up by this EC2NodeClass based on tags, id, name, or an alias. Multiple AMIs may be specified, and Karpenter will choose the newest compatible AMI when spinning up new nodes.&lt;/li>
&lt;li>Some &lt;code>amiSelectorTerm&lt;/code> types are static and always resolve to the same AMI (e.g. &lt;code>id&lt;/code>). However, some are dynamic and may resolve to different AMIs over time. Examples of dynamic types include &lt;code>alias&lt;/code>, &lt;code>tags&lt;/code>, and &lt;code>name&lt;/code> (when using a wildcard). For example, if you specify an &lt;code>amiSelectorTerm&lt;/code> with an &lt;code>alias&lt;/code> set to &lt;code>@latest&lt;/code> (e.g. &lt;code>al2023@latest&lt;/code>, &lt;code>al2@latest&lt;/code>, or &lt;code>bottlerocket@latest&lt;/code>), Karpenter will use the &lt;em>latest&lt;/em> release for that AMI type when spinning up a new node.&lt;/li>
&lt;li>When a node is replaced, Karpenter checks to see if a newer AMI is available based on your &lt;code>amiSelectorTerms&lt;/code>. If a newer AMI is available, Karpenter will automatically use the new AMI to spin up the new node. &lt;strong>In particular, if you are using a dynamic &lt;code>amiSelectorTerm&lt;/code> type, you may get a new AMI deployed to your environment without having properly tested it.&lt;/strong>&lt;/li>
&lt;/ul>
&lt;p>Whenever a node is replaced, the replacement node will be launched using the newest AMI based on your &lt;code>amiSelectorTerms&lt;/code>. Nodes may be replaced due to manual deletion, or any of Karpenter&amp;rsquo;s automated methods:&lt;/p>
&lt;ul>
&lt;li>&lt;a href="/v1.15/concepts/disruption/#expiration">&lt;strong>Expiration&lt;/strong>&lt;/a>: Automatically initiates replacement at a certain time after the node is created.&lt;/li>
&lt;li>&lt;a href="/v1.15/concepts/disruption/#consolidation">&lt;strong>Consolidation&lt;/strong>&lt;/a>: If Karpenter detects that a cheaper node can be used to run the same workloads, Karpenter may replace the current node automatically.&lt;/li>
&lt;li>&lt;a href="/v1.15/concepts/disruption/#drift">&lt;strong>Drift&lt;/strong>&lt;/a>: If a node&amp;rsquo;s state no longer matches the desired state dictated by the &lt;code>NodePool&lt;/code> or &lt;code>EC2NodeClass&lt;/code>, it will be replaced, including if the node&amp;rsquo;s AMI no longer matches the latest AMI selected by the &lt;code>amiSelectorTerms&lt;/code>.&lt;/li>
&lt;li>&lt;a href="/v1.15/concepts/disruption/#interruption">&lt;strong>Interruption&lt;/strong>&lt;/a>: Nodes are sometimes involuntarily disrupted by things like Spot interruption, health changes, and instance events, requiring new nodes to be deployed.&lt;/li>
&lt;/ul>
&lt;p>See &lt;a href="/v1.15/concepts/disruption/#automated-methods">&lt;strong>Automated Methods&lt;/strong>&lt;/a> for details on how Karpenter uses these automated actions to replace nodes.&lt;/p>
&lt;p>The most relevant automated disruption method is &lt;a href="/v1.15/concepts/disruption/#drift">&lt;strong>Drift&lt;/strong>&lt;/a>, since it is initiated when a new AMI is selected-on by your &lt;code>amiSelectorTerms&lt;/code>. This could be due to a manual update (e.g. a new &lt;code>id&lt;/code> term was added), or due to a new AMI being resolved by a dynamic term.&lt;/p>
&lt;p>If you&amp;rsquo;re using an &lt;code>alias&lt;/code> with the &lt;code>latest&lt;/code> pin (e.g. &lt;code>al2023@latest&lt;/code>), Karpenter periodically checks for new AMI releases. Since AMI releases are outside your control, this could result in new AMIs being deployed before they have been properly tested in a lower environment. This is why we &lt;strong>strongly recommend&lt;/strong> using version pins in production environments when using an alias (e.g. &lt;code>al2023@v20240807&lt;/code>).&lt;/p>
&lt;div class="alert alert-warning" role="alert">
&lt;h4 class="alert-heading">Important&lt;/h4>
If you are new to Karpenter, you should know that the behavior described here is different than you get with Managed Node Groups (MNG). MNG will always use the assigned AMI when it creates a new node and will never automatically upgrade to a new AMI when a new node is required. See &lt;a href="https://docs.aws.amazon.com/eks/latest/userguide/update-managed-node-group.html">Updating a Managed Node Group&lt;/a> to see how you would manually update MNG to use new AMIs.
&lt;/div>
&lt;h2 id="controlling-ami-replacement">Controlling AMI Replacement&lt;/h2>
&lt;p>Karpenter&amp;rsquo;s automated node replacement functionality in tandem with the &lt;code>EC2NodeClass&lt;/code> gives you a lot of flexibility to control the desired state of nodes on your cluster. For example, you can opt-in to AMI auto-upgrades using &lt;code>alias&lt;/code> set to &lt;code>@latest&lt;/code>; however, this has to be weighed heavily against the risk of newer versions of an AMI breaking existing applications on your cluster. Alternatively, you can choose to pin your AMIs in your production clusters to avoid the risk of breaking changes; however, this has to be weighed against the management cost of testing new AMIs in pre-production and keeping up with the latest AMI versions.&lt;/p>
&lt;p>Karpenter offers you various controls to ensure you don&amp;rsquo;t take on too much risk as you rollout new versions of AMIs to your production clusters. Below shows how you can use these controls:&lt;/p>
&lt;ul>
&lt;li>&lt;a href="#pinning-amis">Pinning AMIs&lt;/a>: If workloads require a particluar AMI, this control ensures that it is the only AMI used by Karpenter. This can be used in combination with &lt;a href="#testing-amis">Testing AMIs&lt;/a> where you lock down the AMI in production, but allow the newest AMIs in a test cluster while you test your workloads before upgrading production.&lt;/li>
&lt;li>&lt;a href="#testing-amis">Testing AMIs&lt;/a>: The safest way for ensuring that a new AMI doesn&amp;rsquo;t break your workloads is to test it before putting it into production. This takes the most effort on your part, but most effectively models how your workloads will run in production, allowing you to catch issues ahead of time. Note that you can sometimes get different results from your test environment when you roll a new AMI into production, since issues like scale and other factors can elevate problems you might not see in test. Combining this with other controls like &lt;a href="#using-disruption-budgets">Using Disruption Budgets&lt;/a> can allow you to catch problems before they impact your whole cluster.&lt;/li>
&lt;li>&lt;a href="#using-disruption-budgets">Using Disruption Budgets&lt;/a>: This option can be used as a way of mitigating the scope of impact if a new AMI causes problems with your workloads. With Disruption budgets you can slow the pace of upgrades to nodes with new AMIs or make sure that upgrades only happen during selected dates and times (using &lt;code>schedule&lt;/code>). This doesn&amp;rsquo;t prevent a bad AMI from being deployed, but it allows you to control when nodes are upgraded, and gives you more time to respond to rollout issues.&lt;/li>
&lt;/ul>
&lt;h3 id="pinning-amis">Pinning AMIs&lt;/h3>
&lt;p>When you configure the &lt;a href="/v1.15/concepts/nodeclasses/">&lt;strong>EC2NodeClass&lt;/strong>&lt;/a>, you are required to configure which AMIs you want Karpenter to select on using the &lt;code>amiSelectorTerms&lt;/code> field. When pinning to a specific &lt;code>id&lt;/code>, &lt;code>name&lt;/code>, &lt;code>tags&lt;/code> or an &lt;code>alias&lt;/code> that contains a fixed version, Karpenter will only select on a single AMI and won&amp;rsquo;t automatically upgrade your nodes to a new version of an AMI. This prevents a new and potentially untested AMI from replacing existing nodes when those nodes are terminated.
).&lt;/p>
&lt;div class="alert alert-primary" role="alert">
&lt;h4 class="alert-heading">Note&lt;/h4>
Pinning an AMI to an &lt;code>alias&lt;/code> type with a fixed version &lt;em>will&lt;/em> pin the AMI so long as your K8s control plane version doesn&amp;rsquo;t change. Unlike &lt;code>id&lt;/code> and &lt;code>name&lt;/code> types, specifying a version &lt;code>alias&lt;/code> in your &lt;code>amiSelectorTerms&lt;/code> will cause Karpenter to consider the K8s control plane version of your cluster when choosing the AMI. If you upgrade your Kubernetes cluster while using this alias type, Karpenter &lt;em>will&lt;/em> automatically drift your nodes to a new AMI that still matches the AMI version but also matches your new K8s control plane version.
&lt;/div>
&lt;p>These examples show three different ways to identify the same AMI:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#8f5902;font-style:italic"># Using alias&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#8f5902;font-style:italic"># Pinning to this fixed version alias will pull this version of the AMI,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#8f5902;font-style:italic"># matching the K8s control plane version of your cluster&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">amiSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>- &lt;span style="color:#204a87;font-weight:bold">alias&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">al2023@v20240219&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#8f5902;font-style:italic"># Using name&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#8f5902;font-style:italic"># This will only ever select the AMI that contains this exact name&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">amiSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">al2023-ami-2023.3.20240219.0-kernel-6.1-x86_64&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#8f5902;font-style:italic"># Using id&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#8f5902;font-style:italic"># This will only ever select this specific AMI id&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">amiSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>- &lt;span style="color:#204a87;font-weight:bold">id&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">ami-052c9ea013e6e3567&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#8f5902;font-style:italic"># Using tags&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#8f5902;font-style:italic"># You can use a CI/CD system to test newer versions of an AMI&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#8f5902;font-style:italic"># and automatically tag them as you validate that they are safe to upgrade to&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">amiSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>- &lt;span style="color:#204a87;font-weight:bold">tags&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">karpenter.sh/discovery&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;${CLUSTER_NAME}&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">environment&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">prod&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>See the &lt;a href="/v1.15/concepts/nodeclasses/#specamiselectorterms">&lt;strong>spec.amiSelectorTerms&lt;/strong>&lt;/a> section of the NodeClasses page for details.
Keep in mind, that this could prevent you from getting critical security patches when new AMIs are available, but it does give you control over exactly which AMI is running.&lt;/p>
&lt;h3 id="testing-amis">Testing AMIs&lt;/h3>
&lt;p>Instead of avoiding AMI upgrades, you can set up test clusters where you can try out new AMI releases before they are put into production. For example, you could have:&lt;/p>
&lt;ul>
&lt;li>
&lt;p>&lt;strong>Test clusters&lt;/strong>: On lower environment clusters, you can run the latest AMIs e.g. &lt;code>al2023@latest&lt;/code>, &lt;code>al2@latest&lt;/code>, &lt;code>bottlerocket@latest&lt;/code>, for your workloads in a safe environment. This ensures that you get the latest patches for AMIs where downtime to applications isn&amp;rsquo;t as critical and allows you to validate patches to AMIs before they are deployed to production.&lt;/p>
&lt;/li>
&lt;li>
&lt;p>&lt;strong>Production clusters&lt;/strong>: After you&amp;rsquo;ve confirmed that the AMI works in your lower environments, you can pin the latest AMIs to be deployed in your production clusters to roll out the AMI. Refer to &lt;a href="#pinning-amis">Pinning AMIs&lt;/a> for how to choose a particular AMI by &lt;code>alias&lt;/code>, &lt;code>name&lt;/code> or &lt;code>id&lt;/code>. Remember that it is still best practice to gradually roll new AMIs into your cluster, even if they have been tested. So consider implementing that for your production clusters as described in &lt;a href="#using-disruption-budgets">Using Disruption Budgets&lt;/a>.&lt;/p>
&lt;/li>
&lt;/ul>
&lt;h3 id="using-disruption-budgets">Using Disruption Budgets&lt;/h3>
&lt;p>To reduce the risk of entire workloads being immediately degraded when a new AMI is deployed, you can enable Karpenter&amp;rsquo;s &lt;a href="#node-disruption-budgets ">&lt;strong>Node Disruption Budgets&lt;/strong>&lt;/a> as well as ensure that you have &lt;a href="#pod-disruption-budgets ">&lt;strong>Pod Disruption Budgets&lt;/strong>&lt;/a> configured for applications on your cluster. Below provides more details on how to configure each.&lt;/p>
&lt;h4 id="node-disruption-budgets">Node Disruption Budgets&lt;/h4>
&lt;p>&lt;a href="/v1.15/concepts/disruption/#disruption-budgets ">Disruption Budgets&lt;/a> limit when and to what extent nodes can be disrupted. You can prevent disruption based on nodes (a percentage or number of nodes that can be disrupted at a time) and schedule (excluding certain times from disrupting nodes).
You can set Disruption Budgets in a &lt;code>NodePool&lt;/code> spec. Here is an example:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">disruption&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">budgets&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">nodes&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#0000cf;font-weight:bold">15&lt;/span>&lt;span style="color:#000">%&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">nodes&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;3&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">nodes&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;0&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">schedule&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;0 9 * * sat,sun&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">duration&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">24h&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">nodes&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;0&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">schedule&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;0 17 * * mon-fri&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">duration&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">16h&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">reasons&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#000">Drifted&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Settings for budgets in the above example include the following:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Percentage of nodes&lt;/strong>: From the first &lt;code>nodes&lt;/code> setting, only &lt;code>15%&lt;/code> of the NodePool’s nodes can be disrupted at a time.&lt;/li>
&lt;li>&lt;strong>Number of nodes&lt;/strong>: The second &lt;code>nodes&lt;/code> setting limits the number of nodes that can be disrupted at a time to &lt;code>3&lt;/code>.&lt;/li>
&lt;li>&lt;strong>Schedule&lt;/strong>: The third &lt;code>nodes&lt;/code> setting uses schedule to say that zero disruptions (&lt;code>0&lt;/code>) are allowed starting at 9am on Saturday and Sunday and continues for 24 (fully blocking disruptions all day).
The format of the schedule follows the &lt;code>crontab&lt;/code> format for identifying dates and times.
See the &lt;a href="https://man7.org/linux/man-pages/man5/crontab.5.html">crontab&lt;/a> page for information on the supported values for these fields.&lt;/li>
&lt;li>&lt;strong>Reasons&lt;/strong>: The fourth &lt;code>nodes&lt;/code> setting uses &lt;code>reasons&lt;/code> which implies that this budget only applies to the &lt;code>Drifted&lt;/code> disruption condition. This setting uses schedule to say that zero disruptions (&lt;code>0&lt;/code>) are allowed starting at 5pm on Monday, Tuesday, Wednesday, Thursday, and Friday and continues for 16h (effectively blocking rolling nodes due to drift outside of working hours).&lt;/li>
&lt;/ul>
&lt;p>As with all disruption settings, keep in mind that avoiding updated AMIs for your nodes can result in not getting fixes for known security risks and bugs.
You need to balance that with your desire to not risk breaking the workloads on your cluster.&lt;/p>
&lt;h4 id="pod-disruption-budgets">Pod Disruption Budgets&lt;/h4>
&lt;p>&lt;a href="https://kubernetes.io/docs/tasks/run-application/configure-pdb/#specifying-a-poddisruptionbudget">Pod Disruption Budgets&lt;/a> allow you to describe how much disruption an application can tolerate before it begins to become unhealthy. This is critical to configure for Karpenter, since Karpenter uses this information to determine if it can continue to replace nodes. Specifically, if replacing a node would cause a Pod Disruption Budget to be breached (for graceful forms of disruption e.g. Drift or Consolidation), Karpenter will not replace the node.&lt;/p>
&lt;p>In a scenario where a faulty AMI is rolling out and begins causing downtime to your applications, configuring Pod Disruption Budgets is critical since this will tell Karpenter that it must stop replacing nodes until your applications become healthy again. This prevents Karpenter from deploying the faulty AMI throughout your cluster, reduces the imact the AMI has on your production applications, and gives you manually intervene in the cluster to remediate the issue.&lt;/p>
&lt;h2 id="follow-up">Follow-up&lt;/h2>
&lt;p>The Karpenter project continues to add features to give you greater control over AMI upgrades on your clusters.
If you have opinions about features you would like to see to manage AMIs with Karpenter, feel free to enter a Karpenter &lt;a href="https://github.com/aws/karpenter-provider-aws/issues/new/choose">New Issue&lt;/a>.&lt;/p></description></item><item><title>V1.15: Monitoring Amazon EC2 API Usage</title><link>/v1.15/tasks/monitoring-ec2-api-usage/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/v1.15/tasks/monitoring-ec2-api-usage/</guid><description>
&lt;p>AWS throttling can impact Karpenter&amp;rsquo;s ability to manage your cluster. Karpenter calls the Amazon EC2
API to discover infrastructure and to launch and terminate nodes, and Amazon EC2 enforces
per-account, per-Region request-rate limits. When your request rate exceeds a limit, Amazon EC2
rejects the excess requests with the &lt;code>RequestLimitExceeded&lt;/code> error (HTTP 503).&lt;/p>
&lt;p>The volume of these calls is not fixed. It scales with the number of &lt;code>EC2NodeClass&lt;/code>es and clusters
you run, how often your cluster scales up and down, and the Karpenter version you run, so a large
enough fleet can generate enough requests to be throttled. Because this volume can scale, you should
monitor it. This task describes the Amazon EC2 APIs Karpenter calls, how to observe your call volume
and throttling, how to compare that volume against your account&amp;rsquo;s request-rate limits, and other best
practices.&lt;/p>
&lt;h2 id="what-amazon-ec2-apis-karpenter-calls">What Amazon EC2 APIs Karpenter calls&lt;/h2>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>API&lt;/th>
&lt;th>Category&lt;/th>
&lt;th>When Karpenter calls it&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>CreateFleet&lt;/code>&lt;/td>
&lt;td>Launch (hot path)&lt;/td>
&lt;td>Launching nodes to satisfy pending pods&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>CreateLaunchTemplate&lt;/code>&lt;/td>
&lt;td>Launch (hot path)&lt;/td>
&lt;td>Preparing launch configuration for new nodes&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>RunInstances&lt;/code>&lt;/td>
&lt;td>Launch (hot path)&lt;/td>
&lt;td>Launching nodes&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>CreateTags&lt;/code>&lt;/td>
&lt;td>Launch (hot path)&lt;/td>
&lt;td>Tagging instances, fleets, and launch templates as they are created&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>TerminateInstances&lt;/code>&lt;/td>
&lt;td>Terminate&lt;/td>
&lt;td>Removing nodes during consolidation, drift, or expiration&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>RebootInstances&lt;/code>&lt;/td>
&lt;td>Reboot&lt;/td>
&lt;td>Rebooting unhealthy nodes in place during &lt;a href="/v1.15/concepts/disruption/#repair-actions">Node Auto Repair&lt;/a> with the &lt;code>RebootForRepair&lt;/code> AWS feature gate&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>DeleteLaunchTemplate&lt;/code>&lt;/td>
&lt;td>Cleanup&lt;/td>
&lt;td>Removing launch templates Karpenter manages&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>DescribeSubnets&lt;/code>&lt;/td>
&lt;td>Discovery / refresh&lt;/td>
&lt;td>Resolving &lt;code>subnetSelectorTerms&lt;/code> for each &lt;code>EC2NodeClass&lt;/code>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>DescribeSecurityGroups&lt;/code>&lt;/td>
&lt;td>Discovery / refresh&lt;/td>
&lt;td>Resolving &lt;code>securityGroupSelectorTerms&lt;/code> for each &lt;code>EC2NodeClass&lt;/code>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>DescribeImages&lt;/code>&lt;/td>
&lt;td>Discovery / refresh&lt;/td>
&lt;td>Resolving &lt;code>amiSelectorTerms&lt;/code> for each &lt;code>EC2NodeClass&lt;/code>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>DescribeCapacityReservations&lt;/code>&lt;/td>
&lt;td>Discovery / refresh&lt;/td>
&lt;td>Resolving &lt;code>capacityReservationSelectorTerms&lt;/code> (On-Demand Capacity Reservations)&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>DescribeInstanceTypes&lt;/code>, &lt;code>DescribeInstanceTypeOfferings&lt;/code>&lt;/td>
&lt;td>Discovery / refresh&lt;/td>
&lt;td>Resolving available instance types and their offerings&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>DescribeSpotPriceHistory&lt;/code>&lt;/td>
&lt;td>Discovery / refresh&lt;/td>
&lt;td>Determining spot pricing for instance type selection&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>DescribePlacementGroups&lt;/code>&lt;/td>
&lt;td>Discovery&lt;/td>
&lt;td>Resolving placement groups referenced by an &lt;code>EC2NodeClass&lt;/code>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>DescribeInstances&lt;/code>, &lt;code>DescribeInstanceStatus&lt;/code>&lt;/td>
&lt;td>Discovery&lt;/td>
&lt;td>Reconciling the state of launched instances&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>DescribeLaunchTemplates&lt;/code>&lt;/td>
&lt;td>Discovery&lt;/td>
&lt;td>Reconciling the launch templates Karpenter manages&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;h2 id="how-to-observe-call-volume-and-throttling">How to observe call volume and throttling&lt;/h2>
&lt;p>Karpenter exposes Prometheus metrics (by default at &lt;code>:8080/metrics&lt;/code>, configurable via &lt;code>METRICS_PORT&lt;/code>;
see the &lt;a href="/v1.15/reference/metrics/">Metrics reference&lt;/a>). The AWS SDK request metrics are
the most direct measure of Karpenter&amp;rsquo;s Amazon EC2 call volume and throttling. They are labeled by
&lt;code>service&lt;/code> (for example, &lt;code>EC2&lt;/code>), &lt;code>action&lt;/code> (the API operation, for example &lt;code>DescribeSubnets&lt;/code> or
&lt;code>CreateFleet&lt;/code>), and &lt;code>code&lt;/code> (the HTTP status code — &lt;code>200&lt;/code> for success, and &lt;code>503&lt;/code> for the
&lt;code>RequestLimitExceeded&lt;/code> throttling response):&lt;/p>
&lt;ul>
&lt;li>&lt;code>aws_sdk_go_request_total&lt;/code> — total AWS SDK requests, by &lt;code>service&lt;/code>, &lt;code>action&lt;/code>, and &lt;code>code&lt;/code>.&lt;/li>
&lt;li>&lt;code>aws_sdk_go_request_attempt_total&lt;/code> — total request attempts (a single request may make multiple
attempts when retried).&lt;/li>
&lt;li>&lt;code>aws_sdk_go_request_retry_count&lt;/code> — number of retry attempts per request. Sustained retries are an
early indicator of throttling, because the AWS SDK retries throttled requests before they surface
as an error.&lt;/li>
&lt;/ul>
&lt;p>For example, to graph Karpenter&amp;rsquo;s Amazon EC2 request rate by operation:&lt;/p>
&lt;pre tabindex="0">&lt;code>sum by (action) (rate(aws_sdk_go_request_total{service=&amp;#34;EC2&amp;#34;}[5m]))
&lt;/code>&lt;/pre>&lt;p>To graph the throttled fraction of Karpenter&amp;rsquo;s Amazon EC2 requests:&lt;/p>
&lt;pre tabindex="0">&lt;code>sum(rate(aws_sdk_go_request_total{service=&amp;#34;EC2&amp;#34;, code=&amp;#34;503&amp;#34;}[5m]))
/ sum(rate(aws_sdk_go_request_total{service=&amp;#34;EC2&amp;#34;}[5m]))
&lt;/code>&lt;/pre>&lt;h3 id="see-also">See also&lt;/h3>
&lt;ul>
&lt;li>Karpenter sets a User-Agent of &lt;code>karpenter.sh-&amp;lt;version&amp;gt;&lt;/code> on its AWS SDK clients, so you can attribute
Amazon EC2 API events to Karpenter in
&lt;a href="https://docs.aws.amazon.com/awscloudtrail/latest/userguide/cloudtrail-user-guide.html">AWS CloudTrail&lt;/a>
by filtering &lt;code>userAgent&lt;/code> for a value that begins with &lt;code>karpenter.sh-&lt;/code>.&lt;/li>
&lt;/ul>
&lt;h2 id="how-to-compare-against-your-accounts-amazon-ec2-api-request-rate-limits">How to compare against your account&amp;rsquo;s Amazon EC2 API request-rate limits&lt;/h2>
&lt;p>Amazon EC2 API request-rate limits are enforced &lt;strong>per account, per Region&lt;/strong>, and are independent for
different groups of actions (for example, the non-mutating &lt;code>Describe*&lt;/code> actions are limited separately
from mutating actions). Review your account&amp;rsquo;s applied limits with
&lt;a href="https://docs.aws.amazon.com/servicequotas/latest/userguide/intro.html">Service Quotas&lt;/a> for Amazon
EC2 in each Region where you run Karpenter, compare them against the request rate you observe from the
metrics above, and request an increase if your steady-state request rate is close to, or exceeds, your
limits. Because these limits are per account and per Region, the combined request volume of every
cluster in an account competes for the same limits.&lt;/p>
&lt;h2 id="best-practices">Best practices&lt;/h2>
&lt;h3 id="use-a-multi-account-architecture-where-clusters-are-isolated-by-account">Use a multi-account architecture where clusters are isolated by account&lt;/h3>
&lt;p>Because request-rate limits are per account and per Region, concentrating a large number of clusters
in a single account concentrates all of their Amazon EC2 request volume against that one account&amp;rsquo;s
limits. Use a multi-account architecture where clusters are isolated by account to spread the volume
across multiple accounts&amp;rsquo; limits. See the multi-account guidance in the
&lt;a href="https://docs.aws.amazon.com/wellarchitected/latest/framework/welcome.html">AWS Well-Architected Framework&lt;/a>.&lt;/p>
&lt;h2 id="increasing-refresh-intervals-as-a-workaround">Increasing refresh intervals as a workaround&lt;/h2>
&lt;div class="alert alert-warning" role="alert">
&lt;h4 class="alert-heading">Warning&lt;/h4>
Increasing these refresh intervals is a workaround, not a recommended configuration — avoid relying on
it. It reduces &lt;code>Describe*&lt;/code> call volume only by trading away freshness: Karpenter takes longer to
observe changes such as a subnet&amp;rsquo;s available IP capacity or a new AMI, and may act on stale data as a
result. Prefer the best practices above and an appropriately sized request-rate limit, and reach for
this only as a temporary measure while you address the underlying call volume.
&lt;/div>
&lt;p>Karpenter refreshes each &lt;code>EC2NodeClass&lt;/code>&amp;rsquo;s cached subnet, security group, and AMI data from Amazon EC2
on an interval. These intervals are configurable (see the
&lt;a href="/v1.15/reference/settings/">Settings reference&lt;/a>):&lt;/p>
&lt;ul>
&lt;li>&lt;code>SUBNET_REFRESH_INTERVAL&lt;/code> — how often subnet data is refreshed (bounds &lt;code>DescribeSubnets&lt;/code>). Defaults
to &lt;code>1m&lt;/code>.&lt;/li>
&lt;li>&lt;code>SECURITY_GROUP_REFRESH_INTERVAL&lt;/code> — how often security group data is refreshed (bounds
&lt;code>DescribeSecurityGroups&lt;/code>). Defaults to &lt;code>1m&lt;/code>.&lt;/li>
&lt;li>&lt;code>AMI_REFRESH_INTERVAL&lt;/code> — how often AMI data is refreshed (bounds &lt;code>DescribeImages&lt;/code>). Defaults to &lt;code>1m&lt;/code>.&lt;/li>
&lt;/ul>
&lt;p>Increasing an interval reduces that call&amp;rsquo;s steady-state rate proportionally — for example, changing an
interval from &lt;code>1m&lt;/code> to &lt;code>5m&lt;/code> reduces that call&amp;rsquo;s rate by roughly 5x.&lt;/p></description></item><item><title>V1.15: Utilizing GPUs and EFAs with Dynamic Resource Allocation</title><link>/v1.15/tasks/dra/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/v1.15/tasks/dra/</guid><description>
&lt;p>&lt;i class="fa-solid fa-circle-info">&lt;/i> &lt;b>Feature State: &lt;/b> &lt;a href="/v1.15/reference/settings/#aws-specific-feature-gates">Alpha&lt;/a>&lt;/p>
&lt;p>Karpenter can provision nodes for pods that request NVIDIA GPUs and Elastic Fabric Adapters (EFAs) through &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/dynamic-resource-allocation/">Dynamic Resource Allocation&lt;/a> (DRA).
To enable this support, set &lt;a href="#enabling-karpenter-support">&lt;code>settings.enableDRA&lt;/code>&lt;/a> in the Helm chart.&lt;/p>
&lt;div class="alert alert-primary" role="alert">
&lt;h4 class="alert-heading">Note&lt;/h4>
DRA requires Kubernetes 1.34 or later.
&lt;a href="#example-requesting-a-timesliced-gpu">GPU time-slicing&lt;/a> requires Kubernetes 1.36 or later.
&lt;/div>
&lt;h2 id="overview">Overview&lt;/h2>
&lt;p>With device plugins, a pod asks for devices as an integer count of an extended resource, such as &lt;code>nvidia.com/gpu: 1&lt;/code>.
With DRA, a pod references a &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/dynamic-resource-allocation/#resourceclaims-templates">ResourceClaim&lt;/a>, which describes the devices it needs.
A claim can select devices by their attributes using &lt;a href="https://kubernetes.io/docs/reference/using-api/cel/">CEL&lt;/a>, require several devices to share an attribute (for example, a GPU and an EFA on the same PCIe root), and request a portion of a device&amp;rsquo;s capacity.
A DRA driver runs on each node and publishes the node&amp;rsquo;s devices and their attributes as ResourceSlices. The scheduler allocates each claim from the devices in those ResourceSlices.&lt;/p>
&lt;p>The drivers only publish ResourceSlices once a node is running, but Karpenter has to choose an instance type before the node exists.
To close that gap, Karpenter includes metadata for each &lt;a href="#supported-instance-types">supported instance type&lt;/a> that describes the devices the NVIDIA and EFA DRA drivers publish on that instance type.
When a pod with ResourceClaims is pending, Karpenter evaluates its claims against this metadata, along with the ResourceSlices of existing nodes. If no existing node can satisfy the claims, Karpenter launches the cheapest instance type that can.
This works from zero: no GPU node has to exist first, and NodePools and EC2NodeClasses need no DRA-specific configuration.&lt;/p>
&lt;div class="alert alert-primary" role="alert">
&lt;h4 class="alert-heading">Note&lt;/h4>
Static NodePools launch nodes without simulating pods, and the DRA drivers and kube-scheduler handle allocation on those nodes.
DRA does not require explicit Karpenter support for static NodePools.
This page covers dynamic provisioning.
&lt;/div>
&lt;h2 id="capabilities">Capabilities&lt;/h2>
&lt;p>DRA makes requests possible that extended resources can&amp;rsquo;t express. For example:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Select GPUs by their attributes&lt;/strong>, such as model, architecture, CUDA compute capability, or memory. See &lt;a href="#example-requesting-a-specific-gpu-via-dra">Requesting a specific GPU&lt;/a>.&lt;/li>
&lt;li>&lt;strong>Align devices on the same PCIe root.&lt;/strong> For example, request a GPU and an EFA that share a PCIe root for GPUDirect RDMA, or two GPUs that share a PCIe root. See &lt;a href="#example-requesting-a-gpu-and-efa-that-share-a-pcie-root">Requesting a GPU and EFA that share a PCIe root&lt;/a>.&lt;/li>
&lt;li>&lt;strong>Share a GPU between pods&lt;/strong> with consumable capacity, where each claim takes a share of a GPU. See &lt;a href="#example-requesting-a-timesliced-gpu">Requesting a timesliced GPU&lt;/a>.&lt;/li>
&lt;/ul>
&lt;h2 id="installing-the-drivers">Installing the Drivers&lt;/h2>
&lt;p>Install each driver by following its install guide, then set the Helm values below.&lt;/p>
&lt;div class="alert alert-warning" role="alert">
&lt;h4 class="alert-heading">Warning&lt;/h4>
A DRA driver and a device plugin must never manage the same device on the same node. If you also run the NVIDIA or EFA device plugins in your cluster, see &lt;a href="#running-dra-drivers-alongside-device-plugins">Running DRA drivers alongside device plugins&lt;/a>.
&lt;/div>
&lt;h3 id="nvidia">NVIDIA&lt;/h3>
&lt;p>Install the NVIDIA DRA driver by following the &lt;a href="https://dra-driver-nvidia-gpu.sigs.k8s.io/docs/install/">NVIDIA DRA driver install guide&lt;/a> or the &lt;a href="https://docs.aws.amazon.com/eks/latest/userguide/device-management-nvidia-dra-device-plugin.html#eks-nvidia-dra-driver">EKS guide&lt;/a>. Neither value below is the chart&amp;rsquo;s default, so set both in the driver&amp;rsquo;s Helm chart:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Helm value&lt;/th>
&lt;th>Default&lt;/th>
&lt;th>Value&lt;/th>
&lt;th>Why&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>gpuResourcesEnabledOverride&lt;/code>&lt;/td>
&lt;td>&lt;code>false&lt;/code>&lt;/td>
&lt;td>&lt;code>true&lt;/code>&lt;/td>
&lt;td>Required for GPU allocation.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>resources.computeDomains.enabled&lt;/code>&lt;/td>
&lt;td>&lt;code>true&lt;/code>&lt;/td>
&lt;td>&lt;code>false&lt;/code>&lt;/td>
&lt;td>Optional. Disables ComputeDomains, which Karpenter doesn&amp;rsquo;t support provisioning for.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>To share GPUs between pods, you also set the driver&amp;rsquo;s consumable shares values. See &lt;a href="#example-requesting-a-timesliced-gpu">Requesting a timesliced GPU&lt;/a>.&lt;/p>
&lt;p>The chart creates the &lt;code>gpu.nvidia.com&lt;/code> DeviceClass.
The kubelet plugin tolerates the &lt;code>nvidia.com/gpu&lt;/code> taint by default. Its default affinity only schedules it on nodes with a GPU presence label, such as &lt;code>nvidia.com/gpu.present=true&lt;/code>.
The EKS-optimized AL2023 NVIDIA AMI sets that label on GPU nodes, and Karpenter selects it automatically for GPU instance types with the &lt;code>al2023@latest&lt;/code> alias.
If your nodes use another AMI, check that it sets the label. If it doesn&amp;rsquo;t, add the label to your DRA NodePool. Otherwise the driver won&amp;rsquo;t run on the nodes Karpenter launches, and DRA pods stay pending:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.sh/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">NodePool&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu-dra&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">template&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">labels&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">nvidia.com/gpu.present&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;true&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>
&lt;div class="alert alert-primary" role="alert">
&lt;h4 class="alert-heading">Note&lt;/h4>
Karpenter doesn&amp;rsquo;t support dynamic provisioning for &lt;a href="https://dra-driver-nvidia-gpu.sigs.k8s.io/docs/concepts/compute-domains/">ComputeDomains&lt;/a>. It has no metadata for the &lt;code>compute-domain.nvidia.com&lt;/code> driver, so it won&amp;rsquo;t launch nodes for claims on ComputeDomain DeviceClasses.
&lt;/div>
&lt;h3 id="efa">EFA&lt;/h3>
&lt;p>Install the EFA DRA driver (DRANET) by following &lt;a href="https://docs.aws.amazon.com/eks/latest/userguide/device-management-efa.html#efa-dra-driver">Install the EFA DRA driver&lt;/a>. Set this value in the &lt;code>aws-dranet&lt;/code> Helm chart:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Helm value&lt;/th>
&lt;th>Value&lt;/th>
&lt;th>Why&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>tolerations&lt;/code>&lt;/td>
&lt;td>&lt;code>[{key: nvidia.com/gpu, operator: Exists, effect: NoSchedule}]&lt;/code>&lt;/td>
&lt;td>The DaemonSet doesn&amp;rsquo;t tolerate the &lt;code>nvidia.com/gpu&lt;/code> taint by default, so it won&amp;rsquo;t run on tainted GPU nodes without this.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>The chart creates the &lt;code>efa.networking.k8s.aws&lt;/code> DeviceClass. It selects devices whose &lt;code>dra.net/pciDevice&lt;/code> attribute is &lt;code>Elastic Fabric Adapter (EFA)&lt;/code>.
DRANET publishes every network interface on a node, but Karpenter&amp;rsquo;s metadata includes only EFA devices, so request EFAs through this DeviceClass.&lt;/p>
&lt;h2 id="enabling-karpenter-support">Enabling Karpenter Support&lt;/h2>
&lt;p>Set &lt;code>settings.enableDRA: true&lt;/code> in Karpenter&amp;rsquo;s Helm values to enable Karpenter&amp;rsquo;s support for every DRA driver it models, currently the NVIDIA GPU and EFA drivers.
For how to change the values of an existing installation, see the &lt;a href="/v1.15/upgrading/upgrade-guide/">upgrade guide&lt;/a>.&lt;/p>
&lt;p>This enables the &lt;code>DRANVIDIAGPU&lt;/code> and &lt;code>DRAEFA&lt;/code> feature gates and sets &lt;code>settings.ignoreDRARequests&lt;/code> to &lt;code>false&lt;/code>.
The chart always grants Karpenter read access to DeviceClasses, ResourceClaims, and ResourceSlices.&lt;/p>
&lt;p>&lt;code>settings.enableDRA&lt;/code> assumes both drivers run on your DRA nodes. If a driver isn&amp;rsquo;t running, Karpenter launches nodes for its claims that can never serve them.
If you only run one of the drivers, &lt;a href="#enabling-drivers-individually">enable its gate individually&lt;/a> instead.&lt;/p>
&lt;div class="alert alert-primary" role="alert">
&lt;h4 class="alert-heading">Note&lt;/h4>
When you upgrade an existing release with &lt;code>helm upgrade --reuse-values&lt;/code>, Helm also reuses the default values of the chart version you installed before.
Earlier chart versions default &lt;code>settings.ignoreDRARequests&lt;/code> to &lt;code>true&lt;/code>, so setting &lt;code>settings.enableDRA&lt;/code> with &lt;code>--reuse-values&lt;/code> fails with the error above, even if you never set &lt;code>settings.ignoreDRARequests&lt;/code> yourself.
Use &lt;code>--reset-then-reuse-values&lt;/code> (Helm 3.14 or later) instead. If you previously set &lt;code>settings.awsFeatureGates.draNVIDIAGPU&lt;/code> or &lt;code>settings.awsFeatureGates.draEFA&lt;/code>, unset them in the same upgrade, for example with &lt;code>--set settings.awsFeatureGates.draNVIDIAGPU=null&lt;/code>.
&lt;/div>
&lt;h2 id="example-requesting-a-specific-gpu-via-dra">Example: Requesting a specific GPU via DRA&lt;/h2>
&lt;p>This ResourceClaimTemplate requests one GPU that&amp;rsquo;s Hopper or newer (CUDA compute capability 9.0 or later) with at least 100Gi of memory:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">resource.k8s.io/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">ResourceClaimTemplate&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">large-gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">devices&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">requests&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">exactly&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">deviceClassName&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu.nvidia.com&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">count&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#0000cf;font-weight:bold">1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">selectors&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">cel&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">expression&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000;font-weight:bold">|&lt;/span>&lt;span style="color:#8f5902;font-style:italic">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#8f5902;font-style:italic"> device.attributes[&amp;#34;gpu.nvidia.com&amp;#34;].cudaComputeCapability.compareTo(semver(&amp;#34;9.0.0&amp;#34;)) &amp;gt;= 0 &amp;amp;&amp;amp;
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#8f5902;font-style:italic"> device.capacity[&amp;#34;gpu.nvidia.com&amp;#34;].memory.compareTo(quantity(&amp;#34;100Gi&amp;#34;)) &amp;gt;= 0&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#000">---&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">Pod&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">inference&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">tolerations&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">nvidia.com/gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">operator&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">Exists&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">effect&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">NoSchedule&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">resourceClaims&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">resourceClaimTemplateName&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">large-gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">containers&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">model&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">image&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">public.ecr.aws/amazonlinux/amazonlinux:2023-minimal&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">command&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000;font-weight:bold">[&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;/bin/sh&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;-c&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;nvidia-smi &amp;amp;&amp;amp; sleep infinity&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">]&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">resources&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">claims&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Karpenter only considers instance types with a GPU that matches the selector: here, H200 (&lt;code>p5e&lt;/code>, &lt;code>p5en&lt;/code>) and Blackwell (&lt;code>p6-b200&lt;/code>, &lt;code>p6-b300&lt;/code>). It launches the cheapest one the NodePool allows.
To request several GPUs, raise &lt;code>count&lt;/code>. Karpenter checks that the instance type has enough matching GPUs.&lt;/p>
&lt;p>You can only select on attributes that Karpenter knows before launch. See &lt;a href="#supported-attributes-by-driver">Supported Attributes by Driver&lt;/a>.&lt;/p>
&lt;h2 id="example-requesting-a-gpu-and-efa-that-share-a-pcie-root">Example: Requesting a GPU and EFA that share a PCIe Root&lt;/h2>
&lt;p>Enable DRA support with &lt;code>settings.enableDRA&lt;/code> (or both individual gates), and install both drivers.
This ResourceClaimTemplate requests one GPU and one EFA, and requires them to share a PCIe root:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">resource.k8s.io/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">ResourceClaimTemplate&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu-efa-aligned&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">devices&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">requests&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">exactly&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">deviceClassName&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu.nvidia.com&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">count&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#0000cf;font-weight:bold">1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">efa&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">exactly&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">deviceClassName&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">efa.networking.k8s.aws&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">count&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#0000cf;font-weight:bold">1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">constraints&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">requests&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000;font-weight:bold">[&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;gpu&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;efa&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">]&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">matchAttribute&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">resource.kubernetes.io/pcieRoot&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Reference the claim from a pod as in the &lt;a href="#example-requesting-a-specific-gpu-via-dra">previous example&lt;/a>.
Karpenter only launches instance types where a GPU and an EFA share a PCIe root. For example, each of the eight GPUs in a &lt;code>p5.48xlarge&lt;/code> shares a PCIe root with four EFAs.&lt;/p>
&lt;div class="alert alert-warning" role="alert">
&lt;h4 class="alert-heading">Warning&lt;/h4>
Configuring &lt;code>spec.networkInterfaces&lt;/code> on an EC2NodeClass used for EFA DRA workloads is currently unsupported and results in undefined behavior. Leave it unset.
&lt;/div>
&lt;h2 id="example-requesting-a-timesliced-gpu">Example: Requesting a timesliced GPU&lt;/h2>
&lt;div class="alert alert-primary" role="alert">
&lt;h4 class="alert-heading">Note&lt;/h4>
GPU sharing through DRA requires Kubernetes 1.36 or later, where &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/dynamic-resource-allocation/">consumable capacity&lt;/a> is enabled by default.
&lt;/div>
&lt;p>With the NVIDIA DRA driver&amp;rsquo;s &lt;a href="https://dra-driver-nvidia-gpu.sigs.k8s.io/docs/guides/gpu-allocation/consumable-capacity/">consumable shares&lt;/a>, several claims can share one GPU, and the driver time-slices between them.
First, enable consumable shares by setting these values in the NVIDIA DRA driver&amp;rsquo;s Helm chart. This example splits each GPU into four shares:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Helm value&lt;/th>
&lt;th>Value&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>featureGates.ConsumableShares&lt;/code>&lt;/td>
&lt;td>&lt;code>true&lt;/code>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>consumableShares&lt;/code>&lt;/td>
&lt;td>&lt;code>4&lt;/code>&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>For the other modes and how the driver applies them, see the &lt;a href="https://dra-driver-nvidia-gpu.sigs.k8s.io/docs/guides/gpu-allocation/consumable-capacity/">consumable capacity guide&lt;/a>.&lt;/p>
&lt;p>Karpenter can&amp;rsquo;t read the driver&amp;rsquo;s configuration, so set the same mode with the &lt;code>karpenter.k8s.aws/nvidia-consumable-capacity&lt;/code> annotation on the EC2NodeClass:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.k8s.aws/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">EC2NodeClass&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu-dra&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">annotations&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">karpenter.k8s.aws/nvidia-consumable-capacity&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;4&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#8f5902;font-style:italic"># ...&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>The annotation accepts the same values as the driver&amp;rsquo;s &lt;code>consumableShares&lt;/code> setting:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Value&lt;/th>
&lt;th>Karpenter models each GPU as&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>absent or &lt;code>disabled&lt;/code>&lt;/td>
&lt;td>Not shared. Each GPU serves one claim.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>positive integer &lt;code>N&lt;/code>&lt;/td>
&lt;td>Shared, with &lt;code>N&lt;/code> shares. A claim takes one share by default, and its memory request defaults to 0.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>memory&lt;/code>&lt;/td>
&lt;td>Shared by memory. A claim with no memory request takes the GPU&amp;rsquo;s full memory.&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>unlimited&lt;/code>&lt;/td>
&lt;td>Shared without limit. A claim&amp;rsquo;s memory request defaults to 0.&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>Then give each pod its own claim from a template. A plain GPU request takes one share:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">resource.k8s.io/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">ResourceClaimTemplate&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu-share&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">devices&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">requests&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">exactly&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">deviceClassName&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu.nvidia.com&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#000">---&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">apps/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">Deployment&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">shared-inference&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">replicas&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#0000cf;font-weight:bold">8&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">selector&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">matchLabels&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">app&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">shared-inference&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">template&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">labels&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">app&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">shared-inference&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">tolerations&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">nvidia.com/gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">operator&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">Exists&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">effect&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">NoSchedule&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">resourceClaims&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">resourceClaimTemplateName&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu-share&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">containers&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">model&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">image&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">public.ecr.aws/amazonlinux/amazonlinux:2023-minimal&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">command&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000;font-weight:bold">[&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;/bin/sh&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;-c&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;nvidia-smi &amp;amp;&amp;amp; sleep infinity&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">]&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">resources&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">claims&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Karpenter packs up to four of these claims onto each GPU, so the eight replicas need only two GPUs.&lt;/p>
&lt;p>In &lt;code>memory&lt;/code> mode, request &lt;code>memory&lt;/code> instead, for example &lt;code>memory: 10Gi&lt;/code>.
Shares and memory only affect scheduling. They don&amp;rsquo;t limit how much GPU memory a process actually uses.&lt;/p>
&lt;h2 id="appendix">Appendix&lt;/h2>
&lt;h3 id="reference-configuration">Reference configuration&lt;/h3>
&lt;h4 id="ec2nodeclass-and-nodepool">EC2NodeClass and NodePool&lt;/h4>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.k8s.aws/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">EC2NodeClass&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu-dra&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#8f5902;font-style:italic"># Only for GPU sharing. Must match the driver&amp;#39;s consumableShares setting.&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#8f5902;font-style:italic"># annotations:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#8f5902;font-style:italic"># karpenter.k8s.aws/nvidia-consumable-capacity: &amp;#34;4&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">role&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;KarpenterNodeRole-${CLUSTER_NAME}&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">amiSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">alias&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">al2023@latest&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#8f5902;font-style:italic"># Resolves to the NVIDIA variant for GPU instance types&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">subnetSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">tags&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">karpenter.sh/discovery&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;${CLUSTER_NAME}&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">securityGroupSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#8f5902;font-style:italic"># For EFA, include a security group that allows all traffic to and from itself&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">tags&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">karpenter.sh/discovery&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;${CLUSTER_NAME}&amp;#34;&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#8f5902;font-style:italic"># Leave networkInterfaces unset. Karpenter attaches every EFA when a pod is allocated dra.net devices.&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#000">---&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.sh/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">NodePool&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu-dra&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">template&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">nodeClassRef&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">group&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.k8s.aws&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">EC2NodeClass&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu-dra&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">requirements&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.k8s.aws/instance-family&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">operator&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">In&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">values&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000;font-weight:bold">[&lt;/span>&lt;span style="color:#4e9a06">&amp;#34;g6&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;g6e&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;p5&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;p5en&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#34;p6-b200&amp;#34;&lt;/span>&lt;span style="color:#000;font-weight:bold">]&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">taints&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">nvidia.com/gpu&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">effect&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">NoSchedule&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;h4 id="pods">Pods&lt;/h4>
&lt;p>A pod that uses DRA devices needs:&lt;/p>
&lt;ul>
&lt;li>A &lt;code>spec.resourceClaims&lt;/code> entry that references a ResourceClaimTemplate (one claim per pod) or a ResourceClaim (shared by every pod that references it).&lt;/li>
&lt;li>A &lt;code>resources.claims&lt;/code> entry on each container that uses the devices.&lt;/li>
&lt;li>A toleration for the &lt;code>nvidia.com/gpu&lt;/code> taint.&lt;/li>
&lt;/ul>
&lt;p>Karpenter treats a pod as a DRA pod only if it has &lt;code>spec.resourceClaims&lt;/code> or a container lists &lt;code>resources.claims&lt;/code>.
A pod that requests the &lt;code>nvidia.com/gpu&lt;/code> extended resource isn&amp;rsquo;t treated as a DRA pod, even when a DeviceClass serves that resource through DRA.&lt;/p>
&lt;h3 id="running-dra-drivers-alongside-device-plugins">Running DRA drivers alongside device plugins&lt;/h3>
&lt;p>A DRA driver and a device plugin must never manage the same device on the same node. If both run, they can each hand out the same device, oversubscribing it without any error.
If your cluster also runs the &lt;a href="https://docs.aws.amazon.com/eks/latest/userguide/device-management-nvidia-dra-device-plugin.html#eks-nvidia-device-plugin">NVIDIA device plugin&lt;/a> or the &lt;a href="https://docs.aws.amazon.com/eks/latest/userguide/device-management-efa.html#eks-efa-device-plugin">EFA device plugin&lt;/a>, keep each node on one mechanism by giving DRA and device plugin NodePools separate labels.&lt;/p>
&lt;p>Karpenter&amp;rsquo;s DRA metadata applies to every NodePool that allows a supported instance type, including NodePools whose nodes run the device plugins.
If a DRA pod can schedule to both, Karpenter can launch it on a device plugin node, where no DRA driver publishes devices, and the pod stays pending.
Constrain DRA pods to your DRA NodePools as well.&lt;/p>
&lt;ol>
&lt;li>
&lt;p>Label each NodePool with the mechanism its nodes use, for example &lt;code>example.com/device-manager: dra&lt;/code> or &lt;code>example.com/device-manager: device-plugin&lt;/code>:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">apiVersion&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.sh/v1&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">kind&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">NodePool&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">name&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">gpu-dra&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">template&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">metadata&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">labels&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">example.com/device-manager&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">dra&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;/li>
&lt;li>
&lt;p>Set a matching &lt;code>nodeSelector&lt;/code> in each chart&amp;rsquo;s Helm values:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Chart&lt;/th>
&lt;th>Helm value&lt;/th>
&lt;th>Value&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>NVIDIA DRA driver&lt;/td>
&lt;td>&lt;code>kubeletPlugin.nodeSelector&lt;/code>&lt;/td>
&lt;td>&lt;code>example.com/device-manager: dra&lt;/code>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>EFA DRA driver (DRANET)&lt;/td>
&lt;td>&lt;code>nodeSelector&lt;/code>&lt;/td>
&lt;td>&lt;code>example.com/device-manager: dra&lt;/code>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>NVIDIA device plugin&lt;/td>
&lt;td>&lt;code>nodeSelector&lt;/code>&lt;/td>
&lt;td>&lt;code>example.com/device-manager: device-plugin&lt;/code>&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>EFA device plugin&lt;/td>
&lt;td>&lt;code>nodeSelector&lt;/code>&lt;/td>
&lt;td>&lt;code>example.com/device-manager: device-plugin&lt;/code>&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;/li>
&lt;li>
&lt;p>Add a &lt;code>nodeSelector&lt;/code> on &lt;code>example.com/device-manager: dra&lt;/code> to pods that use DRA devices:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">spec&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">nodeSelector&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">example.com/device-manager&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">dra&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;/li>
&lt;/ol>
&lt;h3 id="supported-instance-types">Supported Instance Types&lt;/h3>
&lt;p>Karpenter includes DRA metadata for the instance types below. An instance type with no EFA devices listed doesn&amp;rsquo;t support EFA.
Support for a new instance type requires a new Karpenter release.&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Instance type&lt;/th>
&lt;th>GPU&lt;/th>
&lt;th>GPUs&lt;/th>
&lt;th>EFA devices&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>g4dn.xlarge&lt;/code>&lt;/td>
&lt;td>Tesla T4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g4dn.2xlarge&lt;/code>&lt;/td>
&lt;td>Tesla T4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g4dn.4xlarge&lt;/code>&lt;/td>
&lt;td>Tesla T4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g4dn.8xlarge&lt;/code>&lt;/td>
&lt;td>Tesla T4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g4dn.12xlarge&lt;/code>&lt;/td>
&lt;td>Tesla T4&lt;/td>
&lt;td>4&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g4dn.16xlarge&lt;/code>&lt;/td>
&lt;td>Tesla T4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g4dn.metal&lt;/code>&lt;/td>
&lt;td>Tesla T4&lt;/td>
&lt;td>8&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5.xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA A10G&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5.2xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA A10G&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5.4xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA A10G&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5.8xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA A10G&lt;/td>
&lt;td>1&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5.12xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA A10G&lt;/td>
&lt;td>4&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5.16xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA A10G&lt;/td>
&lt;td>1&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5.24xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA A10G&lt;/td>
&lt;td>4&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5.48xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA A10G&lt;/td>
&lt;td>8&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5g.xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA T4G&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5g.2xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA T4G&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5g.4xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA T4G&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5g.8xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA T4G&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5g.16xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA T4G&lt;/td>
&lt;td>2&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g5g.metal&lt;/code>&lt;/td>
&lt;td>NVIDIA T4G&lt;/td>
&lt;td>2&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6.xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6.2xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6.4xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6.8xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6.12xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4&lt;/td>
&lt;td>4&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6.16xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6.24xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4&lt;/td>
&lt;td>4&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6.48xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4&lt;/td>
&lt;td>8&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6e.xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L40S&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6e.2xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L40S&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6e.4xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L40S&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6e.8xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L40S&lt;/td>
&lt;td>1&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6e.12xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L40S&lt;/td>
&lt;td>4&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6e.16xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L40S&lt;/td>
&lt;td>1&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6e.24xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L40S&lt;/td>
&lt;td>4&lt;/td>
&lt;td>2&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6e.48xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L40S&lt;/td>
&lt;td>8&lt;/td>
&lt;td>4&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6f.large&lt;/code>&lt;/td>
&lt;td>NVIDIA L4-3Q&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6f.xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4-3Q&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6f.2xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4-6Q&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>g6f.4xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4-12Q&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>gr6.4xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>gr6.8xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4&lt;/td>
&lt;td>1&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>gr6f.4xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA L4-12Q&lt;/td>
&lt;td>1&lt;/td>
&lt;td>-&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>p3dn.24xlarge&lt;/code>&lt;/td>
&lt;td>Tesla V100-SXM2-32GB&lt;/td>
&lt;td>8&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>p4d.24xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA A100-SXM4-40GB&lt;/td>
&lt;td>8&lt;/td>
&lt;td>4&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>p4de.24xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA A100-SXM4-80GB&lt;/td>
&lt;td>8&lt;/td>
&lt;td>4&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>p5.4xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA H100 80GB HBM3&lt;/td>
&lt;td>1&lt;/td>
&lt;td>1&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>p5.48xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA H100 80GB HBM3&lt;/td>
&lt;td>8&lt;/td>
&lt;td>32&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>p5e.48xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA H200&lt;/td>
&lt;td>8&lt;/td>
&lt;td>32&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>p5en.48xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA H200&lt;/td>
&lt;td>8&lt;/td>
&lt;td>16&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>p6-b200.48xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA B200&lt;/td>
&lt;td>8&lt;/td>
&lt;td>8&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>p6-b300.48xlarge&lt;/code>&lt;/td>
&lt;td>NVIDIA B300 SXM6 AC&lt;/td>
&lt;td>8&lt;/td>
&lt;td>16&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;h3 id="supported-attributes-by-driver">Supported Attributes by Driver&lt;/h3>
&lt;p>Karpenter can only evaluate a claim against attributes it knows before the node launches. Some attributes can be used in two ways:&lt;/p>
&lt;ul>
&lt;li>&lt;strong>Selector&lt;/strong>: a CEL expression in a claim or DeviceClass that compares the attribute to a value, for example &lt;code>device.attributes[&amp;quot;gpu.nvidia.com&amp;quot;].architecture == &amp;quot;Hopper&amp;quot;&lt;/code>.&lt;/li>
&lt;li>&lt;strong>matchAttribute&lt;/strong>: a constraint that requires every device allocated for a set of requests to have the same value for the attribute, without naming the value.&lt;/li>
&lt;/ul>
&lt;p>Other attributes can only be used with &lt;code>matchAttribute&lt;/code>. They describe how devices relate to each other on a node. Their values are resolved at runtime, or are specific to a platform, so Karpenter can tell whether devices will match but not what the value will be.&lt;/p>
&lt;h4 id="nvidia-gpunvidiacom">NVIDIA (&lt;code>gpu.nvidia.com&lt;/code>)&lt;/h4>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Attribute&lt;/th>
&lt;th>Type&lt;/th>
&lt;th>Example value&lt;/th>
&lt;th>Selector&lt;/th>
&lt;th>matchAttribute&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>productName&lt;/code>&lt;/td>
&lt;td>string&lt;/td>
&lt;td>&lt;code>NVIDIA H100 80GB HBM3&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>architecture&lt;/code>&lt;/td>
&lt;td>string&lt;/td>
&lt;td>&lt;code>Hopper&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>brand&lt;/code>&lt;/td>
&lt;td>string&lt;/td>
&lt;td>&lt;code>Nvidia&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>cudaComputeCapability&lt;/code>&lt;/td>
&lt;td>version&lt;/td>
&lt;td>&lt;code>9.0.0&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>type&lt;/code>&lt;/td>
&lt;td>string&lt;/td>
&lt;td>&lt;code>gpu&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>driverVersion&lt;/code>&lt;/td>
&lt;td>version&lt;/td>
&lt;td>&lt;code>580.82.7&lt;/code>&lt;/td>
&lt;td>&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>cudaDriverVersion&lt;/code>&lt;/td>
&lt;td>version&lt;/td>
&lt;td>&lt;code>13.0.0&lt;/code>&lt;/td>
&lt;td>&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>resource.kubernetes.io/pcieRoot&lt;/code>&lt;/td>
&lt;td>string&lt;/td>
&lt;td>&lt;code>pci0000:10&lt;/code>&lt;/td>
&lt;td>&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Capacity&lt;/th>
&lt;th>Example value&lt;/th>
&lt;th>Selector&lt;/th>
&lt;th>Capacity request&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>memory&lt;/code>&lt;/td>
&lt;td>&lt;code>81152Mi&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓ with consumable capacity&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>shares&lt;/code>&lt;/td>
&lt;td>&lt;code>4&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓ with an integer consumable capacity mode&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;h4 id="efa-dranet">EFA (&lt;code>dra.net&lt;/code>)&lt;/h4>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Attribute&lt;/th>
&lt;th>Type&lt;/th>
&lt;th>Example value&lt;/th>
&lt;th>Selector&lt;/th>
&lt;th>matchAttribute&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>pciDevice&lt;/code>&lt;/td>
&lt;td>string&lt;/td>
&lt;td>&lt;code>Elastic Fabric Adapter (EFA)&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>pciVendor&lt;/code>&lt;/td>
&lt;td>string&lt;/td>
&lt;td>&lt;code>Amazon.com, Inc.&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>pciSubsystem&lt;/code>&lt;/td>
&lt;td>string&lt;/td>
&lt;td>&lt;code>efa1&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>rdma&lt;/code>&lt;/td>
&lt;td>bool&lt;/td>
&lt;td>&lt;code>true&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>numaNode&lt;/code>&lt;/td>
&lt;td>int&lt;/td>
&lt;td>&lt;code>0&lt;/code>&lt;/td>
&lt;td>✓&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>resource.kubernetes.io/pcieRoot&lt;/code>&lt;/td>
&lt;td>string&lt;/td>
&lt;td>&lt;code>pci0000:10&lt;/code>&lt;/td>
&lt;td>&lt;/td>
&lt;td>✓&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table></description></item><item><title>V1.15: Utilizing On-Demand Capacity Reservations and Capacity Blocks</title><link>/v1.15/tasks/odcrs/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>/v1.15/tasks/odcrs/</guid><description>
&lt;p>&lt;i class="fa-solid fa-circle-info">&lt;/i> &lt;b>Feature State: &lt;/b> &lt;a href="/v1.15/reference/settings/#feature-gates">Beta&lt;/a>&lt;/p>
&lt;p>Karpenter introduced native support for &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-capacity-reservations.html">EC2 On-Demand Capacity Reservations&lt;/a> (ODCRs) in &lt;a href="https://github.com/aws/karpenter-provider-aws/releases/tag/v1.3.0">v1.3&lt;/a>, enabling users to select upon and prioritize specific capacity reservations.
In &lt;a href="https://github.com/aws/karpenter-provider-aws/releases/tag/v1.6.0">v1.6&lt;/a>, this support was expanded to include &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-capacity-blocks.html">EC2 Capacity Blocks for ML&lt;/a>. In &lt;a href="https://github.com/aws/karpenter-provider-aws/releases/tag/v1.10.0">v1.10&lt;/a>, this was further extended to support &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/interruptible-capacity-reservations.html">Interruptible Capacity Reservations&lt;/a>.
To enable native ODCR support, ensure the &lt;a href="/v1.15/reference/settings/#feature-gates">&lt;code>ReservedCapacity&lt;/code> feature gate&lt;/a> is enabled.&lt;/p>
&lt;div class="alert alert-primary" role="alert">
&lt;h4 class="alert-heading">Note&lt;/h4>
If you were previously utilizing &lt;code>open&lt;/code> ODCRs using Karpenter, review the &lt;a href="#migrating-from-previous-versions">migration section&lt;/a> of this task before enabling this feature.
&lt;/div>
&lt;h2 id="selecting-capacity-reservations">Selecting Capacity Reservations&lt;/h2>
&lt;p>To configure native ODCR support, you will need to make updates to both your EC2NodeClass and NodePool.
First, you should configure &lt;code>capacityReservationSelectorTerms&lt;/code> on your EC2NodeClass.
Similar to &lt;code>amiSelectorTerms&lt;/code>, you can specify a number of terms which are ANDed together to select ODCRs in your AWS account.
The following example demonstrates how to select all capacity reservations tagged with &lt;code>application: foobar&lt;/code> in addition to &lt;code>cr-56fac701cc1951b03&lt;/code>:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">capacityReservationSelectorTerms&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>- &lt;span style="color:#204a87;font-weight:bold">tags&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">application&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">foobar&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>- &lt;span style="color:#204a87;font-weight:bold">id&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">cr-56fac701cc1951b03&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>
&lt;div class="alert alert-primary" role="alert">
&lt;h4 class="alert-heading">Note&lt;/h4>
Capacity blocks are modeled as on-demand capacity reservations in EC2.
To select capacity blocks, specify them in your &lt;code>capacityReservationSelectorTerms&lt;/code> in the same way you would for a default ODCR.
&lt;/div>
&lt;p>For more information on configuring &lt;code>capacityReservationSelectorTerms&lt;/code>, see the &lt;a href="/v1.15/concepts/nodeclasses/#speccapacityreservationselectorterms">NodeClass docs&lt;/a>.&lt;/p>
&lt;p>Additionally, you will need to update your NodePool to be compatible with ODCRs.
Karpenter doesn&amp;rsquo;t model ODCRs as standard on-demand capacity, and instead uses a dedicated capacity type: &lt;code>reserved&lt;/code>.
For a NodePool to utilize ODCRs, it must be compatible with &lt;code>karpenter.sh/capacity-type: reserved&lt;/code>.
The following example demonstrates how to configure a NodePool to prioritize ODCRs and fallback to on-demand capacity:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">requirements&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.sh/capacity-type&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">operator&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">In&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">values&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000;font-weight:bold">[&lt;/span>&lt;span style="color:#4e9a06">&amp;#39;reserved&amp;#39;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#39;on-demand&amp;#39;&lt;/span>&lt;span style="color:#000;font-weight:bold">]&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>Additionaly, Karpenter supports the following scheduling labels:&lt;/p>
&lt;table>
&lt;thead>
&lt;tr>
&lt;th>Label&lt;/th>
&lt;th>Example&lt;/th>
&lt;th>Description&lt;/th>
&lt;/tr>
&lt;/thead>
&lt;tbody>
&lt;tr>
&lt;td>&lt;code>karpenter.k8s.aws/capacity-reservation-id&lt;/code>&lt;/td>
&lt;td>&lt;code>cr-56fac701cc1951b03&lt;/code>&lt;/td>
&lt;td>The capacity reservation&amp;rsquo;s ID&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>karpenter.k8s.aws/capacity-reservation-type&lt;/code>&lt;/td>
&lt;td>&lt;code>default&lt;/code> or &lt;code>capacity-block&lt;/code>&lt;/td>
&lt;td>The type of capacity reservation&lt;/td>
&lt;/tr>
&lt;tr>
&lt;td>&lt;code>karpenter.k8s.aws/capacity-reservation-interruptible&lt;/code>&lt;/td>
&lt;td>&lt;code>true&lt;/code> or &lt;code>false&lt;/code>&lt;/td>
&lt;td>Whether the capacity reservation is interruptible&lt;/td>
&lt;/tr>
&lt;/tbody>
&lt;/table>
&lt;p>These labels will only be present on reserved nodes.
They are supported as NodePool requirements and as pod scheduling constaints (e.g. &lt;a href="https://kubernetes.io/docs/concepts/scheduling-eviction/assign-pod-node/#node-affinity">node affinity&lt;/a>).&lt;/p>
&lt;div class="alert alert-warning" role="alert">
&lt;h4 class="alert-heading">Warning&lt;/h4>
Karpenter does &lt;strong>not&lt;/strong> support open matching for ODCRs.
This means that all ODCRs you wish to utilize, even those with &lt;code>open&lt;/code> instance eligibility, must be included in your NodeClass&amp;rsquo; &lt;code>spec.capacityReservationSelectorTerms&lt;/code>.
&lt;/div>
&lt;h2 id="prioritization-behavior">Prioritization Behavior&lt;/h2>
&lt;p>NodePools are not limited to a single compatible capacity-type &amp;ndash; they can be compatible with any combination of the available capacity-types (&lt;code>on-demand&lt;/code>, &lt;code>spot&lt;/code>, and &lt;code>reserved&lt;/code>).
Consider the following NodePool requirements:&lt;/p>
&lt;div class="highlight">&lt;pre tabindex="0" style="background-color:#f8f8f8;-moz-tab-size:4;-o-tab-size:4;tab-size:4;">&lt;code class="language-yaml" data-lang="yaml">&lt;span style="display:flex;">&lt;span>&lt;span style="color:#204a87;font-weight:bold">requirements&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline">&lt;/span>- &lt;span style="color:#204a87;font-weight:bold">key&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">karpenter.sh/capacity-type&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">operator&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000">In&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;span style="display:flex;">&lt;span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#204a87;font-weight:bold">values&lt;/span>&lt;span style="color:#000;font-weight:bold">:&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#000;font-weight:bold">[&lt;/span>&lt;span style="color:#4e9a06">&amp;#39;reserved&amp;#39;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#39;on-demand&amp;#39;&lt;/span>&lt;span style="color:#000;font-weight:bold">,&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline"> &lt;/span>&lt;span style="color:#4e9a06">&amp;#39;spot&amp;#39;&lt;/span>&lt;span style="color:#000;font-weight:bold">]&lt;/span>&lt;span style="color:#f8f8f8;text-decoration:underline">
&lt;/span>&lt;/span>&lt;/span>&lt;/code>&lt;/pre>&lt;/div>&lt;p>In this example, the NodePool is compatible with all capacity types.
Karpenter will prioritize ODCRs, but if none are available or none are compatible with the pending workloads it will fallback to spot or on-demand.
Similarly, Karpenter will prioritize reserved capacity during consolidation.
Since ODCRs are pre-paid, Karpenter will model them as free and consolidate spot / on-demand nodes when possible.&lt;/p>
&lt;h2 id="disrupting-nodes-in-full-reservations">Disrupting Nodes in Full Reservations&lt;/h2>
&lt;p>Karpenter pre-spins a replacement before it disrupts a node.
When a node&amp;rsquo;s reservation is full and its pods can&amp;rsquo;t run anywhere else, such as a lower-weight on-demand NodePool, Karpenter can&amp;rsquo;t pre-spin a replacement, and &lt;a href="/v1.15/concepts/disruption/#drift">Drift&lt;/a> and &lt;a href="/v1.15/concepts/disruption/#node-auto-repair">Node Auto Repair&lt;/a> are blocked for that node.
To replace these nodes in place, enable the &lt;code>TerminateFirstDrift&lt;/code> and &lt;code>TerminateFirstRepair&lt;/code> &lt;a href="/v1.15/reference/settings/#feature-gates">feature gates&lt;/a>.
Karpenter will then terminate the node first and launch its replacement into the freed reservation slot.
See &lt;a href="/v1.15/concepts/disruption/#terminate-first-disruption">Terminate-First Disruption&lt;/a> for details.&lt;/p>
&lt;div class="alert alert-warning" role="alert">
&lt;h4 class="alert-heading">Warning&lt;/h4>
Don&amp;rsquo;t enable terminate-first disruption unless this Karpenter installation is the only thing that launches into its capacity reservations.
Anything else that launches into the reservation, such as another Karpenter installation, an Auto Scaling group, or a matching instance consuming an &lt;code>open&lt;/code> reservation, can claim the freed slot first, leaving the terminated node&amp;rsquo;s pods pending indefinitely.
&lt;/div>
&lt;h2 id="expiration">Expiration&lt;/h2>
&lt;p>An instance launched into an ODCR is not necessarily in that ODCR indefinitely.
The ODCR could expire, be cancelled, or the instance could be manually removed from the ODCR.
If any of these occur, and Karpenter detects that the instance no longer belongs to an ODCR, it will update the &lt;code>karpenter.sh/capacity-type&lt;/code> label to &lt;code>on-demand&lt;/code>.&lt;/p>
&lt;h3 id="capacity-blocks">Capacity Blocks&lt;/h3>
&lt;p>Unlike default ODCRs, Capacity Blocks must have an end time.
Additionally, instances launched into a capacity block will be terminated by EC2 ahead of the end time, rather than becoming standard on-demand capacity.&lt;/p>
&lt;p>From the &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-capacity-blocks.html">AWS docs&lt;/a>:&lt;/p>
&lt;blockquote>
&lt;p>You can use all the instances you reserved until 30 minutes (for instance types) or 60 minutes (for UltraServer type) before the end time of the Capacity Block.
With 30 minutes (for instance types) or 60 minutes (for UltraServer types) left in your Capacity Block reservation, we begin terminating any instances that are running in the Capacity Block.
We use this time to clean up your instances before delivering the Capacity Block to the next customer.&lt;/p>
&lt;/blockquote>
&lt;p>Karpenter will preemptively begin draining nodes launched for capacity blocks 10 minutes before EC2 begins termination, ensuring your workloads can gracefully terminate before reclaimation.&lt;/p>
&lt;h3 id="interruptible-capacity-reservations">Interruptible Capacity Reservations&lt;/h3>
&lt;p>Unlike default ODCRs, capacity launched from interruptible ODCRs can be interrupted when capacity is reclaimed back the source ODCR.&lt;/p>
&lt;p>From the &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/interruptible-capacity-reservations.html">AWS docs&lt;/a>, when capacity is reclaimed:&lt;/p>
&lt;blockquote>
&lt;p>Running instances receive a 2-minute interruption warning through EventBridge events.
After the notice period, running instances in the reclaimed capacity enter a shutting down state and get terminated.&lt;/p>
&lt;/blockquote>
&lt;p>Karpenter will begin draining nodes launched for IODCRs when the 2-minute interruption warning is recieved.&lt;/p>
&lt;h2 id="migrating-from-previous-versions">Migrating From Previous Versions&lt;/h2>
&lt;p>Although it was not natively supported, it was possible to utilize ODCRs on previous versions of Karpenter.
If a NodeClaim&amp;rsquo;s requirements happened to be compatible with an open ODCR in the target AWS account, it may have launched an instance into that open ODCR.
This could be ensured by constraining a NodePool such that it was only compatible with the desired open ODCR, and limits could be used to enable fallback to a different NodePool once the ODCR was exhausted.
This behavior is no longer supported when native on-demand capacity support is enabled.&lt;/p>
&lt;p>If you were relying on this behavior, you should configure your &lt;code>EC2NodeClasses&lt;/code> to select the desired ODCRs &lt;strong>before&lt;/strong> enabling the feature gate.
You should also ensure any NodePools which you wish to use with ODCRs are compatible with &lt;code>karpenter.sh/capacity-type: reserved&lt;/code>.
Performing these steps before enabling the feature gate will ensure that Karpenter can immediately continue utilizing your reservations, rather than falling back to on-demand.&lt;/p></description></item></channel></rss>