Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
33 changes: 23 additions & 10 deletions gpu-operator/getting-started.rst
Original file line number Diff line number Diff line change
Expand Up @@ -56,7 +56,7 @@ Prerequisites
For worker nodes or node groups that run CPU workloads only, the nodes can run any operating system because
the GPU Operator does not perform any configuration or management of nodes for CPU-only workloads.

If you are planning to use NVIDIA GPU Driver Custom Resource Definition, you can use a mix of operating system versions on CPU and GPU nodes. Refer to the :doc:`NVIDIA GPU Driver Custom Resource Definition <gpu-driver-configuration>` page for more information.
If you are planning to use NVIDIA GPU Driver Custom Resource Definition, you can use a mix of operating system versions on CPU and GPU nodes. Refer to the :doc:`NVIDIA Driver Custom Resource Definition <nvidia-driver-configuration>` page for more information.

#. Nodes must be configured with a container engine such as CRI-O or containerd.

Expand Down Expand Up @@ -94,16 +94,29 @@ Use ``--set`` options to customize the deployment for your environment.
For installation on Red Hat OpenShift Container Platform,
refer to :external+ocp:doc:`steps-overview`.

#. Add the NVIDIA Helm repository:
#. Choose how to access the GPU Operator Helm chart:

.. code-block:: console
- To install the chart as an OCI artifact, no Helm repository setup is required.

- To use the classic NVIDIA Helm repository, add and update the repository:

$ helm repo add nvidia https://helm.ngc.nvidia.com/nvidia \
&& helm repo update
.. code-block:: console

$ helm repo add nvidia https://helm.ngc.nvidia.com/nvidia \
&& helm repo update

#. Install the GPU Operator.

- Install the Operator with the default configuration:
- Install the Operator from the OCI artifact with the default configuration:

.. code-block:: console

$ helm install --wait gpu-operator \
-n gpu-operator --create-namespace \
oci://nvcr.io/nvidia/cloud-native-charts/gpu-operator \
--version=${version}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we add a note that oci artifacts are available only starting v26.7.0 release and onwards?

@mikemckiernan mikemckiernan Aug 27, 2026 •

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The RN cover the introduction of features and changes rather than including that info in the docs (though, we haven't been 100% consistent on that point, I admit). Thank you for looking over the changes!

- Install the Operator from the classic Helm repository with the default configuration:

.. code-block:: console

Expand All @@ -112,7 +125,7 @@ Use ``--set`` options to customize the deployment for your environment.
nvidia/gpu-operator \
--version=${version}

- Install the Operator and specify configuration options:
- Install the Operator from the classic Helm repository and specify configuration options:

.. code-block:: console

Expand Down Expand Up @@ -470,7 +483,7 @@ To view all the options, run ``helm show values nvidia/gpu-operator``.

* - ``driver.nvidiaDriverCRD.enabled``
- When set to ``true``, the Operator deploys NVIDIA GPU Driver Custom Resource Definition.
Refer to the :doc:`NVIDIA GPU Driver Custom Resource Definition <gpu-driver-configuration>` page for more information.
Refer to the :doc:`NVIDIA Driver Custom Resource Definition <nvidia-driver-configuration>` page for more information.
- ``false``

* - ``driver.repository``
Expand Down Expand Up @@ -540,7 +553,7 @@ To view all the options, run ``helm show values nvidia/gpu-operator``.
When set to ``true``, the GDRCopy Driver runs as a sidecar container in the GPU driver pod.
For information about GDRCopy, refer to the `gdrcopy <https://developer.nvidia.com/gdrcopy>`__ page.

You can enable GDRCopy if you use the :doc:`gpu-driver-configuration`.
You can enable GDRCopy if you use the :doc:`nvidia-driver-configuration`.
- ``false``


Expand Down Expand Up @@ -899,6 +912,6 @@ After verifying the installation, you can configure the GPU Operator for your wo
- :doc:`gpu-operator-mig` — Configure Multi-Instance GPU (MIG) partitioning on supported GPUs.
- :doc:`gpu-operator-rdma` — Enable GPUDirect RDMA for high-performance networking.
- :doc:`dra-intro-install` — Allocate GPUs by using Kubernetes Dynamic Resource Allocation (DRA).
- :doc:`gpu-driver-configuration` — Use the NVIDIA GPU Driver Custom Resource Definition to manage drivers per node.
- :doc:`nvidia-driver-configuration` — Use the NVIDIA GPU Driver Custom Resource Definition to manage drivers per node.
- :doc:`precompiled-drivers` — Speed up driver deployments with precompiled kernel modules.
- :doc:`cdi` — Learn about Container Device Interface (CDI) and NRI Plugin mode.
45 changes: 36 additions & 9 deletions gpu-operator/gpu-operator-kubevirt-dra.rst
Original file line number Diff line number Diff line change
Expand Up @@ -65,7 +65,7 @@ Assumptions, Constraints, and Dependencies
You could also use labels and node selectors to avoid the race condition.
Refer to `Limitations and Considerations <https://dra-driver-nvidia-gpu.sigs.k8s.io/docs/guides/gpu-allocation/kubevirt-vfio-gpu-passthrough/#limitations-and-considerations>`__ in the DRA driver documentation for more information.

* DCGM and DGCM-Exporter cannot be enabled on the KubeVirt nodes.
* DCGM and DCGM Exporter cannot be enabled on the KubeVirt nodes.
* GPU Operator does not install the NVIDIA driver in the guest operating system.

*************
Expand Down Expand Up @@ -172,18 +172,27 @@ Pre-Installed Driver
$ sudo systemctl start nvidia-fabricmanager
$ sudo systemctl status nvidia-fabricmanager

******************************
Disable DCGM and DCGM Exporter
******************************

DCGM and DCGM Exporter are not supported on KubeVirt nodes that use VFIO passthrough.
Both operands keep NVIDIA client connections open and prevent the driver from rebinding a GPU to the VFIO driver.

Disable both operands in the ``GPUCluster`` resource before you enable VFIO support:

.. code-block:: console
:force:

$ kubectl patch gpucluster gpu-cluster \
--type=merge -p '{"spec":{"dcgm":{"enabled":false},"dcgmExporter":{"enabled":false}}}'

****************************
Enable VFIO Support with DRA
****************************

Enable ``PassthroughSupport`` and ``DeviceMetadata`` in the ``GPUCluster`` resource.

#. Isolate GPU consumers that would keep NVIDIA clients open during VFIO rebinding:

.. code-block:: console

$ kubectl patch gpucluster gpu-cluster --type=merge -p '{"spec":{"dcgmExporter":{"enabled":false}}}'

#. Add the feature gates under ``spec.draDriver.featureGates``.
Preserve any other DRA driver settings:

Expand All @@ -197,14 +206,27 @@ Enable ``PassthroughSupport`` and ``DeviceMetadata`` in the ``GPUCluster`` resou
"computeDomains":{"enabled":false},
"featureGates":{
"PassthroughSupport":true,
"DeviceMetadata":true,
"FabricManagerPartitioning":true # For NVLink 5 or later systems with NVSwitch-managed fabric.
"DeviceMetadata":true
}}}}'

``PassthroughSupport`` enables allocation with ``VfioDeviceConfig``.
``DeviceMetadata`` provides KubeVirt with the PCI address for each allocated device and exposes the VFIO API device to the ``virt-launcher`` pod.
GPU Operator continues to manage the DRA driver, DeviceClasses, RBAC, and operand lifecycle.

#. Optional: If you configured Fabric Manager partitioning, also enable the ``FabricManagerPartitioning`` feature gate.
Perform this step only on supported HGX or single-node NVL systems with an NVSwitch-managed fabric:

.. code-block:: console
:force:

$ kubectl patch gpucluster gpu-cluster \
--type=merge -p '{
"spec":{
"draDriver":{
"featureGates":{
"FabricManagerPartitioning":true # For NVLink 5 or later systems with NVSwitch-managed fabric.
}}}}'

#. Wait until GPU Operator finishes reconciling the change:

.. code-block:: console
Expand Down Expand Up @@ -269,6 +291,11 @@ Verify DRA Resources

$ kubectl get resourceslice <slice-name> -o yaml

The driver publishes partition membership through attributes named ``partitionN``,
where ``N`` is ``1``, ``2``, ``4``, or ``8`` and specifies the number of GPUs in the partition.
The attribute value is the Fabric Manager partition ID.
For example, ``partition2: 4`` indicates that the GPU belongs to two-GPU partition 4.

Confirm that devices with the ``gpu.nvidia.com/type`` attribute set to
``vfio`` include a ``gpuModuleID`` attribute and the ``partitionN`` attributes
reported for the hardware. For the two-GPU example in this procedure, at
Expand Down
2 changes: 1 addition & 1 deletion gpu-operator/index.rst
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,7 @@
Outdated Kernels <install-gpu-operator-outdated-kernels.rst>
Custom GPU Driver Parameters <custom-driver-params.rst>
precompiled-drivers.rst
GPU Driver CRD <gpu-driver-configuration.rst>
NVIDIA Driver CRD <nvidia-driver-configuration.rst>
Comment thread
rahulait marked this conversation as resolved.
CDI and NRI Support <cdi.rst>

.. toctree::
Expand Down
12 changes: 6 additions & 6 deletions gpu-operator/life-cycle-policy.rst
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,7 @@ NVIDIA GPU Operator Versioning
NVIDIA GPU Operator is versioned following the calendar versioning convention.

The version follows the pattern ``YY.MM.PP``, such as 23.6.0, 23.6.1, and 23.9.0.
The first two fields, ``YY.MM`` identify a major version and indicates when the major version was initially released.
The first two fields, ``YY.MM``, identify a major version and indicate when the major version was initially released.
The third field, ``PP``, identifies the patch version of the major version.
Patch releases typically include critical bug and CVE fixes, but can include minor features.

Expand All @@ -39,7 +39,7 @@ NVIDIA GPU Operator Life Cycle
******************************

When a new major version of NVIDIA GPU Operator is released, the previous major version enters deprecated support and only receives patch release updates for critical bug and CVE fixes.
All prior major versions enter end of support and are no longer supported and do not receive patch release updates.
All prior major versions reach end of support and no longer receive patch release updates.

The product life cycle and versioning are subject to change in the future.

Expand All @@ -53,7 +53,7 @@ The product life cycle and versioning are subject to change in the future.
* - GPU Operator Version
- Status

* - 26.6.x
* - 26.7.x
- Supported

* - 26.3.x
Expand All @@ -80,7 +80,7 @@ When post-release testing confirms support for newer versions of operands, these
Refer to :ref:`Upgrading the NVIDIA GPU Operator` for more information.

.. note::
All the following components are supported as :ref:`government-ready <install-gpu-operator-gov-ready>` in the NVIDIA GPU Operator v26.3, except for NVIDIA GDS Driver, NVIDIA Confidential Computing Manager, and NVIDIA GDRCopy Driver.
All of the following components are supported as :ref:`government-ready <install-gpu-operator-gov-ready>` in the NVIDIA GPU Operator v26.7, except for NVIDIA GDS Driver, NVIDIA Confidential Computing Manager, and NVIDIA GDRCopy Driver.

**D** = Default driver, **R** = Recommended driver

Expand Down Expand Up @@ -127,10 +127,10 @@ Refer to :ref:`Upgrading the NVIDIA GPU Operator` for more information.
- ${version}

* - NVIDIA KubeVirt GPU Device Plugin
- `v1.5.0 <https://github.com/NVIDIA/kubevirt-gpu-device-plugin>`__
- `v1.6.0 <https://github.com/NVIDIA/kubevirt-gpu-device-plugin>`__

* - NVIDIA vGPU Device Manager
- `v0.4.2 <https://github.com/NVIDIA/vgpu-device-manager>`__
- `v0.5.0 <https://github.com/NVIDIA/vgpu-device-manager>`__

* - NVIDIA GDS Driver |gds|_
- `2.29.4 <https://github.com/NVIDIA/gds-nvidia-fs/releases>`__
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -17,14 +17,18 @@

.. headings (h1/h2/h3/h4/h5) are # * = -

############################################
NVIDIA GPU Driver Custom Resource Definition
############################################
.. _nvidia-gpu-driver-custom-resource-definition:

########################################
NVIDIA Driver Custom Resource Definition
########################################

*****************************************************
Overview of the GPU Driver Custom Resource Definition
*****************************************************

.. _overview-of-the-gpu-driver-custom-resource-definition:

********************************************************
Overview of the NVIDIA Driver Custom Resource Definition
********************************************************

You can create one or more instances of an NVIDIA driver (``NVIDIADriver``) custom resource
to specify the NVIDIA GPU driver type and driver version to configure on specific nodes.
Expand Down Expand Up @@ -161,7 +165,7 @@ Custom Driver Parameters
For more information, refer to :doc:`Customizing NVIDIA GPU Driver Parameters during Installation <custom-driver-params>`.


.. _migrate-clusterpolicy-to-nvidiadriver:
.. _migrating-from-cluster-policy-driver-management:

***********************************************
Migrating from Cluster Policy Driver Management
Expand Down Expand Up @@ -348,8 +352,8 @@ The following table describes some of the fields in the custom resource.

* - ``kernelModuleType``
- Specifies the type of the NVIDIA GPU Kernel modules to use.
Valid values are ``auto`` (default), ``proprietary``, and ``open``.
Valid values are ``auto`` (default), ``proprietary``, and ``open``.

``Auto`` means that the recommended kernel module type is chosen based on the GPU devices on the host and the driver branch used.
- ``auto``

Expand Down Expand Up @@ -384,7 +388,7 @@ The following table describes some of the fields in the custom resource.
- ``nvcr.io/nvidia``

* - ``useOpenKernelModules`` Deprecated.
- This field is deprecated as of v25.3.0 and will be ignored. Use ``kernelModuleType`` instead.
- This field is deprecated as of v25.3.0 and will be ignored. Use ``kernelModuleType`` instead.
Specifies to use the NVIDIA Open GPU Kernel modules.
- ``false``

Expand Down Expand Up @@ -576,10 +580,11 @@ Precompiled Driver Container on Some Nodes


.. _nvd-upgrade:
.. _upgrading-the-nvidia-gpu-driver:

*******************************
Upgrading the NVIDIA GPU Driver
*******************************
***************************
Upgrading the NVIDIA Driver
***************************

To upgrade a driver managed by an NVIDIA driver custom resource, update the ``spec.version`` field.
The upgrade controller applies the resource's ``spec.upgradePolicy`` to the nodes that it manages.
Expand Down
2 changes: 1 addition & 1 deletion gpu-operator/overview.rst
Original file line number Diff line number Diff line change
Expand Up @@ -72,7 +72,7 @@ more information on how to contribute and the release artifacts.
The base images used by the software might include software that is licensed under open-source licenses such as GPL.
The source code for these components is archived on the CUDA opensource `index <https://developer.download.nvidia.com/compute/cuda/opensource/>`_.

The following table identifieis the licenses for the Operator and software components.
The following table identifies the licenses for the Operator and software components.
By installing and using the GPU Operator, you accept the terms and conditions of these licenses.

.. list-table::
Expand Down
15 changes: 11 additions & 4 deletions gpu-operator/platform-support.rst
Original file line number Diff line number Diff line change
Expand Up @@ -331,7 +331,7 @@ Bare Metal / Virtual Machines with GPU Passthrough and NVIDIA vGPU
- 1.33---1.36
- 1.33---1.36
- 1.33---1.36
- 1.33---1.35
- 1.33---1.36
- 2.17

* - Ubuntu 24.04 LTS
Expand All @@ -341,7 +341,7 @@ Bare Metal / Virtual Machines with GPU Passthrough and NVIDIA vGPU
- 1.33---1.36
- 1.33---1.36
- 1.33---1.36
- 1.33---1.35
- 1.33---1.36
- 2.17

* - Ubuntu 22.04 LTS |fn2|_
Expand All @@ -351,7 +351,7 @@ Bare Metal / Virtual Machines with GPU Passthrough and NVIDIA vGPU
- 1.33---1.36
- 1.33---1.36
- 1.33---1.36
- 1.33---1.35
- 1.33---1.36
- 2.15 2.16 2.17

* - Red Hat Core OS
Expand Down Expand Up @@ -483,18 +483,23 @@ Cloud Service Providers
| Kubernetes
- | Google GKE
| Kubernetes
- | Microsoft Azure
| Kubernetes Service

* - Ubuntu 26.04 LTS
- 1.33---1.36
- 1.33---1.36
- 1.33---1.36

* - Ubuntu 24.04 LTS
- 1.33---1.36
- 1.33---1.36
- 1.33---1.36

* - Ubuntu 22.04 LTS
- 1.33---1.36
- 1.33---1.36
- 1.33---1.36

.. _supported-precompiled-drivers:

Expand Down Expand Up @@ -558,14 +563,16 @@ To report an issue with the GPU Operator on one of these configurations, open an

For more details about partners and their supported configurations, refer to the :external+pv:doc:`index` page.

.. _supported-container-runtimes:

****************************
Supported Container Runtimes
****************************

The GPU Operator has been validated for the following container runtimes:

+----------------------------+------------------------+----------------+
| Operating System | Containerd 1.8 - 2.3 | CRI-O |
| Operating System | Containerd 2.0 - 2.3 | CRI-O |
+============================+========================+================+
| Ubuntu 26.04 LTS | Yes | -- |
+----------------------------+------------------------+----------------+
Expand Down
Loading