Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 12 additions & 11 deletions gpu-operator/life-cycle-policy.rst
Original file line number Diff line number Diff line change
Expand Up @@ -91,48 +91,49 @@ Refer to :ref:`Upgrading the NVIDIA GPU Operator` for more information.
- ${version}

* - NVIDIA GPU Driver |ki|_
- | `580.65.06 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-580-65-06/index.html>`_ (recommended)
- | `580.82.07 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-580-82-07/index.html>`_ (default, recommended)
| `580.65.06 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-580-65-06/index.html>`_
| `575.57.08 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-575-57-08/index.html>`_
| `570.172.08 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-570-172-08/index.html>`_ (default)
| `570.172.08 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-570-172-08/index.html>`_
| `570.158.01 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-570-158-01/index.html>`_
| `570.148.08 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-570-148-08/index.html>`_
| `535.261.03 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-535-261-03/index.html>`_
| `550.163.01 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-550-163-01/index.html>`_
| `535.247.01 <https://docs.nvidia.com/datacenter/tesla/tesla-release-notes-535-247-01/index.html>`_

* - NVIDIA Driver Manager for Kubernetes
- `v0.8.0 <https://ngc.nvidia.com/catalog/containers/nvidia:cloud-native:k8s-driver-manager>`__
- `v0.8.1 <https://ngc.nvidia.com/catalog/containers/nvidia:cloud-native:k8s-driver-manager>`__

* - NVIDIA Container Toolkit
- `1.17.8 <https://github.com/NVIDIA/nvidia-container-toolkit/releases>`__

* - NVIDIA Kubernetes Device Plugin
- `0.17.3 <https://github.com/NVIDIA/k8s-device-plugin/releases>`__
- `0.17.4 <https://github.com/NVIDIA/k8s-device-plugin/releases>`__

* - DCGM Exporter
- `4.2.3-4.1.3 <https://github.com/NVIDIA/dcgm-exporter/releases>`__
- `4.3.1-4.4.0 <https://github.com/NVIDIA/dcgm-exporter/releases>`__

* - Node Feature Discovery
- `v0.17.3 <https://github.com/kubernetes-sigs/node-feature-discovery/releases/>`__

* - | NVIDIA GPU Feature Discovery
| for Kubernetes
- `0.17.3 <https://github.com/NVIDIA/k8s-device-plugin/releases>`__
- `0.17.4 <https://github.com/NVIDIA/k8s-device-plugin/releases>`__

* - NVIDIA MIG Manager for Kubernetes
- `0.12.2 <https://github.com/NVIDIA/mig-parted/tree/main/deployments/gpu-operator>`__
- `0.12.3 <https://github.com/NVIDIA/mig-parted/blob/main/CHANGELOG.md>`__

* - DCGM
- `4.2.3 <https://docs.nvidia.com/datacenter/dcgm/latest/release-notes/changelog.html>`__
- `4.3.1 <https://docs.nvidia.com/datacenter/dcgm/latest/release-notes/changelog.html>`__

* - Validator for NVIDIA GPU Operator
- ${version}

* - NVIDIA KubeVirt GPU Device Plugin
- `v1.3.1 <https://github.com/NVIDIA/kubevirt-gpu-device-plugin>`__
- `v1.4.0 <https://github.com/NVIDIA/kubevirt-gpu-device-plugin>`__

* - NVIDIA vGPU Device Manager
- `v0.3.0 <https://github.com/NVIDIA/vgpu-device-manager>`__
- `v0.4.0 <https://github.com/NVIDIA/vgpu-device-manager>`__

* - NVIDIA GDS Driver |gds|_
- `2.20.5 <https://github.com/NVIDIA/gds-nvidia-fs/releases>`__
Expand All @@ -145,7 +146,7 @@ Refer to :ref:`Upgrading the NVIDIA GPU Operator` for more information.
- v0.1.1

* - NVIDIA GDRCopy Driver
- `v2.5.0 <https://github.com/NVIDIA/gdrcopy/releases>`__
- `v2.5.1 <https://github.com/NVIDIA/gdrcopy/releases>`__

.. _known-issue:

Expand Down
11 changes: 9 additions & 2 deletions gpu-operator/platform-support.rst
Original file line number Diff line number Diff line change
Expand Up @@ -148,6 +148,8 @@ The following NVIDIA data center GPUs are supported on x86 based platforms:
| NVIDIA RTX PRO 6000 | NVIDIA Blackwell |
| Blackwell Server Edition| |
+-------------------------+------------------------+
| NVIDIA RTX PRO 6000D | NVIDIA Blackwell |
+-------------------------+------------------------+
| NVIDIA RTX A6000 | NVIDIA Ampere /Ada |
+-------------------------+------------------------+
| NVIDIA RTX A5000 | NVIDIA Ampere |
Expand Down Expand Up @@ -468,6 +470,9 @@ See the :doc:`precompiled-drivers` page for more information about using precomp
| Ubuntu 22.04 | Generic, NVIDIA, Azure | 5.15 | R535, R550, R570 |
| | AWS, Oracle | | |
+----------------------------+------------------------+----------------+---------------------+
| Ubuntu 22.04 | Generic, NVIDIA, Azure | 6.8 | R535, R570 |
| | AWS, Oracle | | |
+----------------------------+------------------------+----------------+---------------------+
| Ubuntu 24.04 | Generic, NVIDIA, Azure | 6.8 | R550, R570 |
| | AWS, Oracle | | |
+----------------------------+------------------------+----------------+---------------------+
Expand Down Expand Up @@ -508,8 +513,8 @@ Operating System Kubernetes KubeVirt OpenShift Virtual
\ \ | GPU vGPU | GPU vGPU
| Passthrough | Passthrough
================ =========== ============= ========= ============= ===========
Ubuntu 20.04 LTS 1.23---1.29 0.36+ 0.59.1+
Ubuntu 22.04 LTS 1.23---1.29 0.36+ 0.59.1+
Ubuntu 20.04 LTS 1.23---1.33 0.36+ 0.59.1+
Ubuntu 22.04 LTS 1.23---1.33 0.36+ 0.59.1+
Red Hat Core OS 4.12---4.19 4.13---4.19
================ =========== ============= ========= ============= ===========

Expand All @@ -524,6 +529,8 @@ Refer to :ref:`GPU Operator with KubeVirt` or :ref:`NVIDIA GPU Operator with Ope

KubeVirt and OpenShift Virtualization with NVIDIA vGPU is supported on the following devices:

- RTX Pro 6000 Blackwell Server Edition

- H200NVL

- H100
Expand Down
32 changes: 32 additions & 0 deletions gpu-operator/release-notes.rst
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,38 @@ See the :ref:`GPU Operator Component Matrix` for a list of software components a

----

.. _v25.3.3:

25.3.3
======

.. _v25.3.3-new-features:

New Features
------------

* Supports these NVIDIA Data Center GPU Driver versions:

- 580.82.07 (default, recommended)

* Added support for additional features:

- RTX Pro 6000 Blackwell Server Edition

- MIG profiles support
- KubeVirt and OpenShift Virtualization: VM with GPU passthrough (Ubuntu 22.04 only)
- KubeVirt and OpenShift Virtualization: VM with time-slice vGPU (Ubuntu 22.04 only)

- RTX Pro 6000D

- KubeVirt and OpenShift Virtualization: VM with GPU passthrough (Ubuntu 22.04 only)

Comment thread
chenopis marked this conversation as resolved.
Fixed Issues
-------------

* Fixed an issue where user-supplied environment variables configured in ClusterPolicy were not getting set in the rendered DaemonSet.
User-supplied environment variables now take precedence over environment variables set by the ClusterPolicy controller.

.. _v25.3.2:

25.3.2
Expand Down
23 changes: 23 additions & 0 deletions gpu-operator/troubleshooting.rst
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,29 @@ If you are facing an issue that is not covered by this page, please file an issu
`NVIDIA GPU Operator GitHub repository <https://github.com/NVIDIA/gpu-operator/issues>`_.


**************************************************
The ``nouveau`` driver fails to initialize the GPU
**************************************************

.. rubric:: Observation
:class: h4

- The GPU driver fails to initialize the GPU with the error ``Failed to enable MSI-X`` in the system journal logs.
- All GPU Operator pods become stuck in the ``init`` state.

.. rubric:: Root Cause
:class: h4

- The ``nouveau`` Linux kernel module is loaded.

.. rubric:: Action
:class: h4

The ``nouveau`` driver must be denylisted when using NVIDIA vGPU.

Follow the instructions in the `NVIDIA AI Enterprise: VMware Deployment Guide <https://docs.nvidia.com/ai-enterprise/deployment/vmware/latest/nouveau.html#disable-nouveau>`_
to disable ``nouveau`` on your OS/distro to resolve this issue.

***********************************
GPU Operator pods are stuck in Init
***********************************
Expand Down
14 changes: 7 additions & 7 deletions gpu-operator/versions.json
Original file line number Diff line number Diff line change
@@ -1,24 +1,24 @@
{
"latest": "24.9.2",
"latest": "25.3.3",
"versions":
[
{
"version": "24.9.2"
"version": "25.3.3"
},
{
"version": "24.9.1"
"version": "25.3.2"
},
{
"version": "24.9.0"
"version": "25.3.1"
},
{
"version": "24.6.2"
"version": "25.3.0"
},
{
"version": "24.6.1"
"version": "24.9.2"
},
{
"version": "24.6.0"
"version": "24.9.1"
}
]
}
10 changes: 5 additions & 5 deletions gpu-operator/versions1.json
Original file line number Diff line number Diff line change
@@ -1,6 +1,10 @@
[
{
"preferred": "true",
"url": "../25.3.3",
"version": "25.3.3"
},
{
"url": "../25.3.2",
"version": "25.3.2"
},
Expand All @@ -19,9 +23,5 @@
{
"url": "../24.9.1",
"version": "24.9.1"
},
{
"url": "../24.9.0",
"version": "24.9.0"
}
]
]
4 changes: 2 additions & 2 deletions repo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -166,8 +166,8 @@ output_format = "linkcheck"
docs_root = "${root}/gpu-operator"
project = "gpu-operator"
name = "NVIDIA GPU Operator"
version = "25.3.2"
source_substitutions = { version = "v25.3.2", recommended = "580.65.06" }
version = "25.3.3"
source_substitutions = { version = "v25.3.3", recommended = "580.82.07" }
copyright_start = 2020
sphinx_exclude_patterns = [
"life-cycle-policy.rst",
Expand Down