From 181da7f60c08f8480d324521e3eaf61845332022 Mon Sep 17 00:00:00 2001 From: Andrew Chen Date: Tue, 9 Sep 2025 16:08:09 -0700 Subject: [PATCH 1/4] add driver container image tag note for OCP 4.19+ Signed-off-by: Andrew Chen --- openshift/gpu-operator-with-precompiled-drivers.rst | 5 +++++ 1 file changed, 5 insertions(+) diff --git a/openshift/gpu-operator-with-precompiled-drivers.rst b/openshift/gpu-operator-with-precompiled-drivers.rst index fa14ea377..1942a1a2c 100644 --- a/openshift/gpu-operator-with-precompiled-drivers.rst +++ b/openshift/gpu-operator-with-precompiled-drivers.rst @@ -121,6 +121,11 @@ Perform the following steps to build a custom driver image for use with Red Hat export DRIVER_VERSION=525.105.17 export OS_TAG=rhcos4.12 + .. note:: The driver container image tag for OpenShift has changed after the ``OCP 4.19`` release. + + - Before OCP 4.19: The driver image tag is formed with the suffix ``-rhcos4.17`` (for example with OCP 4.17). + - Starting OCP 4.19 and onwards: The driver image tag is formed with the suffix ``-rhel9.6`` (for example with OCP 4.19). + #. Build and push the image: .. code-block:: console From c7af1b7d146a7c2851cb9a9874cbe41a47622717 Mon Sep 17 00:00:00 2001 From: Andrew Chen Date: Tue, 9 Sep 2025 16:14:37 -0700 Subject: [PATCH 2/4] formatting and style guide fixes Signed-off-by: Andrew Chen --- openshift/appendix-ocp.rst | 10 ++-- .../gpu-operator-with-precompiled-drivers.rst | 26 ++++----- openshift/install-gpu-ocp.rst | 54 +++++++++---------- openshift/install-nfd.rst | 26 ++++----- openshift/introduction.rst | 4 +- openshift/openshift-virtualization.rst | 10 ++-- openshift/prerequisites.rst | 2 +- 7 files changed, 66 insertions(+), 66 deletions(-) diff --git a/openshift/appendix-ocp.rst b/openshift/appendix-ocp.rst index 22ce929ba..ed38f4f06 100644 --- a/openshift/appendix-ocp.rst +++ b/openshift/appendix-ocp.rst @@ -23,14 +23,14 @@ Introduction If you encounter the :ref:`"broken driver toolkit detected" ` warning on OpenShift 4.10 or later, you should :ref:`troubleshoot ` to find the root cause instead of falling back to entitled driver builds. - If the broken DTK warning is encountered on an older version of OpenShift, refer to the documentation for an older version of the NVIDIA GPU operator to enable entitled builds. Keep in mind that older versions of OpenShift might no longer be supported. + If the broken DTK warning is encountered on an older version of OpenShift, refer to the documentation for an older version of the NVIDIA GPU Operator to enable entitled builds. Keep in mind that older versions of OpenShift might no longer be supported. .. _broken-dtk-troubleshooting: Troubleshooting Broken Driver Toolkit Errors -------------------------------------------- -The most likely reason for the broken DTK message is Node Feature Discovery (NFD) not working correctly. NFD might be disabled, failing, or not updating the kernel version label for other reasons. Another cause might be a missing or incomplete DTK image stream, e.g. because of broken mirroring. +The most likely reason for the broken DTK message is Node Feature Discovery (NFD) not working correctly. NFD might be disabled, failing, or not updating the kernel version label for other reasons. Another cause might be a missing or incomplete DTK image stream, for example, because of broken mirroring. Follow these steps for initial troubleshooting of Node Feature Discovery: @@ -48,7 +48,7 @@ Follow these steps for initial troubleshooting of Node Feature Discovery: $ oc get nodes -o jsonpath='{range .items[*]}{.metadata.name}{":\t"}{.metadata.labels.feature\.node\.kubernetes\.io/kernel-version\.full}{"\n"}{end}' - Ensure nodes have proper kernel version labels that match current OpenShift version of the cluster. + Ensure nodes have proper kernel version labels that match the current OpenShift version of the cluster. #. **Check Driver Toolkit image stream:** @@ -56,12 +56,12 @@ Follow these steps for initial troubleshooting of Node Feature Discovery: $ oc get -n openshift is/driver-toolkit - Verify the driver-toolkit image stream exists and has the correct tags that correspond to current OpenShift version. + Verify the driver-toolkit image stream exists and has the correct tags that correspond to the current OpenShift version. For additional troubleshooting resources: * `Node Feature Discovery documentation `_. * `Red Hat Node Feature Discovery Operator documentation `_ * `OpenShift Driver Toolkit documentation `_ -* `OpenShift Driver Toolkit GihHub repository `_ +* `OpenShift Driver Toolkit GitHub repository `_ * `OpenShift troubleshooting guide `_ diff --git a/openshift/gpu-operator-with-precompiled-drivers.rst b/openshift/gpu-operator-with-precompiled-drivers.rst index 1942a1a2c..bb23acbd2 100644 --- a/openshift/gpu-operator-with-precompiled-drivers.rst +++ b/openshift/gpu-operator-with-precompiled-drivers.rst @@ -17,7 +17,7 @@ About Precompiled Driver Containers *********************************** By default, NVIDIA GPU drivers are built on the cluster nodes when you deploy the GPU Operator. -Driver compilation and packaging is done on every Kubernetes node, which leads to bursts of compute demand, waste of resources, and long provisioning times. +Driver compilation and packaging is done on every Kubernetes node, leading to bursts of compute demand, waste of resources, and long provisioning times. In contrast, using container images with precompiled drivers makes the drivers immediately available on all nodes, resulting in faster provisioning and cost savings in public cloud deployments. *********************************** @@ -43,7 +43,7 @@ Perform the following steps to build a custom driver image for use with Red Hat .. rubric:: Prerequisites -* You have access to a container registry, such as NVIDIA NGC Private Registry, Red Hat Quay, or the OpenShift internal container registry, and can push container images to the registry. +* You have access to a container registry such as NVIDIA NGC Private Registry, Red Hat Quay, or the OpenShift internal container registry and can push container images to the registry. * You have a valid Red Hat subscription with an activation key. @@ -51,11 +51,11 @@ Perform the following steps to build a custom driver image for use with Red Hat * Your build machine has access to the internet to download operating system packages. -* You know a CUDA version, such as ``12.1.0``, that you want to use. +* You know a CUDA version such as ``12.1.0`` that you want to use. - One way to find a supported CUDA version for your operating system is to access the NVIDIA GPU Cloud registry at `CUDA | NVIDIA NGC `_ and view the tags. Use the search field to filter the tags, such as ``base-ubi8`` for RHEL 8 and ``base-ubi9`` for RHEL 9. The filtered results show the CUDA versions, such as ``12.1.0``, ``12.0.1``, ``12.0.0``, and so on. + One way to find a supported CUDA version for your operating system is to access the NVIDIA GPU Cloud registry at `CUDA | NVIDIA NGC `_ and view the tags. Use the search field to filter the tags such as ``base-ubi8`` for RHEL 8 and ``base-ubi9`` for RHEL 9. The filtered results show the CUDA versions such as ``12.1.0``, ``12.0.1``, and ``12.0.0``. -* You know the GPU driver version, such as ``525.105.17``, that you want to use. +* You know the GPU driver version such as ``525.105.17`` that you want to use. .. rubric:: Procedure @@ -65,26 +65,26 @@ Perform the following steps to build a custom driver image for use with Red Hat $ git clone https://gitlab.com/nvidia/container-images/driver -#. Change the directory to ``rhel8/precompiled`` under the cloned repository. You can build precompiled driver images for versions 8 and 9 of RHEL from this directory: +#. Change to the ``rhel8/precompiled`` directory under the cloned repository. You can build precompiled driver images for versions 8 and 9 of RHEL from this directory: .. code-block:: console $ cd driver/rhel8/precompiled -#. Create a Red Hat Customer Portal Activation Key and note your Red Hat Subscription Management (RHSM) organization ID. These are to install packages during a build. Save the values to files, for example, ``$HOME/rhsm_org`` and ``$HOME/rhsm_activationkey``: +#. Create a Red Hat Customer Portal Activation Key and note your Red Hat Subscription Management (RHSM) organization ID. These are to install packages during a build. Save the values to files such as ``$HOME/rhsm_org`` and ``$HOME/rhsm_activationkey``: .. code-block:: console export RHSM_ORG_FILE=$HOME/rhsm_org export RHSM_ACTIVATIONKEY_FILE=$HOME/rhsm_activationkey -#. Download your Red Hat OpenShift pull secret and store it in a file, for example, ``${HOME}/pull-secret``: +#. Download your Red Hat OpenShift pull secret and store it in a file such as ``${HOME}/pull-secret``: .. code-block:: console export PULL_SECRET_FILE=$HOME/pull-secret.txt -#. Set the Red Hat OpenShift version and target architecture of your cluster, for example, ``x86_64``: +#. Set the Red Hat OpenShift version and target architecture of your cluster such as ``x86_64``: .. code-block:: console @@ -123,8 +123,8 @@ Perform the following steps to build a custom driver image for use with Red Hat .. note:: The driver container image tag for OpenShift has changed after the ``OCP 4.19`` release. - - Before OCP 4.19: The driver image tag is formed with the suffix ``-rhcos4.17`` (for example with OCP 4.17). - - Starting OCP 4.19 and onwards: The driver image tag is formed with the suffix ``-rhel9.6`` (for example with OCP 4.19). + - Before OCP 4.19: The driver image tag is formed with the suffix ``-rhcos4.17`` (such as with OCP 4.17). + - Starting OCP 4.19 and later: The driver image tag is formed with the suffix ``-rhel9.6`` (such as with OCP 4.19). #. Build and push the image: @@ -132,9 +132,9 @@ Perform the following steps to build a custom driver image for use with Red Hat make image image-push -Optionally, override the ``IMAGE_REGISTRY``, ``IMAGE_NAME``, and ``CONTAINER_TOOL``. You can also override ``BUILDER_USER`` and ``BUILDER_EMAIL`` if you want, otherwise your Git username and email are used. See the Makefile for all available variables. +Optionally, override the ``IMAGE_REGISTRY``, ``IMAGE_NAME``, and ``CONTAINER_TOOL``. You can also override ``BUILDER_USER`` and ``BUILDER_EMAIL`` if you want. Otherwise, your Git username and email are used. Refer to the Makefile for all available variables. -.. note:: Do not set the ``DRIVER_TYPE``. The only supported value is currently ``passthrough``, which is set by default. +.. note:: Do not set the ``DRIVER_TYPE``. The only supported value is currently ``passthrough``, and this is set by default. ********************************************* Enabling Precompiled Driver Container Support diff --git a/openshift/install-gpu-ocp.rst b/openshift/install-gpu-ocp.rst index 5cfe1af95..9b17ab08e 100644 --- a/openshift/install-gpu-ocp.rst +++ b/openshift/install-gpu-ocp.rst @@ -13,18 +13,18 @@ Installing the NVIDIA GPU Operator by using the web console #. In the OpenShift Container Platform web console from the side menu, navigate to **Operators** > **OperatorHub** and select **All Projects**. -#. In **Operators** > **OperatorHub**, search for the **NVIDIA GPU Operator**. For additional information see the `Red Hat OpenShift Container Platform documentation `_. +#. In **Operators** > **OperatorHub**, search for the **NVIDIA GPU Operator**. For additional information, refer to the `Red Hat OpenShift Container Platform documentation `_. #. Select the **NVIDIA GPU Operator**, click **Install**. In the subsequent screen click **Install**. .. note:: Here, you can select the namespace where you want to deploy the GPU Operator. The suggested namespace to use is the ``nvidia-gpu-operator``. You can choose any existing namespace or create a new namespace under **Select a Namespace**. - If you install in any other namespace other than ``nvidia-gpu-operator``, the GPU Operator will **not** automatically enable namespace monitoring, and metrics and alerts will **not** be collected by Prometheus. - If only trusted operators are installed in this namespace, you can manually enable namespace monitoring with this command: + If you install in any other namespace other than ``nvidia-gpu-operator``, the GPU Operator will **not** automatically enable namespace monitoring, and metrics and alerts will **not** be collected by Prometheus. + If only trusted operators are installed in this namespace, you can manually enable namespace monitoring with this command: - .. code-block:: console + .. code-block:: console - $ oc label ns/$NAMESPACE_NAME openshift.io/cluster-monitoring=true + $ oc label ns/$NAMESPACE_NAME openshift.io/cluster-monitoring=true Proceed to :ref:`Create the cluster policy for the NVIDIA GPU Operator `. @@ -46,13 +46,13 @@ As a cluster administrator, you can install the **NVIDIA GPU Operator** using th name: nvidia-gpu-operator .. note:: The suggested namespace to use is the ``nvidia-gpu-operator``. You can choose any existing namespace or create a new namespace name. - If you install in any other namespace other than ``nvidia-gpu-operator``, the GPU Operator will **not** automatically enable namespace monitoring, and metrics and alerts will **not** be collected by Prometheus. + If you install in any other namespace other than ``nvidia-gpu-operator``, the GPU Operator will **not** automatically enable namespace monitoring, and metrics and alerts will **not** be collected by Prometheus. - If only trusted operators are installed in this namespace, you can manually enable namespace monitoring with this command: + If only trusted operators are installed in this namespace, you can manually enable namespace monitoring with this command: - .. code-block:: console + .. code-block:: console - $ oc label ns/$NAMESPACE_NAME openshift.io/cluster-monitoring=true + $ oc label ns/$NAMESPACE_NAME openshift.io/cluster-monitoring=true #. Create the namespace by running the following command: @@ -89,7 +89,7 @@ As a cluster administrator, you can install the **NVIDIA GPU Operator** using th operatorgroup.operators.coreos.com/nvidia-gpu-operator-group created -#. Run the following command to get the ``channel`` value required for step number 5. +#. Run the following command to get the ``channel`` value required for step 5. .. code-block:: console @@ -101,7 +101,7 @@ As a cluster administrator, you can install the **NVIDIA GPU Operator** using th v22.9 -#. Run the following commands to get the ``startingCSV`` value required for step number 5. +#. Run the following commands to get the ``startingCSV`` value required for step 5. .. code-block:: console @@ -134,7 +134,7 @@ As a cluster administrator, you can install the **NVIDIA GPU Operator** using th sourceNamespace: openshift-marketplace startingCSV: "gpu-operator-certified.v22.9.0" - .. note:: Update the ``channel`` and ``startingCSV`` fields with the information returned in step 3 and 4. + .. note:: Update the ``channel`` and ``startingCSV`` fields with the information returned in steps 3 and 4. #. Create the subscription object by running the following command: @@ -198,7 +198,7 @@ When you install the **NVIDIA GPU Operator** in the OpenShift Container Platform .. note:: If you create a ClusterPolicy that contains an empty specification, such as ``spec{}``, the ClusterPolicy fails to deploy. As a cluster administrator, you can create a ClusterPolicy using the OpenShift Container Platform CLI or the web console. Also, these steps differ -when using **NVIDIA vGPU**. Please refer to appropriate sections below. +when using **NVIDIA vGPU**. Refer to the appropriate sections below. .. _create-cluster-policy-web-console: @@ -209,7 +209,7 @@ Create the cluster policy using the web console #. Select the **ClusterPolicy** tab, then click **Create ClusterPolicy**. The platform assigns the default name *gpu-cluster-policy*. - .. note:: You can use this screen to customize the ClusterPolicy; although, the default values are sufficient to get the GPU configured and running in most cases. + .. note:: You can use this screen to customize the ClusterPolicy. However, the default values are sufficient to get the GPU configured and running in most cases. .. note:: For OpenShift 4.12 with GPU Operator 25.3.1 or later, you must expand the **Driver** section and set the following fields: @@ -219,7 +219,7 @@ Create the cluster policy using the web console #. Click **Create**. - At this point, the GPU Operator proceeds and installs all the required components to set up the NVIDIA GPUs in the OpenShift 4 cluster. Wait at least 10-20 minutes before digging deeper into any form of troubleshooting because this may take a period of time to finish. + At this point, the GPU Operator proceeds and installs all the required components to set up the NVIDIA GPUs in the OpenShift 4 cluster. Wait at least 10 to 20 minutes before troubleshooting because this process can take some time to finish. #. The status of the newly deployed ClusterPolicy *gpu-cluster-policy* for the NVIDIA GPU Operator changes to ``State:ready`` when the installation succeeds. @@ -237,7 +237,7 @@ Create the cluster policy using the CLI $ oc get csv -n nvidia-gpu-operator gpu-operator-certified.v22.9.0 -ojsonpath={.metadata.annotations.alm-examples} | jq .[0] > clusterpolicy.json - .. note:: For OpenShift 4.12 with GPU Operator 25.3.1 or later, modify clusterpolicy.json file to specify ``driver.licensingConfig``, ``driver.repository``, ``driver.image``, ``driver.version`` and ``driver.imagePullSecrets`` (optional). The below snippet is shown as an example, please change values accordingly. Refer to :ref:`operator-release-notes` for recommended driver versions. + .. note:: For OpenShift 4.12 with GPU Operator 25.3.1 or later, modify the clusterpolicy.json file to specify ``driver.licensingConfig``, ``driver.repository``, ``driver.image``, ``driver.version``, and ``driver.imagePullSecrets`` (optional). The following snippet is shown as an example. Change values accordingly. Refer to :ref:`operator-release-notes` for recommended driver versions. .. code-block:: json @@ -262,7 +262,7 @@ Create the ClusterPolicy instance with NVIDIA vGPU Pre-requisites -------------- -* Please refer to :ref:`install-gpu-operator-vgpu` section for pre-requisite steps for using NVIDIA vGPU on RedHat OpenShift. +* Refer to the :ref:`install-gpu-operator-vgpu` section for prerequisite steps for using NVIDIA vGPU on Red Hat OpenShift. Create the cluster policy using the web console ----------------------------------------------- @@ -271,17 +271,17 @@ Create the cluster policy using the web console #. Select the **ClusterPolicy** tab, then click **Create ClusterPolicy**. The platform assigns the default name *gpu-cluster-policy*. -#. Provide name of the licensing ``ConfigMap`` under **Driver** section, this should be created during pre-requsite steps above for NVIDIA vGPU. Refer to below screenshots for example and modify values accordingly. +#. Provide the name of the licensing ``ConfigMap`` under the **Driver** section. This should be created during the prerequisite steps for NVIDIA vGPU. Refer to the following screenshots for examples and modify values accordingly. .. image:: graphics/cluster_policy_vgpu_1.png -#. Specify ``repository`` path, ``image`` name and NVIDIA vGPU driver ``version`` bundled under **Driver** section. If the registry is not public, please specify the ``imagePullSecret`` created during pre-requisite step under **Driver** advanced configurations section. +#. Specify the ``repository`` path, ``image`` name, and NVIDIA vGPU driver ``version`` bundled under the **Driver** section. If the registry is not public, specify the ``imagePullSecret`` created during the prerequisite step under the **Driver** advanced configurations section. .. image:: graphics/cluster_policy_vgpu_2.png #. Click **Create**. - At this point, the GPU Operator proceeds and installs all the required components to set up the NVIDIA GPUs in the OpenShift 4 cluster. Wait at least 10-20 minutes before digging deeper into any form of troubleshooting because this may take a period of time to finish. + At this point, the GPU Operator proceeds and installs all the required components to set up the NVIDIA GPUs in the OpenShift 4 cluster. Wait at least 10 to 20 minutes before troubleshooting because this process can take some time to finish. #. The status of the newly deployed ClusterPolicy *gpu-cluster-policy* for the NVIDIA GPU Operator changes to ``State:ready`` when the installation succeeds. @@ -297,7 +297,7 @@ Create the cluster policy using the CLI $ oc get csv -n nvidia-gpu-operator gpu-operator-certified.v22.9.0 -ojsonpath={.metadata.annotations.alm-examples} | jq .[0] > clusterpolicy.json - Modify clusterpolicy.json file to specify ``driver.licensingConfig``, ``driver.repository``, ``driver.image``, ``driver.version`` and ``driver.imagePullSecrets`` created during pre-requiste steps. Below snippet is shown as an example, please change values accordingly. + Modify the clusterpolicy.json file to specify ``driver.licensingConfig``, ``driver.repository``, ``driver.image``, ``driver.version``, and ``driver.imagePullSecrets`` created during the prerequisite steps. The following snippet is shown as an example. Change values accordingly. .. code-block:: json @@ -372,7 +372,7 @@ The GPU Operator generates GPU performance metrics (DCGM-export), status metrics When the GPU Operator is installed in the suggested ``nvidia-gpu-operator`` namespace, the GPU Operator automatically enables monitoring if the ``openshift.io/cluster-monitoring`` label is not defined. If the label is defined, the GPU Operator will not change its value. -Disable cluster monitoring in the ``nvidia-gpu-operator`` namespace by setting ``openshift.io/cluster-monitoring=false`` as shown: +Disable cluster monitoring in the ``nvidia-gpu-operator`` namespace by setting ``openshift.io/cluster-monitoring=false``: .. code-block:: console @@ -414,7 +414,7 @@ The ``nvidia-driver-daemonset`` pod has two containers. Running a sample GPU Application ************************************************************* -Run a simple CUDA VectorAdd sample, which adds two vectors together to ensure the GPUs have bootstrapped correctly. +Run a simple CUDA VectorAdd sample that adds two vectors together to ensure the GPUs have bootstrapped correctly. #. Run the following: @@ -459,7 +459,7 @@ Run a simple CUDA VectorAdd sample, which adds two vectors together to ensure th Getting information about the GPU ************************************************************* -The ``nvidia-smi`` shows memory usage, GPU utilization, and the temperature of the GPU. Test the GPU access by running the popular ``nvidia-smi`` command within the pod. +The ``nvidia-smi`` command shows memory usage, GPU utilization, and the temperature of the GPU. Test the GPU access by running the popular ``nvidia-smi`` command within the pod. To view GPU utilization, run ``nvidia-smi`` from a pod in the GPU Operator daemonset. @@ -481,7 +481,7 @@ To view GPU utilization, run ``nvidia-smi`` from a pod in the GPU Operator daemo nvidia-driver-daemonset-410.84.202203290245-0-xxgdv 2/2 Running 0 23m 10.130.2.18 ip-10-0-143-147.ec2.internal - .. note:: With the Pod and node name, run the ``nvidia-smi`` on the correct node. + .. note:: With the pod and node name, run the ``nvidia-smi`` command on the correct node. #. Run the ``nvidia-smi`` command within the pod: @@ -513,6 +513,6 @@ To view GPU utilization, run ``nvidia-smi`` from a pod in the GPU Operator daemo | No running processes found | +-----------------------------------------------------------------------------+ - Two tables are generated. The first table reflects the information about all available GPUs (the example shows one GPU). The second table provides details on the processes using the GPUs. + Two tables are generated. The first table reflects the information about all available GPUs (the example shows one GPU). The second table provides details about the processes using the GPUs. - For more information describing the contents of the tables see the man page for ``nvidia-smi``. + For more information describing the contents of the tables, refer to the man page for ``nvidia-smi``. diff --git a/openshift/install-nfd.rst b/openshift/install-nfd.rst index 146c05169..2f424373a 100644 --- a/openshift/install-nfd.rst +++ b/openshift/install-nfd.rst @@ -11,9 +11,9 @@ Installing the Node Feature Discovery Operator on OpenShift Procedure ********* -The Node Feature Discovery (NFD) Operator is a prerequisite for the **NVIDIA GPU Operator**. Install the NFD Operator using the Red Hat OperatorHub catalog in the OpenShift Container Platform web console . +The Node Feature Discovery (NFD) Operator is a prerequisite for the **NVIDIA GPU Operator**. Install the NFD Operator using the Red Hat OperatorHub catalog in the OpenShift Container Platform web console. -#. Follow the Red Hat documentation guidance in `The Node Feature Discovery Operator `_ to install the Node Feature Discovery Operator. +#. Follow the Red Hat documentation guidance in the `Node Feature Discovery Operator guide `_ to install the Node Feature Discovery Operator. #. Verify the Node Feature Discovery Operator is running: @@ -26,19 +26,19 @@ The Node Feature Discovery (NFD) Operator is a prerequisite for the **NVIDIA GPU NAME READY STATUS RESTARTS AGE nfd-controller-manager-7f86ccfb58-nqgxm 2/2 Running 0 11m -#. When the Node Feature Discovery is installed, create an instance of Node Feature Discovery using the **NodeFeatureDiscovery** tab. +#. When the Node Feature Discovery is installed, create an instance of Node Feature Discovery using the **NodeFeatureDiscovery** tab: - #. Click **Operators** > **Installed Operators** from the side menu. + #. Click **Operators** > **Installed Operators** from the side menu. - #. Find the **Node Feature Discovery** entry. + #. Find the **Node Feature Discovery** entry. - #. Click **NodeFeatureDiscovery** under the **Provided APIs** field. + #. Click **NodeFeatureDiscovery** under the **Provided APIs** field. - #. Click **Create NodeFeatureDiscovery**. + #. Click **Create NodeFeatureDiscovery**. - #. In the subsequent screen click **Create**. This starts the Node Feature Discovery Operator that proceeds to label the nodes in the cluster that have GPUs. + #. In the following screen, click **Create**. This starts the Node Feature Discovery Operator that proceeds to label the nodes in the cluster that have GPUs. - .. note:: The values pre-populated by the OperatorHub are valid for the GPU Operator. + .. note:: The values prepopulated by the OperatorHub are valid for the GPU Operator. ************************************************************************* Verify that the Node Feature Discovery Operator is functioning correctly @@ -49,19 +49,19 @@ The Node Feature Discovery Operator uses vendor PCI IDs to identify hardware in #. In the OpenShift Container Platform web console, click **Compute** > **Nodes** from the side menu. -#. Select a worker node that you know contains a GPU. +#. Select a worker node that contains a GPU. #. Click the **Details** tab. -#. Under **Node labels** verify that the following label is present: +#. Under **Node Labels**, verify that the following label is present: .. code-block:: console feature.node.kubernetes.io/pci-10de.present=true - .. note:: ``0x10de`` is the PCI vendor ID that is assigned to NVIDIA. + .. note:: ``0x10de`` is the PCI vendor ID assigned to NVIDIA. -#. Verify the GPU device (``pci-10de``) is discovered on the GPU node: +#. Verify that the GPU device (``pci-10de``) is discovered on the GPU node: .. code-block:: console diff --git a/openshift/introduction.rst b/openshift/introduction.rst index c929ce5a0..be25e3929 100644 --- a/openshift/introduction.rst +++ b/openshift/introduction.rst @@ -13,11 +13,11 @@ Introduction to NVIDIA GPU Operator on OpenShift Kubernetes is an open-source platform for automating the deployment, scaling, and managing of containerized applications. Red Hat OpenShift Container Platform is a security-centric and enterprise-grade hardened Kubernetes platform for deploying and managing Kubernetes clusters at scale, developed and supported by Red Hat. -Red Hat OpenShift Container Platform includes enhancements to Kubernetes so users can easily configure and use GPU resources for accelerating workloads such as deep learning. +Red Hat OpenShift Container Platform includes enhancements to Kubernetes so users can easily configure and use GPU resources for accelerating workloads like deep learning. The NVIDIA GPU Operator uses the operator framework within Kubernetes to automate the management of all NVIDIA software components needed to provision GPU. These components include the NVIDIA drivers (to enable CUDA), Kubernetes device plugin for GPUs, the `NVIDIA Container Toolkit `_, -automatic node labelling using `GFD `_, `DCGM `_ based monitoring and others. +automatic node labeling using `GFD `_, `DCGM `_-based monitoring, and others. For guidance on the specific NVIDIA support entitlement needs, refer |essug|_ if you have an NVIDIA AI Enterprise entitlement. diff --git a/openshift/openshift-virtualization.rst b/openshift/openshift-virtualization.rst index 4334463c2..25e0e8017 100644 --- a/openshift/openshift-virtualization.rst +++ b/openshift/openshift-virtualization.rst @@ -22,7 +22,7 @@ Red Hat OpenShift Virtualization is an OpenShift feature to run virtual machines In addition to the GPU Operator being able to provision worker nodes for running GPU-accelerated containers, the GPU Operator can also be used to provision worker nodes for running GPU-accelerated virtual machines. -There are some different prerequisites required virtual machines with GPU(s) than running containers with GPU(s). +There are some different prerequisites required when running virtual machines with GPUs compared to running containers with GPUs. The primary difference is the drivers required. For example, the datacenter driver is needed for containers, the vfio-pci driver is needed for GPU passthrough, and the `NVIDIA vGPU Manager `_ is needed for creating vGPU devices. @@ -87,7 +87,7 @@ Assumptions, constraints, and dependencies * A worker node can run GPU-accelerated containers, or GPU accelerated VMs with GPU passthrough, or GPU accelerated-VMs with vGPU, but not a combination of any of them. -* The cluster admin or developer has knowledge about their cluster ahead of time, and can properly label nodes to indicate what types of GPU workloads they will run. +* The cluster admin or developer has knowledge about their cluster ahead of time and can properly label nodes to indicate what types of GPU workloads they will run. * Worker nodes running GPU accelerated VMs (with pGPU or vGPU) are assumed to be bare metal. @@ -118,7 +118,7 @@ Prerequisites hyperconverged.hco.kubevirt.io/kubevirt-hyperconverged patched -* If planning to use NVIDIA vGPU, SR-IOV must be enabled in the BIOS if your GPUs are based on the NVIDIA Ampere architecture or later. Refer to the `NVIDIA vGPU Documentation `_ to ensure you have met all of the prerequisites for using NVIDIA vGPU. +* If planning to use NVIDIA vGPU, SR-IOV must be enabled in the BIOS if your GPUs are based on the NVIDIA Ampere architecture or later. Refer to the `NVIDIA vGPU Documentation `_ to ensure you have met all the prerequisites for using NVIDIA vGPU. *********************************************************** Configure NVIDIA GPU Operator with OpenShift Virtualization @@ -365,7 +365,7 @@ Create the cluster policy using the CLI: In general, the flag ``sandboxWorkloads.enabled`` in ``ClusterPolicy`` controls whether the GPU Operator can provision GPU worker nodes for virtual machine workloads, in addition to container workloads. This flag is disabled by default, meaning all nodes get provisioned with the same software which enables container workloads, and the ``nvidia.com/gpu.workload.config`` node label is not used. - The term ``sandboxing`` refers to running software in a separate isolated environment, typically for added security (i.e. a virtual machine). We use the term ``sandbox workloads`` to signify workloads that run in a virtual machine, irrespective of the virtualization technology used. + The term ``sandboxing`` refers to running software in a separate isolated environment, typically for added security (that is, a virtual machine). We use the term ``sandbox workloads`` to signify workloads that run in a virtual machine, irrespective of the virtualization technology used. #. Apply the changes: @@ -406,7 +406,7 @@ As a cluster administrator, you can create a ClusterPolicy using the OpenShift C In general, when sandbox workloads are enabled, ``ClusterPolicy`` controls whether the GPU Operator can provision GPU worker nodes for virtual machine workloads, in addition to container workloads. This flag is disabled by default, meaning all nodes get provisioned with the same software which enables container workloads, and the ``nvidia.com/gpu.workload.config`` node label is not used. - The term ``sandboxing`` refers to running software in a separate isolated environment, typically for added security (i.e. a virtual machine). We use the term ``sandbox workloads`` to signify workloads that run in a virtual machine, irrespective of the virtualization technology used. + The term ``sandboxing`` refers to running software in a separate isolated environment, typically for added security (that is, a virtual machine). We use the term ``sandbox workloads`` to signify workloads that run in a virtual machine, irrespective of the virtualization technology used. * Click **Create** to create the ClusterPolicy. .. image:: graphics/cluster_policy_enable_sandbox_workloads.png diff --git a/openshift/prerequisites.rst b/openshift/prerequisites.rst index 1623c36c1..f78c4ca0d 100644 --- a/openshift/prerequisites.rst +++ b/openshift/prerequisites.rst @@ -7,7 +7,7 @@ Prerequisites for GPU Operator on OpenShift Before following the steps in this guide, ensure that your environment has: -* A working OpenShift cluster up and running with a GPU worker node. See `OpenShift Container Platform installation `_ for guidance on installing. +* A working OpenShift cluster up and running with a GPU worker node. Refer to the `OpenShift Container Platform installation `_ guide for installation guidance. Refer to :external+gpuop:ref:`Container Platforms ` for the support matrix of the GPU Operator releases and the supported container platforms for more information. * Access to the OpenShift cluster as a ``cluster-admin`` to perform the necessary steps. * OpenShift CLI (``oc``) installed. From 81e27908780b9a88d3ce2be001551d3b36c42bbf Mon Sep 17 00:00:00 2001 From: Andrew Chen Date: Wed, 10 Sep 2025 13:45:23 -0700 Subject: [PATCH 3/4] add reference links to RH docs Signed-off-by: Andrew Chen --- openshift/gpu-operator-with-precompiled-drivers.rst | 4 ++++ 1 file changed, 4 insertions(+) diff --git a/openshift/gpu-operator-with-precompiled-drivers.rst b/openshift/gpu-operator-with-precompiled-drivers.rst index bb23acbd2..747709d60 100644 --- a/openshift/gpu-operator-with-precompiled-drivers.rst +++ b/openshift/gpu-operator-with-precompiled-drivers.rst @@ -126,6 +126,10 @@ Perform the following steps to build a custom driver image for use with Red Hat - Before OCP 4.19: The driver image tag is formed with the suffix ``-rhcos4.17`` (such as with OCP 4.17). - Starting OCP 4.19 and later: The driver image tag is formed with the suffix ``-rhel9.6`` (such as with OCP 4.19). + Refer to `RHEL Versions Utilized by RHEL CoreOS and OCP `_ + and `Split RHCOS into layers: /etc/os-release `_ + for more information. + #. Build and push the image: .. code-block:: console From ca8577a9a845a63b3a6da5fb3679cd9d390b002d Mon Sep 17 00:00:00 2001 From: Andrew Chen Date: Thu, 18 Sep 2025 10:55:33 -0700 Subject: [PATCH 4/4] accept suggestions Signed-off-by: Andrew Chen --- openshift/gpu-operator-with-precompiled-drivers.rst | 2 +- openshift/install-gpu-ocp.rst | 4 ++-- openshift/openshift-virtualization.rst | 2 +- 3 files changed, 4 insertions(+), 4 deletions(-) diff --git a/openshift/gpu-operator-with-precompiled-drivers.rst b/openshift/gpu-operator-with-precompiled-drivers.rst index 747709d60..3777a4ed4 100644 --- a/openshift/gpu-operator-with-precompiled-drivers.rst +++ b/openshift/gpu-operator-with-precompiled-drivers.rst @@ -121,7 +121,7 @@ Perform the following steps to build a custom driver image for use with Red Hat export DRIVER_VERSION=525.105.17 export OS_TAG=rhcos4.12 - .. note:: The driver container image tag for OpenShift has changed after the ``OCP 4.19`` release. + .. note:: The driver container image tag for OpenShift has changed after the OCP 4.19 release. - Before OCP 4.19: The driver image tag is formed with the suffix ``-rhcos4.17`` (such as with OCP 4.17). - Starting OCP 4.19 and later: The driver image tag is formed with the suffix ``-rhel9.6`` (such as with OCP 4.19). diff --git a/openshift/install-gpu-ocp.rst b/openshift/install-gpu-ocp.rst index 9b17ab08e..bf86b7f67 100644 --- a/openshift/install-gpu-ocp.rst +++ b/openshift/install-gpu-ocp.rst @@ -19,7 +19,7 @@ Installing the NVIDIA GPU Operator by using the web console .. note:: Here, you can select the namespace where you want to deploy the GPU Operator. The suggested namespace to use is the ``nvidia-gpu-operator``. You can choose any existing namespace or create a new namespace under **Select a Namespace**. - If you install in any other namespace other than ``nvidia-gpu-operator``, the GPU Operator will **not** automatically enable namespace monitoring, and metrics and alerts will **not** be collected by Prometheus. + If you install in any other namespace other than ``nvidia-gpu-operator``, the GPU Operator does **not** automatically enable namespace monitoring, and metrics and alerts are **not** collected by Prometheus. If only trusted operators are installed in this namespace, you can manually enable namespace monitoring with this command: .. code-block:: console @@ -297,7 +297,7 @@ Create the cluster policy using the CLI $ oc get csv -n nvidia-gpu-operator gpu-operator-certified.v22.9.0 -ojsonpath={.metadata.annotations.alm-examples} | jq .[0] > clusterpolicy.json - Modify the clusterpolicy.json file to specify ``driver.licensingConfig``, ``driver.repository``, ``driver.image``, ``driver.version``, and ``driver.imagePullSecrets`` created during the prerequisite steps. The following snippet is shown as an example. Change values accordingly. + Modify the ``clusterpolicy.json`` file to specify ``driver.licensingConfig``, ``driver.repository``, ``driver.image``, ``driver.version``, and ``driver.imagePullSecrets`` created during the prerequisite steps. The following snippet is shown as an example. Change values accordingly. .. code-block:: json diff --git a/openshift/openshift-virtualization.rst b/openshift/openshift-virtualization.rst index 25e0e8017..7cf5313f7 100644 --- a/openshift/openshift-virtualization.rst +++ b/openshift/openshift-virtualization.rst @@ -365,7 +365,7 @@ Create the cluster policy using the CLI: In general, the flag ``sandboxWorkloads.enabled`` in ``ClusterPolicy`` controls whether the GPU Operator can provision GPU worker nodes for virtual machine workloads, in addition to container workloads. This flag is disabled by default, meaning all nodes get provisioned with the same software which enables container workloads, and the ``nvidia.com/gpu.workload.config`` node label is not used. - The term ``sandboxing`` refers to running software in a separate isolated environment, typically for added security (that is, a virtual machine). We use the term ``sandbox workloads`` to signify workloads that run in a virtual machine, irrespective of the virtualization technology used. + The term *sandboxing* refers to running software in a separate isolated environment, typically for added security (that is, a virtual machine). We use the term ``sandbox workloads`` to signify workloads that run in a virtual machine, irrespective of the virtualization technology used. #. Apply the changes: