From c3f61be6efb74af6c0932fb6141d0d75b9eb8842 Mon Sep 17 00:00:00 2001 From: =?UTF-8?q?Efe=20G=C3=B6kdemir?= Date: Tue, 22 Sep 2026 01:59:36 +0300 Subject: [PATCH] docs(gpu-operator): clarify MPS support MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Signed-off-by: Efe Gökdemir --- gpu-operator/getting-started.rst | 5 +++-- gpu-operator/gpu-sharing.rst | 11 +++++++++++ 2 files changed, 14 insertions(+), 2 deletions(-) diff --git a/gpu-operator/getting-started.rst b/gpu-operator/getting-started.rst index cd0428c9f..519eafbe0 100644 --- a/gpu-operator/getting-started.rst +++ b/gpu-operator/getting-started.rst @@ -917,10 +917,11 @@ Next Steps After verifying the installation, you can configure the GPU Operator for your workloads: -- :doc:`gpu-sharing` — Share a single GPU across multiple pods using time-slicing or MPS. +- :doc:`gpu-sharing` — Share a single GPU across multiple pods using time-slicing. - :doc:`gpu-operator-mig` — Configure Multi-Instance GPU (MIG) partitioning on supported GPUs. - :doc:`gpu-operator-rdma` — Enable GPUDirect RDMA for high-performance networking. -- :doc:`dra-intro-install` — Allocate GPUs by using Kubernetes Dynamic Resource Allocation (DRA). +- :doc:`dra-intro-install` — Allocate GPUs by using Kubernetes Dynamic Resource Allocation (DRA), + including experimental MPS sharing. - :doc:`nvidia-driver-configuration` — Use the NVIDIA GPU Driver Custom Resource Definition to manage drivers per node. - :doc:`precompiled-drivers` — Speed up driver deployments with precompiled kernel modules. - :doc:`cdi` — Learn about Container Device Interface (CDI) and NRI Plugin mode. diff --git a/gpu-operator/gpu-sharing.rst b/gpu-operator/gpu-sharing.rst index 0c2ff5882..fcf744500 100644 --- a/gpu-operator/gpu-sharing.rst +++ b/gpu-operator/gpu-sharing.rst @@ -19,6 +19,17 @@ of extended options for the `NVIDIA Kubernetes Device Plugin `__ + for application considerations. Experimental MPS support through the DRA + driver requires the Alpha ``MPSSupport`` feature gate. Refer to + :doc:`dra-intro-install` for the DRA feature gates and their compatibility + constraints. + This mechanism for enabling *time-slicing* of GPUs in Kubernetes enables a system administrator to define a set of *replicas* for a GPU, each of which can be handed out independently to a