diff --git a/gpu-operator/getting-started.rst b/gpu-operator/getting-started.rst index cd0428c9f..519eafbe0 100644 --- a/gpu-operator/getting-started.rst +++ b/gpu-operator/getting-started.rst @@ -917,10 +917,11 @@ Next Steps After verifying the installation, you can configure the GPU Operator for your workloads: -- :doc:`gpu-sharing` — Share a single GPU across multiple pods using time-slicing or MPS. +- :doc:`gpu-sharing` — Share a single GPU across multiple pods using time-slicing. - :doc:`gpu-operator-mig` — Configure Multi-Instance GPU (MIG) partitioning on supported GPUs. - :doc:`gpu-operator-rdma` — Enable GPUDirect RDMA for high-performance networking. -- :doc:`dra-intro-install` — Allocate GPUs by using Kubernetes Dynamic Resource Allocation (DRA). +- :doc:`dra-intro-install` — Allocate GPUs by using Kubernetes Dynamic Resource Allocation (DRA), + including experimental MPS sharing. - :doc:`nvidia-driver-configuration` — Use the NVIDIA GPU Driver Custom Resource Definition to manage drivers per node. - :doc:`precompiled-drivers` — Speed up driver deployments with precompiled kernel modules. - :doc:`cdi` — Learn about Container Device Interface (CDI) and NRI Plugin mode. diff --git a/gpu-operator/gpu-sharing.rst b/gpu-operator/gpu-sharing.rst index 0c2ff5882..fcf744500 100644 --- a/gpu-operator/gpu-sharing.rst +++ b/gpu-operator/gpu-sharing.rst @@ -19,6 +19,17 @@ of extended options for the `NVIDIA Kubernetes Device Plugin `__ + for application considerations. Experimental MPS support through the DRA + driver requires the Alpha ``MPSSupport`` feature gate. Refer to + :doc:`dra-intro-install` for the DRA feature gates and their compatibility + constraints. + This mechanism for enabling *time-slicing* of GPUs in Kubernetes enables a system administrator to define a set of *replicas* for a GPU, each of which can be handed out independently to a