Thanks for publishing the Qwen3.8-Flash-Next community image and Docker Compose recipe.
Would it be possible to support multiple independent TP1 Qwen3.8-Flash-Next replicas on the same node sharing one CPU-offloaded PLE N-gram table?
The use case is multiple NVIDIA RTX PRO 6000 GPUs, with one complete model replica on each GPU. This is different from TP2: TP2 splits one model across two GPUs, while we would like separately schedulable TP1 replicas for throughput and availability.
With VLLM_PLE_CPU_OFFLOAD=1, each TP1 replica currently needs its own CUDA-mapped host-memory PLE allocation. Sharing the immutable table between same-node replicas, while keeping each GPU's model weights and KV cache independent, could make this configuration much more memory-efficient.
The current recipe is Docker Compose based. It would also be very useful if any solution could work cleanly for Kubernetes deployments, where multiple TP1 replicas may run as separate pods pinned to GPUs on the same node.
Is this technically feasible, and is it something you would consider for a future recipe/runtime update?
Thank you.
Thanks for publishing the Qwen3.8-Flash-Next community image and Docker Compose recipe.
Would it be possible to support multiple independent TP1 Qwen3.8-Flash-Next replicas on the same node sharing one CPU-offloaded PLE N-gram table?
The use case is multiple NVIDIA RTX PRO 6000 GPUs, with one complete model replica on each GPU. This is different from TP2: TP2 splits one model across two GPUs, while we would like separately schedulable TP1 replicas for throughput and availability.
With
VLLM_PLE_CPU_OFFLOAD=1, each TP1 replica currently needs its own CUDA-mapped host-memory PLE allocation. Sharing the immutable table between same-node replicas, while keeping each GPU's model weights and KV cache independent, could make this configuration much more memory-efficient.The current recipe is Docker Compose based. It would also be very useful if any solution could work cleanly for Kubernetes deployments, where multiple TP1 replicas may run as separate pods pinned to GPUs on the same node.
Is this technically feasible, and is it something you would consider for a future recipe/runtime update?
Thank you.