Skip to content

Qwen3.8-Flash-Next: could multiple TP1 replicas share the CPU-offloaded PLE table on one node? #102

Description

@renehonig

Thanks for publishing the Qwen3.8-Flash-Next community image and Docker Compose recipe.

Would it be possible to support multiple independent TP1 Qwen3.8-Flash-Next replicas on the same node sharing one CPU-offloaded PLE N-gram table?

The use case is multiple NVIDIA RTX PRO 6000 GPUs, with one complete model replica on each GPU. This is different from TP2: TP2 splits one model across two GPUs, while we would like separately schedulable TP1 replicas for throughput and availability.

With VLLM_PLE_CPU_OFFLOAD=1, each TP1 replica currently needs its own CUDA-mapped host-memory PLE allocation. Sharing the immutable table between same-node replicas, while keeping each GPU's model weights and KV cache independent, could make this configuration much more memory-efficient.

The current recipe is Docker Compose based. It would also be very useful if any solution could work cleanly for Kubernetes deployments, where multiple TP1 replicas may run as separate pods pinned to GPUs on the same node.

Is this technically feasible, and is it something you would consider for a future recipe/runtime update?

Thank you.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions