Skip to content

[Serve][LLM][SGLang] RL weight sync and inference-trainer coordination #62794

Description

@eicherseiji

The ingress mixins for RL (collective_rpc, sleepable, pausable) call engine-protocol methods that SGLangServer doesn't implement. /collective_rpc, /sleep, /pause fail on SGLang deployments. This issue tracks implementing those methods plus an ObjectRef-based weight-transfer path and session affinity for multi-turn rollouts.

Subitems

  • SGLangServer implements collective_rpc / sleep / wakeup / is_sleeping / pause / resume / is_paused.
  • ObjectRef-based weight-transfer wrapper for SGLang.
  • Session / multi-turn affinity for SGLang: routing pins follow-up requests to the same replica, exposed through a small SGLang-scoped session API.
  • SGLang-specific reference RL driver exercising deploy → open session → stream → pause → sync → resume.
  • SGLang RL weight-sync release test: bounded sync latency, no requests stuck across a sync boundary.

Open questions

  • Whether the inference-trainer handoff becomes a documented first-class contract for RL frameworks, or stays implicit in the mixin surface.
  • Where session affinity lives: Serve routing, engine-level session manager, or both.
  • Weight-sync latency budget. Drives whether broadcast-and-reload is acceptable or delta sync is needed.
  • Composition with PD disaggregation (separate rollout-decode, shared prefill pool): natural fallout or dedicated design pass?

Upstream coordination

  • update_weights_from_object_ref-style API accepting Ray references directly.
  • abort_request acknowledgement so cancellation is deterministic.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    llmserveRay Serve Related Issue

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions