The ingress mixins for RL (collective_rpc, sleepable, pausable) call engine-protocol methods that SGLangServer doesn't implement. /collective_rpc, /sleep, /pause fail on SGLang deployments. This issue tracks implementing those methods plus an ObjectRef-based weight-transfer path and session affinity for multi-turn rollouts.
Subitems
Open questions
- Whether the inference-trainer handoff becomes a documented first-class contract for RL frameworks, or stays implicit in the mixin surface.
- Where session affinity lives: Serve routing, engine-level session manager, or both.
- Weight-sync latency budget. Drives whether broadcast-and-reload is acceptable or delta sync is needed.
- Composition with PD disaggregation (separate rollout-decode, shared prefill pool): natural fallout or dedicated design pass?
Upstream coordination
update_weights_from_object_ref-style API accepting Ray references directly.
abort_request acknowledgement so cancellation is deterministic.
The ingress mixins for RL (
collective_rpc,sleepable,pausable) call engine-protocol methods thatSGLangServerdoesn't implement./collective_rpc,/sleep,/pausefail on SGLang deployments. This issue tracks implementing those methods plus an ObjectRef-based weight-transfer path and session affinity for multi-turn rollouts.Subitems
SGLangServerimplementscollective_rpc/sleep/wakeup/is_sleeping/pause/resume/is_paused.Open questions
Upstream coordination
update_weights_from_object_ref-style API accepting Ray references directly.abort_requestacknowledgement so cancellation is deterministic.