Skip to content

[Serve LLM] Fix SGLang sleep state after empty-tag wakeup - #65969

Open
0z5a wants to merge 1 commit into
ray-project:masterfrom
0z5a:fix/sglang-empty-wakeup-state
Open

0z5a wants to merge 1 commit into
ray-project:masterfrom
0z5a:fix/sglang-empty-wakeup-state

Conversation

@0z5a

@0z5a 0z5a commented Sep 7, 2026 •

Copy link
Copy Markdown

Incremental review: Files changed against master.

Description

Clear Ray's tracked sleep state after an acknowledged wakeup(tags=[]).
SGLang treats both None and an empty list as all memory regions, but Ray
previously cleared _sleeping_tags only for None. This left is_sleeping()
true after a successful full wakeup.

The production change is the all-tags condition plus its config documentation.
The existing SGLang release suite gains acknowledgement/error/cancellation
contracts and nine omitted/None/empty sleep-wakeup cycles. The GPU fixture now
configures TorchMemorySaver through the public configure_subprocess() helper
and Ray actor runtime_env, enabling actual release/restoration checks.

Related issues

Related to #62794; follows the control-plane implementation merged in #63021.
This does not implement ObjectRef weight transfer, sessions, an RL driver, or
the separate P/D builder in #63741.

Non-duplication: reread #62794's comments and searched open PRs for
62794 in:body
and sglang wakeup
immediately before submission; no same-fix PR was found. The bounded scope and
validation were already described in
the issue discussion.

Additional information

Tests

CPU contract selection (real Ray class/SGLang request types; only backend
tokenizer-manager calls mocked):

python -m pytest release/llm_tests/serve/test_llm_serve_sglang.py \
  -k TestSGLangSleepState -q --tb=short
  • Original: 4 failed / 20 passed. Fixed, final-source Linux: 24 passed, zero skips.
  • Real single-L20, Qwen2.5-0.5B-Instruct controls: original 4 failed / 10 passed;
    fixed 14 passed, zero skips. Three original failures are empty-tag wakeups;
    the fourth is the subsequent selective-tag test observing their stale state.
  • Expanded final-source GPU selection: 19 passed, zero skips:
python -m pytest release/llm_tests/serve/test_llm_serve_sglang.py \
  -k '(sleep or pause or reset or test_sglang_serve_e2e or streaming or tokenize or batched) and not TestSGLangSleepState and not multi and not pipeline' \
  -q --timeout=600

The GPU harness initializes Ray with one GPU and substitutes a verified offline
checkpoint path before invoking this selection. All nine lifecycle cases check
at least 256 MiB release/restoration and exact greedy-output equality. Recorded
same-GPU process totals were approximately 38,384 -> 1,732 -> 38,384 MiB; these
are not allocator-level measurements. No backend calls are mocked in this run.

pre-commit run --files \
  python/ray/llm/_internal/serve/engines/sglang/sglang_engine.py \
  release/llm_tests/serve/test_llm_serve_sglang.py

All applicable hooks passed on the final files. Validation uses Ray's documented
Python-only development setup: core wheel dede511b61fbf6383f5b03cdcc1fa2c0128efafa
with the checkout's LLM Python source, not a full Ray C++ source build. GPU
environment: Torch 2.13.0+cu130, SGLang 0.5.19, Transformers 5.12.1; model revision
7ae557604adf67be50417f59c2c2f167def9a775. Driver/container LD_PRELOAD was unset;
the fixture configures the actor environment. Initial dependency/preload setup
failures and an aborted macOS import-I/O rerun are not counted as passing tests
or bug reproductions. Multi-GPU/RL/performance qualification is outside scope.

AI assistance and human accountability

OpenAI Codex assisted with the audit, implementation, tests, and submission.
The human submitter confirmed reviewing every changed line and personally
running the relevant tests for commit 1c5ea3ee38ed10ec4eb84becc81e9cbeb9293135
before this PR was opened. The commit's earlier pending-review wording records
its preparation-time status; the human confirmation was provided afterward.

@0z5a
0z5a requested a review from a team as a code owner September 7, 2026 05:59

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the SGLang engine to treat both omitted tags and empty lists as targeting all components during sleep and wakeup operations. It also enhances the release tests by configuring a memory saver subprocess, parametrizing the sleep/wakeup tests to verify memory usage, and adding unit tests for state contracts using mocked components. The reviewer suggested safely accessing the LD_PRELOAD environment variable using os.environ.get to avoid potential KeyError exceptions.


# Ray launches scheduler actors outside the engine's subprocess context.
with configure_subprocess():
memory_saver_env = {"LD_PRELOAD": os.environ["LD_PRELOAD"]}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Accessing os.environ["LD_PRELOAD"] directly can raise a KeyError if the environment variable is not set (for example, on non-Linux platforms or if configure_subprocess() does not set it). It is safer to use os.environ.get("LD_PRELOAD") and only populate memory_saver_env if it is present.

Suggested change
memory_saver_env = {"LD_PRELOAD": os.environ["LD_PRELOAD"]}
ld_preload = os.environ.get("LD_PRELOAD")
memory_saver_env = {"LD_PRELOAD": ld_preload} if ld_preload else {}

@ray-gardener ray-gardener Bot added serve Ray Serve Related Issue llm community-contribution Contributed by the community labels Sep 7, 2026
@github-actions

Copy link
Copy Markdown

This pull request has been automatically marked as stale because it has not had
any activity for 14 days. It will be closed in another 14 days if no further activity occurs.
Thank you for your contributions.

You can always ask for help on our discussion forum or Ray's public slack channel.

If you'd like to keep this open, just leave any comment, and the stale label will be removed.

@github-actions github-actions Bot added stale The issue is stale. It will be closed within 7 days unless there are further conversation unstale A PR that has been marked unstale. It will not get marked stale again if this label is on it. and removed stale The issue is stale. It will be closed within 7 days unless there are further conversation labels Sep 21, 2026
@0z5a
0z5a force-pushed the fix/sglang-empty-wakeup-state branch from 1c5ea3e to bb597ff Compare September 24, 2026 05:54
Keep Ray state consistent with acknowledged SGLang all-region wakeups. Cover CPU acknowledgement/failure contracts and real single-GPU sleep cycles with memory release, restoration, and deterministic output checks. The fixture configures TorchMemorySaver for Ray actors. AI-assisted implementation; human review and human-run tests remain required before requesting review.

Signed-off-by: 0z5a <0z5a@users.noreply.github.com>
Signed-off-by: 0z5a <dezhen.lu@student.uni-tuebingen.de>
@0z5a
0z5a force-pushed the fix/sglang-empty-wakeup-state branch from bb597ff to 7f68ce3 Compare September 24, 2026 07:08

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

community-contribution Contributed by the community llm serve Ray Serve Related Issue unstale A PR that has been marked unstale. It will not get marked stale again if this label is on it.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant