[core] Back off failed placement group recovery retries - #65970
hikaru-212 wants to merge 1 commit into
Conversation
There was a problem hiding this comment.
Code Review
This pull request modifies the placement group rescheduling logic in GcsPlacementGroupManager to apply a backoff delay to failed recovery retries instead of always scheduling them with the highest priority (rank 0). This prevents failed retries from starving other pending placement groups. Additionally, the test suite is updated to thoroughly verify this backoff behavior and ensure that normal placement groups can be scheduled during a recovery cooldown. I have no feedback to provide.
|
This pull request has been automatically marked as stale because it has not had You can always ask for help on our discussion forum or Ray's public slack channel. If you'd like to keep this open, just leave any comment, and the stale label will be removed. |
Signed-off-by: Yen-Hua Chen <226400984+hikaru-212@users.noreply.github.com>
676def3 to
cd65fc9
Compare
|
Thanks for the contribution! But I think the original code was designed this way on purpose: always letting the rescheduled PG be scheduled with top priority Your change does let other PGs get scheduled, but it also makes the rescheduled PG less likely to get scheduled, so it's actually a tradeoff |
|
Thanks, I agree this is a tradeoff. |
Description
This PR addresses a scheduler-liveness issue observed while investigating #65147.
Fresh placement-group recovery triggered by node loss is intentionally scheduled at the highest priority. However, when a feasible
RESCHEDULINGattempt fails, the current path requeues the placement group at rank0again with fresh retry state.Repeated recovery failures can therefore continue taking the highest-priority scheduling turn and prevent unrelated, already-due pending placement groups from reaching the scheduler.
This change preserves rank
0for the initial recovery attempt, but reuses the existingExponentialBackoffafter a feasibleRESCHEDULINGattempt fails.As a result:
The focused regression exercises the full manager-level transition from a successfully created placement group through node loss and
RESCHEDULING. It verifies that the initial recovery attempt wins over a normal pending placement group, but after recovery fails, the pending placement group receives a scheduling attempt while recovery is in backoff.The same regression also verifies that recovery becomes eligible again when its retry time is reached and that a subsequent failure receives a larger retry delay.
This does not resolve the broader complementary partial-placement-group resource cycle described in #65147. It does not introduce victim selection, disruption, cycle detection, or a rescheduler policy.
Related issues
Related to #65147
Additional information
The focused regression was validated as a RED → GREEN test:
RESCHEDULINGattempt still hadhighest_retry_delay_ms == 0;Local validation:
GcsPlacementGroupManagerMockTest.PendingQueuePriorityReschedule: 10/10 runs passed.test_placement_group_failovertarget: 6 tests passed.For the source-built validation, the Ray Python package and native artifacts were generated from the corresponding checkout. The tests used the checkout-specific GCS server, raylet,
_raylet, and Redis artifacts.