Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 17 additions & 0 deletions doc/source/serve/advanced-guides/performance.md
Original file line number Diff line number Diff line change
Expand Up @@ -98,6 +98,23 @@ Ray Serve allows you to fine-tune the backoff behavior of the request router, wh
- `RAY_SERVE_ROUTER_RETRY_BACKOFF_MULTIPLIER`: The multiplier applied to the backoff time after each retry. Default is `2`.
- `RAY_SERVE_ROUTER_RETRY_MAX_BACKOFF_S`: The maximum backoff time (in seconds) between retries. Default is `0.5`.

### Set a timeout for choosing a replica

By default, a request waits in the router until a replica accepts it. If no replica becomes available, for example because a deployment scaled to zero can't schedule a new replica, the request waits indefinitely. Set `request_routing_timeout_s` in the deployment's `request_router_config` to bound this wait:

```python
from ray import serve
from ray.serve.config import RequestRouterConfig

@serve.deployment(
request_router_config=RequestRouterConfig(request_routing_timeout_s=10),
)
class Model:
...
```

A request that isn't assigned to a replica within the timeout fails with a `TimeoutError`. The timeout doesn't cover the time the replica takes to process the request. Unlike `request_timeout_s`, it also applies to `DeploymentHandle` calls.

### Set timeouts while probing replicas for queue length

Ray Serve's request router probes replicas for their queue lengths to make intelligent load balancing decisions. You can tune the following environment variables to optimize this behavior for your workload:
Expand Down
26 changes: 25 additions & 1 deletion python/ray/serve/_private/request_router/request_router.py
Original file line number Diff line number Diff line change
Expand Up @@ -495,6 +495,7 @@ def __init__(
initial_backoff_s: float = RAY_SERVE_ROUTER_RETRY_INITIAL_BACKOFF_S,
backoff_multiplier: float = RAY_SERVE_ROUTER_RETRY_BACKOFF_MULTIPLIER,
max_backoff_s: float = RAY_SERVE_ROUTER_RETRY_MAX_BACKOFF_S,
request_routing_timeout_s: Optional[float] = None,
*args,
**kwargs,
):
Expand All @@ -510,6 +511,8 @@ def __init__(
self.backoff_multiplier = backoff_multiplier
self.max_backoff_s = max_backoff_s

self.request_routing_timeout_s = request_routing_timeout_s

# Current replicas available to be routed.
# Updated via `update_replicas`.
self._replica_id_set: Set[ReplicaID] = set()
Expand Down Expand Up @@ -1405,13 +1408,34 @@ async def _choose_replica_for_request(

self._add_pending_request_to_indices(pending_request)
self._maybe_start_routing_tasks()
replica = await pending_request.future
if self.request_routing_timeout_s is None:
replica = await pending_request.future
else:
replica = await asyncio.wait_for(
pending_request.future, self.request_routing_timeout_s
)
except asyncio.CancelledError as e:
pending_request.future.cancel()
self._cancel_routing_task_for_pending_request(pending_request)
self._remove_pending_request_from_indices(pending_request)

raise e from None
except asyncio.TimeoutError:
self._cancel_routing_task_for_pending_request(pending_request)
self._remove_pending_request_from_indices(pending_request)
# No routing task runs while the deployment has no replicas, so the
# lazy cleanup in `_fulfill_pending_requests` can't drop the request.
for queue in (
self._pending_requests_to_fulfill,
self._pending_requests_to_route,
):
while queue and queue[0].future.done():
queue.popleft()
Comment on lines +1428 to +1433

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The current cleanup logic only removes done/timed-out requests from the front of the queues (while queue and queue[0].future.done(): queue.popleft()). If there is a pending request at the front of the queue that has not timed out (e.g., because it has no timeout set), any timed-out requests behind it will remain in the queue indefinitely. This head-of-line blocking leads to a memory leak of timed-out requests and their associated arguments/metadata.

To prevent this memory leak, we should filter the queues in-place to remove all done/timed-out requests regardless of their position in the queue.

Suggested change
for queue in (
self._pending_requests_to_fulfill,
self._pending_requests_to_route,
):
while queue and queue[0].future.done():
queue.popleft()
for queue in (
self._pending_requests_to_fulfill,
self._pending_requests_to_route,
):
active_requests = [r for r in queue if not r.future.done()]
queue.clear()
queue.extend(active_requests)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With a timeout set, requests mostly expire in FIFO order, and each expiry pops the done prefix. 50 staggered timeouts with no replicas leave both deques empty. Entries can sit behind a pending head that has a later deadline: a retry (re-inserted by created_at with a fresh timeout), or a head from before the timeout was lowered. They stay only until that head resolves or expires. They're held indefinitely only if the head was enqueued before any timeout was set, and that head waits forever regardless. Cancelled requests already get this lazy cleanup today. Rebuilding both deques on every timeout would be O(n) per timeout under overload, so I kept the lazy pop.


raise TimeoutError(
f"Failed to route request to a replica of {self._deployment_id} "
f"within {self.request_routing_timeout_s}s."
) from None
Comment on lines +1435 to +1438

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

In Python versions prior to 3.11, asyncio.TimeoutError and the built-in TimeoutError are distinct exception classes (with different inheritance hierarchies). Since this is an asynchronous API and callers typically expect asyncio.TimeoutError when awaiting async operations, raising asyncio.TimeoutError is more idiomatic and ensures backward compatibility for callers catching asyncio.TimeoutError on Python 3.8 - 3.10.

Suggested change
raise TimeoutError(
f"Failed to route request to a replica of {self._deployment_id} "
f"within {self.request_routing_timeout_s}s."
) from None
raise asyncio.TimeoutError(
f"Failed to route request to a replica of {self._deployment_id} "
f"within {self.request_routing_timeout_s}s."
) from None

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Builtin TimeoutError is intentional. The HTTP/gRPC proxies map isinstance(exc, TimeoutError) to 408/DEADLINE_EXCEEDED (http_util.py, grpc_util.py), and handle.py re-raises asyncio.TimeoutError as the builtin for the same reason. On 3.10, asyncio.TimeoutError isn't a subclass of the builtin, so switching would turn the 408 into a 500.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Timeout path leaves pending request future

High Severity

The TimeoutError handler never cancels pending_request.future, unlike the CancelledError path. Queue cleanup only drops entries whose futures are already done, so a still-pending timed-out request stays in _pending_requests_to_fulfill and _pending_requests_to_route. A later routing task can still assign a replica to that abandoned request, and the request args remain queued. On Python 3.11+, asyncio.wait_for no longer directly cancels a bare Future; it cancels the calling task via asyncio.timeout, so this is not guaranteed.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit ecff197. Configure here.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

wait_for already cancels the awaited future on timeout. 3.10/3.11 call fut.cancel(). On 3.12+ asyncio.timeout cancels the task, and Task.cancel() cancels its _fut_waiter, which is this future. Checked on 3.10-3.13: future.cancelled() is True after the timeout. Routing tasks skip done futures, so a later replica goes to a live request. test_request_routing_timeout_no_replicas relies on this when it asserts the route queue is empty.


return replica

Expand Down
14 changes: 13 additions & 1 deletion python/ray/serve/_private/router.py
Original file line number Diff line number Diff line change
Expand Up @@ -667,6 +667,7 @@ def __init__(
self._initial_backoff_s: Optional[float] = None
self._backoff_multiplier: Optional[float] = None
self._max_backoff_s: Optional[float] = None
self._request_routing_timeout_s: Optional[float] = None

# Initializing `self._metrics_manager` before `self.long_poll_client` is
# necessary to avoid race condition where `self.update_deployment_config()`
Expand Down Expand Up @@ -781,6 +782,10 @@ def request_router(self) -> Optional[RequestRouter]:
backoff_kwargs["backoff_multiplier"] = self._backoff_multiplier
if self._max_backoff_s is not None:
backoff_kwargs["max_backoff_s"] = self._max_backoff_s
if self._request_routing_timeout_s is not None:
backoff_kwargs[
"request_routing_timeout_s"
] = self._request_routing_timeout_s

request_router = self._request_router_class(
deployment_id=self.deployment_id,
Expand Down Expand Up @@ -855,13 +860,19 @@ def update_deployment_config(self, deployment_config: DeploymentConfig):
deployment_config.request_router_config.backoff_multiplier
)
self._max_backoff_s = deployment_config.request_router_config.max_backoff_s
self._request_routing_timeout_s = (
deployment_config.request_router_config.request_routing_timeout_s
)

if self._request_router:
self._request_router.update_backoff_params(
initial_backoff_s=self._initial_backoff_s,
backoff_multiplier=self._backoff_multiplier,
max_backoff_s=self._max_backoff_s,
)
self._request_router.request_routing_timeout_s = (
self._request_routing_timeout_s
)

# Guard against the case where request_router is None (e.g., when
# request_router_class is None and lazy initialization has not yet
Expand Down Expand Up @@ -1139,7 +1150,8 @@ async def route_and_send_request(
"""Choose a replica for the request and send it.

This will block indefinitely if no replicas are available to handle the
request, so it's up to the caller to time out or cancel the request.
request and `request_routing_timeout_s` is unset, so it's up to the
caller to time out or cancel the request.
"""
# Wait for the router to be initialized before sending the request.
await self._request_router_initialized.wait()
Expand Down
12 changes: 12 additions & 0 deletions python/ray/serve/config.py
Original file line number Diff line number Diff line change
Expand Up @@ -274,6 +274,16 @@ class RequestRouterConfig(BaseModel):
),
)

request_routing_timeout_s: Optional[PositiveFloat] = Field(
default=None,
description=(
"Maximum duration in seconds that a request waits to be assigned a "
"replica before it fails with a TimeoutError. This bounds routing "
"only, not the time the replica takes to handle the request. "
"Defaults to None, meaning a request waits indefinitely."
),
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

None timeout breaks proto serialization

High Severity

DeploymentConfig.to_proto dumps request_routing_timeout_s as None by default and passes it into RequestRouterConfigProto. The same function already pops retry_after_s when it is None so an optional proto field is left unset. Without that, default configs can raise TypeError during proto construction, which breaks deployment config broadcast for every deployment that leaves the timeout unset.

Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit ecff197. Configure here.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The proto field is optional, and the protobuf constructor treats None as unset. HasField is False on protobuf 3.20.3 (Ray's floor, cpp and python) and on 6.31.1. It doesn't raise. test_request_routing_timeout_proto_round_trip[None] covers this path.


@field_validator("request_router_kwargs")
@classmethod
def request_router_kwargs_json_serializable(cls, v):
Expand Down Expand Up @@ -308,6 +318,7 @@ def __eq__(self, other):
and self.initial_backoff_s == other.initial_backoff_s
and self.backoff_multiplier == other.backoff_multiplier
and self.max_backoff_s == other.max_backoff_s
and self.request_routing_timeout_s == other.request_routing_timeout_s
)

def __hash__(self):
Expand All @@ -329,6 +340,7 @@ def __hash__(self):
self.initial_backoff_s,
self.backoff_multiplier,
self.max_backoff_s,
self.request_routing_timeout_s,
)
)

Expand Down
19 changes: 19 additions & 0 deletions python/ray/serve/tests/unit/test_config.py
Original file line number Diff line number Diff line change
Expand Up @@ -411,6 +411,25 @@ def test_backoff_params_declarative_schema(self):
assert schema.request_router_config.backoff_multiplier == 3.0
assert schema.request_router_config.max_backoff_s == 2.0

@pytest.mark.parametrize("request_routing_timeout_s", [None, 1.5])
def test_request_routing_timeout_proto_round_trip(self, request_routing_timeout_s):
"""None must not become 0.0 on the way through the proto."""
config = DeploymentConfig.from_default(
request_router_config=RequestRouterConfig(
request_routing_timeout_s=request_routing_timeout_s
)
)
round_tripped = DeploymentConfig.from_proto_bytes(config.to_proto_bytes())
assert (
round_tripped.request_router_config.request_routing_timeout_s
== request_routing_timeout_s
)

def test_request_routing_timeout_changes_equality(self):
"""A timeout-only change must be broadcast to routers."""
config = RequestRouterConfig(request_routing_timeout_s=1.5)
assert config != RequestRouterConfig()

def test_deployment_actors_config(self):
"""Test deployment_actors config and proto roundtrip."""

Expand Down
36 changes: 36 additions & 0 deletions python/ray/serve/tests/unit/test_pow_2_request_router.py
Original file line number Diff line number Diff line change
Expand Up @@ -2232,5 +2232,41 @@ def test_compute_backoff_s_does_not_overflow():
assert router._compute_backoff_s(2048) == router.max_backoff_s


@pytest.mark.asyncio
async def test_request_routing_timeout_no_replicas(pow_2_router):
"""A request that isn't assigned a replica within the timeout fails."""
s = pow_2_router
s.request_routing_timeout_s = 0.05
loop = get_or_create_event_loop()

task = loop.create_task(s._choose_replica_for_request(fake_pending_request()))
with pytest.raises(TimeoutError, match="Failed to route request to a replica"):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Update the test to expect asyncio.TimeoutError to match the exception raised by the request router.

Suggested change
with pytest.raises(TimeoutError, match="Failed to route request to a replica"):
with pytest.raises(asyncio.TimeoutError, match="Failed to route request to a replica"):

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The router raises the builtin TimeoutError on purpose (see the thread on request_router.py), so the test stays as is.

await asyncio.wait_for(task, timeout=10)

assert s.num_pending_requests == 0
assert len(s._pending_requests_to_route) == 0


@pytest.mark.asyncio
async def test_request_routing_timeout_when_replicas_maxed(pow_2_router):
"""A request that times out stops its routing task."""
s = pow_2_router
s.request_routing_timeout_s = 0.05
loop = get_or_create_event_loop()

r1 = FakeRunningReplica("r1")
r1.set_queue_len_response(DEFAULT_MAX_ONGOING_REQUESTS)
s.update_replicas([r1])

task = loop.create_task(s._choose_replica_for_request(fake_pending_request()))
with pytest.raises(TimeoutError, match="Failed to route request to a replica"):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Update the test to expect asyncio.TimeoutError to match the exception raised by the request router.

Suggested change
with pytest.raises(TimeoutError, match="Failed to route request to a replica"):
with pytest.raises(asyncio.TimeoutError, match="Failed to route request to a replica"):

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same as above: the builtin TimeoutError is intentional, so the test stays as is.

await asyncio.wait_for(task, timeout=10)

await async_wait_for_condition(
lambda: s.curr_num_routing_tasks == 0, retry_interval_ms=1
)
assert s.num_pending_requests == 0


if __name__ == "__main__":
sys.exit(pytest.main(["-v", "-s", __file__]))
2 changes: 2 additions & 0 deletions python/ray/serve/tests/unit/test_router.py
Original file line number Diff line number Diff line change
Expand Up @@ -3172,6 +3172,7 @@ async def test_update_deployment_config_sets_backoff_params(self):
initial_backoff_s=custom_initial_backoff,
backoff_multiplier=custom_multiplier,
max_backoff_s=custom_max_backoff,
request_routing_timeout_s=1.5,
)
)

Expand All @@ -3187,6 +3188,7 @@ async def test_update_deployment_config_sets_backoff_params(self):
assert fake_request_router.initial_backoff_s == custom_initial_backoff
assert fake_request_router.backoff_multiplier == custom_multiplier
assert fake_request_router.max_backoff_s == custom_max_backoff
assert fake_request_router.request_routing_timeout_s == 1.5


class TestOnRequestCompleted:
Expand Down
3 changes: 3 additions & 0 deletions src/ray/protobuf/serve.proto
Original file line number Diff line number Diff line change
Expand Up @@ -135,6 +135,9 @@ message RequestRouterConfig {

// Maximum backoff time (in seconds) between retries.
double max_backoff_s = 8;

// Maximum time (in seconds) a request waits to be assigned a replica.
optional double request_routing_timeout_s = 9;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Proto change needs fault-tolerance review

Low Severity

⚠️ This PR modifies one or more .proto files.
Please review the RPC fault-tolerance & idempotency standards guide here:
https://github.com/ray-project/ray/tree/master/doc/source/ray-core/internals/rpc-fault-tolerance.rst

This is required by the RPC Fault Tolerance Standards Guide rule because serve.proto is in the change set.

Fix in Cursor Fix in Web

Triggered by project rule: Bugbot Rules

Reviewed by Cursor Bugbot for commit 0d9909b. Configure here.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No RPC is added or changed. The PR only adds an optional config field to the RequestRouterConfig message. An unset field reads back as None, and older readers ignore it.

}
//[End] ROUTING CONFIG

Expand Down
Loading