Skip to content

relay: bounded queues, rooms, documents and rates - #377

Merged
mishan merged 2 commits into
masterfrom
relay-limits
Oct 7, 2026
Merged

mishan merged 2 commits into
masterfrom
relay-limits

Conversation

@mishan

@mishan mishan commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

Under production container limits (512m, a 384 MB heap) a single client
could crash the relay several ways -- a socket that stopped reading, a
flood of rooms, garbage or crafted document updates, oversized run logs
and starts -- and CPU gave out well below what the caps allowed. Each is
bounded.

Send queues: 16 MiB per socket and 128 MiB across the relay; past its
own a socket is cut `slow', past the relay's the one holding the most
bytes goes. One catch-up or document sync per socket is held apart from
its cap and counted once. Fan-out and catch-ups go as shared Buffers;
each socket is corked for a loop turn and flushed past 64 KiB.

Rooms: made only by a welcomed hello, at most 4096 and 128 MiB charged in
all, 16 MiB each, rooms empty longest evicted to make room; creation and
joins rate-limited per address and /48; 64 peers a room, a page
rejoining with a live ticket displacing its old peer; sockets that never
say hello capped per address and cut after 5 s. Empty rooms are kept
10 minutes.

Documents: a room is charged for its document's real structure --
structs, clients, content and roots, counted after every transaction
-- and every update is bounded before it is applied by the most it
could add, deletions credited and subdocuments refused, so a room at
its cap is still joined and read and only growth is refused. What Yjs
holds pending is charged and bounded per socket. Every charge goes
through Room.charge, which is a no-op once the room is gone. Run logs
are kept as JSON strings and charged, 8 MiB a run; starts are rebuilt
from the fields pages use; cursor states keep only user and cursor, at
1 KiB.

Rates: a token bucket per socket for every room message type, chat
included; inbound byte budgets charged before a frame is parsed; per
peer and per address budgets for document updates before they are
decoded; to' lists deduplicated and capped; byte budgets for fan-out and for run commands. Over a limit a message is dropped with a refused' line; a sustained flood cuts the socket (`flood').

Memory backstop: every 250 ms the relay reads its heap (and, with
MEMORY_MAX_BYTES, heap plus external) and past 75% forces a GC; while
past 60% it sheds empty rooms, then the largest runs' logs, then a
room, telling its pages `memory' and terminating its sockets, and
waits a look after each shed for the memory to show freed.

The page: an edit the relay refuses closes only the document socket,
the editor going read only and saying why; a page joins again by itself,
with backoff and no sooner than retryMs, after a refusal that passes
(rooms, full, slow, flood, memory, accounts); a refused catch-up ends
its wait at once.

One room at 112 players: p99 forward 6.7 s -> 4.0 ms; the loop passes
20 ms at ~192 players (was 96); peak output ~95k -> ~460k frames/s. A
slow reader, a room flood, delete-splitting, padded varints, oversized
logs, starts, subdocuments and cursor floods all leave the relay up.

relaytest covers each limit, cut and charge, and the reproductions of
each crash; jamtest a refused edit, a cut page joining again, and a
refused Play.

Under production container limits (512m, a 384 MB heap) a single client
could crash the relay several ways -- a socket that stopped reading, a
flood of rooms, garbage or crafted document updates, oversized run logs
and starts -- and CPU gave out well below what the caps allowed. Each is
bounded.

Send queues: 16 MiB per socket and 128 MiB across the relay; past its
own a socket is cut `slow', past the relay's the one holding the most
bytes goes. One catch-up or document sync per socket is held apart from
its cap and counted once. Fan-out and catch-ups go as shared Buffers;
each socket is corked for a loop turn and flushed past 64 KiB.

Rooms: made only by a welcomed hello, at most 4096 and 128 MiB charged in
all, 16 MiB each, rooms empty longest evicted to make room; creation and
joins rate-limited per address and /48; 64 peers a room, a page
rejoining with a live ticket displacing its old peer; sockets that never
say hello capped per address and cut after 5 s. Empty rooms are kept
10 minutes.

Documents: a room is charged for its document's real structure --
structs, clients, content and roots, counted after every transaction
-- and every update is bounded before it is applied by the most it
could add, deletions credited and subdocuments refused, so a room at
its cap is still joined and read and only growth is refused. What Yjs
holds pending is charged and bounded per socket. Every charge goes
through Room.charge, which is a no-op once the room is gone. Run logs
are kept as JSON strings and charged, 8 MiB a run; starts are rebuilt
from the fields pages use; cursor states keep only user and cursor, at
1 KiB.

Rates: a token bucket per socket for every room message type, chat
included; inbound byte budgets charged before a frame is parsed; per
peer and per address budgets for document updates before they are
decoded; `to' lists deduplicated and capped; byte budgets for fan-out
and for run commands. Over a limit a message is dropped with a
`refused' line; a sustained flood cuts the socket (`flood').

Memory backstop: every 250 ms the relay reads its heap (and, with
MEMORY_MAX_BYTES, heap plus external) and past 75% forces a GC; while
past 60% it sheds empty rooms, then the largest runs' logs, then a
room, telling its pages `memory' and terminating its sockets, and
waits a look after each shed for the memory to show freed.

The page: an edit the relay refuses closes only the document socket,
the editor going read only and saying why; a page joins again by itself,
with backoff and no sooner than retryMs, after a refusal that passes
(rooms, full, slow, flood, memory, accounts); a refused catch-up ends
its wait at once.

One room at 112 players: p99 forward 6.7 s -> 4.0 ms; the loop passes
20 ms at ~192 players (was 96); peak output ~95k -> ~460k frames/s. A
slow reader, a room flood, delete-splitting, padded varints, oversized
logs, starts, subdocuments and cursor floods all leave the relay up.

relaytest covers each limit, cut and charge, and the reproductions of
each crash; jamtest a refused edit, a cut page joining again, and a
refused Play.
@mishan mishan changed the title relay: bounded queues, rooms and rates relay: bounded queues, rooms, documents and rates Oct 6, 2026
The relay-wide total check set the total to 1 MiB and sent three
gestures of 900 KiB to two readers at once: until the readers drained
them the relay held 5 MB, and on a slow runner the 250 ms trim rightly
cut a reader before the edit it was meant to catch came. The total is
now above one gesture to both, and each is read before the next.

The malformed document frame checks opened three document sockets on
one ticket one after another; a close can reach the client before the
relay has counted that socket gone, and the third was refused past
DOCS_PER_PEER with an error nothing listened for. Each open now retries
until the relay lets it in.

Both failed with relaytest run on one core shared with three spinning
processes, and pass there now.
@mishan
mishan merged commit a6c99a3 into master Oct 7, 2026
18 checks passed
@mishan
mishan deleted the relay-limits branch October 7, 2026 02:17
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant