Skip to content

mesh: no stuck notes, a late joiner's mesh, and pairs that go back to direct - #392

Merged
mishan merged 13 commits into
masterfrom
mesh-fixes
Oct 7, 2026
Merged

mishan merged 13 commits into
masterfrom
mesh-fixes

Conversation

@mishan

@mishan mishan commented Oct 7, 2026

Copy link
Copy Markdown
Owner

Stacked on #391, whose harness measured the problems and these fixes.

Keys and one-shot commands on a reliable channel

  • Each peer connection gets a second channel, keys, which is ordered and reliable.
  • Every command except a knob move goes on both keys and the unreliable gestures channel; the first copy to arrive wins and Dedupe drops the other. Knobs stay on gestures only.
  • A direct key that arrives behind a later key of the same note is dropped rather than left sounding. A stamped key in time is played at its stamp, whatever it arrived behind.
  • Dedupe tells a duplicate apart by sequence distance: anything within 4096 of the sender's newest is remembered. That also catches duplicate sequencer clicks.

A late joiner's mesh opens. The mesh is opened right after the welcome, before the document sync, so an offer from an existing peer is no longer dropped.

Pairs go back to direct

  • A fallback no longer closes the connection.
  • disconnected is treated as temporary.
  • A failed connection is left to its ICE restart, which keeps the reliable channel's queue.
  • While a pair is relayed, both sides check back with backoff (5 s doubling to 60 s), and the offering side rebuilds a connection that is gone or has lost a channel.
  • Signals carry a generation number.

Before / after (nettest on a load machine, 40 s runs at 50 ms round trip)

before after
Stuck keys, 5% loss 15–29 0
Stuck keys, 10% loss 25–31 0
Keys arriving, ahead mode, 10% loss 259–264 of 266 all
Late joiner's pairs direct about 1 in 2 always
Pairs direct again after a 10 s or 25 s blackhole never yes
  • Median key delay is unchanged.
  • Under loss, keys that used to be lost arrive late instead: worst case about 450 ms at 5% loss and about 700 ms at 10%.

Compatibility. PROTOCOL is unchanged.

  • An old page that offers creates gestures only, and a new page answering it uses that one channel.
  • A new page offering to an old one sends both copies; the old page's Dedupe drops one.
  • An old page's answer without a generation is taken as the current one.

Tests

  • jamtest: the late joiner's log never says it went through the relay (it fails 3 of 3 with the open-before-sync change reverted), and a key still arrives when everything on the unreliable channel is lost.
  • protocoltest: a duplicate 1000 commands later is still dropped.
  • protocoltest, relaytest, mirrortest and jamtest pass.

mishan added 13 commits October 7, 2026 10:41
Each pair gets a second data channel, ordered and reliable. Every
command but a knob goes on it as well as on the unordered channel with
no retransmission. The first copy to arrive is applied and Dedupe drops
the other, so a lost packet costs a key only the wait for the other
copy, and the reliable channel's head-of-line wait holds up only what
the unreliable one also lost. Knobs stay on the unreliable channel
alone, since a later knob supersedes a lost one.

A press can now arrive after its own release: its unreliable copy is
lost and the reliable one is held for a retransmission. jam.js drops a
key older than the last one applied for the same peer, seat and note,
so the press cannot sound until the next release.

The relay fanning out its log copy would have been the other way. The
log holds only replayable commands, so direct keys are not in it, and
each key would cost the relay a second send and the receiver an extra
hop.

A page from before this offers only the unreliable channel, and the
answering side sends everything on that. An older page that answers an
offer of both takes the reliable one for a second unreliable one,
receives both copies, and Dedupe drops one. PROTOCOL is unchanged: the
relay sees nothing new.

nettest mesh, rtt 50 ms, 2 runs per condition; keys arrived of sent,
keys left stuck, and key delay as median / p99 / worst:

  loss             before              after
  1%               523/534, 5 stuck    534/534, 0; 51 / 53-56 / 54-154 ms
  5%               472/534, 29         531/532, 0; 48-50 / 184-301 / 326-453
  5% from 12 s     496/534, 15         531/534, 0; 49-53 / 51-56 / 57-149
  10% from 12 s    468/534, 25         528/536, 0; 49-51 / 400-451 / 451-699

The median is 48-53 ms either way. Before, a key arrived on time or not
at all; now the ones that were lost come late. With the reliable
channel alone the worst at 10% from 12 s was 552-1355 ms, and at 5%
from 12 s 80-1073 ms over three runs.

jamtest: a key pressed and let go while the sender's unreliable channel
drops everything still reaches the other page.
join() made the mesh only after the document had synced, and room.js
drops a signal that has no handler. A peer already in the room offers
the moment it hears `joined', so a joiner whose id sorts after the
offerer's lost the offer, and that pair fell back to relayed at
OPEN_WITHIN. The mesh is now made straight after the welcome, with
nothing awaited in between, so every signal finds it. room.js's other
events have their handlers from openRoom, before the socket connects,
so they are not affected.

Buffering signals in room.js until a handler is attached was the other
way. Moving one call is smaller and leaves no queue to bound.

nettest, 2 runs per condition; how many runs ended with every pair
direct:

  condition                      before    after
  join, rtt 20 ms                1 of 2    2 of 2
  join, rtt 100 ms               1 of 2    2 of 2
  join, rtt 300 ms               2 of 2    2 of 2
  clock, rtt 50, ±10 ms, 5%      1 of 2    2 of 2
  clock, rtt 300, ±10 ms, 5%     0 of 2    2 of 2

The clock runs fell back at the start for the same reason: the second
page joins while the first is already there. clock with rtt 20,300 was
direct in both runs before and after.

jamtest holds the late joiner's document socket back a second, then
checks that the joiner's pairs are direct. With random ids the joiner
offers to any peer whose id sorts after its own, so the check catches
the old order unless the joiner has the smallest id.
A fallback closed the RTCPeerConnection, and `disconnected' was a
fallback, so a pair that missed OPEN_WITHIN, or a page whose network
went away for a while, stayed relayed for the rest of the session.

The connection now stays. A pair is direct again once both channels
are open and the connection is `connected' again. While it is
`disconnected' the pair is relayed, and ICE finds its own way back. On
`failed', the offering side calls restartIce(), and the answering side
sends it a `restart' signal asking for one. While relayed, the offering
side checks back with backoff, 5 s doubling to 60 s, and makes the
connection again if it is new, failed or closed, or has lost a
channel. Signals carry a generation, so an answer or candidate for a
replaced connection is not applied to the next one. An older page
sends no generation, which reads as the first.

A key relayed and its release sent direct, or the other way round,
can arrive out of order across a switch. The overtaken-key check in
jam.js drops the stale press.

nettest blackhole, 3 pages, p2 cut off from 8 s, 2 runs per hole; when
every pair was direct again:

  hole                   before          after
  10 s, back at 18 s     never           18.4, 18.9 s
  25 s, back at 33 s     never           43.1, 43.3 s
  60 s, back at 68 s     77.1, 80.6 s    79.1, 79.4 s

After 60 s the relay has cut p2's socket and it rejoins, which gives it
a new mesh either way. After 25 s the restart offer reaches p2 only
once its relay socket's TCP backoff lets it through.

Without impairment (clock, 30 s, 2 runs) nothing changes: direct, no
late commands.
A key behind a later one of its note was dropped in every mode, but a
quantised or ahead key is played at its stamp, not on arrival, and the
order it came in says nothing about the order it sounds in. With loss
on the mesh, an ahead key whose unreliable copy was lost and whose
reliable one came after the next key of its note was dropped two
seconds early, and a quick repeat of one key became one long note on
the page that dropped it. Only a key that would be played now -- a
direct one, or a stamped one already late -- is held to its order.
Dedupe kept the last 128 to 256 seqs of each sender and took anything
older as new. A knob drag is a command per slider tick, so a reliable
copy of a key or a cell click held up behind one by a second or two was
applied again: a second attack on a poly seat, or a cell toggled back.
It now remembers every seq within 4096 of the sender's newest, about a
minute of a slider moving sixty times a second, and takes anything
older as seen.
window.jam.dropped is the last of the commands applyOne let go without
applying: a key overtaken by a later one of its note, a command of a run
since replaced, or one that came before Start. nettest counts them
apart from commands that never arrived.
connect sets the connection, its generation, its channels and the
candidates held for it, before anything reads them.
A page from before generations sends a signal with none, which was read
as generation 0, so its answer to a pair made again was dropped and the
pair stayed relayed. Such a page has one connection, and what it sends
is for whichever this side has.
The answering side took a link with no `keys' channel for one with a
page from before it, which has none, and called the pair direct as soon
as the unreliable channel opened: for a moment after each connection,
keys and one-shot commands went on that channel alone. An offer with a
generation is from a page that opens both, and the pair waits for both.
A failed connection both restarted ICE and, at the next backoff tick,
was made again, closing the channels mid-restart and throwing away what
was queued on the reliable one: a release sent between an outage and
its detection was lost, and the note left sounding. The restart is left
to it. If the connection is truly gone its channels close, and the pair
is made again then.
A key on the reliable channel is not lost while the connection lives,
an ICE restart included, but what is queued on it when a connection
comes apart for good goes with the connection.
A fallback with no change of state behind it -- a signal that failed to
apply, or an offer that could not be made, on a connection that was
fine -- left the pair relayed for good: nothing would call it direct
again, and the backoff, seeing a connection that was neither new nor
failed nor closed, left it alone. Each backoff tick now first asks
whether the pair is direct after all, on both sides, the answering side
included, which never ticked before; only the offering side makes the
connection again. And an offer that failed for a connection already
replaced is not a fallback of the one that replaced it.
The check that the late joiner's mesh is direct passed with the joiner
missing its first offer, since the pair is made again later and is
direct by the end. The joiner's log says whether a pair went through
the relay at all, which nothing after can hide. And the joiner joins
again, under another name, until another page's id is the smaller, so
it is that page that offers while the joiner's document syncs; five
joins without one fails as not tested.

The key sent with the unreliable channel losing everything is direct
when it is let go, too, not only when it is pressed.
Base automatically changed from nettest to master October 7, 2026 22:06
@mishan
mishan merged commit e19edd6 into master Oct 7, 2026
18 of 28 checks passed
@mishan
mishan deleted the mesh-fixes branch October 7, 2026 22:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant