Skip to content

[fix] Re-probe reachability while the relay is firewalled - #16

Merged
ok merged 1 commit into
mainfrom
ok/reprobe-reachability
Sep 20, 2026
Merged

ok merged 1 commit into
mainfrom
ok/reprobe-reachability

Conversation

@ok

@ok ok commented Sep 20, 2026

Copy link
Copy Markdown
Owner

The bug

On a live StartOS box the relay came up firewalled after an update and stayed that way, with the port forward, outbound gateway and public address all intact. The same box had been firewalled for 5.5 hours on 0.3.0 three days earlier, and both times a plain restart fixed it.

The cause is in dht-rpc. It probes reachability once at bootstrap (_checkIfFirewalled: up to 5 nodes are asked to ping the server port, and fewer than 3 replies counts as firewalled). Its periodic re-check is then guarded by:

if (this._lastHost === this._nat.host) {
  // do not recheck the same network...
  this._stableTicks = MORE_STABLE_TICKS
}

so while the public host is unchanged it never probes again. One lost round in the first seconds after a restart is permanent until the next one.

The fix

src/reprobe.js re-runs dht._updateNetworkState() — the same call dht-rpc's own tick makes, minus the same-host guard — while the node is firewalled: after 1 minute, doubling to a 15-minute ceiling, and stopping for good once a probe passes. It updates relay_dht_firewalled, which was previously only set at startup, and logs reachability re-probe passed.

It is not started under MIRALL_RELAY_ASSUME_REACHABLE. The hook is dht-rpc internals, so its absence degrades to a single warning and no self-heal; nothing else depends on it.

Testing

  • test/unit/reprobe.test.js: never probes a reachable node; retries until a probe passes, then schedules nothing; backoff sequence and ceiling; a throwing probe does not end the retries; stop() cancels pending and in-flight work; a missing hook is a no-op with one warning.
  • Run against the real pinned dht-rpc 6.27.0 on a NAT'd machine: the hook exists, each re-probe performs a real ~4 s ping round, and the node correctly stays firewalled where there is no forward.
  • Not shown: a firewalled-to-reachable flip on a real forwarded host. That needs the StartOS box to land in the bad state again; the log line above is what to look for.
  • npm run lint clean, npm test 408/408.

dht-rpc probes once at bootstrap and skips its periodic re-check while
the public host is unchanged, so one lost probe round left a correctly
forwarded relay firewalled until restart. Re-run the probe while
firewalled, from 1 minute backing off to 15, and keep the firewalled
gauge current.
@ok
ok merged commit 8efdd73 into main Sep 20, 2026
4 checks passed
@ok
ok deleted the ok/reprobe-reachability branch September 20, 2026 16:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant