Skip to content

fix(bindings): Bluetooth peers find each other in busy rooms, and Android survives Bluetooth toggles and stack crashes - #511

Open
kivtxs wants to merge 2 commits into
mainfrom
fix/android-ble-busy-air
Open

kivtxs wants to merge 2 commits into
mainfrom
fix/android-ble-busy-air

Conversation

@kivtxs

@kivtxs kivtxs commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Found in the v0.28 device test. Two Android phones a metre apart, both advertising the mesh service (a Mac scan saw both adverts), never connected over Bluetooth for fifteen minutes after Bluetooth was switched off and on. This PR fixes three causes (the third found while testing the first two). Both predate v0.28. Neither is about Wi-Fi Direct, but Bluetooth is the fallback carrier v0.28 leans on, so they belong in the release.

1. The dense-mesh filters counted every Bluetooth device in range (iOS and Android)

estimatedVisiblePeerCount counts scan callbacks from every advertiser in range: televisions, earbuds, watches. Above 10, the RSSI filter, the rate limit and shouldProbabilisticallySkip engage. The skip passes over up to 80% of peers, chosen by hash(address) alone, so the same peers are passed over on every advert for as long as their address lasts. On iOS it is peripheral.hashValue, which is seeded per process.

Measured on the phones: the estimate read 32, giving a skip share of 0.44. The Seeker's address hashes to 0.091 and the Infinix's to 0.284, both under 0.44. In the logs, each phone matches the other's service UUID every 30 s and then silently drops it before Discovered device.

Fix: BleDensityPolicy, the same on both platforms.

  • What counts: density is the number of distinct mesh candidates (addresses the discovery gate admits) in the last 5 s. The RSSI filter, the dense rate limit and the skip use it. The unknown-device bootstrap keeps the all-advert count, because "how busy the air is" is exactly what that policy needs.
  • How the skip draws: per address, per one-minute slot, mixed with the murmur3 finalizer. Without the mix, Java's String.hashCode put consecutive slots of one address in neighbouring buckets; the new test caught that. A skipped peer is reconsidered in the next minute.
  • Cross-platform agreement: Android and iOS compute the same bucket for the same id. Both test suites pin the same values.
  • Where the count is taken: after the gate and before the filters, the same place as the iOS restoration sighting. The Rust guard react_native_ios_restoration_commands_wait_for_powered_on now pins the new call and explains the new reason.

2. Android kept dead links after Bluetooth went off

When the adapter goes off, Android delivers no disconnect for most links. The facade reported BLE unavailable but kept every link, so:

  • the dead links counted against the connection cap;
  • peers stayed mapped to addresses they no longer use;
  • each BluetoothGatt still held a client registration the stack had forgotten (the Infinix showed five, and BtGatt.ContextMap: Context not found for ID 6 repeated several times a second).

Separately, a peer probed while its radio was still coming back serves no mesh service yet, so it landed in verifiedNonMeshDevices for five minutes. On the phones, the Seeker probed the Infinix 17 s before the Infinix's recovery rebuilt its GATT service.

Fix: dropLinksAfterRadioLoss(), the Android equivalent of iOS's function of the same name. When startScanning first sees the adapter off (once per outage, on the BLE thread), it:

  • reports each identified peer lost;
  • closes every BluetoothGatt;
  • clears the link, MTU, handshake and discovery caches, including the non-mesh cache.

3. Android never heard Bluetooth go off, come back, or crash (7c4811f3)

Found while validating 1 and 2. As the Infinix dialled the Seeker, Android's own Bluetooth stack crashed (SIGSEGV in libbluetooth_jni.so, connection_manager::on_connection_complete, Infinix X670 on Android 13). The stack restarted in under a second and took the app's GATT server, advertiser and pending connect with it.

The facade only learns about the adapter by polling isEnabled from the 60 s scan watchdog. That reads "enabled" throughout a crash, so nothing was rebuilt:

  • the pending connect pinned the Seeker's address as "connecting";
  • the Seeker could not verify the Infinix until the app restarted.

The same polling is why a plain toggle was noticed up to a minute late, and recovered 12 to 29 s after Bluetooth returned.

Fix: listen for ACTION_STATE_CHANGED on the BLE handler. Android reports a crash as the same ON → TURNING_OFF transition a toggle produces.

  • TURNING_OFF / OFF: drop the links (section 2), follow the dead scan locally, report BLE unavailable.
  • ON after an outage: rebuild the scan, GATT server and advertising at once.
  • AdapterStateTransition maps the states (3 JVM tests). A Rust guard pins the registration after RUNNING, the unregistration behind the shutdown barrier, and the radio-lost arm.

On the phones (Infinix NOTE 12 / Android 13, Seeker / Android 15; build = #509's tree with this branch)

Scenario Before After
Bluetooth only, app start 2 s 2 s
Both phones toggle Bluetooth off 20 s, 3 rounds never relinked (15 min watched) neighbor_lost exactly 1 per phone, headers 0 while off; relinked in 7, 21, 4 s; GATT client registrations stayed 1 to 4
One phone toggles n/a relinked in 10 s; 4/4 each way after
Wi-Fi Direct + Bluetooth, then Bluetooth off n/a no neighbor_lost; messages over Wi-Fi Direct p50 50 to 69 ms
…then Wi-Fi off (no carrier left) phantom neighbour kept (fixed with #505's be01882a) neighbor_lost 1 per phone, headers 0; Bluetooth back on relinked in 13 s
Whole session, toggles included n/a 43/43 each way: sent = received = receipts

Known and not changed here (pre-existing; I'll file issues):

  • After a reconnect, writes from a central whose address the server hasn't mapped to a peer wait for the resolution dial. That costs seconds to about 40 s; the core's retries deliver them meanwhile.
  • When the other phone switches Bluetooth off, its clean disconnect is treated as transient, so this phone lists it for over a minute.

Related

#505 has the core half (be01882a): ble_status_changed(false) now ends the core's Bluetooth peers. Before that, the stale entry suppressed a later Wi-Fi Direct neighbor_lost. Each PR is correct without the other.

Tests

Suite Result
BleDensityPolicyTest (JVM) 9 tests
BleDensityPolicyTests (Swift) 8 tests
StaleAddressRegistryTest new deviceIds test
Android JVM suite 581 pass
swift test 456 pass
iOS bridge typecheck passes
cargo test --workspace --lib passes
Packaging script tests pass

I could not run an iPhone. The iOS change is typechecked and unit-tested, and mirrors the Android logic line for line.

…droid drops dead links when Bluetooth goes off

Found in the v0.28 device test: two Android phones a metre apart, both
advertising the mesh service (a Mac scan saw both), never connected over
Bluetooth for fifteen minutes.

- Density counted every advert in range as mesh density. The estimate read 32
  in a house (televisions, earbuds, watches), so the dense-mesh filters passed
  over 44% of peers, chosen by a hash of the address alone: each phone's
  address hashed under the line on the other (0.091 and 0.284), so each was
  passed over on every advert. Density now counts distinct mesh candidates
  (devices the discovery gate admits) in the last five seconds, and the
  pass-over is drawn per address per one-minute slot, mixed with the murmur3
  finalizer so consecutive slots are independent draws. The unknown-device
  bootstrap keeps the all-advert count, which is what it is about. Same
  change on iOS, where the hash was Swift's per-process seed; both platforms
  compute the same bucket for the same id (pinned by tests on both sides).
- Android: Bluetooth switched off under a running transport left every link
  in place, since the stack delivers no disconnect for most of them. The dead
  links counted against the connection cap, kept peers mapped to old
  addresses, and held GATT client registrations the stack had forgotten. The
  transport now reports each identified peer lost and clears the link state
  when it first sees the adapter off, mirroring iOS's dropLinksAfterRadioLoss,
  including the non-mesh cache that had marked a peer probed mid-recovery as
  not a mesh device for five minutes.

BleDensityPolicy on both platforms (9 JVM tests, 8 Swift tests), a registry
test, and the iOS restoration guard updated for the new call.
…uding a stack crash

Found on the phones while validating the previous commit: the Infinix's
Bluetooth stack crashed (SIGSEGV in libbluetooth_jni.so,
connection_manager::on_connection_complete) as the app dialled the other
phone. The stack restarted in under a second and took the app's GATT server,
advertiser and pending connect with it. The transport polls the adapter once
a minute from the scan watchdog, which read "enabled" throughout, so nothing
was rebuilt: the pending connect pinned the peer's address as "connecting",
and the other phone could not verify this one until the app restarted.

Android reports a crash as an ordinary ON -> TURNING_OFF broadcast. The
facade now registers for ACTION_STATE_CHANGED on the BLE handler once it
runs, and unregisters behind the shutdown barrier:

- TURNING_OFF / OFF: drop the dead links (once per outage), follow the dead
  scan locally, report BLE unavailable, arm recovery.
- ON after an outage: run the recovery now (scan, GATT server, advertising)
  instead of waiting for the ladder's next rung (12 to 29 s measured).

AdapterStateTransition maps the states (3 JVM tests); a Rust guard pins the
registration after RUNNING, the unregistration after the barrier, and the
radio-lost arm dropping links.
@kivtxs kivtxs changed the title fix(bindings): Bluetooth peers in a busy room find each other, and Android drops dead links when Bluetooth goes off fix(bindings): Bluetooth peers find each other in busy rooms, and Android survives Bluetooth toggles and stack crashes Oct 6, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant