Skip to content

chore(release): 0.28.0 - #509

Open
kivtxs wants to merge 10 commits into
mainfrom
chore/release-0.28.0
Open

kivtxs wants to merge 10 commits into
mainfrom
chore/release-0.28.0

Conversation

@kivtxs

@kivtxs kivtxs commented Oct 5, 2026 •

Copy link
Copy Markdown
Member

The 0.28.0 cut, following CONTRIBUTING's "Cutting a Release". It touches the same file set as #460 (0.27.0).

Merge order

Merge this last. It renames ## [Unreleased] to ## [0.28.0] — 2026-10-06, so anything that lands after it would be described under a release it isn't in.

Already on main: #504 (Android library artifact id) and #505 (Android Wi-Fi Direct fixes). #508 (demo app) was closed. Still to merge, in any order:

This branch currently stacks the remaining three (main + merges of each + the release commit), so its CI tests exactly the tree 0.28.0 would ship. You can also merge it alone: it carries their commits unchanged.

Then rebase this branch; I'll do it if you'd like. The only expected conflict is the changelog heading line. Update the date to the day you tag.

docs/UPGRADING.md §26 already describes #505, #506, #507 and #511's change to neighbor_lost. If any of them doesn't make 0.28.0, drop its paragraph; I can do that too.

What changed

  • Versions to 0.28.0:
    • Cargo.toml: [workspace.package] plus the 11 internal deps;
    • the four local deps in offline-protocol-sealed and offline-protocol-leaf;
    • Cargo.lock, tools/embedded-footprint/Cargo.lock, tools/mls-interop/Cargo.lock (refreshed with cargo metadata; only our own crates moved);
    • bindings/python/pyproject.toml;
    • bindings/react-native/package.json and its lockfile (npm version 0.28.0 --no-git-tag-version).
  • THIRD-PARTY-NOTICES.md (×3): regenerated with scripts/generate-third-party-notices.sh and cargo-about 0.9.1, the pinned version.
  • CHANGELOG.md: Unreleased becomes 0.28.0. The 0.27.0 section moves to the new docs/changelog/0.27.md, with its title and release table, and a row in both archive tables.
  • Moved links: ](docs/…) becomes ](../…) (four of them).
  • The move is lossless: re-extracting the section from HEAD:CHANGELOG.md and applying the same rewrite gives the archived text exactly, apart from one trailing blank line.
  • SECURITY.md: the current line is 0.28.x; earlier lines are ≤ 0.27.x.
  • docs/UPGRADING.md:
    • the current line is v0.28.x;
    • seven releases since 0.17 can now break a build. 0.28.0's two breaks: iOS peer streams move to Network framework, so an iPhone on 0.28 doesn't see one on ≤ 0.27 over that slot, and Python's InternetManager now requires app_id=;
    • a 0.28.0 summary paragraph;
    • a new §26, "Behaviour that changes without a compile error (v0.28.0)".

Verification

  • cargo test -p offline-protocol-core -p offline-protocol-leaf --lib local_dep_versions (the version guards): pass.
  • cargo check --workspace --locked: clean.
  • cargo test -p offline-protocol-uniffi --lib (the bridge guards): 189 passed.
  • Markdown links: every relative ](path#anchor) in every tracked file, resolved against real headings. No new breakage; the only two flagged are the illustrative ](docs/…) and ](path#anchor) in CONTRIBUTING itself.

Device validation of this tree (two Android phones, Infinix NOTE 12 on Android 13 and Seeker on Android 15)

Area Result
Wi-Fi Direct group formation (#507), Bluetooth off, location not granted 8/8 clean starts, 18 to 225 s (median about a minute), no system dialog
Wi-Fi Direct chat 30/30, p50 73 to 82 ms; 40-round soak 80/80 received, 79/80 receipts (see below)
App restart / 30 s Wi-Fi outage / group chat stream re-proved in 3 s / 3 of 3 queued messages delivered after Wi-Fi returned / 3/3 each way
Bluetooth only discovery 2 s; 30-round soak 30/30 each way, with receipts
Bluetooth toggled on both phones, 3 rounds (#511) one neighbor_lost per phone; relinked in 7, 21, 4 s (never, before)
Both carriers, then Bluetooth off, then Wi-Fi off (#505 be01882a) no false loss while Wi-Fi Direct holds; one neighbor_lost per phone when nothing is left
DORS with both carriers routes Wi-Fi Direct (p50 59 to 88 ms), falls back to Bluetooth at once when Wi-Fi drops
Clocks 40 days apart (#506) KEY_PACKAGE_OUTSIDE_VALIDITY_WINDOW once per phone; demo alert shown; messaging after the clock is fixed
App in background 2/2 received
Mesh services ping.v1 discovered, pong received
Whole final session, toggles included 43/43 each way: sent = received = receipts

Seen and not fixed in 0.28.0:

  • Event gaps: twice, one message in a 30-message soak arrived with a receipt but no message_received in the receiver's JS log, and once a receipt was missing in a 40-round soak. Not reproduced in the 86 messages after native-side event logging was added; I'll keep watching.
  • Pre-existing Android BLE behaviour, issues to follow:
    • after a reconnect, writes from an address the server hasn't mapped to a peer wait for a resolution dial (seconds to about 40 s; retries deliver);
    • a peer that switches Bluetooth off stays listed for over a minute on the other phone (a clean disconnect is treated as transient).
  • Flaky relay test: mesh relay: a message can die inside a dense cluster when no fan-out ever picks the one bridge node (mesh_forwarding flake, ~1.5%) #510.

After merge (from CONTRIBUTING)

  1. Optional rehearsal: push v0.28.0-rc.1, or a workflow_dispatch with dry_run: true, version: 0.28.0. That checks npm OIDC before anything is permanent.
  2. Push v0.28.0.
  3. gh workflow run publish.yml -R Offline-Protocol/offline-protocol-swift -f version=0.28.0. This is the first release with a Swift package.
  4. Maven Central and PyPI stay behind their switches (MAVEN_CENTRAL_PUBLISH, PYPI_PUBLISH), unchanged here.

For the release run

One flaky test appeared in CI today: mesh_forwarding::everyone_hearing_everyone_does_not_multiply_the_traffic (#510, about 1.5%). It failed on #508, a demo-only change, and on an earlier run on main (36897622339); the failure is non-delivery (far's inbox empty). A flake in the tag's run means a re-run, not a new patch version, but it's worth knowing before you push the tag.

A peer's key package is valid from an hour before it was minted, judged by
the receiver's clock. Between two devices whose clocks disagree by more
than that, no session forms and messages and connection requests stay
pending forever, and the only trace was a debug line. Two Android phones
that had never been online (one set to February 2024, one to October 2025)
reproduced it during the v0.28 smoke test.

- MlsError::KeyPackageOutsideValidityWindow, from OpenMLS's
  KeyPackageVerifyError::InvalidLifetime, on every route that admits a
  peer's package (import, cache read, add_group_member) through one helper.
- SecurityWarningCode::KeyPackageOutsideValidityWindow
  (KEY_PACKAGE_OUTSIDE_VALIDITY_WINDOW), raised by establish_secure_session
  once per peer through the control-gate throttle: the package has not
  proved its sender when its window is checked.
- The window itself is unchanged.
@kivtxs
kivtxs requested a review from bahdotsh October 5, 2026 22:39
@kivtxs
kivtxs force-pushed the chore/release-0.28.0 branch from 129b946 to b10b3ce Compare October 5, 2026 23:22
…droid drops dead links when Bluetooth goes off

Found in the v0.28 device test: two Android phones a metre apart, both
advertising the mesh service (a Mac scan saw both), never connected over
Bluetooth for fifteen minutes.

- Density counted every advert in range as mesh density. The estimate read 32
  in a house (televisions, earbuds, watches), so the dense-mesh filters passed
  over 44% of peers, chosen by a hash of the address alone: each phone's
  address hashed under the line on the other (0.091 and 0.284), so each was
  passed over on every advert. Density now counts distinct mesh candidates
  (devices the discovery gate admits) in the last five seconds, and the
  pass-over is drawn per address per one-minute slot, mixed with the murmur3
  finalizer so consecutive slots are independent draws. The unknown-device
  bootstrap keeps the all-advert count, which is what it is about. Same
  change on iOS, where the hash was Swift's per-process seed; both platforms
  compute the same bucket for the same id (pinned by tests on both sides).
- Android: Bluetooth switched off under a running transport left every link
  in place, since the stack delivers no disconnect for most of them. The dead
  links counted against the connection cap, kept peers mapped to old
  addresses, and held GATT client registrations the stack had forgotten. The
  transport now reports each identified peer lost and clears the link state
  when it first sees the adapter off, mirroring iOS's dropLinksAfterRadioLoss,
  including the non-mesh cache that had marked a peer probed mid-recovery as
  not a mesh device for five minutes.

BleDensityPolicy on both platforms (9 JVM tests, 8 Swift tests), a registry
test, and the iOS restoration guard updated for the new call.
…uding a stack crash

Found on the phones while validating the previous commit: the Infinix's
Bluetooth stack crashed (SIGSEGV in libbluetooth_jni.so,
connection_manager::on_connection_complete) as the app dialled the other
phone. The stack restarted in under a second and took the app's GATT server,
advertiser and pending connect with it. The transport polls the adapter once
a minute from the scan watchdog, which read "enabled" throughout, so nothing
was rebuilt: the pending connect pinned the peer's address as "connecting",
and the other phone could not verify this one until the app restarted.

Android reports a crash as an ordinary ON -> TURNING_OFF broadcast. The
facade now registers for ACTION_STATE_CHANGED on the BLE handler once it
runs, and unregisters behind the shutdown barrier:

- TURNING_OFF / OFF: drop the dead links (once per outage), follow the dead
  scan locally, report BLE unavailable, arm recovery.
- ON after an outage: run the recovery now (scan, GATT server, advertising)
  instead of waiting for the ladder's next rung (12 to 29 s measured).

AdapterStateTransition maps the states (3 JVM tests); a Rust guard pins the
registration after RUNNING, the unregistration after the barrier, and the
radio-lost arm dropping links.
kivtxs added 7 commits October 6, 2026 00:35
…a system dialog

Closes the residual #465 left open ("the Android manager does not form a
group itself"). With wifiDirect.autoAccept on Android 10+:

- Devices of one app find each other over Wi-Fi P2P DNS-SD: the stream
  chapter's _offlineprotocol._tcp record (txtvers, addr) plus `app` (a tag of
  the app id) and `net` (the group name, while owning one).
- The lowest address creates a group under a name derived from its address;
  the others join it with a passphrase derived from the app id. Joining by
  credentials shows no dialog on either phone, unlike connect() by device
  address, whose invitation a phone in a pocket never answers.
- The query names the instance, because a query by type returns only the PTR
  record and never reaches the TXT listener. A group owner answers no service
  discovery query, so joiners target the name the lowest peer will create,
  and an owner whose group stays empty for 45 s dissolves it so it can be
  found again.

The rules are a pure object (WifiDirectGroupFormation) with 20 JVM tests. Off
by default; groupOwnerIntent stays unused. Measured on an Android 13 and an
Android 15 phone: launch to proved stream in 26-55 s over five clean starts,
chat 3/3 and 4/4 each way at ~100 ms.
Turning Wi-Fi P2P off reported the slot down; turning it on reported
nothing, restarted no discovery and, with group formation on, left the
framework's dropped service request and record unregistered, so every
later discoverServices failed with NO_SERVICE_REQUESTS. An app started
with Wi-Fi off never got Wi-Fi Direct.

The manager tracks whether it last reported the layer up, reports it up
again when P2P returns (a repeated broadcast is a no-op), restarts peer
discovery, re-registers formation's service request and record, and
re-registers the request on NO_SERVICE_REQUESTS. Formation pauses while
the layer is down instead of failing BUSY.

On the phones: app started with Wi-Fi off, then Wi-Fi on: group formed
and streams proved in 73 s; Wi-Fi off and on mid-session on one phone:
re-formed in 19 s.
…tials

The first formation design named the group after its owner, so a joiner had
to hear the owner's DNS-SD record before it could join. On two phones that
failed: a device that owns or is joining a group answers no service
discovery query, and the two often heard each other minutes apart or never.
Of nine clean starts, two never formed and two took about two minutes.

- The group's name and passphrase derive from the app id alone, so a joiner
  needs nothing from the owner. The `net` record entry is gone.
- The lowest address heard creates; others join, and take over creating
  after three joins found no group (the lower peer may never hear them).
- A device that heard no record joins only as a paced probe (once a minute)
  when a group owner is nearby: Wi-Fi Direct televisions are owners too, and
  joining on every step kept both phones deaf.
- A join attempt is cancelled after 15 s (it must scan, and the supplicant
  retries a rejected association) and attempts are 30 s apart, so a joiner
  stays discoverable half the time.
- An owner whose group stays empty for a randomised 30-60 s dissolves it and
  joins twice before it may create again, so groups formed at once merge.

Eight clean starts on an Android 13 and an Android 15 phone next to two
Wi-Fi Direct televisions: all formed, 18-225 s, median about a minute, no
system dialog. 18 JVM tests cover the rules, including the asymmetric cases.
@kivtxs
kivtxs force-pushed the chore/release-0.28.0 branch from 5720309 to dcb5e2f Compare October 6, 2026 04:38

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant