The working map of how a silo is built: the key hierarchy, the persisted state, the operation log, and the invariants every change has to preserve. Written for whoever changes this code next, human or tool.
This repository holds the domain crates, the persisted formats and the standalone extraction tool. It knows nothing about a user interface. The desktop application lives in silentsilo/desktop, and the parts of the picture that belong to it, the sync pass order and the session and lock lifecycle, are in that repository's own architecture map.
The other documents each own one slice: FORMATS.md owns the persisted bytes and CRYPTO.md owns the cryptography as an auditor reads it. This page owns the moving parts and the reasoning.
Keep it true. A change that alters anything described here updates this page in the same commit. A stale map is worse than none: it answers with confidence and it answers wrong.
SilentSilo is a local-first encrypted vault. There is no server. Every change to the tree is an immutable, encrypted operation record appended to a log; devices converge by exchanging records through dumb storage the user owns (S3, WebDAV, SFTP, a folder, OneDrive, Dropbox, Google Drive). File content lives beside the log as encrypted blobs, each under a key of its own. Everything a device shows is a disposable cache rebuilt from the log; the log and the blobs are the only things that matter, and the whole design bends around never losing either.
flowchart TD
subgraph domain["domain crates"]
VFS["silentsilo-vfs<br/>oplog, tree, names, snapshots"]
VAULT["silentsilo-vault<br/>sessions, keys on disk, cache"]
SYNC["silentsilo-sync<br/>transport: bucket layout, passes"]
CRYPTO["silentsilo-crypto<br/>seal/unseal, blob format, keys"]
end
subgraph edge["edge crates"]
STORE["silentsilo-store<br/>ObjectStore: folder, S3, WebDAV, SFTP"]
CLOUD["silentsilo-cloud<br/>OneDrive, Dropbox, Google Drive:<br/>sign-in, tokens, stores"]
S3C["silentsilo-s3"]
FIDO["silentsilo-fido"]
AUDIT["silentsilo-audit<br/>activity log: sealed events, segments"]
CORE["silentsilo-core<br/>shared types"]
end
APP["silentsilo-app<br/>sessions, sync pass order (being moved in)"]
EXTRACT["silentsilo-extract<br/>standalone recovery binary"]
FIXTURE["silentsilo-fixture<br/>format compatibility corpus"]
TESTKIT["silentsilo-testkit<br/>dev-only: hostile conditions, skip detector"]
CLIENT["client applications<br/>(silentsilo/desktop, silentsilo/mobile)"]
CLIENT --> APP & VFS & VAULT & SYNC & FIDO
APP --> SYNC & VFS & VAULT & STORE
SYNC --> VFS & VAULT & CRYPTO & STORE & AUDIT
AUDIT --> CRYPTO
VFS --> VAULT & CRYPTO & CORE
VAULT --> CRYPTO & CLOUD
CLOUD --> STORE & S3C
STORE --> S3C
EXTRACT --> SYNC & VFS & VAULT & CRYPTO & STORE
FIXTURE --> SYNC & VFS & VAULT
VAULT & STORE & SYNC & S3C -.dev.-> TESTKIT
Dependency direction is the rule worth defending: nothing here knows about a
user interface, silentsilo-sync is transport and knows no UI,
silentsilo-vfs owns the model and knows no network, and
silentsilo-crypto knows nothing above bytes. Nothing in this repository may
depend on a client application or on its OS integration crate. CI enforces
that with the no-ui-deps job.
silentsilo-app is the application logic every client shares, moving in
from the desktop's command layer one area at a time: so far the session map,
closing a silo, the sync pass, the recovery and device key flows, and file
previews. A client gives it a Host for events, diagnostics and its saved
storage settings. The order and invariants of the pass are described in the
desktop repository's docs/ARCHITECTURE.md until the move is done, with one
two steps that exist only here so far: before pushing, the pass reconciles
the enrolled keys with keys/ (silentsilo-sync/key_sync.rs, FORMATS.md),
and after pulling it imports the inbox (below). After its own records it
delivers the activity log (deliver_audit): it reads the log's policy every
ten minutes and pins the key the policy names, closes a segment once the
oldest queued event is fifteen minutes old, and sends what waits to every
copy. A segment leaves the device only once every configured copy holds it,
as a record does, so a copy in a drawer keeps it queued; a failure warns
and never stops the sync. A policy that later names another key than an
organisation's pinned one is not followed: the device goes on sealing to
the key it pinned, and says so. A personal log follows the newest policy.
An organisation's log (audit_admin) is started, given another reading
key, has its retention changed and its old segments removed only after a
touch of one of its keys, which the client asks for; core takes the wrap
key that touch gave.
A personal log is turned on and off on the device (AppState::set_audit_log),
copies or not: the queue keeps the key and the policy, and the pass writes
the newest policy, by changed_at, to every copy that lacks it
(settle_audit_policy), so turning it on never waits for storage.
Gathering is the caller's and reading is shared: the app and the extract
tool both hand their segments to silentsilo_audit::reading::read_log.
Reading (audit_read::read_audit_log) takes every segment from the cache
beside the silo, this device's outbox and its pending records, and every
copy it can reach, keeping what it fetched. It checks each device's chain
and event count from the oldest segment present (a retention may have
removed what came before) and names what is missing, records that do not
open with the log's key, and copies it could not read. A personal silo's
log opens with the content key; an organisation's only with an
organisation key's wrap key.
The device's queue (Spool) is held under a file lock while it is open, so
a command recording an event and the pass closing a segment take turns
instead of numbering two events alike. The pass opens it for each local step
and never across a network call, which keeps a recording command's wait
short.
The extract binary deliberately reuses the same crates rather than reimplementing the read path: a second interpretation of the log is a second thing that can be wrong.
flowchart TD
FIDO2["FIDO2 hmac-secret<br/>(per enrolled key)"] -->|"wraps"| DEK
SE["Secure Enclave ECDH / Android Keystore AES<br/>(macOS, iOS, Android, per enrolled key)"] -->|"wraps"| DEK
DS["device secret<br/>(keyring, per device)<br/>Argon2id + vault.salt"] -->|"wraps, until a key is enrolled"| DEK
RC["recovery code<br/>(160-bit, on paper)<br/>Argon2id + salt in envelope"] -->|"wraps"| DEK
DEK["vault DEK (32B)<br/>one per silo, shared by devices"]
DEK -->|"seals"| OPS["operation records (ops/*.op)"]
DEK -->|"seals"| SNAP["snapshots (snapshots/*.snap)"]
DEK -->|"seals"| DB["vault.db.enc (+ .bak, .next)"]
DEK -->|"seals"| KEKENV["keys/content.kek"]
KEKENV --> KEK["content KEK (32B)<br/>one per silo, never rotates"]
KEK -->|"wraps"| CK["per-blob content keys<br/>carried inside records / entry JSON"]
KEK -->|"seals"| PW["password entry JSON<br/>(rows + UpsertPassword records)"]
CK -->|"AES-256-GCM, AAD = blob_id"| BLOB["blobs/*.sslo"]
Why two layers under the DEK: a record's fingerprint covers its body, so nothing inside a record can ever be rewritten. Content keys therefore hide behind the KEK, and rotating the vault key re-wraps one KEK envelope plus the DEK envelopes, touching no record and no blob. That is what makes rotation affordable on a terabyte. The KEK itself never rotates; the DEK does. Password entries seal under the KEK for the same reason: the ciphertext travels inside records.
A security key's wrap key is BLAKE3 of its hmac-secret output for the
silo's salt. hmac-secret keeps two secrets per credential, one for
assertions verified with the key's PIN and one without, and Windows verifies
every ceremony on a key that has a PIN. So every platform does the same: a
key with a PIN is asked for it, one without never is. Windows does this in
its own dialog; Linux and macOS speak CTAP2 over USB HID through
silentsilo_fido::ctap2 (ceremony.rs, the code Android uses over NFC and
USB) and ask through the prompt the client registers with
set_pin_prompt. A silo made by a Linux build before core 1.9.0 on a key
with a PIN holds the unverified secret, and has to enrol that key again.
Rotation state machine (silentsilo-vault/rotation.rs, driven by the
client application):
stateDiagram-v2
[*] --> Staged: stage_rotation writes .next key files
Staged --> Resealed: reseal_under_new_key per deletable target (idempotent, resumable)
Resealed --> Committed: commit_rotation_with (keys and recovery staged, KEK renamed, the rest renamed)
Staged --> Staged: crash → resume with any enrolled key
Committed --> [*]: new recovery code shown, silos locked
commit_rotation_with takes the re-wrapped fido.json and the new
recovery.json and writes them as .next beside their targets before it
moves the KEK. Moving the KEK is the point of no return; everything after is
renames, and finish_interrupted_commit (called by load_fido_keys and
rotation_pending) finishes them with no key. Writing the keys after the
commit, as the app used to, left a window where the new key was in force and
nothing on disk opened it. Past the KEK the commit reports success whatever
the renames do: an error there hid the new recovery code while the old one
no longer worked. A rename refused because something holds the file falls
back to writing it in place, as every other key file does.
A rotation reaches every target or does not start: a target whose settings
will not open stops it before the first touch, since one skipped keeps the
old key readable there. An object that opens under neither key does not stop
it (ResealOutcome.unreadable): nobody could read it before either, and one
corrupt record otherwise blocked every rotation and resume for good. The KEK
envelope is the exception and still fails it.
The order is forced: the staged key is durable on disk before any object in
storage is re-sealed, because an object under a key that existed only in
memory is one nothing opens. Going backwards is never attempted; a second
rotation cannot be staged over a pending one. A device that was not kept in
the rotation is detected on its next pass, before it pushes: the published
KEK envelope only opens under the current DEK, and kek_envelope_state turns
that into needs_rejoin. Without that gate, the stale device pushed records
nobody could read and overwrote the rotated KEK envelope with its own.
The same check answers a second question, because that envelope alone cannot
tell a rotation from someone putting an old copy of it back. The reseal works
through ops/ and snapshots/ and writes keys/content.kek last, so an
envelope under a newer key always has records under that key beside it. A
device that cannot open the envelope therefore reads the newest records: when
they open under its own key, no rotation happened and the object was replaced
(KekState::Replaced), reported as storage having been written to rather
than as needs_rejoin. Rejoining reads the same object, so the old answer
sent a whole fleet round a loop with no way out. With nothing readable to
compare against the answer stays Rotated. The join flows say the same two
things (flows::kek_refusal).
With several copies, every copy that answers is asked and the gravest answer
decides (rotated, then replaced, then current): a copy that missed a rotation
still says current, and letting the first answer decide pushed its stale
envelope over the rotated one. Never-delete copies vote only on a silo that
has no working copy: a rotation leaves them on the old key, so a device on
the new key reads them as rotated, and letting them decide sent it to rejoin
in a loop, or whenever the working copy was unplugged. Such a copy, one that
reads as rotated beside working copies, is retired: the pass leaves it out
with RETIRED_COPY as its status and does not count it as a copy to reach.
Counting it held the inbox, the sweep and compaction for as long as it stayed
configured, since it can never be reached under the new key.
Three distinct places, and the boundary between them is a security property:
- The silo folder (user-chosen, portable, safe in a synced directory):
only ciphertext.
vault.db.encand.bak(the index, sealed under the DEK),vault.db.enc.next(mid-rotation only),blobs/*.sslo,vault.salt,master.dek.enc,keys/(fido.json, recovery.json). Every irreplaceable file here is written atomically, temp then sync then rename (workdir::write_private), and a blob is synced to disk before the import returns: the record naming it can reach the bucket within seconds, and a truncated key file or a hollow blob after a power cut is a lockout or a permanently unopenable file. - The local secrets (outside the folder): device credentials, storage
settings and the silo list. Kept in the OS keyring where there is one,
else in files. On Windows those files are DPAPI-wrapped; elsewhere a client
may register a
LocalProtector(Android: a Keystore key), and without one they are plaintext, private to the user or app. Backup targets of a kind core 1.7.2 does not know live intargets.more.config.json, never in the list or the single slot an older release reads (FORMATS.md). Test runs use the keyring servicecom.silentsilo.test, never the app's (keychain::service). The refresh token of each OneDrive, Dropbox or Google Drive target is kept the same way, undercloud-token:<target id>(FORMATS.md). The access token is never stored. - The machine workdir (keyed by silo path, outside the folder): the
working copy
vault.sqlcipherwith its WAL, ciphered by SQLCipher under a random page key, that key sealed under the DEK asvault.key(andvault.key.nextmid-rotation), and decrypted files the user opened (open/). A lock and the sweeps remove the plaintext and keep the ciphered copy and its key (workdir::KEPT_ACROSS_LOCKS); the next unlock reuses the copy while it still stands forvault.db.enc, and exports a fresh one otherwise. Only removing the silo deletes the copy. Unlock and snapshots move the decrypted index through memory only (CRYPTO.md, "Local metadata"). Beside it, kept across locks, the cache directory:cache.db(blob bookkeeping), and the protected folders' list and import ledger (protected.enc,protected.sqlcipherand its keyprotected.key), which between them name every mirrored file by its full local path and are therefore encrypted, under the content KEK so a rotation cannot strand them. Nothing here may ever land in the silo folder. - The bucket (per target):
vault.json(the only plaintext object, one random UUID),ops/,blobs/,snapshots/,keys/*.env,keys/content.kek,recovery.env. Layout and versions are FORMATS.md's jurisdiction.
Object keys sort meaningfully: op keys are
ops/{lamport:020}-{device_id}-{op_id}.op, so a plain listing is already in
apply order and both the Lamport value and the op id can be read without
downloading. Snapshot keys are the zero-padded horizon. Only blobs/ keys
are immutable-by-key; everything else may be rewritten in place (rotation
re-seals, envelopes re-wrap), which is why seeding size-skips blobs alone.
The app seeds with seed_target_checked: records, snapshots and the KEK
envelope are copied only when they open under the current key, and key
envelopes and the manifest only where the destination has none, so an
append-only copy that missed a rotation cannot undo it on another target.
Key envelopes go only for keys this device holds as in use: a never-delete
copy keeps the envelope of a removed key, and once the removal's tombstone
is gone nothing would delete it from the copy it was seeded into. The pass
after the seed publishes every key this device knows.
State is a pure function of the record set. The total order is
(lamport, device_id, op_id); replay sorts, so arrival order is
irrelevant. Alongside the Lamport value every record carries seq and
prev: a per-device hash chain that makes silently dropped or replaced
records detectable (verify_chains; see the gotchas for why it is not
wired). emit runs local changes through the same apply_op as remote
ones, inside one transaction with the Lamport reservation and the log row,
so the write path cannot drift from the replay path and a crash cannot
leave an effect without its record.
replay takes one transaction for the whole batch, with a savepoint around
each record. A record still stands or falls on its own, and everything
applied before a record this build refuses is committed before the refusal
goes up, so a device keeps what it could read rather than meeting the same
wall from further back every pass. When the caller already holds a
transaction, that one decides and replay adds nothing.
Name resolution is the subtle part. Uniqueness is per folder,
case-insensitive and Unicode-composed (names::fold). The name_claims
table records which operation claimed which name; ranks within a claim
group assign name, name (2), and so on, as a pure function of the
record set. A suffixed name that another entry in the folder asked for
outright is skipped, so report (2).pdf given to a second report.pdf
never meets a file really called that. Subtree queries use GLOB, not
LIKE: LIKE folds ASCII case and would treat x (2) and X (2) as one
subtree. Typed names are validated (names::check) and NFC-composed at
the boundary; replayed names are repaired (names::sanitize) because a
record can never be refused. Concurrent edits of one file resolve by total
order, with the loser preserved as a deterministic conflict copy (id
derived from the losing record via UUIDv5, carrying the losing content's
wrapped key).
A snapshot at a chosen Lamport horizon stands in for every record at or
below it. choose_horizon keeps a time margin (30 days, from untrusted
timestamps) and a count margin (500 records, which holds when a clock
lies), never splits a Lamport value, and only the device that has just
synced everything may compact. Order, one way only: capture, publish the
snapshot to every target, read it back and refuse to prune unless the bytes
in storage are the bytes written (it is about to become the only copy of
everything below the horizon), delete covered ops from targets that allow
it, prune locally (unpushed records are never dropped and are reported as
stranded). A snapshot is captured from this device's base plus its log,
never from the log alone: after a rebuild or an earlier compaction the log
starts at the base, and a snapshot of the log alone left out everything
below it. A device that falls below the horizon gets needs_rebuild and
comes back via the snapshot with a fresh device id, because its old chain
positions died with the log. Append-only targets keep their whole log and
still receive the snapshot, which is what a joining device replays from.
Whether a device is behind is judged against the lowest horizon across the
targets that answer. A target below the highest horizon counts only when its
records reach that horizon: a drive left in a drawer since before the others
were compacted has a low horizon because it is stale, and it does not hold
what they pruned. A folder target whose root is gone (an unplugged drive)
is Unreachable for every call, never an empty listing and never a write
that recreates the root and fills it as a new copy; a missing prefix under
an existing root is still an empty listing. Only check, which runs while
the user adds or tests the place, creates the root.
The received mark alone cannot see one case. A device that wrote offline for
a month sends records whose Lamport values the others passed long ago; they
are older than the policy's month by their own clock, so the first
compaction folds them in and prunes them, possibly in the same pass that
sent them. A device whose mark is above the horizon then never fetches them
and never counts as behind. So once per new horizon each device replays its
own log to that horizon (vfs::state_at) and holds it against the snapshot
(vfs::holds_more): anything the snapshot has that it lacks, or has
differently, sends it to rebuild. Only that direction counts, so a snapshot
short of state, as 1.6.1 published after a rebuild, is not rebuilt from. The
horizon that agreed is kept as snapshot_checked_through. When no target answers, the pass marks each
one failed and backs off instead of stopping with an error.
stateDiagram-v2
[*] --> LocalOnly: import (encrypt_file, record_blob_present)
LocalOnly --> Delivered: put_from_file per target, record_blob_delivered
Delivered --> Synced: settle_blob_delivery (every configured target has it)
Synced --> Evicted: cache limit (LRU, never full-copy silos, never unsynced)
Evicted --> LocalOnly: fetch_blob on open/export (from any copy)
Synced --> Candidate: sweep sees it unreferenced
Candidate --> Deleted: still unreferenced on a later sweep, 30 days on
Candidate --> Synced: a record referencing it arrives
Deleted --> Synced: a row points at it and a device holds it (restored)
The referenced set is files.blob_id (trash included, a restore needs the
bytes) plus password attachment blobs parsed out of the decrypted entries,
because attachments have no row anywhere; the sealed entry is their only
reference (Vfs::referenced_blobs_with_attachments). Purge and attachment
removal clean the local cache only, and a purge leaves content no copy holds
yet until a push has sent it (files::release_purged_blobs): a device that
has not seen the purge may keep an edit pointing at it. The pass drops such
content once every copy has it (drop_delivered_unreferenced). The bucket
copy is always the sweep's to delete, because the sweep re-asks after the ops have converged, and a
row written concurrently on another device may still need the bytes. All
transfers stream through disk (put_from_file/get_to_file); nothing
holds a whole blob in memory.
Every backend also answers the same two transfers with a byte count as it
goes: put_from_file_reporting and get_to_file_reporting take a callback
that receives the bytes moved since the last call and answers whether to
carry on. A ControlFlow::Break stops the transfer and the call returns
StoreError::Cancelled, which SyncError carries through as Cancelled.
The default implementations move the whole file and report it once at the
end, so a backend that has not been taught to stream still answers. That is
what lets a seed put a number on one large blob instead of a counter that
sits still for minutes, and what lets Stop land inside an object rather than
after it.
An ordinary pass says the same thing: push_blobs_reporting names each blob
before it looks at it, then reports how much of it has gone up, at most every
250 ms, ending with all of it. The pacing is the point as much as the number:
a client turns each report into an event, and a gigabyte at 512 KiB a report
would be two thousand of them. Nothing stops a pass inside a blob, because no
client has a stop for a pass to answer; that callback always says carry on,
and a cancel would be one argument on the call.
The sweep keeps a candidate for 30 days after this device first saw it
unreferenced (snapshot::gc_first_seen). A move records the file again over
the content its device holds, and a device that has not synced for days can
make one on top of an edit or a purge it has not received: without the grace
the others had already deleted that content. The same listing puts back
content a row here points at and the target lacks, from the local cache or
another copy (sync::restore_missing_blobs). A 1.0.0 device sweeps after
two sightings and has no row for some content this build keeps (below), so
beside one the content goes and comes back until it updates.
A device that cannot open the silo sends content to inbox/, and an
unlocked device imports it (silentsilo-sync/src/inbox.rs, formats in
FORMATS.md). The import copies the item into blobs/ before it records
the file and deletes the item only after. Nothing waiting may live under
blobs/: the sweep would delete it after two passes, on any version.
In silentsilo-app the import runs after the pull and before housekeeping.
An item recorded in one pass leaves the inbox in a later pass, and only when
that pass reached every configured target: the record is pushed at the start
of the pass after the one that wrote it, so the item is never gone from
storage while the only trace of it is one device's database. An archive
target keeps its items; they are skipped as already known. With more than one
target the importer fetches the content down, so its next push spreads it to
the targets the phone did not send to.
Several unlocked devices import the same items at once, and nothing stops that. What keeps them agreeing:
- the file id is the item id, and the folders the import creates take ids
derived from parent and name (
Vfs::ensure_folder_path), so both devices write the same file into the same folder. Random folder ids made a "(2)" folder, and each device kept the files in its own. Where that id is taken (a trashed or purged folder of that name) the id is derived from the item as well, the same on both devices; a fresh id split the items again. A device whose base hides the purge reuses the plain id, and replay keeps that folder: a creation is dropped only when a purge of its id sorts after it (purged_after), since a purge never deletes what its author did not name. Two imports of one item that still name two folders are settled by order: the creation first in the order places the file on every device, while no rename has claimed it since; - content already in
blobs/at the signed size is not copied again; - an item another device finished between the listing and the read is skipped, not an error that stops the scan;
- before an item leaves the inbox its content is checked in
blobs/and copied again when missing. A device that recorded an item and stayed locked for days can find that copy swept by another device, which does not know the file yet; - a sender never sends an item again once its envelope is in the inbox. A
resend used to seal the same item under a new blob id at the same size
and overwrite the content, so an import between scan and copy recorded
bytes whose header named another blob. The import also reads the header
(
get_prefix, 82 bytes) before and after the copy and refuses content sealed for another item or blob, which covers senders on older builds.
What gets someone out of which hole, all of it built from the same pieces
(fetch_join_plan picks snapshot-plus-tail or whole log by asking the
bucket):
| Situation | Way out |
|---|---|
| Lost the security key, machine fine | vault_unlock_with_recovery (code, local or bucket envelope) |
| Machine gone entirely | vault_join_with_recovery, or a key on a new machine via vault_join_from_storage |
| Machine gone, project gone | silentsilo-extract: files, trash to one side, passwords as CSV, attachments |
| Local index corrupt beyond both snapshots | vault_repair_from_storage, offered by the unlock screen, in place, blobs kept |
| Device below the compaction horizon | vault_rebuild_from_snapshot (any copy that holds one), fresh device id |
| Device rotated away | needs_rejoin: remove and rejoin with a current credential |
| Bit rot in a copy | vault_verify deep read finds and names it; nothing replaces it yet. restore_missing_blobs puts back only objects that are gone: a damaged one has the right key and size, so every pass takes it as present |
Read this before "fixing" any of it.
verify_chainsis never called. Wiring it naively false-positives on every rebootstrapped device, whose new chain legitimately starts above zero from the fetcher's point of view. It waits for a checkpoint design.derivationon key envelopes decides more than it looks. It shipped before anything varied, as groundwork for platform authenticators, and the macOS build is the first to use it:ecdh-p256-hkdf-sha256-v1alongsidehmac-secret-v1. Every device reads every other device's envelopes, and a device skips the ones whose derivation it cannot perform. Mobile will add to the list rather than change it.- Purge does not delete bucket blobs. The sweep does, two-pass, after convergence and a 30-day grace. Deleting at purge time froze "referenced" at what one device knew and destroyed content another device still pointed at, and so did the sweep without a grace, for a move made on a device that had not synced.
- A pass uploads content it did not create. The sweep puts back what a row points at and storage lost. It never sends content no row here references, so emptying the trash is not undone by a device that still holds the bytes in its cache.
- Compaction and the sweep skip a partial pass. A record held back, an
unreadable object or a copy that could not be read means this device's
idea of what is referenced is short, so neither runs. The same complete
passes record
received_through(vfs::snapshot), and the horizon check compares that, not the highest local Lamport value, which counts this device's own offline writes. - An opened record under another record's name is skipped. The name is where the order and the horizon filter come from; storage copying a genuine record to a later name would otherwise replay it there. Snapshots likewise count only when their sealed horizon matches the name.
- Arrival order decides nothing, across passes too.
replaysorts a batch, but records reach a device over many passes in whatever order storage and other devices deliver them. So the derived state is worked out from kept records rather than from what happened to be applied first: trash state from every trash and restore record (trash_events), a file's content and conflict copies from every edit (content_versions), password edits and renames guarded by total order (password_order, the claim's own order), purged ids remembered (purged_ids), and each purge with its place in the order (purges). A conflict copy an edit has built on is retired unless a record touched it (trash, rename, edit, star); a record aimed at a file not here yet waits inpending_touchesand applies when the file appears, and a copy with such a record stands even once superseded. Before this a device that retired the copy first dropped the record, and a device rebuilt from the log did too.tests/arrival_order.rsapplies records one at a time in random causal orders;silentsilo-app/tests/fleet.rsruns three devices on one storage. - A purge never deletes what its author did not name. Entries another
device put in a purged folder meanwhile move to the top of the silo, as do
entries a later record creates there (
purged_ids). A rebuild writes again, on top, only what this device wrote and no copy took (vfs::undelivered_own_ops,sync::apply_rebuild). Notpushed = 0: that flag waits for every copy, so beside an unplugged drive or a never-delete copy under a replaced key it is unset on everything, and a rebuild that took it wrote old history again as new changes, purged folders included. - A purge keeps the edits its author could not have seen. An edit to a
purged file made on a device that had not received the purge becomes a
file of its own at the top of the silo, named as a conflict copy of the
purged file and holding exactly that content; the purged file stays gone
(
settle_kept_edits). What the author had seen: a purge this build writes names, for each file, the conflict copy id of every content record it held and a marker saying the list is complete (purge_groups,purge_marker), one file's group never split across records. An edit it lists, or one a listed edit was written on top of, was seen. A copy it does not list came from an edit it never had, so every edit to that copy is unseen. A purge without the marker, from an earlier build, is read by the order: its author's clock was past every record it had received, so an edit from another device at the purge's Lamport value or above was unseen, and one below counts as seen. The purging device's own later edits count as seen. Every unseen edit is kept, not only the newest, and one purge that missed it is enough: the newest edit and the set of purges both change with arrival, and a kept file that went away again would take later work on it along. What a purge saw only grows as its records arrive, so an edit kept while the edit the purge listed (or one written on it) was missing is taken back when that arrives, unless a record touched the file it became: trash, rename, edit or star (retire_kept_edit). A record aimed at a kept file not made here yet waits inpending_touchesand keeps it. A conflict copy counts as purged with its file (copy_origins). The purge leaves the file'scontent_versionsin place and records what it named inpurges; the kept file's id derives from the file id and the edit's op id (kept_edit_id), and its name follows the purged file's last name (purged_names), so every arrival order and a rebuild give the same file. A snapshot carries all of it (Snapshot::purged), with the versions and copies of every edited file: a device rebuilt from one judges a later purge against the same edits as a device that replayed them. A 1.0.0 device ignores the markers and drops the edit; the kept file references the content, so the sweep here keeps it andrestore_missing_blobsputs back what a 1.0.0 sweep deleted. - A large purge is several records. Readers refuse records over their size ceiling, so a purge is split files first, then folders deepest first, and each record stands on its own for a 1.0.0 reader.
- This build writes nothing a 1.0.0 device refuses from what it holds.
A 1.0.0 replay stops for good at a record it cannot apply, so the
Vfsauthors around 1.0.0's rules with records every version applies (oplog::plan_claim_for_1_0_0,touches_after_purge,purge_groups). 1.0.0 ranks name groups with no gaps; while no group needs a suffix another group asked for, both rankings agree, so a record joining a group asks for the name shown where they would not, and asking for a taken suffix renames its holders first. Emptying the trash moves live entries out of trashed folders before the purge, a purge names the conflict copies it takes along, and the last entry of each group a purge left gaps in is renamed to the name it already asked for, because 1.0.0 does not rank a group again on purge. The extra renames look redundant here and are not. Changes made at once on two devices can still meet on 1.0.0.silentsilo-fixture/tests/refused_by_1_0_0.rsreplays random histories into a 1.0.0 database;mixed_fleet.rsfails when 1.0.0 refuses a record this build wrote given only what its author held. - Seeding size-skips only
blobs/. Everything else is rewritten in place at identical length by rotation, so "same key, same size" would skip the one write that matters. Which side is current is decided by the key (seed_target_checked), not by size or time. - A seeded object counts its size once, not twice. It moves twice, down
from the source and up to the destination, and
SeedProgress.bytes_totalis the size of the silo: each leg credits half the object, and whatever is left is credited when the object is done with, skipped and failed ones included. So the bar ends where the object count does. Counting the bytes that actually cross this machine would show a total twice the size of the thing being copied, which is a number nobody can check against their storage. - A stop on the last progress report still stops. Whether the final chunk falls inside a backend's loop or after it is an accident of that backend, and an answer that depends on it would make Stop work on SFTP and not on S3. An upload stopped that way leaves the key untouched wherever the backend writes through a temporary name (a folder, SFTP) and may leave the whole object where it does not (WebDAV, a small S3 PUT). Never a short one: the seed re-copies it on the next run either way.
- S3 uploads any file over 16 MiB in parts, a sync pass as well as a seed. One PUT reports nothing until it ends, cannot be stopped, and S3 refuses a single PUT over 5 GiB, which capped what a silo could hold at that per file. Parts are sequential, and a stop or a failure aborts the upload rather than leaving parts that are billed and invisible. Below the threshold it stays one PUT, which is every record, every snapshot and most blobs.
- A multipart upload first aborts every unfinished upload of its key.
A killed process aborts nothing, and the retry that always follows starts
a second upload, so the first one's parts would stay billed and invisible
for good. Another device sending the same key at that moment (two devices
putting back the same missing blob) loses its upload and retries on a
later pass; the bytes are the same either way. The daily blob sweep also
aborts unfinished uploads older than 24 hours under
blobs/,snapshots/andinbox/, for content deleted before anything retried it. Never younger: that may be another device mid-upload. Never from the bucket root: another program's uploads may live there. MinIO lists unfinished uploads only for a whole key, so there the sweep finds nothing, and MinIO expires them itself. - S3 HEAD treats 403 as absent. A prefix-scoped credential gets 403 for a missing key; callers use HEAD to decide whether to write, writes are idempotent, and a genuinely bad credential fails loudly on PUT.
init_opensslbefore every connection. SQLCipher's vendored OpenSSL was built with the build machine's path as OPENSSLDIR and would readopenssl.cnffrom there, or fromOPENSSL_CONF, when SQLite first initialises; a config can load a provider library. Every place here that opens a connection initialises OpenSSL first with no config (silentsilo-vault/openssl.rs). A client that opens its own connection has to call it too.tests/openssl_config.rsproves it in a child process.- S3 brings its own HTTP client on Android only. The SDK's client reads
root certificates from files, and a phone has none where it looks, so every
handshake failed. It also takes no custom verifier. On Android
silentsilo-s3hands the SDK a small hyper client (https.rs) with the platform verifier fromtls.rs; desktop keeps the SDK's client and native roots. WebDAV uses the sametls::client_configeverywhere. That verifier panics on Android until the app callsinit_android_tls, so it sits behind a readiness check that refuses the handshake instead (README, "HTTPS on Android"). SetDeviceLabelandAnnounceDeviceare separate ops so a machine re-announcing its hostname can never overwrite a name a person typed. Empty label means "no label", deliberately.settle_delivery([])marks everything unpushed. No targets means no copy holds anything, and compaction must not prune on the strength of copies that no longer exist. Same shape for blobs.fetch_blobdoes not mark the blob evictable even though it plainly came from a target: eviction needs every target to hold it and this call knows only one. The next pass settles it.is_skippableis a per-variant decision, carried on the record, because it is all an older build has to go on when it meets an operation from a newer one. Decoration skips; structure refuses.MAX_OP_BYTESrejects from the listing, before download: storage is untrusted and an object sized to exhaust memory must never be fetched. Same posture as the Argon2 parameter ceiling onrecovery.env. Key envelopes, revocation markers, inbox keys and senders, the manifest, the KEK envelope andrecovery.envgetMAX_SMALL_OBJECT_BYTES(64 KiB) the same way, from the listing or a HEAD. A server that answers the HEAD small and the GET huge still gets its bytes read:ObjectStore::gethas no limit.- The working copy is not called
vault.db, and its key sits beside it. An older release after a downgrade decrypts the snapshot over anything namedvault.db, with the ciphered WAL still beside it; under its own name the copy is never touched. The page key is sealed under the DEK on disk rather than held in memory because a key that dies with the process takes the copy with it: a crashed session's changes, and the next unlock's reuse.stage_local_backupseals it under the new DEK too, so a crash after a rotation commits still adopts.commit_rotationmoves it overvault.keyright away (or deletesvault.keywhen there is none), so the retired DEK opens nothing here after the commit. - The working copy outlives the lock. Exporting a 12 MB index into a
fresh SQLCipher copy took 600 ms of every unlock. Reuse is decided by
fingerprint, not by timestamps: each snapshot write records the BLAKE3 of
the sealed
vault.db.encin the copy (working_copy_state), after the image was taken, so a recorded fingerprint always means "this copy holds that snapshot or more". An unlock reuses the copy only when the file on disk matches one (a rotation's staged.nextincluded), or when the snapshot is damaged and the shadow copy matches. A missingvault.db.encis put back from the shadow copy (or a staged snapshot) on unlock, andVaultPaths::existsandSiloEntry::is_presentcount those too: a sync client that removed the snapshot left a silo shown as unplugged. Anything else that wrote the snapshot (a rotation elsewhere, a repair, a restored file, another release) gets a fresh export. A lock also recordslocked; a copy that has it is used as is, with noquick_check(130 ms on 12 MB), because the lock just read every table through SQLCipher's page HMACs. A copy without it never locked, so it is checked and the snapshot catches up, as after a crash. The table is dropped from every image, so no release sees it. A lock with nothing written since the snapshot keeps it rather than sealing it again. "Nothing written" istotal_changes()plus the schema cookie, noted in a TEMP table of the same connection when the snapshot is written, exported or adopted. Never a timestamp or the revision, which sync writes do not bump; the note dies with the connection it counts. - Cloud targets open through a hook the vault installs.
StoreConfig::openlives insilentsilo-store, which cannot reach the keyring, so the three cloud kinds open throughset_cloud_opener, set bysilentsilo_vault::install_cloud()at startup. A client that never calls it refuses those targets with a message instead of opening them without a token. OneTokenSourceper target for the life of the process, so a pass that opens a store again does not refresh again, and one refresh at a time: two transfers must not both spend a refresh token Microsoft rotates. - A sign-in's tokens wait in memory until a target adopts them. The UI
gets the account to show and an id. The target is built from the account
that sign-in reached, never from what the UI sent, checked with the pending
tokens, and only stored under the target id once the list is saved: a
failed check leaves no token behind. A reconnect must reach the same
account, or the target would point at an empty folder. A sign-in nobody
adopts goes when its dialog closes (
cancel_cloud_sign_in), when the silos lock (forget_cloud_sign_ins) or after 30 minutes, and a Dropbox one is revoked as it goes. - Adopting a sign-in keeps its token source. The store that was checked with it goes on using the same source, so a rotation from then on is written under the target. Each target's source carries a generation, and only the current one writes: a pass still holding the source a reconnect replaced would otherwise put an older token over the new one, and one holding the source of a removed target would write a token back for it.
- On a phone the code is traded after the app is back. The browser is
answered as soon as the redirect arrives, and the exchange waits for the
readya client passes (sign_in_when,cloud_sign_in_when); the phone passes "the app is on screen again". Android keeps an app off the network while another app is in front, and a request made then fails its certificate revocation check, after which the platform verifier answers "revoked" for that host for about half a minute. The exchange also retries a request that never got out, for up to a minute: never one that did, since the code may be spent. - Only Dropbox sign-ins are revoked on removal. Google's revocation ends the whole grant, every other computer's sign-in to that account included, and Microsoft has none for a personal account's token. Removing a target forgets its token here either way.
- Tokens never follow a redirect. None of the HTTP clients follow one;
the OneDrive download's 302 is followed by hand with the plain client, and
upload session addresses (OneDrive, Google) and OneDrive's
nextLinkare checked before use. Graph refuses a token on an upload URL anyway. - OneDrive is personal accounts only (
consumers).Files.ReadWrite.AppFolderexists only there; a work account would have to grant every file it holds. A drive that answers as anything butpersonalis refused. - Google Drive may hold two files with one name. Drive addresses by id,
and two devices can create the same key or the same folder at once. The
newest file by (modifiedTime, id) is the object; a writer removes only the
older copies, a delete removes them all, and a listing names each key once.
Twin folders are read as one: a read that misses asks Drive again rather
than trusting the folder cache, because the twin another device made after
the cache was filled holds files a read must see. The layout is real
folders, so a folder downloaded from drive.google.com opens with
silentsilo-extractlike a local copy. - A Google Drive store remembers what it found missing. Folder paths
and keys found absent, and folders looked up fresh, are kept for the life
of the store, which is one sync pass: every new record and blob is asked
about before it is written, and on an account with several silos the twin
search found the others'
vault.jsonandrecovery.envevery time. Another device's write in the meantime is seen by the next pass, as with a listing, and a duplicate it leaves is one the rule above settles. - Small objects are read in one request where the backend can.
get_smallreturns the bytes or says the object is absent, and refuses one over the limit unread: the default asks for the size first; OneDrive and Dropbox read the download and stop past the limit, Google Drive knows the size from its lookup. probeis a folder nobody writes. Asking a provider for the account or the silo folders goes through a store built with that placeholder name, which keeps one HTTP path per provider instead of two.- Recovery codes map O→0, I/L→1, U→V on input. Crockford's alphabet excludes those on output precisely because handwriting confuses them; strict parsing would reject correct codes.
The checklist, in order:
-
Does it touch a persisted format? The list is in FORMATS.md, and the question to answer is what a client on the previous release does with the new bytes. Only two answers are acceptable, and they must be true by test: it ignores the new thing safely, or it refuses explicitly and tells the user to update. Anything else needs a version field first. There is an installed base from 1.0.0 onward, so this is never theoretical.
-
Does it change what replay produces? Then every device must produce the same result from the same records, in any order, and a test proving both orders belongs next to the change. The derived tables are dropped and rebuilt from the local oplog, which is what makes schema changes free: never add an in-place SQLite migration.
-
Does it touch delivery, eviction, sweeping or compaction? State the invariant it preserves: nothing is marked sent before storage confirms; nothing local is dropped unless every copy holds it; nothing in a bucket is deleted unless it is covered by a snapshot or unreferenced across two sweeps and for 30 days; nothing is put back that no row references.
-
Run the whole CI sequence locally, from the workspace root, in the order
.github/workflows/ci.ymlruns it. The cargo commands need--allfrom the root or they silently skip crates. Addcargo test -p silentsilo-fixturewhen replay or formats moved: a fixture whose decoded output changes is a break, not a fixture to update.Anything touching a storage backend goes through
scripts/test-local.ps1, which brings up MinIO (built from source), WebDAV and SFTP (in containers) and runs the same sequence with the endpoints set. Those suites skip themselves without an endpoint, so on a developer machine they otherwise never run at all. That script setsSILENTSILO_TEST_REQUIRE_BACKENDS, which turnssilentsilo_testkit::skip_or_failfrom a printed line into a failure: a suite that skips itself during a run that asked for it is a hole, not a note. -
Does it change a write the app cannot afford to lose? Then it belongs in
silentsilo-vault/tests/hostile_environment.rs, which runs each of those writes with something holding the destination open and with the temp path blocked. A clean temporary directory is not the machine the app runs on: the worst defect found so far was a temp-then-rename that a scanner's open handle refused, and no test in the suite held anything open. -
Does a client have to change with it? A release here is a tag, and
silentsilo/desktoppins one. Say so in the release notes; the client moves its pin deliberately, which is the moment the two are tested together. -
Update this page and FORMATS.md in the same commit when behavior they describe moves.