Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .dagger/src/index.ts
Original file line number Diff line number Diff line change
Expand Up @@ -1505,6 +1505,8 @@ export class AtomicServer {
`--test-threads ${this.hostKnobs.nextestTestThreads} ` +
`--retries ${this.hostKnobs.nextestRetries}`,
])
// Select the core alone so server dependencies cannot mask missing features.
.withExec(['sh', 'lib/tests/check-instance-checkpoint.sh'])
.stdout()
);
}
Expand Down
6 changes: 6 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -58,6 +58,12 @@ See [STATUS.md](server/STATUS.md) to learn more about which features will remain
switch. `Db::init_redb_file` now owns the 100ms durable-flush tick for every
binding, and an idle tick no longer writes anything.
- Fix remaining `clippy` warnings in `wasm/src/lib.rs` blocking `develop`'s pre-commit hook ([#1508](https://github.com/ontola/atomic-server/issues/1508)).
- Instance checkpoint capture, verification and offline restore are available in
`atomic_lib` through the optional native `backup` feature, without server/Actix.
- Add opt-in full-instance backups: `--backup-dir`, local authenticated backup
control, `backup`/`restore` CLI commands, checksummed ZIP64 archives and an
offline restore guard. Writes and incoming sync application pause during
capture and resume before compression. See [instance backups](docs/src/instance-backups.md).

## [v0.41.0-beta.7] - 2026-09-12

Expand Down
3 changes: 3 additions & 0 deletions Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

24 changes: 24 additions & 0 deletions TESTING_COVERAGE.md
Original file line number Diff line number Diff line change
Expand Up @@ -180,6 +180,30 @@ duplicate issues/comments. Live proxy OAuth,
GitHub writes and a guided uncertain-write recovery UI remain unverified/unbuilt;
proxy v40 CORS and browser OAuth are verified, but its GitHub credential returns 404 for the private sandbox.

Instance backup: `lib/src/backup.rs` covers full redb/file round-trip including
Loro historical checkout, blobs, envelopes and configuration files; future byte tables;
concurrent writer exclusion; buffered-batch refusal; capture error/panic recovery;
unsafe paths, hash mismatch and corrupt-database refusal.
`server/src/backup.rs` covers loopback/token authorization and adapter setup.
`lib/src/db/maintenance.rs` covers nested admission during drain,
overlapping pause refusal, queued work and cancellation recovery. The shared
sync-engine test proves a paused import is neither acknowledged nor dropped.
`server/tests/it/instance_backup.rs` starts a real server process, replicates a
second node over WebSockets, invokes the backup CLI, replicates a later change,
restores the earlier checkpoint and verifies accidental startup is refused.

`lib/tests/check-instance-checkpoint.sh` selects only `atomic_lib` with
`backup,config`, rejects server/Actix dependencies, runs the core backup tests,
and tests an originless capture/verify/restore/reopen with persisted identity in
`lib/tests/instance_checkpoint.rs`. Dagger runs this separately after workspace tests.

Remaining backup coverage gaps: desktop VFS staging drain and native UI restore,
OS-level disk-full/power-loss injection,
large (>4 GiB) ZIP64 fixtures and pause-duration benchmarks, and a dedicated
Iroh disconnect-during-capture test. The gate is shared by both transports;
these tests do not assert a globally synchronized checkpoint or lifetime history
retention. See `docs/src/instance-backups.md` for the operational limits.

What is tested, at which layer, and — the part that matters — **what is not**.

This exists because the protocol is far better tested than the glue around it,
Expand Down
1 change: 1 addition & 0 deletions docs/src/SUMMARY.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,7 @@
- [When (not) to use it](atomicserver/when-to-use.md)
- [Installation](atomicserver/installation.md)
- [Browser peer sync](browser-peer-sync.md)
- [Full-instance backups](instance-backups.md)
- [Using the GUI](atomicserver/gui.md)
- [Tables](atomicserver/gui/tables.md)
- [AI and Atomic Assistant](atomicserver/gui/ai-and-atomic-assistant.md)
Expand Down
141 changes: 141 additions & 0 deletions docs/src/instance-backups.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,141 @@
# Full-instance backups

Run a dedicated replica with an output directory outside its data and config
folders. Data and config must be separate, non-overlapping directories; keep
`--cache-dir` outside both too:

```sh
atomic-server --data-dir /srv/atomic/data --config-dir /srv/atomic/config \
--backup-dir /srv/atomic-backups
```

Request a backup on that same machine:

```sh
atomic-server --config-dir /srv/atomic/config backup --server http://127.0.0.1:9883
```

The command prints the archive path after verification, or exits nonzero on
failure. The server creates `backup.token` in its config directory with private
permissions. Backup control requires that token and a loopback connection; drive
write access does not grant instance backup access. Do not share the token.
The `--backup-dir` option also accepts `ATOMIC_BACKUP_DIR`.

The server drains admitted operations (up to 30 seconds), temporarily gates new
HTTP requests with `503` and `Retry-After`, and pauses incoming sync application.
It copies every redb table through one read transaction into a fresh database,
while a storage write transaction prevents all writers from crossing the capture
boundary. This includes Loro histories, retained signed envelopes, blobs, node
identity and internal metadata. Associated data/config files are captured under
the same barrier. Symlinks and special files fail the backup rather than escaping
the source directories. Operator edits to these directories should wait until the
capture finishes.

Capture time grows with database and file size; multi-GB pause duration has not
yet been benchmarked. Normal writes and queued sync resume before the staged
files are compressed.
Sync connections need not be disconnected: pending operations wait for admission
and then continue; a timed-out connection uses the existing reconnect/reconcile
path. No maintenance update is acknowledged and discarded. The output is a
ZIP64 archive named `atomic-backup-<UTC timestamp>-<process id>.zip`. File sizes
and BLAKE3 hashes, package version, Git revision, redb version, capture time, retention policy and source path
mapping are recorded in `manifest.json`. The ZIP is read back and hashes checked
before it receives its final name. Partial work uses private temporary paths.

This is a checkpoint of **this instance**, not proof that it received every
change from every peer. Replication freshness is explicitly reported as unknown.
It cannot recover Loro history or envelopes that were previously deleted or
never replicated. Cached vector indexes are outside the data/config roots and
are rebuilt when needed. External services, environment variables, separately
mounted storage and an exact executable are not bundled; retain your binary and
deployment configuration alongside the archive. A ZIP contains private data and
credentials; keep an independent protected copy off-machine.

## Restore

Use the same server version and a destination that does not exist:

```sh
atomic-server restore --archive /srv/atomic-backups/atomic-backup-EXAMPLE.zip \
--target /srv/atomic-restored
```

Restore validates the manifest, paths, all file hashes and the redb database
before promoting staged files. It does not boot the node or start networking.
The result contains `data/`, `config/` and `manifest.json`. A marker in `data/`
prevents accidentally starting a normal server with copied node identities.
Inspect the restored files first. For actual recovery, stop the original instance
before explicitly activating the restored one:

```sh
atomic-server --data-dir /srv/atomic-restored/data \
--config-dir /srv/atomic-restored/config --activate-restored
```

Reapply the original deployment settings, including the envelope-retention
policy recorded in the manifest. Activation can reconnect existing peers and
integrations. Reconnecting an old
checkpoint can merge newer remote changes back into it. Experimental branches
need a separate, network-isolated environment; the restore command deliberately
does not start one automatically. Do not activate both copies with the same node
identity on the same network.

## Nightly scheduling

`scripts/backup-instance.sh` accepts `ATOMIC_SERVER_BIN`, `ATOMIC_BACKUP_SERVER`
and the required `ATOMIC_BACKUP_TOKEN_FILE`. For example, a cron entry:

```cron
0 2 * * * ATOMIC_BACKUP_TOKEN_FILE=/srv/atomic/config/backup.token ATOMIC_SERVER_BIN=/usr/local/bin/atomic-server /path/to/atomic-server/scripts/backup-instance.sh >> /srv/atomic-backup.log 2>&1
```

On macOS, use the same script and environment with a launchd job whose
`StartCalendarInterval` has `Hour=2` and `Minute=0`, and set explicit
`StandardOutPath`/`StandardErrorPath`. The server rejects overlapping jobs. The
CLI waits for its job, detects server restarts or replaced status and exits
nonzero on failure. Check logs and disk capacity; nothing prunes old backups.

For status, use authenticated `GET /__atomic/backup` from loopback. An
unauthenticated request cannot see filesystem paths or backup errors. `POST` to
that route starts a job and returns its ID. Client disconnect does not cancel a
capture. Process crashes may leave `.atomic-backup-*` staging directories, but
never make a partial ZIP appear complete; remove stale staging only while no
backup job is running. Restore promotion errors may leave a partial destination;
inspect it and choose a fresh destination for retry.

Test a restore before relying on the first archive, then repeat periodically.
Take an extra checkpoint before risky experiments. Nightly retention accepts up
to a day of data loss; it does not retain every intermediate edit.

## Shared native checkpoint API

Capture, ZIP verification, and offline restore live in `atomic_lib::backup`,
behind the optional native `backup` feature. `CheckpointOptions` supplies explicit
data, configuration and output roots plus the caller's build revision. The server
provides only operator authentication, HTTP/CLI control, status, and server cache
configuration checks. Native callers invoke the same `create`, `verify`, `restore`
and `check_restore_activation` functions without an HTTP listener or Actix.
The v1 layout remains `data/store/atomic.redb` plus data files and
`config/config.toml`. The legacy manifest key `server_version` now records the
core package version; v1 archives remain readable at that same package version.

These mechanisms serve different recovery needs:

| Mechanism | Purpose | Restore behavior |
| --- | --- | --- |
| Desktop virtual filesystem (`desktop/src/vfs.rs`) | Read/write projection of graph folders and files; writes become ordinary signed commits | Filesystem operations edit live graph state; the projection does not contain all instance metadata or history |
| Encrypted vault (`atomic_lib::vault`) | Portable drive-level encrypted segments using `VaultObjectStore`, including its filesystem backend | Imports and merges drive data into a node |
| Instance checkpoint (`atomic_lib::backup`) | Consistent persisted database plus local configuration and files | Restores a complete local instance into a new offline directory |

Checkpoint creation uses the existing `Db` and its maintenance gate; it does not
introduce a second sync engine. A checkpoint is a streamed full-instance ZIP, so
it does not use the vault's in-memory sealed-object interface or its merge format.
A future destination abstraction should preserve streaming and these distinct
restore semantics.

The checkpoint boundary covers **persisted state**. Desktop VFS writes are staged
in memory before becoming commits. A desktop adapter must stop new staging,
flush pending writes before `create`, and coordinate external file writers with
`Db::maintenance`. It must check the restored-offline marker before starting sync
or integrations. This PR provides the native API, not a desktop backup UI or a
verified desktop staging/restore integration.
7 changes: 7 additions & 0 deletions lib/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -72,6 +72,11 @@ wasm-bindgen = { version = "0.2.122", optional = true }
wasm-bindgen-futures = { version = "0.4.72", optional = true }
web-sys = { version = "0.3.99", optional = true, features = ["DomException", "FileSystemDirectoryHandle", "FileSystemFileHandle", "FileSystemGetFileOptions", "FileSystemSyncAccessHandle", "FileSystemReadWriteOptions", "StorageManager", "WorkerGlobalScope", "WorkerNavigator"] }

[target.'cfg(not(target_arch = "wasm32"))'.dependencies]
chrono = { version = "0.4.44", optional = true }
tempfile = { version = "3", optional = true }
zip = { version = "8.6.0", default-features = false, features = ["deflate"], optional = true }

[dev-dependencies]
criterion = { version = "0.8.2", features = ["async_tokio"] }
iai = "0.1"
Expand All @@ -88,6 +93,8 @@ optional = true
version = "0.32"

[features]
## Native full-instance checkpoints; no hosted server or Actix dependency.
backup = ["db-redb", "dep:chrono", "dep:tempfile", "dep:zip", "tokio/time"]
config = ["directories", "toml"]
## Core Db with encoding support (BTreeMapStore included).
## Does NOT pull in sled — works in WASM.
Expand Down
Loading
Loading