Skip to content

Multi-machine sync correctness: transcript conflicts, history merge, debris, desktop records - #77

Open
sjalife wants to merge 7 commits into
tawanorg:mainfrom
sjalife:upstream-pr/jsonl-conflict-resolution
Open

sjalife wants to merge 7 commits into
tawanorg:mainfrom
sjalife:upstream-pr/jsonl-conflict-resolution

Conversation

@sjalife

@sjalife sjalife commented Aug 2, 2026 •

Copy link
Copy Markdown

Fixes #69, #70, #71, #72, #76.

This supersedes and includes #78, which I've closed. Everything here came out of running claude-sync across three machines for a few weeks and then investigating what went wrong; each issue has the full evidence, and this is the code.

Seven self-contained commits, one per concern — reviewable and revertable individually. If you'd rather take a subset, say which and I'll split; commits 1–5 are the data-safety fixes and 6–7 are additive.

# Commit Issue
1 fix(sync) resolve append-only transcript conflicts by comparing bytes #69, #70
2 fix(sync) never sync per-machine debris (.lock and .conflict artifacts) #71
3 feat(storage) typed ErrNotFound prerequisite for 4
4 fix(sync) union-merge history.jsonl instead of last-writer-wins #72
5 feat(sync) union-merge genuinely diverged session transcripts completes #69
6 feat(sync) sync Claude Code Desktop session records #76
7 fix(cli) report per-file sync errors to stderr in quiet mode —

1–2: transcripts stop losing data

Pull declares a conflict from metadata alone — remote newer than our last upload, local hash differs from state. For append-only transcripts both are routinely true with nothing diverging, because a live session appends between two syncs. Content is never compared, so a fast-forward is indistinguishable from a fork.

Three machines produced 92 .conflict artifacts in about a week. 91 held no unique data. The 92nd was inverted: local was a truncated 41-line copy and the .conflict file was the complete 251-line transcript, so the rule kept the damaged copy and hid the intact one under a name Claude Code never reads.

Pull now classifies the byte relation first: equal → refresh state; remote ahead → fast-forward by appending the missing tail (never a rewrite, so a concurrent appender can't lose lines); local ahead → keep, no .conflict; diverged → commit 5. Push gets the mirror guard so a truncated local can't clobber a fuller remote, paid only when a file shrank, failing open on any fetch error.

Commit 2 stops the tool replicating its own debris: zero-byte .lock files, and the .conflict artifacts themselves — which self-replicate, since handleConflict writes them inside a synced directory. The .conflict pattern is anchored on the exact timestamp format the tool generates, because excluding an already-tracked path prunes it from the bucket; a loose *.conflict.* would remotely delete a user's notes.conflict.md.

3–4: history.jsonl stops losing prompts

One file appended by every machine, synced as an opaque blob, so pushes are last-writer-wins. #63 recovers the loss after the fact; this prevents it. Push GETs the remote, unions, uploads; pull appends missing lines rather than overwriting or writing .conflict.

Three properties worth keeping in any implementation: merges only ever append (a live session can't lose lines to a stale-read rewrite), they're idempotent by raw-line dedupe, and a fetch error aborts rather than uploading a local-only file over a union that couldn't be read. That last one needs ErrNotFound (commit 3) to tell "no remote yet, first push" from "storage unreachable" — adapters previously returned provider-specific errors and the caller couldn't distinguish them.

5: the diverged case

Before writing this I checked 1,180 real transcripts. Two findings pull in opposite directions:

  • Unioning is structurally safe. 782 (66%) already contain a fork — a parentUuid with two children — and Claude Code reads them fine. Records carry their own parent pointers, so a uuid-keyed union yields a valid DAG regardless of append order.
  • But transcripts aren't uniformly append-only. Six record types (mode, custom-title, ai-title, permission-mode, last-prompt, worktree-state) are mutable last-wins state with no uuid and — across all 12,679 of them — no timestamp. A naive line-union can therefore place an older state record last and silently resurrect a stale mode or title.

So events union by uuid, and each side's final state record per type is compared as a whole, with file-level times breaking the tie (LastModified is upload time and lags the write, so near-ties resolve in the remote's favour — bounded and self-correcting). Merge failure degrades to the legacy .conflict path.

6: desktop sidebar

The desktop app renders its sidebar from per-machine records at %APPDATA%\Claude\claude-code-sessions\<install-id>\<org-id>\local_*.json, not from ~/.claude/projects. So a machine can receive every transcript, resume any of them from the CLI, and still show an empty sidebar. Records sync under _ccd-sessions/<org-id>/ with the machine-specific install-id deliberately left out of the key, merged last-writer-wins by the record's own lastActivityAt.

Details that cost me time and are easy to miss: MSIX-packaged installs redirect the Roaming path somewhere os.Stat can't traverse, so the Local\Packages glob is also probed; trust state is deliberately not synced because it gates hook execution and a compromised bucket must not be able to pre-trust projects; records are normalised in transit (sorted keys, machine-local permission fields stripped) so they don't ping-pong; writes are atomic because the app may be reading. Machines without the desktop app skip silently.

This is the most separable commit — happy to pull it out if you'd rather land the data-safety fixes alone.

Notes on the diff

Two existing tests were retargeted, not weakened. TestPullDetectsConflicts and TestConflictCreatesConflictFile used history.jsonl as the vehicle for the generic conflict machinery, which by design no longer conflicts after commit 4. Every assertion is byte-for-byte unchanged; only the file under test moved to settings.json, with a comment explaining why.

DetectChanges skips the CCD prefix: those records are tracked in state but live outside claudeDir, so without it the first push would report them as deletions and prune them from the bucket.

Scope: 18 files, +2472/-35. Nothing outside internal/, plus 16 lines in cmd/claude-sync/main.go for commit 7. No changes to config defaults, crypto, or the installer.

Tests: go vet ./... and go test ./... green on Linux/ext4 across all 12 packages, with roughly 40 new cases. I don't think CI has run here — the workflow doesn't appear to trigger on fork PRs without approval — so that's local verification, not GitHub's.

Pull declares a conflict whenever the remote changed after our last upload
AND the local hash differs from state. For append-only session transcripts
(projects/**/*.jsonl) both conditions are routinely true without any real
divergence: a live Claude Code session appends between two syncs. The content
is never compared, so a fast-forward is indistinguishable from a fork.

Across three synced machines this produced 92 .conflict artifacts in about a
week. Byte analysis showed 91 held no unique data. The 92nd was the dangerous
inversion: the local file was a truncated 41-line copy and its .conflict
counterpart was the complete 251-line transcript, so the conflict rule kept
the damaged file and hid the only intact one under a name Claude Code never
reads.

Pull now classifies the byte relation before declaring a conflict:

  equal        -> refresh state so it stops re-triggering every pull
  remote ahead -> fast-forward by APPENDING the missing tail (never a
                  rewrite, so a concurrent appender cannot lose lines)
  local ahead  -> keep local, write no .conflict, leave state stale so the
                  next push publishes it
  diverged     -> unchanged: keep local, write .conflict

Push gets the mirror guard. It previously uploaded any file whose hash
differed from state, so a truncated local transcript would overwrite the
fuller bucket copy and propagate the loss to every machine. When a session
transcript has shrunk relative to state, the remote is fetched and the upload
is skipped if local is a prefix of it. The remote round-trip is paid only in
the shrunk case; a non-prefix rewrite still uploads; any fetch error falls
through to a normal upload so the guard can never fail a push. A skip counts
as neither uploaded nor errored.

fetchDecoded is factored out of downloadFile so both paths can inspect remote
plaintext without writing it, and so comparisons are like-for-like (the
remote is resolved to local path form first). The pull path adds no network
cost: the conflict branch it replaces already downloaded the remote to write
the .conflict file.

Refs tawanorg#69, tawanorg#70

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
sjalife and others added 6 commits August 2, 2026 01:50
Two classes of local-only file are currently uploaded and replicated to every
machine.

.lock: Claude Code writes a zero-byte lock per task directory
(tasks/<id>/.lock). A lock represents a process holding a resource on one
machine; replicated elsewhere it is indistinguishable from a lock genuinely
held there, and it long outlives the process that made it. Observed: 12 of
them, the oldest about four weeks old, syncing across three machines.

.conflict.<timestamp>: handleConflict writes the remote copy next to the
original, which is inside a synced directory, so the recovery artifact is
itself uploaded. Each replica can then be re-detected on another machine and
spawn further artifacts. Observed: 92 accumulated in about a week, including
five copies of one snapshot minted by successive pulls.

Both are now excluded unconditionally, ahead of the user's exclude patterns,
because neither has any meaning on another machine under any configuration.

The .conflict pattern is anchored on the exact <yyyymmdd>-<hhmmss> format this
tool generates rather than matching .conflict. loosely. That matters: an
already-tracked file becoming excluded is reported as a deletion and pruned
from the bucket on the next push, so a loose pattern would silently delete a
user file named e.g. notes.conflict.md from remote storage. Tests pin both
directions.

Migration note for existing users: the first push after upgrading prunes any
already-uploaded locks and conflict artifacts from the bucket. Local copies
are untouched.

Fixes tawanorg#71

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…rom unreachable

Head returned the provider's own error verbatim in every adapter, so a caller
had no portable way to tell "the object is not in the bucket" from "the bucket
is unreachable right now" (timeout, throttle, 5xx). Both surfaced as an opaque
error, and the natural reading — treat any Head failure as absent — is exactly
the dangerous one.

storage.ErrNotFound plus storage.IsNotFound give callers that distinction:
each adapter now wraps its own not-found signal (types.NotFound for R2/S3,
storage.ErrObjectNotExist for GCS, HTTP 404 for WebDAV) with the sentinel, and
leaves every other error untouched.

The history merge added in the next commit depends on this: on push it must
merge the remote copy into the upload, so a genuinely absent object means
"first push, nothing to merge" while a transient failure must abort the upload
rather than clobber the bucket's union with local-only content.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
history.jsonl is one shared file that every machine appends to. Syncing it
like any other file makes it last-writer-wins: whichever machine pushes last
replaces the bucket copy, silently dropping every prompt the other machines
recorded since, and the /resume picker degrades to whatever that one machine
happened to hold. On pull the same divergence surfaced as a .conflict file,
parking the other machines' history where nothing reads it.

Push now GETs the remote copy, unions it with the local payload, and uploads
the union. Pull appends the lines the local file is missing instead of
overwriting it or declaring a conflict.

Properties that make this safe on a file a live session is writing to:

  - Merges only ever APPEND. The local bytes are the verbatim prefix of the
    result, so a rewrite from a stale read can never destroy lines appended
    during the merge's network round-trip. Unparseable lines (torn writes)
    and unknown record shapes survive untouched.
  - Idempotent by raw-line dedupe. Remote lines already present byte-for-byte
    are dropped, so blank submissions and unknown-shape records — which have
    no sessionId+display signature — do not multiply as content cycles
    between machines and the bucket. Prompts additionally dedupe by
    sessionId+display within historyDedupeWindowMs, the same rule
    RebuildHistory uses, with the local line winning so its pastedContents
    are kept.
  - Aborts rather than clobbers. A missing remote object means first push and
    merges nothing, but a transient Head failure or an undecodable remote
    aborts the history upload — uploading local-only content would erase the
    union. This is what storage.ErrNotFound in the previous commit is for.
  - A local delete no longer deletes the bucket copy: the union only grows,
    and the next pull restores the file locally.

Builds on the history helpers introduced in PR tawanorg#63 (HistoryEntry, forEachLine,
withinWindow, historyDedupeWindowMs) — tawanorg#63 recovers lost prompts after the
fact, this stops them being lost in the first place.

TestPullDetectsConflicts and TestConflictCreatesConflictFile used
history.jsonl as the vehicle for the generic conflict machinery, which now
never conflicts. They exercise the same assertions against settings.json.

Fixes tawanorg#72

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Once the byte-prefix fast-forward has taken the easy cases, the only apparent
conflict left on a session transcript is a true fork: the same session
advanced on two machines. Writing a .conflict file for that keeps whichever
copy this machine happens to hold and hides the other where Claude Code never
reads it, even though both sides are recoverable.

MergeSessionPayloads unions the two payloads, returning only the lines to
append. The local file is never rewritten, so a live session's concurrent
appends survive and applying the merge is idempotent.

Transcripts mix two record classes with different write semantics:

  - uuid-keyed events are immutable DAG nodes linked by parentUuid, so they
    union safely. Append order is not a correctness concern: Claude Code
    already renders forked DAGs — 782 of 1,180 real transcripts contain
    forks — and threads render from the parent links, not line order.
  - uuid-less records of type mode, custom-title, ai-title, permission-mode,
    last-prompt and worktree-state are MUTABLE state where the last
    occurrence in the file wins. Across 12,679 such records in those same
    transcripts, not one carries a timestamp, so per-record ordering between
    two machines is impossible and cross-machine ordering falls back to
    file-level times. Only each side's final value per type is compared, so a
    stale intermediate remote record cannot resurrect, and a type the local
    file never had is always preserved.
  - anything else uuid-less unions by exact-raw-line multiset, the rule that
    makes the history merge idempotent.

A merge failure degrades to the legacy keep-local-plus-.conflict path rather
than blocking the pull.

TestPullDivergedJSONLStillConflicts asserted the .conflict outcome this commit
deliberately replaces; it now asserts the merge outcome instead — no
conflicts, local content preserved as a verbatim prefix, the remote-only event
present, and a second pull a byte-for-byte no-op.

Completes tawanorg#69: with the fast-forward and this merge, an append-only transcript
no longer produces .conflict files at all.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The desktop app renders its sidebar from per-machine records at
%APPDATA%\Claude\claude-code-sessions\<install-id>\<org-id>\local_*.json,
not from ~/.claude/projects — so synced transcripts are resumable from the
CLI but invisible in the app.

Records now sync under a reserved `_ccd-sessions/<org-id>/` prefix with the
machine-specific install-id deliberately excluded from the key, merged
last-writer-wins by the record's own lastActivityAt.

Notes:
- MSIX-packaged installs redirect the Roaming path (probed via the
  Local\Packages glob).
- Trust state is deliberately NOT synced: it gates hook execution.
- Records are normalized in transit (sorted keys, machine-local permission
  fields stripped).
- Writes are atomic (temp+rename) — the app reads these files live.
- Machines without the desktop app skip the feature silently.

Fixes tawanorg#76

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
`-q` suppresses all output, including per-file errors, so a hook-driven sync
could fail silently on individual files while appearing to succeed. Errors
now go to stderr even in quiet mode; normal output is unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sjalife sjalife changed the title fix(sync): resolve append-only transcript conflicts by comparing bytes Multi-machine sync correctness: transcript conflicts, history merge, debris, desktop records Aug 2, 2026
@tawanorg

tawanorg commented Oct 6, 2026

Copy link
Copy Markdown
Owner

Thanks for this — the substance here (append-only transcript conflict resolution, union-merging history.jsonl instead of last-writer-wins, not syncing per-machine debris, typed ErrNotFound) is the right set of problems to be solving, and this is the only one of the three desktop-sync PRs that still merges clean.

Not merging it in this pass, and I want to be straight about why: it isn't a judgement on the code, which I haven't reviewed line by line yet. It's that this PR restructures the most data-sensitive logic in the repo — conflict resolution, history merge, session merge, state tracking — across 18 files, and that deserves a focused review rather than riding along in a backlog sweep. For calibration, #82 landed earlier today: 4 files, 324 lines, and it still took three review rounds and seven confirmed defects, two of which would have corrupted users' transcripts.

One thing you should know, since it affects sequencing: #86 contains this PR's seven commits verbatim, by SHA, plus 16 of its own. So #86 is this work extended rather than a competing approach. The plan is to land #77 first and have #86 rebase onto it — that keeps attribution with you and shrinks #86 to something reviewable.

Tracking the sequencing and the review in #91.

Nothing needed from you right now; this stays open and is next in line for a proper review. If you'd like to make that easier, splitting the storage ErrNotFound commit (f84476ed) out as its own PR would let that land independently, since it's additive and low-risk on its own.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

needs-decision Blocked on a maintainer decision about direction, not on code

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Pull declares false conflicts for append-only session files (one side is a strict prefix of the other)

2 participants