Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
54 commits
Select commit Hold shift + click to select a range
3333c6b
fix(scrapers): index Claude Code tool results as tool, not user; rend…
fstubner Oct 1, 2026
37ab3cb
fix(mcp): session detail returns the end of the session, within a siz…
fstubner Oct 1, 2026
41b7fe4
fix(mcp,hook): label transcript text as untrusted in JSON output and …
fstubner Oct 1, 2026
56b376f
fix(scrapers): re-read claude-code transcripts once on upgrade to cor…
fstubner Oct 1, 2026
bbff004
fix(cursor): default store path follows the platform, not ~/.cursor
fstubner Oct 1, 2026
182cf93
fix(cursor): attribute conversations by composerHeaders, not by guess…
fstubner Oct 1, 2026
9cd93b4
fix(copilot): replay a journal splice as truncate-then-push, not insert
fstubner Oct 1, 2026
2a2b1d0
fix(copilot): render response items by kind, use per-request times, r…
fstubner Oct 1, 2026
e8b7108
fix(copilot-cli): file subagent output as tool output, skip the syste…
fstubner Oct 1, 2026
38e5e9b
test(drift): fingerprint the current chatSessions journal format
fstubner Oct 1, 2026
6469d84
fix(cursor): mark subagent conversations and stop presenting the pare…
fstubner Oct 1, 2026
31fbbf3
fix(cursor): leave a one-line trace for each tool call
fstubner Oct 1, 2026
5889e77
fix(copilot): skip the older progressTask item kind without a drift w…
fstubner Oct 1, 2026
75fc3bc
test(drift): record composerHeaders in the cursor format fingerprint
fstubner Oct 1, 2026
da5d01c
fix(opencode): leave a one-line trace for each tool call
fstubner Oct 1, 2026
258fbbe
fix(opencode): re-read a session whole when any of its rows changed s…
fstubner Oct 1, 2026
53aef65
test(drift): cursor mutation fixture includes composerHeaders
fstubner Oct 1, 2026
e822e59
feat(embeddings): install the local model on demand instead of shippi…
fstubner Oct 1, 2026
41c1d55
feat(embeddings): semantic search is off until enabled, and off is a …
fstubner Oct 1, 2026
6f5a5e1
docs: semantic search is an optional add-on, with the install size me…
fstubner Oct 1, 2026
109c577
chore(review): put embeddings-runtime/ in the build-config review layer
fstubner Oct 1, 2026
3670a56
fix(scrapers): re-read cursor and opencode once on upgrade to correct…
fstubner Oct 1, 2026
16e4221
fix(setup): pin the hook and MCP configs to the version that ran setup
fstubner Oct 1, 2026
694c3a1
docs(readme): say what setup actually puts in front of the agent
fstubner Oct 1, 2026
88696ac
fix(scrapers): pick up rewritten and late-stamped history in JSONL st…
fstubner Oct 1, 2026
ec19984
fix(index): never prune rows a scan did not see; refuse cursors the i…
fstubner Oct 1, 2026
f8eedc9
fix(index): one scanner per project at a time, across servers
fstubner Oct 1, 2026
38a8ee0
fix(index): rebuild every missing or stale retrieval window, not four…
fstubner Oct 1, 2026
f64e4d7
feat(setup): shrink the managed instruction block to what an agent needs
fstubner Oct 1, 2026
97129eb
fix(index): migrate older schemas in place instead of setting them aside
fstubner Oct 1, 2026
341311d
fix(mcp): label the session-list preview as untrusted transcript text
fstubner Oct 1, 2026
8ef2e7c
fix(index): let tool calls and shutdown in while a scan runs
fstubner Oct 1, 2026
2e74256
docs(changelog): add an Unreleased section and teach the release step…
fstubner Oct 1, 2026
8074a4c
docs(release): describe how the changelog's Unreleased section is han…
fstubner Oct 1, 2026
cb1cf08
test(setup): pin that the report says updated for files that already …
fstubner Oct 1, 2026
2eac02c
fix(index): carry sessions forward from a set-aside index after the r…
fstubner Oct 1, 2026
1cc184f
feat(cli): xtctx export and xtctx import
fstubner Oct 1, 2026
ec411b7
feat(status): say how many sessions exist only in the index
fstubner Oct 1, 2026
4a8016e
fix(status): say 'back it up' when one session exists only in the index
fstubner Oct 1, 2026
38ddda2
docs: the index is derived for sessions on disk and the only copy of …
fstubner Oct 1, 2026
02ba519
merge: one scanner per project, prune bounded by scan start, cursors …
fstubner Oct 1, 2026
3704181
merge: migrate the index in place, carry set-aside sessions forward, …
fstubner Oct 1, 2026
7e88b5a
merge: Claude Code tool results indexed as role tool, tool calls rend…
fstubner Oct 1, 2026
ac150b1
test: a scraper-version re-read under the lease prunes old rows and w…
fstubner Oct 1, 2026
92d8acd
merge: Copilot journal replay, response items by kind, per-request ti…
fstubner Oct 1, 2026
6df3619
merge: Cursor composer attribution, subagents and tool lines; opencod…
fstubner Oct 1, 2026
00504b5
merge: the ML runtime is optional; xtctx embeddings enable/disable; s…
fstubner Oct 1, 2026
67c2e66
fix: reconcile index-durability with optional-embeddings
fstubner Oct 1, 2026
b893bc1
merge: pinned npx version in generated configs, shorter instruction b…
fstubner Oct 1, 2026
56b8091
docs: changelog lines for the merged fixes; setup's npx command is pi…
fstubner Oct 1, 2026
f863a23
merge: main (dependency advisories, #398)
fstubner Oct 2, 2026
28b28f7
Merge branch 'main' into fix/audit-fixes
fstubner Oct 2, 2026
5897818
test(scan): lower the scan-duration floor in the event-loop test to 5…
fstubner Oct 2, 2026
805acf0
Merge branch 'main' into fix/audit-fixes
fstubner Oct 2, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 13 additions & 2 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -203,9 +203,20 @@ jobs:
# Insert above the previous newest entry, keeping the file's header.
# Release Please owned this file before; its format is preserved so
# the existing entries stay readable alongside the new ones.
first_entry=$(grep -n '^## ' CHANGELOG.md | head -n1 | cut -d: -f1 || true)
#
# An `## [Unreleased]` section, when present, stays at the top and
# the new entry goes beneath it: inserting above the first `## `
# heading would put the release above Unreleased and leave the
# section describing work that has now shipped. Its hand-written
# body is replaced by the generated notes (which cover the same
# commits), so the heading is kept and the body dropped.
first_entry=$(grep -n '^## ' CHANGELOG.md | grep -v ':## \[Unreleased\]' | head -n1 | cut -d: -f1 || true)
unreleased=$(grep -n '^## \[Unreleased\]' CHANGELOG.md | head -n1 | cut -d: -f1 || true)
if [ -n "${first_entry:-}" ]; then
head -n "$((first_entry - 1))" CHANGELOG.md > /tmp/new.md
head -n "$(( ${unreleased:-$first_entry} - 1 ))" CHANGELOG.md > /tmp/new.md
if [ -n "${unreleased:-}" ]; then
printf '## [Unreleased]\n\n' >> /tmp/new.md
fi
cat /tmp/entry.md >> /tmp/new.md
tail -n +"${first_entry}" CHANGELOG.md >> /tmp/new.md
else
Expand Down
48 changes: 38 additions & 10 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,8 @@ boundaries between them, and what each part is allowed to trust.

## Parts

- **CLI** (`src/cli/`) — `setup`, `status`, `scan`, `calibrate`,
`disconnect`, and the internal `--hook session-start` entry point. Bare `xtctx` on a non-TTY stdio pair
- **CLI** (`src/cli/`) — `setup`, `status`, `scan`, `export`, `import`,
`calibrate`, `disconnect`, and the internal `--hook session-start` entry point. Bare `xtctx` on a non-TTY stdio pair
starts the MCP server.
- **MCP server** (`src/mcp/`) — stdio JSON-RPC server exposing exactly five
read-only tools. Spawned by coding agents via `npx -y xtctx`.
Expand All @@ -20,9 +20,16 @@ boundaries between them, and what each part is allowed to trust.
(`.xtctx/state/xtctx.db`, WAL, schema-versioned) holding sessions,
messages, retrieval windows, FTS index, and embedding vectors. Refreshed
on demand from the scrapers and at MCP server start. Derived from the
transcripts, but not disposable: the index keeps sessions whose transcripts have since been deleted (Claude Code deletes them after 30 days by default), so for those it is the only copy. A database that will not
open (corruption, schema mismatch) is moved aside to
`xtctx.db.set-aside-<time>`, never deleted, and a new one is built.
transcripts for every session still on disk, and the only copy of the
sessions whose transcripts have since been deleted (Claude Code deletes
them after 30 days by default); deleting it loses those. An index from an
older schema is migrated in place. One that is corrupt, or older in a shape
no migration recognises, is moved aside to `xtctx.db.set-aside-<time>`,
never deleted; a new one is built from the transcripts, and the first full
scan copies every session it lacks back out of the set-aside file. One from
a newer schema is refused. `xtctx export` writes the project's sessions and
messages to a JSON Lines file and `xtctx import` merges one back, without
duplicates.
- **Drift log** (`src/scrapers/drift-log.ts`) — per-tool record of the
places another tool's transcripts did not match what the scraper expected,
summarised once per scan and kept in `.xtctx/state/<tool>-drift.json`.
Expand Down Expand Up @@ -93,6 +100,21 @@ evidence they have, and one with no vector is treated as unknown similarity
rather than none, because scoring it zero penalises it for its position in a
queue.

**Semantic search is an add-on, off until enabled.** The default install has no
ML runtime: `@huggingface/transformers` and the ONNX runtimes under it were
about 550 MB on disk and were fetched before `npx -y xtctx` could answer, which
is longer than an MCP client waits. `optionalDependencies` would not help, as
npm installs those by default, so the library is not a dependency at all.
`xtctx embeddings enable` runs `npm ci` against a pinned manifest and lockfile
shipped in `embeddings-runtime/`, into `~/.xtctx/embeddings`, and
`handoff/embedding-runtime.ts` loads it from there by path. Until then the
provider is `NullEmbeddingProvider`, which carries a `semanticOff` reason: search
answers from keyword without calling it, nothing counts as a backlog, vectors an
earlier install built are kept (the placeholder model identity must not read as
"another model" to `dropVectorsFromOtherModels`), and `xtctx status` and
`xtctx_continuity_status` say which mode is active and the command to change it.
A remote OpenAI-compatible endpoint needs no local runtime and is unaffected.

**Bounded, so a tool call always returns.** Scanning gets four seconds,
vectorizing six, and an indexed view is treated as current for thirty. Work
left over resumes on the next call. A scan also warms the embedding model and
Expand Down Expand Up @@ -150,11 +172,17 @@ the transcripts remain authoritative.
- **`.xtctx/config.yaml` is semi-trusted.** It is repo-committable, so a
cloned repo can point `storePath` anywhere on disk. Store paths are used
read-only, but treat overrides in a foreign repo as a risk surface.
- **The index is trusted state, and not disposable.** It is built from the
transcripts but keeps sessions whose transcripts have since been deleted,
so it is never deleted: a corrupt index or one from an older schema is
moved to `xtctx.db.set-aside-<time>` and rebuilt, one from a newer schema
is refused, and `setup --repair` leaves it alone.
- **The index is trusted state, and not disposable.** It is derived from the
transcripts still on disk but is the only copy of sessions whose
transcripts have since been deleted, so it is never deleted: an older
schema is migrated in place, a corrupt index (or an older one no migration
fits) is moved to `xtctx.db.set-aside-<time>` and rebuilt with the set-aside
file's missing sessions carried forward, one from a newer schema is
refused, and `setup --repair` leaves it alone.
- **An export file is untrusted input**, like the transcripts it came from:
`xtctx import` checks the header and every line before writing it, binds
every value as a SQL parameter, and its content reaches agents through the
same fenced MCP output as any other transcript text.
- **The registry and npm supply chain** are trusted at install time; CI
pins action SHAs and publishes via OIDC with provenance, no long-lived
tokens.
50 changes: 50 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,56 @@ All notable changes to this project are documented in this file.
The format is based on Keep a Changelog, and this project follows Semantic Versioning.
Entries are written by the `release` workflow when a release is cut by hand.

## [Unreleased]

Work on `main` since 0.21.8 that has not been released. The release workflow
leaves this heading in place and empties it when a version is cut.

### Features

* **scan:** add `--embed`, so semantic search can cover a real history ([#359](https://github.com/fstubner/xtctx/issues/359))
* **search:** a literal mode that answers without the index ([#343](https://github.com/fstubner/xtctx/issues/343))
* **hook:** background scan on session start; take the transcript location from the tool instead of deriving it ([#322](https://github.com/fstubner/xtctx/issues/322), [#300](https://github.com/fstubner/xtctx/issues/300))
* **mcp:** name an unconfigured project instead of answering with silence ([#315](https://github.com/fstubner/xtctx/issues/315))
* **index:** `xtctx export` and `xtctx import`; migrate an older index in place and carry sessions forward from one set aside; `xtctx status` counts sessions that exist only in the index
* **embeddings:** semantic search is an optional add-on, off until `xtctx embeddings enable` installs the local model; keyword-only is a working state that status reports

### Bug Fixes

* **index:** stop deleting the index; status and setup say what is true ([#388](https://github.com/fstubner/xtctx/issues/388), [#389](https://github.com/fstubner/xtctx/issues/389))
* **index:** stop a re-read leaving behind the rows it replaced ([#374](https://github.com/fstubner/xtctx/issues/374))
* **config:** stop setup and disconnect destroying files the user wrote ([#367](https://github.com/fstubner/xtctx/issues/367), [#380](https://github.com/fstubner/xtctx/issues/380))
* **setup:** grant the xtctx tools (including under the plugin's server name) and say what setup cannot grant ([#316](https://github.com/fstubner/xtctx/issues/316), [#326](https://github.com/fstubner/xtctx/issues/326))
* **setup:** authenticate the self-hosted branch, and stop mangling flags ([#307](https://github.com/fstubner/xtctx/issues/307))
* **status:** report the MCP command the configs name, and make the embedding estimate describe the run that is happening ([#358](https://github.com/fstubner/xtctx/issues/358), [#361](https://github.com/fstubner/xtctx/issues/361))
* **search:** point a match at where it actually is ([#368](https://github.com/fstubner/xtctx/issues/368))
* **hook:** stop stdin choosing which directory is a project's transcript store ([#370](https://github.com/fstubner/xtctx/issues/370))
* **scope:** close project-boundary leaks, and scope search and status to the project ([#293](https://github.com/fstubner/xtctx/issues/293), [#311](https://github.com/fstubner/xtctx/issues/311), [#313](https://github.com/fstubner/xtctx/issues/313))
* **codex:** read the human turns Codex writes now; stop a resumed scan serving another project's turns; report an oversized record instead of dropping it ([#371](https://github.com/fstubner/xtctx/issues/371), [#309](https://github.com/fstubner/xtctx/issues/309), [#355](https://github.com/fstubner/xtctx/issues/355))
* **claude-code:** collapse the dots and underscores its store directories collapse ([#357](https://github.com/fstubner/xtctx/issues/357))
* **scrapers:** see a project opened through WSL ([#375](https://github.com/fstubner/xtctx/issues/375))
* **security:** scrub every unfenced field and fail closed on an undecided resume ([#312](https://github.com/fstubner/xtctx/issues/312), [#314](https://github.com/fstubner/xtctx/issues/314))
* **release:** publish as its own run so npm trusted publishing accepts it ([#390](https://github.com/fstubner/xtctx/issues/390))
* **index:** one scanner per project across servers, a prune that never deletes rows its scan did not see, cursors refused when the index lost their rows, and rewritten or late-stamped history read again
* **claude-code:** index tool results as tool output rather than the user, keep a one-line trace of each tool call, strip terminal colour codes, and correct already-indexed rows once on upgrade; session detail returns the end of a session first
* **copilot:** replay chat journals as truncate-then-push, stamp each request with its own time and render response items by kind; file Copilot CLI subagent output as tool output; correct already-indexed rows once on upgrade
* **cursor, opencode:** attribute Cursor conversations by composer headers, mark subagents, find the store on macOS and Linux, trace tool calls, and re-read an opencode session whenever it changes
* **setup:** pin the hook and MCP configs to the version that ran setup, shrink the managed instruction block, and label session previews as untrusted transcript text

### Performance

* **scrapers:** resume Codex, Claude Code and Copilot CLI transcripts from a byte offset instead of re-reading them ([#302](https://github.com/fstubner/xtctx/issues/302), [#303](https://github.com/fstubner/xtctx/issues/303))
* bound search memory, skip unchanged Copilot files, batch FTS deletes ([#298](https://github.com/fstubner/xtctx/issues/298))

### Documentation

* the plugin route does not answer in an unconfigured project; design a configurable embedding provider; write down what indexing throughput costs ([#325](https://github.com/fstubner/xtctx/issues/325), [#381](https://github.com/fstubner/xtctx/issues/381), [#382](https://github.com/fstubner/xtctx/issues/382))

### Internal

* Releases are manual: one workflow, run on request, replaces the automatic pipeline ([#296](https://github.com/fstubner/xtctx/issues/296))
* Large module splits (index, scrapers, config) and mutation-sweep test additions ([#328](https://github.com/fstubner/xtctx/issues/328) to [#356](https://github.com/fstubner/xtctx/issues/356))

## [0.21.8](https://github.com/fstubner/xtctx/compare/xtctx-v0.21.7...xtctx-v0.21.8) (2026-08-31)

> **Not on npm.** 0.20.0 through 0.21.8 were tagged and given GitHub
Expand Down
30 changes: 20 additions & 10 deletions PRODUCT.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,10 +9,14 @@ transcript store into a per-project SQLite index and serves it back — to any
of those tools — through a small read-only MCP server, so the next agent can
pick up where the last one left off.

Raw local transcripts are authoritative. xtctx never summarizes and never
persists derived "memory". The index is built from the transcripts, but it
keeps sessions whose transcripts have since been deleted, so it is not
disposable and xtctx never deletes it. It sends transcript content nowhere
Raw local transcripts are authoritative while they exist. xtctx never
summarizes and never persists derived "memory". The index is derived from
the transcripts for every session still on disk, and is the only copy of the
older ones whose transcripts have since been deleted (Claude Code deletes
them after 30 days by default). Deleting the index loses those, so xtctx
never deletes it: schema upgrades migrate it in place, a corrupt one is set
aside and its sessions carried into the rebuilt one, and `xtctx export` /
`xtctx import` keep a copy elsewhere. It sends transcript content nowhere
unless a project opts in to an external embedding endpoint, written into
`.xtctx/config.yaml` by hand, trusted by the user in
`XTCTX_TRUSTED_EMBEDDING_ENDPOINTS` (a repository cannot set that), and
Expand Down Expand Up @@ -55,14 +59,16 @@ Single-user, single-machine. There is no team, sync, or server component.
- Scrapers for the seven supported tools, project-scoped, incremental, and
tolerant of upstream schema drift (warn, never silently drop).
- One per-project SQLite index (`.xtctx/state/xtctx.db`) with keyword (FTS5)
and semantic (local bge-small embeddings) search over chronological windows.
and, once the optional add-on is enabled with `xtctx embeddings enable`,
semantic (local bge-small embeddings) search over chronological windows.
- Five read-only MCP tools: recent sessions, session detail, search,
continuity status, handoff manifest.
- CLI: `setup` (wire MCP config, managed instruction blocks, skills, and the
Claude Code SessionStart hook), `status`, `scan` (read the stores into the
index now, `--embed` to finish vectorizing too), `calibrate` (time the
embedding model on this machine's devices and use the fastest),
`disconnect`.
index now, `--embed` to finish vectorizing too), `embeddings enable|disable`
(install or remove the optional local model; keyword search needs none of
it), `calibrate` (time the embedding model on this machine's devices and use
the fastest), `disconnect`.

Out of scope (deliberately, and documented everywhere the product speaks):
no daemon, no API server, no dashboard, no generated summaries or briefs,
Expand All @@ -71,7 +77,10 @@ no durable memory, no write-back tools, no cloud anything.
## Constraints

- Node ≥ 24, distributed via npm (`npx -y xtctx`); no install step beyond
what a coding agent's MCP config can express.
what a coding agent's MCP config can express. The default install carries no
ML runtime (about 55 MB on disk against 550 MB with it, measured), because
an MCP client will not wait minutes for `npx` to fetch one: the local model
is an add-on, installed by `xtctx embeddings enable`.
- Transcript stores belong to other tools: all reads are read-only
(`readonly` + `fileMustExist` for SQLite stores) and must survive those
tools changing their formats — drift is detected by tests, committed format
Expand All @@ -83,7 +92,8 @@ no durable memory, no write-back tools, no cloud anything.
fences it and never grows write capabilities.
- Everything runs local by default. Three network dependencies exist. Two are
unavoidable and narrow: the one-time embedding-model download from Hugging
Face, and loopback-only HTTPS calls to Antigravity's local language server
Face (and the runtime from npm), made only when the user runs
`xtctx embeddings enable`, and loopback-only HTTPS calls to Antigravity's local language server
(127.0.0.1, exact-PID + CSRF matched; certificate verification is off
because the server is self-signed). The third is opt-in and is the only one
that carries transcript text: an OpenAI-compatible embedding endpoint named
Expand Down
Loading
Loading