fix: audit findings — scrapers, concurrent scans, index durability, optional embeddings - #403
Merged
Merged
Conversation
…er tool calls; strip ANSI message.role is the API role: Claude Code writes tool results back as user records, so 11,434 of a real index's messages were 'user' tool output. A record carrying only tool blocks is now role tool; tool_use blocks render as one short line (ran Bash: ..., edit <path>) instead of being dropped.
…e budget xtctx_session_detail returned the oldest 50 messages with no total cap; the first page of real recent sessions measured 75k-206k characters and a 7,996-message session needed offset=7946 to reach its end. It now returns the newest messages when offset is omitted, caps a response at 40,000 characters (2.5x the per-message cap), excerpts tool output to 1,500 characters with the original length noted, and prints the offset for earlier messages. An explicit offset still counts from the start; from_end=true counts back from the newest. Messages carry a 0-based position. Also adds the untrusted marker to the JSON payloads built in the same file (see next commit).
…the hook preview JSON output of the session tools and the manifest carried raw transcript text in bare fields; markdown fences it and says so. JSON payloads now carry untrusted: true and a notice. The SessionStart hook's 'Opened with' line is labelled on the line itself, and its pointer now says detail returns the most recent messages, not the full turn history.
…rect indexed roles The resume cursor sits past every finished session, so rows indexed with the old roles would stay wrong until a transcript grew. Scraper state now records a scraperVersion; a stored version below the scraper's resets the cutoff and cursors for one scan, and the normal re-read path plus pruneRereadSessions replaces the old rows. The version is saved only when the read ran to the end. Sessions whose transcripts are gone keep their rows.
Without APPDATA the default was ~/.cursor/workspaceStorage on every platform, where Cursor keeps extensions rather than conversations. macOS and Linux now use Application Support and ~/.config, computed the way the VS Code path is.
…ing from file paths Current Cursor no longer lists conversations in each workspace; globalStorage's composerHeaders table records which workspace owns each one. Reading it first stops conversations that touched a second project's files being filed under it. A conversation the header cannot place (no row, or a workspace that resolves to no folder) still falls back to the recorded-path match. With the table present the unlisted-conversation search is a lookup by id plus a key-only range, so the LIKE over every stored value (9-28s on a 6.9GB store) only runs when the table is missing, which is also now reported as drift.
A kind-2 record with an index means "cut the array back to length i, then push v". Applying it as an insert left the old copy of a rewritten request beside the new one, which duplicated the first question and attached answers to the wrong requests. The timestamp sort that hid the resulting misordering is removed; array order is conversation order once the truncate is applied. An unknown record kind now raises a drift warning instead of being skipped, a record with an index and no values is a bare truncate, and an index past the end of the array warns and appends rather than padding it with holes.
…e-read on upgrade
Assistant text used to join the .value of every response item, which dropped
inline file references ("files at:\n- \n- "), dropped tool invocations, and let
thinking text in. Items are now selected by kind: markdown is kept, inline
references render as their path, thinking is excluded, and each tool
invocation becomes a one-line tool chunk. Unrecognised kinds raise a drift
warning.
Each turn takes its request's own timestamp, falling back to the session
creation date. A plain-string response (2023 format) is accepted, and a
cancelled request keeps its prompt.
Indexed rows keep the old output because finished sessions are never read
again, so the scraper records COPILOT_SCRAPER_VERSION in its state; a stored
version below it makes one scan ignore its cutoff, and the version is saved
only after the read completes. scraperVersion is added to ScraperState
(same change as 56b376f on fix/claude-code-roles).
…m prompt, record tool runs
Events carrying data.parentToolCallId are output from a subagent the main
assistant launched; they were indexed as ordinary assistant turns. They are
now role "tool" with metadata { subagent: true, parentToolCallId }.
system.message (the CLI's own system prompt) is skipped. Each
tool.execution_start becomes a one-line tool chunk naming the tool and its
target (path, command, pattern, query, description or name), through the same
project-scope guard as other records; arguments are never dumped whole, since
an edit's would carry file contents.
Re-reads once on upgrade through COPILOT_CLI_SCRAPER_VERSION, with the same
mechanism as the Copilot VS Code and Claude Code scrapers: a stored version
below it ignores the cutoff and the cursors for one scan, and the version is
saved only after the read completes.
capture:formats only fingerprinted the legacy interactive.sessions blob, so the journal VS Code writes now (chatSessions/*.jsonl, records filed by kind since they have no type) had nothing to drift against. The capture script now records it as copilot-chat-sessions.json. The committed file is derived from the fixture journal, not a real store; running capture:formats -- --write on a real machine merges real shapes into it. The drift suite checks that the fixture writes no field the fingerprint does not know and that the fingerprint holds no record kind the replay does not apply.
…nt's prompt as the user's composerHeaders flags conversations a parent agent started. They are indexed with metadata.subagent and subagentType, and the first prompt, which the parent wrote, gets role 'tool' instead of 'user'.
Tool-call bubbles carry no text, so every edit, command and search an agent made was dropped (2,200 bubbles became 278 chunks on a real store). Each now yields a 'tool' chunk naming the tool and its target; the arguments are not indexed, and thinking bubbles stay out.
…arning Seen 34 times in real 2025 session files (chatSessions/*.json); it carries no conversation text.
Attribution now depends on the table's columns, so the committed fingerprint names them and a test ties it to the columns the scraper reads.
Only text parts were kept, so an assistant turn made of tool calls produced no chunk at all. Tool parts now yield a 'tool' chunk with a line per call (tool and target); outputs are not indexed and reasoning stays out.
…ince the cursor The cursor is one timestamp across all sessions, and a message was read only if it was created after it, so a message still streaming at the last scan while another session moved the cursor on was never read again. A session whose own, message or message-recorded time is past the cursor is now read from its first message, which also lets the index replace the partial row instead of keeping it beside the final one. time_updated is used where the schema has it.
The scraper now reports a missing table, which the battery's baseline-must-not-warn check treats as a failure.
…ng it `npx -y xtctx` downloaded ~633 MB (onnxruntime-node 212, onnxruntime-web 161, @huggingface/transformers 174) before it could answer, plus a ~106 MB model on first server start. A cold first --help took 97 s and MCP server starts through npx 17.8-147.9 s, past what an MCP client waits. Mechanism chosen: the runtime is not a dependency at all. `xtctx embeddings enable` copies a pinned package.json + package-lock.json shipped in embeddings-runtime/ into ~/.xtctx/embeddings, runs `npm ci --ignore-scripts` there, fetches the model, and writes a marker last. embeddings.ts and the calibration worker load the library from that directory by file URL. Why not the alternatives: - optionalDependencies: npm installs those by default, so `npx -y xtctx` (and the plugin's MCP command) would still fetch all of it. - A companion package: still needs an install step the user triggers, but adds a second published artifact, version-skew handling and a release to keep in lockstep. A lockfile in this package gives the same pin with none of that. - `npm install --prefix` with a version range: resolves on the registry's state that day. `npm ci` from the shipped lockfile pins the version and checks every package's integrity hash (all 76 entries carry one; the lock was resolved with --before so nothing in it is younger than 24 h, the youngest being 91 h at the time, and `npm audit` on it reports 0). --ignore-scripts is deliberate: the only install script in the tree is onnxruntime-node's, which on Linux x64 downloads CUDA binaries xtctx never uses (it times cpu, dml and webgpu). The model now downloads into ~/.xtctx/embeddings, which also survives npx cache eviction (it used to land inside the npx-cached package and be refetched). @huggingface/transformers stays a devDependency for the scripts, the eval and the smoke test; XTCTX_EMBEDDING_RUNTIME_DIR points a process at any directory that already has it (the repo root, in those). A test fails if any file under src/ imports the library in any form, or if it reappears in dependencies/optionalDependencies/peerDependencies.
…working state With no local runtime installed, the provider is NullEmbeddingProvider carrying a `semanticOff` reason (not_enabled, or disabled_by_env for XTCTX_DISABLE_EMBEDDINGS=1). The index reads it and never calls embed: - hybrid search answers from keyword directly: no stderr line, no `embedding_error`, no "N windows not yet vectorized" or "model still loading" note. An explicit `vector` search is told to run `xtctx embeddings enable`. - nothing counts as a vectorizing backlog while it is off. - vectors an earlier install built are kept. Previously the placeholder model name "null" read as "another model" to dropVectorsFromOtherModels and deleted every vector on first open without the model; this matters now because every existing user starts in that state until they enable. - `xtctx status` and `xtctx_continuity_status` (markdown and JSON: semantic_search, semantic_off_reason) name the mode and the command, and report kept vectors as kept rather than as an error. - calibration (server start, `scan --embed`, `xtctx calibrate`) only runs when the local model is configured and installed. - `scan --embed` says nothing was embedded and exits 1 when it is off. - the remote OpenAI-compatible provider needs no local runtime and is untouched, whether or not the add-on is installed. Tests that fail on the old code: no backlog/error/stderr noise with it off, vectors survive a reopen without the model, provider choice by environment, status text, calibration gated, `scan --embed` message.
…asured README, PRODUCT, ARCHITECTURE, the embedding docs, the testing notes and the landing page's status example and limits FAQ now say that the default install is keyword-only (about 55 MB on disk) and that `xtctx embeddings enable` adds the local model (about 540 MB on disk).
It is what xtctx embeddings enable installs from; check-review-coverage failed on its two files.
… indexed rows The cursor sits past every conversation already read, so rows indexed before the tool-line, subagent and attribution fixes would stay as they were until a conversation changed. Scraper state records a scraperVersion (same field the claude-code scraper uses); a stored version below the scraper's resets the cutoff for one scan, and the normal re-read path plus the index's prune replace the old rows. The version is saved only when the read ran to the end. Conversations whose sources are gone keep their rows.
Unpinned `npx -y xtctx` cost a registry round-trip on every agent start and let the SessionStart hook and the MCP server resolve different versions. Generated commands now name xtctx@<version>; re-running setup moves the pin, and `xtctx status` shows it and says when it differs from the running version. The plugin's own .mcp.json stays unpinned.
…ores Changes to history the index already held never reached it, in a single process with nothing concurrent: - An appended line stamped earlier than the index's saved lastTimestamp was skipped while the file's byte cursor moved past it, so no scan read it again. - A file rewritten in place was re-read from the top, but every record was filtered against that same lastTimestamp, so a rewritten early turn kept its old text. A rewrite past the first kilobyte was not noticed at all: the head hash covered only that kilobyte, so the read resumed at the old offset into different content. The Claude Code, Codex and Copilot CLI scrapers no longer default to the global timestamp; the per-file byte cursor decides what is new. Cursors also record a hash of the kilobyte before their offset, and a mismatch sends the file back to a full read. Re-read rows upsert by deterministic id, and the existing prune replaces the rows a rewrite changed.
…ndex lost Two servers scanning one growing session lost its newest rows for good. A scan that reads a session from the top deletes the rows it did not produce, so a server whose read predated an append deleted the rows a second server had just inserted for it, and that server's cursor already sat past those lines. Reproduced with three servers over a 10,000-message corpus while sessions grew: 70, 72 and 74 rows lost in three runs, still missing after two later rescans. - The prune only considers rows indexed no later than the scan began. Rows another scanner inserted while this one ran are not its to delete. - JSONL cursors record the last chunk their file yielded, and the scan hands the scrapers a probe to check it is still in the index. A cursor the index no longer backs is refused and the file is read again. A cursor from before this field existed is refused once, which re-reads those files a single time and restores rows an index already lost.
Every agent session starts its own xtctx server, and every server scanned the same transcript stores into the same index at once. With three servers started together on a 10,000-message corpus each read all of it (3.8-5.5 s of CPU apiece), and the overlap is what let one server's prune delete another's fresh rows. A scan now takes a lease first: a row in `settings`, claimed inside BEGIN IMMEDIATE, renewed every 5 s while the scan runs and expiring after 30 s without renewal. A holder whose process is gone from this machine loses it at once, so a server killed mid-scan by its host does not hold up the next one. A server that finds the lease taken waits for it and then scans, unless a scan that began after it asked has finished in the meantime, in which case that scan already read everything it would have. Waiting rather than skipping keeps the reason servers scan on every start: another server's scan in progress may have passed a store before this session's tool wrote to it. Callers still wait on the scan only up to the refresh budget and read what the holder has indexed so far.
… a scan A first scan killed between saving its cursor and building windows left the messages indexed and most of them unsearchable: the repair rebuilt four sessions per later scan. Measured over a 100-session corpus with the server killed at that point, 96 of 2,400 windows came back per later session. - The start-of-scan repair is bounded by the refresh budget instead of a count, and the scan runs it again at the end, unbounded, for whatever the first pass left. Each session commits on its own. - A scan marks every session it writes to before writing, and the rebuild clears the mark in the same transaction as the windows. A scan cut off in between leaves the mark, which finds a turn replaced at a position the windows already reach; the old coverage check could not.
The block in CLAUDE.md, AGENTS.md, GEMINI.md, copilot-instructions.md and the Cursor rule was ~30 lines listing every tool, the MCP command and a notes section. It is now what xtctx is, when to call xtctx_recent_sessions then xtctx_session_detail, that results are untrusted transcript text, and the skill pointer. It no longer carries the project path or a command, so it is identical wherever setup ran. The begin/end markers are unchanged.
An index from an older schema version was set aside and rebuilt from the transcripts still on disk. The index is the only copy of sessions whose transcripts have been cleaned up (Claude Code deletes them after 30 days by default), so every schema bump dropped those sessions from retrieval. Every step in the schema history is now a migration (0->1 drops the unread messages_fts table and the single-key vector table, 1->2 adds the git columns, 2->3 canonicalises project_root), run in one BEGIN IMMEDIATE transaction that re-reads the version so concurrent servers do not both migrate. The result is checked against a freshly created schema before the version is stamped. A migration also clears the scraper cursors (through a setting, so a crash in between still re-reads) so sessions still on disk are refreshed the way a rebuild used to refresh them. Set-aside is kept for corruption and for an older file in a shape no step recognises; a newer schema is still refused.
The scan runs on the thread that answers tool calls, over synchronous SQLite, and its longest stretch (building windows for every session it touched) never awaited anything. On a 15,000-message corpus a request needing no index at all waited up to 3.8 s behind it. A server told to shut down mid-scan sat out its 2 s grace window and was then killed by its own timer, so it never closed the index; with several servers open, write-ahead logs of up to 30MB stayed behind after they had all exited. - The scan awaits a checkpoint after every chunk and every session's windows. It yields the event loop when the scan has held it for 20 ms, renews the lease, and stops the scan when the index is closing or the lease was taken over. A stopped scan behaves like a killed one: what was written stays, no cursor moves past it, and touched sessions stay marked for the windows they did not get. - close() therefore returns at the next checkpoint instead of after the whole scan, skips vector warming, and checkpoints the write-ahead log with TRUNCATE when this process scanned and no other holds the lease.
…ebuild A corrupt index is set aside and rebuilt from the transcripts still on disk, but nothing read the set-aside file, so every session whose transcript had been cleaned up dropped out of retrieval anyway. After each full scan, every set-aside file beside the index that has not been read yet is opened (read-only where possible, leaving it as found), and every session the index lacks is copied in with its messages; windows are rebuilt by the normal per-session path. Each session is read in its own try, so a damaged page costs only the sessions on it, and the counts are recorded under carried_forward:<file> and reported on stderr. Writes to the new index stay outside that try, so a lock is retried on the next scan rather than recorded as unreadable. Files are found by listing the directory, so one left by a crash or by an earlier version is picked up too, and none is deleted.
The index is the only copy of sessions whose transcripts have been cleaned up, and there was no way to keep a copy of it anywhere else. `xtctx export [--out <file>]` writes this project's sessions and messages as JSON Lines (format version 1, documented in src/handoff/export-file.ts): a header line, one line per session carrying its messages, and an end line that counts them so a file cut short is recognised. It reads the index as it stands, never scans or touches a transcript, and never overwrites an existing file; a file it could not finish is removed. `--out -` writes to stdout. `xtctx import <file>` merges an export into this project's index, one transaction per session. Message ids are content hashes, so importing twice adds nothing and a session already present gains only what it lacks. Windows are rebuilt for every session that changed. A file that is not an export, or is from a newer format, is refused before anything is written; invalid lines and a missing end line are reported and exit nonzero.
Once Claude Code has cleaned up a transcript (30 days by default), the index is the only copy of that session, and nothing in `xtctx status` said so. Scrapers can now list the session ids in their store without reading them (optional `listSessionIds`); Claude Code implements it as a listing of the same directories it reads. Status compares that with the sessions indexed for the project and, when any are gone from disk, prints: Backup N sessions exist only in this index; back them up with `xtctx export` A store that cannot be listed counts nothing, so the line is never a guess. The count is also on HandoffStatus as index_only_sessions.
…older ones PRODUCT.md, ARCHITECTURE.md, README.md and docs/architecture.md now say the same thing: for every session whose transcript is still on disk the index is derived data; for the older ones whose transcripts have been cleaned up it is the only copy, and deleting it loses them. They describe what keeps those sessions (in-place schema migration, carry-forward from a set-aside index) and the export/import commands, and the README no longer calls the index a cache. The backup design note records what was built and how its open questions were answered. Two comments in the CLI entry point that justified exiting mid-scan with "the index is derived data" now give only the reason that is true: every chunk is committed as it is written and an interrupted scan resumes.
…checked against the index, rewritten history re-read
…xtctx export/import, status backup line # Conflicts: # src/cli/index.ts # src/handoff/sqlite-index.ts # src/types/scraper.ts
…ered, ANSI stripped; detail is tail-first; transcript text labelled untrusted # Conflicts: # src/scrapers/claude-code.ts
…rites checked cursors The version bump (claude-code roles), the cursor refusal for cursors without lastEmitted, and the prune bounded by scan start all act on the first scan after an upgrade. Pins that together they give one correct read: rows indexed before the scan are pruned, the new cursors carry lastEmitted, the next scan trusts them, and a read cut short leaves the version unsaved.
…mestamps, Copilot CLI subagents as role tool # Conflicts: # src/scrapers/copilot-cli.ts
…e re-reads changed sessions; macOS/Linux store paths
…tatus says when semantic search is off # Conflicts: # src/handoff/sqlite-index.ts
The carry-forward tests made an old index fail its open by damaging the vector table, which the open read through dropVectorsFromOtherModels. With semantic search off that call is skipped, the open never touches the vectors, and the damaged file was used instead of set aside. Damage the settings table instead: the open always reads it (the migration marker), and carry-forward reads only sessions and messages.
…lock, README claim fix, untrusted Preview label, CHANGELOG [Unreleased]
…nned One [Unreleased] line per merged fix branch, in the file's groups. The README and docs/architecture.md still showed setup writing an unpinned `npx -y xtctx`; since the version pin it writes `xtctx@<version>`, and only the plugin runs the unpinned package.
This was referenced Oct 2, 2026
…00ms It guards that the corpus is big enough for the gap ratio to mean anything, not speed; a fast ubuntu runner finished in 994ms and failed it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the issues found in the 2026-10-01 whole-product audit, as seven branches merged with
--no-ffso each keeps its own history.What changes for users
xtctx export/xtctx import;xtctx statussays how many sessions exist only in the index.npxinstall 550 MB → 55 MB measured);xtctx embeddings enableadds semantic search. Fixes vectors being deleted when the index was opened without the model.xtctx@<version>; 8-line instruction block with no project path; untrusted-text labels on JSON output, the hook preview and the recent-sessions preview; README/PRODUCT/ARCHITECTURE agree; CHANGELOG has[Unreleased], and the release workflow now writes beneath it.Already-indexed sessions are re-read once on upgrade so they pick up the corrected roles and content; sessions whose transcripts are gone keep their old rows.
Verification
npm run lint,npm run typecheck,npm test(963),test:security,test:smoke,test:drift,test:integration,test:eval,security:checklist,check:review-coverage,check:workflows,build,npm pack --dry-run,demo:public,smoke:cli,audit:production(clean after deps: clear the advisories blocking the release gate #398).Known gaps