Skip to content

fix: audit findings — scrapers, concurrent scans, index durability, optional embeddings - #403

Merged
fstubner merged 54 commits into
mainfrom
fix/audit-fixes
Oct 2, 2026
Merged

fstubner merged 54 commits into
mainfrom
fix/audit-fixes

Conversation

@fstubner

@fstubner fstubner commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Fixes the issues found in the 2026-10-01 whole-product audit, as seven branches merged with --no-ff so each keeps its own history.

What changes for users

  • Claude Code sessions read correctly. Tool results were indexed as the user speaking (11,434 "user" vs 5,325 "assistant" rows in a real index); tool calls were dropped; ANSI codes were kept. Session detail now returns the end of a session within a 40k-character budget.
  • Copilot, Cursor and opencode sessions read correctly. Copilot's chat journal was replayed as an insert, duplicating a question and attaching answers to the wrong requests; subagent prompts in Cursor and Copilot CLI were filed as the user; Cursor attributed conversations by guessing from file paths (3 of 100 landed in a second project) and its default path was wrong on macOS/Linux; tool calls were dropped in Cursor and opencode.
  • No more lost rows when several agents run at once. A cross-process scan lease, a prune bounded by scan start, and cursors that are refused when their rows are missing. Measured on a synthetic corpus: 70–120 rows permanently lost per run before, 0 after; 5 servers idle in ~11 s instead of 21–24 s. Edits to already-indexed history are now picked up; interrupted scans no longer leave search mostly empty; tool calls no longer stall for seconds during a scan.
  • The index keeps sessions it is the only copy of. Schema upgrades migrate in place; sessions in a set-aside index are carried forward; xtctx export / xtctx import; xtctx status says how many sessions exist only in the index.
  • Small default install. The local embedding runtime is no longer a dependency (npx install 550 MB → 55 MB measured); xtctx embeddings enable adds semantic search. Fixes vectors being deleted when the index was opened without the model.
  • Setup and docs. Generated configs pin xtctx@<version>; 8-line instruction block with no project path; untrusted-text labels on JSON output, the hook preview and the recent-sessions preview; README/PRODUCT/ARCHITECTURE agree; CHANGELOG has [Unreleased], and the release workflow now writes beneath it.

Already-indexed sessions are re-read once on upgrade so they pick up the corrected roles and content; sessions whose transcripts are gone keep their old rows.

Verification

  • npm run lint, npm run typecheck, npm test (963), test:security, test:smoke, test:drift, test:integration, test:eval, security:checklist, check:review-coverage, check:workflows, build, npm pack --dry-run, demo:public, smoke:cli, audit:production (clean after deps: clear the advisories blocking the release gate #398).
  • Every new test was shown to fail on the old code.
  • End-to-end with the built CLI in a throwaway project and redirected home, three runs: three servers at once, roles, tail-first detail, no lost rows vs a single-process scan, status lines, export → delete → import.
  • The new Copilot and Cursor parsing was checked read-only against real stores on the author's machine (Cursor scan for this repo: 18.4 s → 1.0 s).

Known gaps

  • A process killed in the moment between saving cursors and pruning after an upgrade re-read keeps old rows next to new ones (same window as before).
  • Rewrite detection samples the first and last kilobyte before the read offset; a same-length edit in the middle of a long file is not caught.

…er tool calls; strip ANSI

message.role is the API role: Claude Code writes tool results back as user
records, so 11,434 of a real index's messages were 'user' tool output. A
record carrying only tool blocks is now role tool; tool_use blocks render as
one short line (ran Bash: ..., edit <path>) instead of being dropped.
…e budget

xtctx_session_detail returned the oldest 50 messages with no total cap; the
first page of real recent sessions measured 75k-206k characters and a
7,996-message session needed offset=7946 to reach its end. It now returns the
newest messages when offset is omitted, caps a response at 40,000 characters
(2.5x the per-message cap), excerpts tool output to 1,500 characters with the
original length noted, and prints the offset for earlier messages. An explicit
offset still counts from the start; from_end=true counts back from the newest.
Messages carry a 0-based position. Also adds the untrusted marker to the JSON
payloads built in the same file (see next commit).
…the hook preview

JSON output of the session tools and the manifest carried raw transcript text
in bare fields; markdown fences it and says so. JSON payloads now carry
untrusted: true and a notice. The SessionStart hook's 'Opened with' line is
labelled on the line itself, and its pointer now says detail returns the most
recent messages, not the full turn history.
…rect indexed roles

The resume cursor sits past every finished session, so rows indexed with the
old roles would stay wrong until a transcript grew. Scraper state now records
a scraperVersion; a stored version below the scraper's resets the cutoff and
cursors for one scan, and the normal re-read path plus pruneRereadSessions
replaces the old rows. The version is saved only when the read ran to the end.
Sessions whose transcripts are gone keep their rows.
Without APPDATA the default was ~/.cursor/workspaceStorage on every platform,
where Cursor keeps extensions rather than conversations. macOS and Linux now
use Application Support and ~/.config, computed the way the VS Code path is.
…ing from file paths

Current Cursor no longer lists conversations in each workspace; globalStorage's
composerHeaders table records which workspace owns each one. Reading it first
stops conversations that touched a second project's files being filed under it.
A conversation the header cannot place (no row, or a workspace that resolves to
no folder) still falls back to the recorded-path match.

With the table present the unlisted-conversation search is a lookup by id plus a
key-only range, so the LIKE over every stored value (9-28s on a 6.9GB store)
only runs when the table is missing, which is also now reported as drift.
A kind-2 record with an index means "cut the array back to length i, then
push v". Applying it as an insert left the old copy of a rewritten request
beside the new one, which duplicated the first question and attached answers
to the wrong requests. The timestamp sort that hid the resulting misordering
is removed; array order is conversation order once the truncate is applied.

An unknown record kind now raises a drift warning instead of being skipped,
a record with an index and no values is a bare truncate, and an index past the
end of the array warns and appends rather than padding it with holes.
…e-read on upgrade

Assistant text used to join the .value of every response item, which dropped
inline file references ("files at:\n- \n- "), dropped tool invocations, and let
thinking text in. Items are now selected by kind: markdown is kept, inline
references render as their path, thinking is excluded, and each tool
invocation becomes a one-line tool chunk. Unrecognised kinds raise a drift
warning.

Each turn takes its request's own timestamp, falling back to the session
creation date. A plain-string response (2023 format) is accepted, and a
cancelled request keeps its prompt.

Indexed rows keep the old output because finished sessions are never read
again, so the scraper records COPILOT_SCRAPER_VERSION in its state; a stored
version below it makes one scan ignore its cutoff, and the version is saved
only after the read completes. scraperVersion is added to ScraperState
(same change as 56b376f on fix/claude-code-roles).
…m prompt, record tool runs

Events carrying data.parentToolCallId are output from a subagent the main
assistant launched; they were indexed as ordinary assistant turns. They are
now role "tool" with metadata { subagent: true, parentToolCallId }.
system.message (the CLI's own system prompt) is skipped. Each
tool.execution_start becomes a one-line tool chunk naming the tool and its
target (path, command, pattern, query, description or name), through the same
project-scope guard as other records; arguments are never dumped whole, since
an edit's would carry file contents.

Re-reads once on upgrade through COPILOT_CLI_SCRAPER_VERSION, with the same
mechanism as the Copilot VS Code and Claude Code scrapers: a stored version
below it ignores the cutoff and the cursors for one scan, and the version is
saved only after the read completes.
capture:formats only fingerprinted the legacy interactive.sessions blob, so
the journal VS Code writes now (chatSessions/*.jsonl, records filed by kind
since they have no type) had nothing to drift against. The capture script now
records it as copilot-chat-sessions.json.

The committed file is derived from the fixture journal, not a real store;
running capture:formats -- --write on a real machine merges real shapes into
it. The drift suite checks that the fixture writes no field the fingerprint
does not know and that the fingerprint holds no record kind the replay does
not apply.
…nt's prompt as the user's

composerHeaders flags conversations a parent agent started. They are indexed with
metadata.subagent and subagentType, and the first prompt, which the parent wrote,
gets role 'tool' instead of 'user'.
Tool-call bubbles carry no text, so every edit, command and search an agent made
was dropped (2,200 bubbles became 278 chunks on a real store). Each now yields a
'tool' chunk naming the tool and its target; the arguments are not indexed, and
thinking bubbles stay out.
…arning

Seen 34 times in real 2025 session files (chatSessions/*.json); it carries
no conversation text.
Attribution now depends on the table's columns, so the committed fingerprint
names them and a test ties it to the columns the scraper reads.
Only text parts were kept, so an assistant turn made of tool calls produced no
chunk at all. Tool parts now yield a 'tool' chunk with a line per call (tool and
target); outputs are not indexed and reasoning stays out.
…ince the cursor

The cursor is one timestamp across all sessions, and a message was read only if
it was created after it, so a message still streaming at the last scan while
another session moved the cursor on was never read again. A session whose own,
message or message-recorded time is past the cursor is now read from its first
message, which also lets the index replace the partial row instead of keeping it
beside the final one. time_updated is used where the schema has it.
The scraper now reports a missing table, which the battery's baseline-must-not-warn check treats as a failure.
…ng it

`npx -y xtctx` downloaded ~633 MB (onnxruntime-node 212, onnxruntime-web
161, @huggingface/transformers 174) before it could answer, plus a ~106 MB
model on first server start. A cold first --help took 97 s and MCP server
starts through npx 17.8-147.9 s, past what an MCP client waits.

Mechanism chosen: the runtime is not a dependency at all. `xtctx embeddings
enable` copies a pinned package.json + package-lock.json shipped in
embeddings-runtime/ into ~/.xtctx/embeddings, runs `npm ci --ignore-scripts`
there, fetches the model, and writes a marker last. embeddings.ts and the
calibration worker load the library from that directory by file URL.

Why not the alternatives:
- optionalDependencies: npm installs those by default, so `npx -y xtctx`
  (and the plugin's MCP command) would still fetch all of it.
- A companion package: still needs an install step the user triggers, but
  adds a second published artifact, version-skew handling and a release to
  keep in lockstep. A lockfile in this package gives the same pin with none
  of that.
- `npm install --prefix` with a version range: resolves on the registry's
  state that day. `npm ci` from the shipped lockfile pins the version and
  checks every package's integrity hash (all 76 entries carry one; the lock
  was resolved with --before so nothing in it is younger than 24 h, the
  youngest being 91 h at the time, and `npm audit` on it reports 0).

--ignore-scripts is deliberate: the only install script in the tree is
onnxruntime-node's, which on Linux x64 downloads CUDA binaries xtctx never
uses (it times cpu, dml and webgpu).

The model now downloads into ~/.xtctx/embeddings, which also survives npx
cache eviction (it used to land inside the npx-cached package and be
refetched).

@huggingface/transformers stays a devDependency for the scripts, the eval
and the smoke test; XTCTX_EMBEDDING_RUNTIME_DIR points a process at any
directory that already has it (the repo root, in those).

A test fails if any file under src/ imports the library in any form, or if
it reappears in dependencies/optionalDependencies/peerDependencies.
…working state

With no local runtime installed, the provider is NullEmbeddingProvider
carrying a `semanticOff` reason (not_enabled, or disabled_by_env for
XTCTX_DISABLE_EMBEDDINGS=1). The index reads it and never calls embed:

- hybrid search answers from keyword directly: no stderr line, no
  `embedding_error`, no "N windows not yet vectorized" or "model still
  loading" note. An explicit `vector` search is told to run
  `xtctx embeddings enable`.
- nothing counts as a vectorizing backlog while it is off.
- vectors an earlier install built are kept. Previously the placeholder
  model name "null" read as "another model" to dropVectorsFromOtherModels
  and deleted every vector on first open without the model; this matters
  now because every existing user starts in that state until they enable.
- `xtctx status` and `xtctx_continuity_status` (markdown and JSON:
  semantic_search, semantic_off_reason) name the mode and the command, and
  report kept vectors as kept rather than as an error.
- calibration (server start, `scan --embed`, `xtctx calibrate`) only runs
  when the local model is configured and installed.
- `scan --embed` says nothing was embedded and exits 1 when it is off.
- the remote OpenAI-compatible provider needs no local runtime and is
  untouched, whether or not the add-on is installed.

Tests that fail on the old code: no backlog/error/stderr noise with it off,
vectors survive a reopen without the model, provider choice by environment,
status text, calibration gated, `scan --embed` message.
…asured

README, PRODUCT, ARCHITECTURE, the embedding docs, the testing notes and the
landing page's status example and limits FAQ now say that the default
install is keyword-only (about 55 MB on disk) and that `xtctx embeddings
enable` adds the local model (about 540 MB on disk).
It is what xtctx embeddings enable installs from; check-review-coverage
failed on its two files.
… indexed rows

The cursor sits past every conversation already read, so rows indexed before the
tool-line, subagent and attribution fixes would stay as they were until a
conversation changed. Scraper state records a scraperVersion (same field the
claude-code scraper uses); a stored version below the scraper's resets the cutoff
for one scan, and the normal re-read path plus the index's prune replace the old
rows. The version is saved only when the read ran to the end. Conversations whose
sources are gone keep their rows.
Unpinned `npx -y xtctx` cost a registry round-trip on every agent start
and let the SessionStart hook and the MCP server resolve different
versions. Generated commands now name xtctx@<version>; re-running setup
moves the pin, and `xtctx status` shows it and says when it differs from
the running version. The plugin's own .mcp.json stays unpinned.
…ores

Changes to history the index already held never reached it, in a single
process with nothing concurrent:

- An appended line stamped earlier than the index's saved lastTimestamp
  was skipped while the file's byte cursor moved past it, so no scan read
  it again.
- A file rewritten in place was re-read from the top, but every record
  was filtered against that same lastTimestamp, so a rewritten early turn
  kept its old text. A rewrite past the first kilobyte was not noticed at
  all: the head hash covered only that kilobyte, so the read resumed at
  the old offset into different content.

The Claude Code, Codex and Copilot CLI scrapers no longer default to the
global timestamp; the per-file byte cursor decides what is new. Cursors
also record a hash of the kilobyte before their offset, and a mismatch
sends the file back to a full read. Re-read rows upsert by deterministic
id, and the existing prune replaces the rows a rewrite changed.
…ndex lost

Two servers scanning one growing session lost its newest rows for good.
A scan that reads a session from the top deletes the rows it did not
produce, so a server whose read predated an append deleted the rows a
second server had just inserted for it, and that server's cursor already
sat past those lines. Reproduced with three servers over a 10,000-message
corpus while sessions grew: 70, 72 and 74 rows lost in three runs, still
missing after two later rescans.

- The prune only considers rows indexed no later than the scan began.
  Rows another scanner inserted while this one ran are not its to delete.
- JSONL cursors record the last chunk their file yielded, and the scan
  hands the scrapers a probe to check it is still in the index. A cursor
  the index no longer backs is refused and the file is read again. A
  cursor from before this field existed is refused once, which re-reads
  those files a single time and restores rows an index already lost.
Every agent session starts its own xtctx server, and every server scanned
the same transcript stores into the same index at once. With three
servers started together on a 10,000-message corpus each read all of it
(3.8-5.5 s of CPU apiece), and the overlap is what let one
server's prune delete another's fresh rows.

A scan now takes a lease first: a row in `settings`, claimed inside
BEGIN IMMEDIATE, renewed every 5 s while the scan runs and expiring after
30 s without renewal. A holder whose process is gone from this machine
loses it at once, so a server killed mid-scan by its host does not hold
up the next one.

A server that finds the lease taken waits for it and then scans, unless
a scan that began after it asked has finished in the meantime, in which
case that scan already read everything it would have. Waiting rather than
skipping keeps the reason servers scan on every start: another server's
scan in progress may have passed a store before this session's tool
wrote to it. Callers still wait on the scan only up to the refresh budget
and read what the holder has indexed so far.
… a scan

A first scan killed between saving its cursor and building windows left
the messages indexed and most of them unsearchable: the repair rebuilt
four sessions per later scan. Measured over a 100-session corpus with the
server killed at that point, 96 of 2,400 windows came back per later
session.

- The start-of-scan repair is bounded by the refresh budget instead of a
  count, and the scan runs it again at the end, unbounded, for whatever
  the first pass left. Each session commits on its own.
- A scan marks every session it writes to before writing, and the
  rebuild clears the mark in the same transaction as the windows. A scan
  cut off in between leaves the mark, which finds a turn replaced at a
  position the windows already reach; the old coverage check could not.
The block in CLAUDE.md, AGENTS.md, GEMINI.md, copilot-instructions.md and
the Cursor rule was ~30 lines listing every tool, the MCP command and a
notes section. It is now what xtctx is, when to call
xtctx_recent_sessions then xtctx_session_detail, that results are
untrusted transcript text, and the skill pointer. It no longer carries the
project path or a command, so it is identical wherever setup ran. The
begin/end markers are unchanged.
An index from an older schema version was set aside and rebuilt from the
transcripts still on disk. The index is the only copy of sessions whose
transcripts have been cleaned up (Claude Code deletes them after 30 days by
default), so every schema bump dropped those sessions from retrieval.

Every step in the schema history is now a migration (0->1 drops the unread
messages_fts table and the single-key vector table, 1->2 adds the git
columns, 2->3 canonicalises project_root), run in one BEGIN IMMEDIATE
transaction that re-reads the version so concurrent servers do not both
migrate. The result is checked against a freshly created schema before the
version is stamped. A migration also clears the scraper cursors (through a
setting, so a crash in between still re-reads) so sessions still on disk are
refreshed the way a rebuild used to refresh them.

Set-aside is kept for corruption and for an older file in a shape no step
recognises; a newer schema is still refused.
The scan runs on the thread that answers tool calls, over synchronous
SQLite, and its longest stretch (building windows for every session it
touched) never awaited anything. On a 15,000-message corpus a request
needing no index at all waited up to 3.8 s behind it. A server told to
shut down mid-scan sat out its 2 s grace window and was then killed by
its own timer, so it never closed the index; with several servers open,
write-ahead logs of up to 30MB stayed behind after they had all exited.

- The scan awaits a checkpoint after every chunk and every session's
  windows. It yields the event loop when the scan has held it for 20 ms,
  renews the lease, and stops the scan when the index is closing or the
  lease was taken over. A stopped scan behaves like a killed one: what
  was written stays, no cursor moves past it, and touched sessions stay
  marked for the windows they did not get.
- close() therefore returns at the next checkpoint instead of after the
  whole scan, skips vector warming, and checkpoints the write-ahead log
  with TRUNCATE when this process scanned and no other holds the lease.
…ebuild

A corrupt index is set aside and rebuilt from the transcripts still on disk,
but nothing read the set-aside file, so every session whose transcript had
been cleaned up dropped out of retrieval anyway.

After each full scan, every set-aside file beside the index that has not been
read yet is opened (read-only where possible, leaving it as found), and every
session the index lacks is copied in with its messages; windows are rebuilt
by the normal per-session path. Each session is read in its own try, so a
damaged page costs only the sessions on it, and the counts are recorded under
carried_forward:<file> and reported on stderr. Writes to the new index stay
outside that try, so a lock is retried on the next scan rather than recorded
as unreadable. Files are found by listing the directory, so one left by a
crash or by an earlier version is picked up too, and none is deleted.
The index is the only copy of sessions whose transcripts have been cleaned
up, and there was no way to keep a copy of it anywhere else.

`xtctx export [--out <file>]` writes this project's sessions and messages
as JSON Lines (format version 1, documented in src/handoff/export-file.ts):
a header line, one line per session carrying its messages, and an end line
that counts them so a file cut short is recognised. It reads the index as it
stands, never scans or touches a transcript, and never overwrites an
existing file; a file it could not finish is removed. `--out -` writes to
stdout.

`xtctx import <file>` merges an export into this project's index, one
transaction per session. Message ids are content hashes, so importing twice
adds nothing and a session already present gains only what it lacks.
Windows are rebuilt for every session that changed. A file that is not an
export, or is from a newer format, is refused before anything is written;
invalid lines and a missing end line are reported and exit nonzero.
Once Claude Code has cleaned up a transcript (30 days by default), the index
is the only copy of that session, and nothing in `xtctx status` said so.

Scrapers can now list the session ids in their store without reading them
(optional `listSessionIds`); Claude Code implements it as a listing of the
same directories it reads. Status compares that with the sessions indexed
for the project and, when any are gone from disk, prints:

  Backup   N sessions exist only in this index; back them up with `xtctx export`

A store that cannot be listed counts nothing, so the line is never a guess.
The count is also on HandoffStatus as index_only_sessions.
…older ones

PRODUCT.md, ARCHITECTURE.md, README.md and docs/architecture.md now say the
same thing: for every session whose transcript is still on disk the index is
derived data; for the older ones whose transcripts have been cleaned up it is
the only copy, and deleting it loses them. They describe what keeps those
sessions (in-place schema migration, carry-forward from a set-aside index)
and the export/import commands, and the README no longer calls the index a
cache. The backup design note records what was built and how its open
questions were answered.

Two comments in the CLI entry point that justified exiting mid-scan with
"the index is derived data" now give only the reason that is true: every
chunk is committed as it is written and an interrupted scan resumes.
…checked against the index, rewritten history re-read
…xtctx export/import, status backup line

# Conflicts:
#	src/cli/index.ts
#	src/handoff/sqlite-index.ts
#	src/types/scraper.ts
…ered, ANSI stripped; detail is tail-first; transcript text labelled untrusted

# Conflicts:
#	src/scrapers/claude-code.ts
…rites checked cursors

The version bump (claude-code roles), the cursor refusal for cursors
without lastEmitted, and the prune bounded by scan start all act on the
first scan after an upgrade. Pins that together they give one correct
read: rows indexed before the scan are pruned, the new cursors carry
lastEmitted, the next scan trusts them, and a read cut short leaves the
version unsaved.
…mestamps, Copilot CLI subagents as role tool

# Conflicts:
#	src/scrapers/copilot-cli.ts
…e re-reads changed sessions; macOS/Linux store paths
…tatus says when semantic search is off

# Conflicts:
#	src/handoff/sqlite-index.ts
The carry-forward tests made an old index fail its open by damaging the
vector table, which the open read through dropVectorsFromOtherModels.
With semantic search off that call is skipped, the open never touches
the vectors, and the damaged file was used instead of set aside. Damage
the settings table instead: the open always reads it (the migration
marker), and carry-forward reads only sessions and messages.
…lock, README claim fix, untrusted Preview label, CHANGELOG [Unreleased]
…nned

One [Unreleased] line per merged fix branch, in the file's groups. The
README and docs/architecture.md still showed setup writing an unpinned
`npx -y xtctx`; since the version pin it writes `xtctx@<version>`, and
only the plugin runs the unpinned package.
…00ms

It guards that the corpus is big enough for the gap ratio to mean anything,
not speed; a fast ubuntu runner finished in 994ms and failed it.
@fstubner
fstubner merged commit ccd4244 into main Oct 2, 2026
5 checks passed
@fstubner
fstubner deleted the fix/audit-fixes branch October 2, 2026 22:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant