Repository navigation
fix: close the five findings left open by the audit - #383
Merged
Merged
Conversation
All four code defects came out of the same audit as #380 and share its shape: a read degrades correctly, and the degraded value is then used as the base for a write or a filter. Antigravity dropped every step whose timestamp fell back to the epoch sentinel. `steps.ts` returns `new Date(0)` when neither the step metadata nor the summary carries a parseable time, `fullSync` reads with `since = new Date(0)`, and `0 <= 0` is true -- so those steps were lost on every path including the rebuild that exists to recover them, silently. Every other scraper guards this; claude-code even spells out why. `disconnect` stripped comments from a TOML config that `setup` refuses to touch. The write path checks `tomlHasComments` and declines, telling the user to edit by hand and keep their comments; the remove path re-serialised and deleted them all, which made following that advice pointless. `disconnect` rewrote an unparseable `.xtctx/config.yaml` from `{}`. Degrading the read is right and deliberate -- disconnect is what someone reaches for when things are already wrong. Using `{}` as the base for the write is not: the file came back holding nothing but `tools:`, and any `storePath` override went with it. Everything else still disconnects; only this rewrite is skipped, with a warning naming the file. Cursor renumbered a composer after a pruned bubble. Two skips did not advance `messageIndex` while the two below them did, and `scan.ts` hashes that index into the row id -- so a bubble Cursor pruned between scans shifted every later turn and re-inserted it under a new id beside the row already stored. The fifth finding is not a fix, because it could not be reproduced. The smoke suite failed four of five cases once under load and passed on every attempt since, including under deliberate contention. What made that unfixable is that nothing drained the spawned server's stderr, so the only evidence was "initialize did not answer within 60s" -- a symptom, with the cause piped into a buffer nobody read. That text is now attached to the failure, and an exited child fails its pending call immediately instead of waiting out the timeout. Each test fails against the previous behaviour: the antigravity one on an empty chunk list, cursor on index 1 where 2 belongs, and both disconnect ones on the destroyed file.
Neither area had ever been reviewed. Both had defects, which is the answer to whether the first audit had found everything. `detail_offset` pointed past the end of the session it was meant to point into. The `target` CTE had no `LIMIT 1` and the query cross-joins `messages` against it, so a session with duplicate `message_index` values multiplied the count by however many duplicates there were. Those duplicates are not hypothetical: the comment directly above the statement records 828 of them in one real session. Reproduced at six messages with three sharing an index -- offset 12 into a six-row session, so `getSessionDetail` paged past the end and returned nothing. That is the "match points somewhere unrelated" failure this statement exists to prevent, in its worst form. `literalUnreadableTools` never left `buildIndexProgress`. The field was not declared on `ProgressInputs` and the caller passes its inputs by spread, so excess-property checking could not catch it and the value was dropped in silence. The "this store cannot be read" branch in `sessions.ts` was therefore unreachable, and every unreadable store was reported as a search that stopped at its limit -- advising the caller to narrow a query, which against an unreadable store returns the same nothing forever. The comment explaining why that advice is wrong was already there; the code that acts on it could never run. `session_refs` widened silently. `uniqueStrings` answered `[]` for a non-array, and `[]` routes to the branch returning recent sessions, so a caller asking for three specific sessions got arbitrary recent ones with `missing_session_refs` empty. `validatedFilter` exists to stop exactly this and its docstring names this handler; the fix had reached `tool_filter` and `branch_filter` and missed this one. Also capped, since each ref is a synchronous sqlite `get` on the event-loop thread. Two leaks on the continuity surface, both one branch away from code that already did it right: `embedding_error` was `inlineSafe`d but not `sanitizeErrorMessage`d, though it carries an absolute path on a model load failure and `tool.last_error` two lines below gets both; and `redirected_tools` was joined raw, though its entries are arbitrary keys from a committable config file while `tool.tool` beside it -- a scraper literal -- was escaped. `loadProjectConfig` treated every read failure as a missing config. `present: false` makes every tool tell the agent to run setup, and setup rewrites config.yaml with every tool enabled, so a locked or busy file would end with the user's `enabled: false` undone. The parse-failure branch beside it distinguishes this case carefully; the read-failure branch did not.
… to see The scripts and the test suite had never been audited. Both were hiding failures rather than catching them, which is the worse half of the thirteen defects found earlier today: some of those were watched by tests that could not see them. `audit:production` reported a clean audit of a tree it had not read. With `--json`, an npm that cannot reach the registry prints an error object and exits 0, so the body parses, carries no `vulnerabilities`, and every check found nothing wrong. Adding `--json` is what inverted it -- bare `npm audit` exits 1 on the same failure. Reproduced with `npm_config_registry=http://127.0.0.1:9`: "audit clean", exit 0, and `verify:release` would have passed on it. A security gate has to fail closed; not knowing is not the same as knowing there is nothing. `shapeOf` could not see inside an array. The docstring said "one level of arrays" and the code did none, so `{content:[{type,text}]}` and `{content:[{kind,body}]}` produced byte-identical shapes. Claude Code's entire assistant payload is `message.content[]`, so the mechanism that exists to catch upstream format drift was blind to the place drift would appear -- in `capture:formats` and in `fixture-fidelity`, which runs in CI and in `verify:release`. Elements are now shaped under `path[]`, keeping the key redaction that stops private paths reaching a committed fingerprint. The security checklist could not see an indented unchecked control -- the matcher was anchored at column 0 while the file uses nested bullets -- and counted a citation pointing at a directory as a verified test file. Two tests could not fail, both proven by breaking the source and watching them stay green. `capSegments` swapped for `segments.slice(0, limit)`: the test asserted first element, last element and ascending order, all of which truncation satisfies. `detail_offset` hardcoded to `0`: the test compared `getSessionDetail(ref, offset, 1)` against the same call sliced at the same offset, so both sides moved together and it proved paging was self-consistent, not that the pointer was right. That is the same feature found broken earlier today by the cross-join in `messageOffsetInSession` -- the test was there and could not see it. The backlog test now asserts the backlog is zero, not only that two numbers agree; it held with one window of twenty-four embedded. Not changed, deliberately: the segment cap stays at 16. Its test only restates the constant, which is a floor against lowering it rather than the percentile property the heading claimed, and that is now written down. Raising it to 17 is a cost decision about embedding time, not a test fix.
…o see A fourth audit, this time asking only one question of the suite: which tests would still pass if the code they cover were broken. Six could not fail, and one of them was hiding a live defect. Drift locations from an incremental scan were wrong. The counter started at zero on every pass while a resumed read starts at `startAt` BYTES -- `readJsonlLines` yields only lines past the cursor -- so a record appended as line 101 reported as `path:1`. A drift location is the only pointer anyone has when chasing an upstream format break, and every one produced after a resume pointed somewhere else. It is now the byte offset the reader already tracks, written `path@offset`, which means the same thing on a first read and a resumed one. Nothing caught that, and nothing could: no test in the suite asserted a location any scraper computed. The two that assert `firstLocation` pass the string to `recordDrift` themselves, so they pin the drift log's plumbing. Changing the format of every location in two scrapers left 700 tests green. Now pinned, including after a resume, where the old behaviour fails with "expected 1 to be greater than 100". The manifest's markdown branch had no positive assertion anywhere in the repository. Returning `""` from it passed all 697 tests: every assertion on that surface was a `not.toMatch`, and an empty string forges nothing. An entire public MCP output branch -- heading, project line, per-session block, retrieve hint, missing-refs list -- was undefended. Each scrub case now pins the value it scrubbed as well as the structure it refused to forge, because a scrub that works by rendering nothing is not a scrub. `fenceFor` was pinned by `>= 4` tildes, which a constant four-tilde fence satisfies -- and content holding four tildes then closes it from the inside. Now parameterised over three run lengths and asserting the delimiter is strictly longer than the content's longest run. The codex cursor test compared `offset` to `size` from the same written record, so recording the file's size instead of the boundary read passed it. That is the permanent silent loss the other two readers pin. Now compared against the real file, with the partial-trailing-line case that actually distinguishes them. Also: the shutdown test now asserts the callback has NOT run before stdin ends, so wiring that tears the server down at boot is distinguishable from wiring that works; and the opencode ordering test no longer sorts the very property it is named for.
None was a live defect. All nine were tests that could not see the thing
they were named for, which is the class the last two audits were looking
for and the reason the earlier defects went unnoticed.
The landing test pinned `v9.astro` -- a design draft carrying
`noindex,nofollow`. The page visitors are actually sent to, `index.astro`
and the components it composes, was pinned by nothing, so the product
claims and the punctuation rule were enforced only on a page nobody
reaches. Retargeting found two em dashes in live hero and install copy
that the draft-only check had never looked at; both sentences are
reworded. The test states plainly that it reads source rather than
rendered output, and what that cannot catch.
Dropped one assertion rather than ship it: banning the words "daemon" and
"dashboard" flagged the page's own copy saying it has neither. Asserting
the local-only commitment is present survives rewording; a substring ban
on correct copy does not.
The integration fixture's `searchSessions` took no parameters, so the
query never had to reach it -- the payload echoes the handler's own local
variable. Passing `""` instead of the caller's query left it green. It
now records what it was asked for.
The continuity disclosure test had two negative assertions over a fixture
whose only path lived in `store_paths`, so both tested the same omission
and returning `{sessions, messages}` passed. It now asserts the
diagnostic is still a diagnostic, and a second case gives a tool a
path-bearing `last_error` -- reaching both `sanitizeErrorMessage` calls on
that surface, which no test had executed.
The setup plan asserted only `kind` strings, so every path in the notice
the user confirms could have been wrong. Paths are now pinned, and a new
test compares the plan against what setup actually writes, since the two
were independent lists that nothing reconciled.
Also: `git_commit` now anchors on the scrubbed value, so a neutraliser is
distinguishable from a deletion; the truncation test pins the real 16,000
boundary and the reported character count instead of a 20,000 bound with
4KB of slack; the Antigravity warning test asserts the global file was
written and that a no-op rerun does not warn again; the opencode
role-less skip pins the surviving turn's index; and the hostile-payload
smoke test asserts nothing was recorded, not only that the exit code was
zero.
The landing page's central pitch and the published plugin's skill text both said the same false thing, in five places between them: install the plugin and retrieval works anywhere, no setup required. It does not. A project with no `.xtctx/config.yaml` gets `present: false`, and `createToolHandlers` then points all five tools at `notConfigured()`. README.md has always said so in its own table -- "Retrieval in an unconfigured project: no (offers setup)" -- so the two surfaces a new user actually meets were the two that disagreed with it. The skill matters most: it is instructions to an agent, and it told them a missing config affected only instruction blocks and the retrieval tools were "worth calling anyway". An agent following that calls five tools that will not answer and reports the project has no history. The hero terminal demo showed the same impossible sequence -- install plugin, cd into a project, get sessions back -- and now runs setup in between. Tests pin both surfaces against the claim returning. Found in the same pass, all verified against the code they contradict: - The Copy button on the three feature cards copied sample OUTPUT as if it were commands, so a user pasting it ran `updated`, `configured` and `├──` in their shell. Those blocks no longer offer copy. - The cache diagram named three tables that do not exist (`retrieval_windows`, `vectors`, `fts_index`); the real ones are `retrieval_units`, `retrieval_units_fts`, `retrieval_unit_vectors`. - The IDE mock labelled an `AGENTS.md` managed block `Tool: cursor`. AGENTS.md is Codex's file; Cursor's block goes to `.cursor/rules/`. - `v9.astro` carried `noindex,nofollow` and then `index,follow` in the same head, so the draft could be indexed as a near-duplicate of the homepage. - `Schematic.astro` was imported by nothing, and was wrong where it could be read at all: "Antigravity" twice, and tool names (`mcp_recent_sessions`) that have never existed. Deleted. Also fixed, from a line-by-line read of the CLI and index: - A literal search waited out the refresh budget for a scan it never reads. That route exists to answer while the index is still filling, and on the defaults it sat four seconds behind a scan before starting its own five. It now starts the scan and moves on; measured 2025ms before, under 1500ms after. - `readStdinWithin` removed its listeners but left stdin flowing, so the hang it exists to prevent still held the process open -- measured at 6012ms against a pipe held for six seconds, with the function itself resolving at 250ms. - `CLAUDE_CONFIG_DIR` was named in a comment as the containment root and read by nothing, so with it set the real transcript path was refused and the scraper searched a `~/.claude` holding nothing. - A relative `storePath` resolved against the process's working directory, so `xtctx status -p X` from elsewhere read a different store than the server does. - `scan --embed` exited 0 while reporting it had not finished. - The release job left the admin PAT in `.git/config` for the whole of `npm ci` and `verify:release`; it now reaches git only at the push. - The upstream watcher deduped on a title with no version in it, searched across all states, so one issue ever silenced it permanently -- and a failing `gh` left the count empty, which also read as "already reported".
Five audits each found a layer the previous one had not looked at: code defects, then tests that could not fail, then release gates that could not fail, then documentation and a landing page claiming things the code refuses. None of those passes was shallow. Each was simply pointed somewhere the last had not pointed, and nothing could say where that was, because the list of places lived in whoever was reviewing. So the list lives in the repository now. Every tracked file must belong to exactly one of seven review layers; a file matching none fails the check, and so does a file matching two, because then neither layer owns it. Adding a directory forces a decision about who reviews it. The layers carry a QUESTION, which is the load-bearing part. Asking "is this correct?" of a test suite finds nothing — a test that cannot fail is perfectly correct — and asking "can this fail?" of a landing page finds nothing either. The two passes that found the most were the two that asked something other than "is there a bug". It earned itself on the first run, which is why it is in `verify:release` rather than in a document: three tracked files belonged to no layer and no pass had opened any of them — `.github/copilot-instructions.md` (written by xtctx, read by an agent), `dependabot.yml`, and `upstream-versions.json`, the baseline the upstream watcher compares against. What it does not do is check that a review happened or was any good. It checks that the map has no blank regions, which is the specific failure that recurred five times. One thing worth recording: the first version of this script exited 0 while doing nothing, because it compared `import.meta.url` against a hand-built `file://` path that a Windows drive letter does not match. A gate that cannot fail, in the script whose `gates` layer asks exactly that question.
…stale copy The sixth pass ran the method the fifth one worked out: enumerate the layers from `check-review-coverage`, ask each its own question, and for the test layer break the source rather than read it. **Test layer: 20 of 21 mutations killed, no survivors.** Semantic and keyword weights, window stride, project-root case folding, the literal match cap, drift scrubbing, the resume offset, CRLF detection, the vector backlog, the stale-vector delete, both cosine thresholds, the tool registry, the FTS indexing split. The method that found six blind tests two passes ago now finds none, which is the first clean result any layer has returned. **Claims layer: a stale skill copy nothing could see.** Setup writes the built-in skill into `.xtctx/skills/` once, and everything afterwards compares the synced targets against THAT copy — so a project set up before the text changed keeps the old wording, every target agrees with it, and status reports `ok`. This repository was in exactly that state: its own skill still told agents "no xtctx setup is required" hours after the claim was corrected at the source. `inspectSkillStatus` now compares the built-in copy against `builtInHandoffSkill()` and status says `stale` with the command to refresh. Two numbers in the source were contradicted by the committed baseline. The model-choice table in `embeddings.ts` and the threshold sweep in `ranking.ts` both quote MiniLM at mrr 0.598; the baseline reads 0.333. Both figures are real — the sweep predates #318, which rebuilt the eval corpus because it "has been measuring a world that does not exist" — so they are annotated rather than deleted, with the note that 0.36 is inherited rather than re-derived against the current corpus. The release-incident count disagreed with itself: 54 versions in an hour in two documents, 112 in a third. Measured from the history: 76 release commits over four days, 49 on 2026-08-29, 12 in the busiest hour. All three corrected. Also corrected, each against the code that contradicts it: the confidence floor (0.4 -> 0.36); "never crosses the project boundary", which a committed `storePath` can cross and status reports rather than blocks; "the index refreshes on the first call that needs data" and "the model loads lazily", both untrue since the server scans and warms at startup; "indexing is on-demand", which is in the managed block every agent reads; Release Please in AGENTS.md, gone since #296; conventional commits "for release automation", which nothing reads; the codex oversized-line defect, fixed in c9beb48 and still described as open; integration tests "against a real index", which use a fixture; the bundled model, which downloads on first use; and `disconnect --all` deleting `.xtctx/skills`, which the README did not mention. `docs/embedding-performance.md` now opens by saying which of its figures are reproducible: one, the eval baseline. Everything else was measured in a session with scratch scripts that are not in the repository. That is stated because the alternative is how "18ms per embed" survived long enough to drive a model change that had to be reverted. Source maps referenced a `src/` the package did not ship, so every map in the published tarball pointed at nothing. `files` now includes it: 359 files, 1.93MB unpacked, 0.50MB packed. One gap that only a release closes, now written down: the plugin installs its skill from this repository while its server comes from `npx -y xtctx` on npm, so a plugin install describes behaviour the server it runs may not have yet. Pinning the manifest to a version npm does not serve would be worse.
The GPU result in docs/embedding-performance.md — ~6x, numerically identical vectors, already installed — has been blocked on one thing since it was measured: it came from one machine. DirectML is Windows-only, WebGPU is the portable candidate, and WebGPU had never run anywhere but this desk. WSL on this host has /dev/dxg but no /dev/dri, so it exercises the fallback rather than a Linux GPU. Implementing device selection on a sample of one is how a fallback path ships broken: it works for whoever wrote it and either throws or, worse, returns different vectors everywhere else. The probe asks two things per device, and only the first is about speed: does it initialise and embed at all, and are its vectors the same as the CPU's. The second decides whether a fallback can be silent — identical vectors mean a machine that falls back is running the same index as one that does not, with no re-embed and no threshold re-sweep implied. Each device runs in its own process, because two providers in one process fail intermittently with `bad allocation`. Locally (win32/x64, 48 segments of 1000 chars): cpu 22.0 ms/segment 10.6 cores baseline dml 5.1 ms/segment 0.8 cores worst cosine 1.000000 webgpu 6.4 ms/segment 0.9 cores worst cosine 0.999999 auto 23.1 ms/segment 10.1 cores worst cosine 0.999999 `auto` is what the product passes today, and it matches cpu — which confirms the GPU is currently unused rather than merely underused. The workflow runs on ubuntu, macos and windows. ubuntu is the important one: no GPU at all, so it is the machine that has to fall back cleanly, and the one this project has no evidence about. It triggers on changes to the probe because a workflow_dispatch workflow is only dispatchable from the default branch, so a manual-only probe could not be run on the PR that adds it.
…ystems The probe ran on macos-latest, windows-latest and ubuntu-latest. Two results matter more than the speed column. `auto` — what the product passes today — measures identical to `cpu` on all three runners, so the GPU is unused rather than underused. WebGPU on the GPU-less Windows runner is 1784.6 ms/segment against the CPU's 35.4: fifty times slower, and it does NOT fail. It found a software adapter, initialised cleanly and returned correct vectors. Linux with no GPU threw instead, which is what a fallback chain assumes. "Try WebGPU, fall back on error" is therefore not implementable — on the one configuration where it is catastrophic there is no error to fall back from, and the symptom is a vector backlog that never drains rather than anything that looks like a failure. macOS arm64 is the case the feature is for: WebGPU 21.4 ms/segment against 125.5 on CPU, 5.9x, at 0.3 cores. Vectors agree everywhere — worst pair 0.999999 across every device on every runner — so device choice does not enter vector identity, and a machine that ends up on CPU shares an index with one that does not. This does not rule out the feature; it rules out the cheap shape of it. A device has to be chosen on evidence: a short timed probe on first use, or an explicit opt-in.
…ding-providers.md Implements the design merged in #381. `EmbeddingProvider` was already an interface taken by injection, so this is a second implementation rather than a new seam. Opt-in per project and never inferred: an OPENAI_API_KEY that happens to be set in the environment does not opt a project into uploading transcript text. The key is read from the env var named by `embedding.apiKeyEnv` — a literal `apiKey` in `.xtctx/config.yaml` is rejected with an error saying why, because that file is committed and would publish the credential to everyone who clones the repository. Vector identity for a remote provider is `openai:{baseUrl}:{model}`, because two services both serving `text-embedding-3-small` produce vectors in different spaces while sharing a model name, and `retrieval_unit_vectors` is keyed on that name. Two deviations from the design, both deliberate: The local identity stays the bare HuggingFace id rather than becoming `local:Xenova/all-MiniLM-L6-v2`. Every remote identity is `openai:`-prefixed and cannot collide with a HuggingFace id, so the prefix buys nothing — while renaming it makes `dropVectorsFromOtherModels` discard every vector in every existing project on the first open after upgrading. That is tens of minutes of keyword-only search bought for no gain. Remote vectors are scaled to unit length on receipt, matching the local pipeline's `normalize: true`. Not for the cosine, which divides by both norms either way — for `poolVectors`, which mean-pools a window's segments and would otherwise weight them by magnitude instead of equally. OpenAI returns unit vectors; Ollama and LM Studio do not promise to. And one defect fixed in review: vectors are placed by the `index` each response element declares, not by position in `data`. A batch endpoint is not required to answer in request order, which is why that field exists. Reading positionally pairs every window with another window's vector — search keeps returning results, ranked against the wrong text, and nothing fails anywhere. Both are covered by tests confirmed to fail against the behaviour they replace.
…chine Measured end to end on this desktop, via `xtctx scan --embed` before and after: 551.9 ms/window on the CPU, 50.7 ms/window on DirectML. The remaining embed backlog for this project went from about 24 minutes to about 1.5. `xtctx calibrate` times the model on every execution provider the platform offers, each in its own process, and remembers the fastest in `~/.xtctx/device.json`. It is a command, not something that happens on its own: it costs real seconds and loads the model once per device, and nothing should spawn processes behind an agent's tool call. A machine that never runs it stays on the CPU — which is what every machine did before this existed, since passing no device at all measured identical to `cpu` on all three operating systems. It is a measurement rather than a fallback chain because a fallback chain cannot be written correctly here. On a Windows machine with no real GPU, WebGPU does not fail: it finds a software adapter, initialises cleanly, returns numerically correct vectors, and runs fifty times slower than the CPU. There is no error to fall back from and no symptom but a backlog that never drains. Device identity would not have solved it either, and this was checked rather than assumed: `onnxruntime-node` exposes no adapter information (`env.webgpu` carries only `powerPreference`, there is no `navigator.gpu`, and the verbose log names every kernel dispatch but never the device), Bun 1.4.2 has no `navigator.gpu` at all, and requiring Deno to run an npm package is not a trade worth making. More to the point, a vendor string answers the wrong question — a laptop iGPU against a 24-core CPU is a real GPU that still loses. Every ambiguous case resolves to the CPU: a GPU inside the noise, a GPU that would not initialise, and a run where the CPU baseline itself failed and there is therefore nothing to compare against. `xtctx status` reports the device, read off the provider rather than off the cache. "A verdict was written" and "the indexer is using it" are different facts, and reporting the first while meaning the second is how a wiring bug hides behind a green check. The worker is TypeScript, not plain JavaScript, because the build is `tsc` over the src tree and copies nothing: a `.js` worker would be absent from `dist` and every device would report "no result" for the published package while every test here passed. A test asserts the resolved path exists.
…long embed Two corrections, both prompted by asking why calibration was neither automatic nor simply "pick the fastest". A single timed pass is not a measurement. The warmup batch was two segments, which is enough for the CPU and nowhere near enough for a GPU — graph compilation and buffer allocation stayed inside the timed window. Same desktop, same 16 segments, only the warmup differing: 2 segments, 1 pass cpu 20.1 dml 6.0 webgpu 10.1 -> dml 2 segments, best of 3 cpu 21.0 dml 4.6 webgpu 3.4 -> webgpu full pass, best of 3 cpu 18.6 dml 3.1 webgpu 3.4 -> dml The short warmup was not measuring a slower device, it was measuring a device still starting up, and it ranked WebGPU as the slower GPU when it is within 7% of the faster one. The last row is three consecutive runs varying by 0.0 on DirectML and 0.1 on WebGPU. The 5.1ms/segment DirectML figure in the docs was warmup-contaminated the same way; the real number on this machine is ~3.1, so the GPU win is ~6x rather than ~4x. With a measurement that repeats, the margin over the CPU drops from 1.3x to 1.1x. It was 1.3 to cover the noise in a single pass; taking the fastest of three passes removes most of that, since other load on the machine only ever makes a pass slower. The rule is "pick the fastest device", and this is how close to 1.0 that rule can honestly get. And it now runs automatically inside `xtctx scan --embed`, which is the one path where the arithmetic is not close: that command is about to spend hours, and a minute spent finding out the GPU is ten times faster is repaid inside the first one. `--no-calibrate` skips it. Still not automatic anywhere else, and for a stated reason rather than caution. The MCP server answers a tool call inside a four-second budget and must not spawn three model-loading processes behind it. `setup` would become an 86MB model download nobody asked for.
… knows Asked why `xtctx scan --embed` has to be run by hand, the answer turned out to be that nothing else ever finishes the job. Searches vectorize about sixteen windows per call. By this repository's own measurement that leaves a 9,232-window project needing on the order of 570 searches to cover its history — so semantic search was keyword-only in practice on any real project, while still paying six seconds a call for the privilege. The only way out was knowing to run a command. Three comments said the session-start hook launched a detached scan that did this. It does not and never has: the hook does a deliberate no-scan read, and nothing under `src` spawned a process except the Antigravity client. The comments are corrected rather than left describing a mechanism that was never built. The MCP server already warms the index at session start. It now also drains the vector backlog there, bounded by this machine's own measured rate: the remainder has to fit in fifteen minutes or it is left alone. The bound is the point — unconditional background embedding is what would make a large project on a CPU unusable, since that path uses nine to eleven of twenty-four cores. One threshold covers both cases because the estimate is built from the measured rate: a large history is about eight minutes on a calibrated GPU and about eighty-five on the CPU. Device calibration is what made the affordable case the common one, which is why this lands now and not before. Verified against the built CLI by starting the server the way a host does: this project went from 1,340 windows outstanding to 3,964 of 3,964 vectorized, in the background, with no command run by hand. `estimateVectorBacklog` gains `etaMs` so the budget check and the status line share one estimate rather than reimplementing it.
docs/embedding-performance.md recorded bge-small beating MiniLM and said
plainly that the rows could not be reproduced: they came from scratch scripts
against a temporarily patched constant, and the eval harness "runs one fixed
provider and has no way to select a model or vary a threshold". The strongest
result in the file was the least checkable thing in it.
scripts/embedding-bakeoff.ts is that missing harness. It reuses the eval's own
corpus and scoring, and its MiniLM row at 0.15/0.36 reproduces
tests/eval/results/ranking-baseline.json exactly — which is what establishes
that the two measure the same thing rather than merely agreeing. On that
footing it reproduced every previously-unreproducible row in the doc.
Sweeping finer than the original then found a better pair than the 0.55/0.65
the doc recommended. The confidence floor, not the semantic floor, is what was
destroying vector mode: at 0.55/0.65 vector scores 0.325, at 0.55/0.64 it
scores 0.385, and the false-positive rate is zero at both.
At 0.62/0.64, sixty queries, both models at a false-positive rate of zero:
hybrid vector
mrr recall@5 top1 mrr recall@5 top1
MiniLM 0.333 0.533 0.183 0.246 0.350 0.183
bge-small 0.417 0.583 0.267 0.398 0.533 0.317
Vector mode gains 62% on mrr, 52% on recall and 73% on top-1. The committed
baseline regenerated to exactly those numbers from ranking.eval.test.ts, which
is a second harness agreeing rather than the same one repeated.
Thresholds do not transfer between models and that is the trap this closes: at
MiniLM's 0.15/0.36, bge-small scores a false-positive rate of 1.00 — every
deliberately unanswerable query, gibberish included, returns something. It is
not worse there; it places its cosine values higher, so a floor tuned to
MiniLM's distribution excludes nothing.
What changed to allow this is cost. bge-small indexes the corpus in 27.5s
against 15.6s, about 1.8x, and that ratio is exactly why it was rejected on
2026-09-20. Device calibration made embedding ~6x faster on a machine with a
GPU and the server now drains the backlog in the background, so 1.8x of a much
smaller number stopped being the deciding term.
The caveat the earlier measurement carried still stands and is not claimed
away: the corpus is synthetic and sixty queries. What is stronger is that the
margin holds across three modes and every threshold pair swept.
Both models are 384 dimensions, so the schema is unchanged.
dropVectorsFromOtherModels keys on the model name, so upgrading discards every
existing vector and re-embeds — paid once per project, and much cheaper than
it was this morning.
One test moved with the threshold: a score-reporting test used a stub cosine
of 0.6, which cleared MiniLM's 0.36 floor and not bge-small's 0.64, so it
failed for a reason unrelated to what it asserts. It now derives its value
from MIN_CONFIDENT_COSINE.
The MCP server drains the vector backlog at session start only while the estimate fits its budget. Above that nothing is working on it, and the status line said the same thing either way: a number of windows outstanding and a time estimate, identical in shape whether it finishes in four minutes or never. A large history on a CPU-only machine lands there the moment a model change invalidates every vector it had — which this session's switch to bge-small does to every existing project, since dropVectorsFromOtherModels deletes them all on the first open. Verified on this project's own index: 3,974 vectors to 0 in a single open, not gradually. So status now names the command when the backlog is past the budget, and the budget constant moved to utils/duration.ts beside the estimate it is compared against, because the server and status both need it and they must not drift. Covered at the layer it renders from: a test stubs the two rates measured on this project — 551.9ms/window on the CPU and 50.7ms on DirectML, the same 9,232 windows — and asserts the advisory appears for one and not the other. Confirmed to fail when the comparison is removed.
Hybrid search answers from keyword while the embedding model loads, because a cold cache takes minutes and nobody should wait for it holding a tool call open. That branch asked one question — `isReady()` — and a model still downloading and a model that will never download both answer false. So a failed load produced, on every call and for the life of the project: keyword results, the note "embedding model still loading, so this answer is keyword-only — ask again shortly for more", and an `xtctx status` with no semantic-unavailable line. The advice could never come true, and the failure was recorded nowhere, because the only code that writes `last_error:embeddings` is a catch that this early return skips. `warm()` swallowed the reason. The provider now keeps its last load error and the search path records it, so the failure reaches `xtctx status` and `xtctx_continuity_status` the same way a mid-search embedding exception already did. Loading is still retried every call and the error cleared on success, because the usual cause is a cold cache behind a flaky network that works on the next attempt — refusing to retry would turn a transient fault into a permanent one. Found by a delegated review of the failure-and-degradation paths, then confirmed against the code rather than taken on trust. Covered by a test using a provider that never loads, confirmed to fail when the recording is removed.
`tool_filter` reaches SQLite as `WHERE tool IN (...)`, so an id naming no tool
matched nothing and the answer was "No matching sessions found." — which an
agent reports to its user as "there is no Claude Code history in this
project". A wrong id and an empty index were indistinguishable.
The ids are not guessable either. They are `claude-code` and `antigravity`,
while the natural guesses from a tool's own name are `claude` and `gemini`,
and the MCP schema advertised only `items: { type: "string" }` with the
description "Optional tool ids to include".
Both halves are fixed: the schema now enumerates the ids from the tool
registry and lists them in the description, and `validatedFilter` rejects an
unrecognised one with an error naming the valid set. An agent that gets that
back can correct its own call; one that got an empty result could not tell a
typo from an empty history.
Found by a delegated review of the agent-facing surface, then confirmed
against the schema and the query path rather than taken on trust.
…setup A review flagged that `xtctx setup` followed by `xtctx disconnect --all` leaves xtctx wired into Antigravity for every project on the machine, and proposed making disconnect remove it by default. Implementing that broke three tests, and reading why they exist is the point of this commit. They record a live incident: disconnecting ONE project used to empty the machine-global Antigravity and Copilot CLI configs, taking xtctx away from every other project on the machine. Those files hold no per-project entry, so there is nothing project-scoped to remove from them — only the whole wiring for every project at once. The asymmetry is deliberate and the proposed fix would have reintroduced a known regression. What the review was right about is the documentation. The README said to pass `--global-mcp` "(as with `setup`)", which is false: `setup` writes Antigravity's config WITHOUT the flag, because Antigravity has no project-scoped MCP file. A reader who believed that sentence would assume `disconnect --all` had undone what `setup` did. So the flag is now documented as explicitly not symmetric, with the reason, and the README gains the full-removal sequence it never had — including that `.xtctx` survives on purpose and holds the indexed transcripts.
…ement of mine A delegated review of prose-against-code found that most of what is now wrong in the docs was made wrong by this session's own changes. The serious one is the privacy claim. README and PRODUCT.md said xtctx "never sends transcript content anywhere" and "nothing is ever sent off the machine". Those stopped being true when the opt-in OpenAI-compatible embedding endpoint shipped a few hours ago, and a claim about where a user's transcripts go is not one to leave stale. The review also caught that the qualification `docs/embedding-providers.md` specifies — "`xtctx status` reports the endpoint whenever one is configured" — is itself false, because no such line exists. Writing the qualification without the line would have replaced one false claim with another, so the line is added first: `xtctx status` now prints the endpoint in full whenever the vector identity is a remote one. "Am I uploading my transcripts, and where to" must never require opening a config file. Also corrected: ARCHITECTURE.md gave the semantic floors as 0.15/0.36 and PRODUCT.md named MiniLM, both superseded today. The thresholds now say they belong to the model and are swept per model, because that is the part that will go stale again otherwise. And a false statement of my own, from two commits ago. I wrote that no code had ever launched a detached scan from the session-start hook. It had: `launchDetachedScan` was added in #322 and removed by #323 the same day, 2026-09-02. The comment now says what is actually true — nothing launches one now, and the stale comments outlived the code by nineteen days — and records the correction rather than quietly rewriting it.
…pply Four findings from the UX pass, all confirmed by running the built CLI in a sandboxed first-run rather than by reading the source. `setup` ended on the last of eighteen file paths. The two things a user needs next are invisible from there: MCP clients read their config at launch, so an agent that was already open has no xtctx tools and reads as a broken install; and nothing is indexed until an agent calls a tool, so `xtctx status` — the obvious way to check setup worked — reports `Scan never` and `0 sessions`, which is the shape of a failure. Both are now stated, with the restart first. `setup` also wires all seven supported tools whether or not they are installed: measured in a clean environment with zero detected, eighteen files and eleven new top-level entries including `GEMINI.md` and `opencode.json`. The behaviour stays — detection reads a transcript store that does not exist until a tool has been used, so wiring only what is detected would silently skip whatever the user installs tomorrow, and a silently unwired tool is worse than an unwanted file because nothing reports it — but the reason is now printed, with `xtctx disconnect <tool>` for anything unwanted. `status` had no branch for an unreadable config. It fell through to "Ask a configured agent to call xtctx_recent_sessions", which cannot work because nothing is scanned at all while the config will not parse — and with an index left from before the file broke it reported "Handoff is wired" six lines under "UNREADABLE ... No transcripts are being read until this is fixed." The last line is the one people act on, so it now names the file to fix. And literal-search advice stopped following unrelated calls. `literalSearchStoppedEarly` and `literalUnreadableTools` were set by a literal pass and never cleared, while `getIndexProgress` — which every tool calls — reports them, so one truncated literal search attached "Narrow the query or raise `limit`" to every later `xtctx_recent_sessions` and `xtctx_session_detail` answer. Those calls carry no query. An agent either follows advice that cannot apply or learns to ignore the notes, which costs the ones that matter. Both behaviour fixes have tests confirmed to fail against the old behaviour.
…s shouting Two more from the UX pass. `.xtctx/config.yaml` records which transcript stores the user allowed to be read, so a file that will not parse yields zero scrapers rather than falling back to defaults. That is the right call and it had a consequence nobody had looked at: every MCP tool then answered "No matching sessions found.", and an agent reads that as "this project has no cross-tool history" and tells the user so. The real answer is that xtctx is reading nothing at all until one file is fixed. `xtctx status` has printed UNREADABLE for a while; agents never run the CLI, and MCP is the only surface they see. Every tool now answers with the file, the parse error, and the fact that this is an unread history rather than an absent one — returned rather than thrown, because it is a state the user can fix and an agent can pass that on. It also tells the agent not to edit the file unprompted, since rewriting it would mean widening its own read access. Kept separate from the not-configured notice on purpose: "nobody opted this directory in" and "somebody did and the file is broken" are opposite situations that had become the same empty answer. And `xtctx status` in a project that has not been set up stops printing the Skills and Managed-files sections. Every line of them reads `missing` there — about thirty, each with an absolute path — between the reader and the one sentence that matters. Someone running `status` to see what this thing does met a wall of faults describing the absence of a thing they had not asked for yet. Eighteen lines now, ending in `Run: xtctx setup`. Both covered by tests confirmed to fail against the old behaviour.
…th it A delegated review checked every factual claim in the docs against the code and found most of what is wrong was made wrong by this session. docs/embedding-performance.md now says up front that it is written in two voices: everything through "Model bake-off" was written on 2026-09-20 when none of it had changed the code and every conclusion was "blocked on", and three of those shipped the next day. Sections are annotated where they are now false rather than rewritten, because how a wrong conclusion was reached is the part worth keeping — but the file said the committed baseline was MiniLM's 0.333/0.533/0.183 when that file now holds bge-small's 0.417/0.583/0.267, said the bge rows could not be reproduced when a harness for them exists, said the `device` option "is simply not passed today" when it is, and still listed bge-small as blocked on a confirmation while it was the default. It also recommended 0.55/0.65 where 0.62/0.64 shipped, and now says why. docs/embedding-providers.md opened with "Nothing here is built yet" for a design that was built hours later. It now records what was built, the two deliberate deviations (the local identity keeps its bare HuggingFace id; remote vectors are normalized on receipt) and the two parts deliberately left out (the unswept-threshold warning, and any sweep against a remote provider), and says the code is current wherever the two disagree. "Semantic search is lazy" and "vector creation is lazy" described behaviour that changed when the server started draining the backlog in the background. ARCHITECTURE.md and PRODUCT.md were missing `scan` and `calibrate` from their CLI lists entirely, and the README documented neither `--embed` nor `calibrate` — a command that makes indexing roughly six times faster was discoverable only from a performance note. Also checked and NOT changed: `xtctx_session_detail` returning nothing for a session literal search found. The empty answer already carries the progress note naming the tools not yet read, so an agent is told to ask again rather than being left with a bare miss.
The first thing `xtctx --help` said was how to start the MCP server from a non-interactive stdio pair and which environment variable suppresses that — scripting advice, before any statement of what the tool is for or what to run first. It now says what xtctx does in two sentences and names the first command, with the stdio note kept below for the people who need it. `--hook` and `--tool` are hidden rather than removed. They are how a tool's hook re-enters this CLI, never something a person types, and listing them beside `--project` and `--version` made them read as options a newcomer is expected to understand. Every command description said what the code does instead of what the user gets. "Configure MCP, hooks, managed handoff instructions, and synced skills" requires knowing what all four are before it says anything; "Set this project up so agents can read each other's history here" says why you would run it. The same for the other four — a stranger could not previously tell from `--help` when they would ever want `scan` or `calibrate`, and `calibrate` makes indexing roughly six times faster.
`npm run verify:release` failed on five type errors that `npx tsc --noEmit` does not see. The two are not the same check: the default project compiles `src`, while `npm run typecheck` uses `tsconfig.test.json` and compiles the tests with it. I had been running the first and calling it a type check all session. Four are test stubs of `HandoffStatus` that predate `vector_device`, and one is a `deviceCandidates` call passing a bare `string` where the parameter is `NodeJS.Platform`. None affects shipped behaviour, which is exactly why they survived a suite that was green the whole time. `npm run verify:release` now exits 0.
…mark lied `MAX_BATCH_SIZE` moves from 32 to 16, worth about 13% of indexing time. The number is not the interesting part. I wrote `scripts/probe-batch-size.mjs` first. It said the GPU wanted the largest batch available — 4.0ms/segment at 128 against 5.3 at 32, a 1.37x win that reproduced cleanly across repeated runs — and that the CPU wanted the smallest. I implemented a device-dependent constant on the strength of it. Then I measured the real path: `xtctx scan --embed` over this project's own index from an empty vector table, 150 seconds per size, on DirectML. batch ms/window 8 123.4 16 107.8, 107.9 32 123.4 128 542.9 The benchmark's recommendation was a 5x regression. Its segments are all exactly 1000 characters. Real ones are not — windows hold a median of 4 segments and a 95th percentile of 17, of varying size — and a batch is padded to its longest member, so a wide batch of mixed lengths spends most of its work on padding. Uniform inputs hide the dominant cost of the real workload, which is why the benchmark was confidently, reproducibly wrong. That is the same failure as the "18ms per embed" figure that drove a model change and had to be reverted: a measurement whose inputs do not have the shape of the real ones is not weak evidence, it is evidence for a different question. So the device-dependence is gone too — it existed only because of the benchmark. One constant, 16, which is also what the independent CPU sweep in docs/embedding-performance.md found on real index segments, at the same margin. The script keeps a warning at the top saying what it is and is not evidence for. Verified end to end at 101.9 ms/window on the shipped build. `npm run verify:release` exits 0.
… I did not `scan --embed` began calibrating automatically earlier today. Calibration loads the embedding model in a child process per device, and it did not check `XTCTX_DISABLE_EMBEDDINGS=1` — a switch whose entire meaning is that the model is never loaded. So a scan told not to touch a model spent minutes doing exactly that. It timed out two tests in `tests/cli/scan-embed.test.ts` on ubuntu-latest at 60s. They passed here every time, because this machine has a cached device verdict: with one present calibration is skipped regardless, so the missing guard was never reached locally. A fresh machine — CI, or any user — has no verdict, which is precisely the case the guard is for. `xtctx calibrate` run directly now says why it did nothing rather than silently doing nothing, since being asked for explicitly is different from being triggered. The test is fixed as well as the code. It redirected nothing, so it would have kept passing here with the guard removed; it now points HOME and USERPROFILE at an empty temp directory for the duration, and fails in 13s when the guard is taken out. Passing for the wrong reason is the same defect the product bug had. Also corrected the file's header comment, which still described a session-start hook launching a detached scan. That code was removed on 2026-09-02.
…s a trap Asked why `xtctx calibrate` needs to exist at all, the honest answer was that it mostly should not — it was papering over a hole in the automatic path, and the hole was self-perpetuating. The MCP server drains the vector backlog in the background only when the estimate fits a fifteen-minute budget. That estimate is computed from the rate of whatever device is in use, which is the CPU until something calibrates. So on a large history: the CPU estimate exceeds the budget, the drain is skipped, and the drain was the only thing that would have made the project fast. Nothing else calibrates except `scan --embed`, which most people never run — so the machine stays on the CPU permanently. The user who loses most is the one with the largest history, which is exactly who the mechanism is for. The server now calibrates when it finds a backlog and no verdict, alongside a background drain that already takes minutes. Not in front of a tool call, which remains the line. It takes effect NEXT session, and deliberately so. The provider was constructed with the device known at startup and keeps using it, so this session's drain still runs at the old speed and is still judged by the old estimate. Relaxing the gate on the strength of a verdict the running provider is not using would start an hour of CPU work on the promise of a GPU that is not attached yet — the same shape of mistake as trusting a benchmark over the real path. It also honours `XTCTX_DISABLE_EMBEDDINGS=1`, which is the bug fixed in the previous commit, reintroduced in a second place an hour later and caught before it shipped. Calibration loads a model per device; that switch means no model is ever loaded. `xtctx calibrate` stays as an escape hatch — `--force` after a hardware change, and a way to see the measurements — rather than something a user is expected to discover.
Asked whether this could be automated entirely rather than left as an escape hatch, the answer was yes, and the thing blocking it was smaller than I said. I had claimed the verdict could only apply to the next session, because the provider is constructed before calibration runs. But the provider loads its model LAZILY — `extractor` and `loading` are both null until something embeds — so until then, pointing it at a different device costs nothing. `retargetDevice` does that, and refuses once the model is loaded or loading. The refusal is the careful half: swapping the device under a loaded pipeline means discarding it and paying the load again, possibly while a tool call waits on it, to speed up work already running. The caller is told it did not apply rather than left to assume it did. That closes the loop the previous commit left open. The server now measures, retargets the not-yet-loaded provider, and re-judges affordability using the per-segment rate calibration just measured on the chosen device — not the stale CPU rate, which would have skipped the drain on exactly the machines the measurement had made fast. That substitution is not a projection: calibration measures milliseconds per segment, and the estimate multiplies a per-segment rate by the segment backlog. Same arithmetic, fresher measurement of the same quantity. So `calibrate` is no longer something a user needs to discover. Both automatic paths cover it, and the command remains only for re-measuring after a hardware change and for showing the numbers. The README now leads with the automatic behaviour instead of the command, and says plainly that it is not a setup step. Still never behind an agent's tool call — that line was always about where the cost lands, not about whether it is automatic, and both automatic callers are background work that was already going to take minutes.
…branch_filter
release.yml was invalid on this branch. A comment inside a `run:` block
mentioned the Actions expression syntax as an empty `${{ }}` pair, and Actions
evaluates expressions in `run:` strings, shell comments included: "An
expression was expected". Merging would have left the only workflow that cuts
releases impossible to dispatch. It was also the cause of the zero-job failed
`release` run on every push to this branch, which I twice dismissed as noise.
CI stayed green because CI never reads release.yml.
`scripts/check-workflows.mjs` now checks every workflow parses and that every
expression opened in a string value is closed and non-empty — YAML comments
excluded, since Actions never evaluates those, which is exactly the
distinction that was missed. Confirmed to flag the broken file at f28ec43.
It runs in CI's checks job alongside `check:review-coverage`, which had also
only ever run inside `verify:release`.
`branch_filter` rejected every real branch name. The unknown-tool-id check
added two days ago went into `validatedFilter`, which both filters share, so
`branch_filter: ["main"]` failed with "unknown tool id \"main\"". The only
branch test passed a bare string, which the array check rejects for its own
reason, so the suite stayed green. The id check is now `validatedToolFilter`,
called only for `tool_filter`, and a test passes a real branch array through
the manifest handler — confirmed to fail when branch_filter is routed back
through the id check.
Also drops `prepack` (npm runs `prepare` before pack and publish anyway) and
the explicit build before `npm pack --dry-run`, which together built the
package three times per `verify:release`.
…d make it safe The audit found that same-session calibration — which I described as done — never applied in the MCP server. Every scan ends by calling `warm()`, so by the time calibration finished the model was already loading and `retargetDevice` refused. The verdict only ever reached the NEXT session, while the README and two docstrings said otherwise. The unit test called the provider in isolation and never exercised that ordering. The provider now defers its model load until the device is known (`deferDeviceUntil`), and the server starts calibration before it scans. A deferral cannot lose the race: whoever asks for the model first — the warm scan, a hybrid search's `warm()`, an explicit vector search — gets one load, on the chosen device. Hybrid search is unaffected in practice because it already answers from keyword while the model is not ready; only an explicit vector search waits, and only on a machine's first run. Because calibration now finishes before the scan, the scan's own vectorizing pass records its per-segment rate on the chosen device — a real measurement on the real path — and the drain budget is judged by that. The previous version substituted calibration's own rate, the fastest of three passes over uniform 1000-character segments, which is systematically quicker than real windows. Made safe for several agents starting at once on a fresh machine: - a machine-wide lock (`~/.xtctx/device.json.lock`, abandoned after ten minutes), so one server calibrates and the rest use the default for that session instead of all loading the model together and contaminating each other's timings; - the verdict written to a temp file and renamed, so a concurrent reader sees the old file or the new one, never half of one; - calibration workers killed when the server exits. It leaves by `process.exit` two seconds after its client disconnects; on Windows a child outlives its parent, and the timeout that would have stopped it lived in the parent. The background path moved out of `cli/index.ts` into `runtime/background.ts`, because the CLI runs `main()` on import and nothing in it could be tested — which is why three mutations to this logic survived the whole suite. It now has tests for the ordering, the embeddings-disabled guard, the budget in both directions, the lock, and failure reporting. A failed background drain is now logged to stderr and recorded in `embedding_error` instead of vanishing into a bare `catch`. Two more test gaps from the same sweep, one of which was hiding a bug: - `formatDuration` rounded minutes and seconds separately and printed "1m 60s" for 119.6 seconds. Swapping its rounding for flooring changed nothing any test could see. Now rounded once and split. - Hybrid mode's keyword half could be zeroed without failing anything; a ranking contract now needs it. The batch width is pinned by behaviour (sixteen per forward pass), through a pipeline-factory seam that lets tests see what the model is asked to do without downloading it. The remaining sweep survivor, the tool-level `limit` cap, is an equivalent mutant: the index caps every path at 100 independently, so no behaviour depends on it.
The rest of the 2026-09-23 audit. Claims made false this week, all corrected: - docs/embedding-performance.md said the MCP server "reads the verdict and nothing more"; it now calibrates in the background, and the paragraph says how and records why an earlier version did not apply. - The landing page's FAQ said indexing happened "on demand" and vectors were "created lazily", and called xtctx local-only without the opt-in endpoint. Its `xtctx status` mock used an output format the command has never printed; it now shows the real labels. - This repository's own four managed blocks still said indexing was on-demand. The generator had been updated; the committed copies had not. - ux-walkthrough.md said the tools work in a project without setup. They name `xtctx setup` and do nothing else. - ranking.ts presented MiniLM's 0.36 sweep as the live threshold. - CHANGELOG now says, inside the 0.21.8 section so the release workflow's insertion point does not move it, that 0.20.0–0.21.8 were never on npm. Lows: - The OpenAI provider now releases the body of a response it gives up on, not only one it retries; and a 429 waits for Retry-After (clamped to 5s, default 1s) instead of retrying instantly into a second 429. - The segment backlog is scoped to the project, like the window count printed beside it; in a shared index the duration described more work than the count. - The unconfigured and unreadable-config notices honour `format: json` — and default to it for `xtctx_handoff_manifest`, whose documented contract is JSON. An orchestrator used to get prose it could not parse. - `scripts/*.ts` is now typechecked and linted; neither covered it. The device probe measured MiniLM for two days after the default changed. - The drift canary FAILS when it cannot run for lack of a key. It passed with a warning so as not to turn a nightly job red, and there is no nightly job: a dispatched check that checked nothing should not look like a pass. - The `v9` and `concepts` design drafts no longer deploy. The site builds one page instead of eight. And a test for the calibration lock added in the previous commit, which had none: a fresh lock refuses, an abandoned one is taken over. Verified with a mutation harness that counts a run only if vitest printed its own summary: all nine mutations — the six survivors from the audit and one per new fix — are killed, and the working tree was byte-identical before and after the sweep. `npm run verify:release` exits 0 and builds the package once rather than three times.
…awn no workers CI on #383 failed two ways. On ubuntu the first explicit vector search waited for a device measurement it did not need and hit the 60s limit; embedBatch now abandons the deferral, so only the background warm-up waits. On windows the calibration-lock tests timed real devices through tsx, and killing tsx orphans its ONNX child, starving the 2-core runner until time-budgeted scans in other tests came back empty. The lock tests now pass devices: [].
On Windows the killed server's handles on the index outlive its exit by a moment, so the first rmdir failed with EBUSY on GitHub's windows runner and locally. The server also ran with the real home, so it indexed and calibrated there; it now uses the demo's temp home.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The four code defects come from the same audit as #380 and share its shape: a read degrades correctly, and the degraded value is then used as the base for a write or a filter.
disconnectsetupexplicitly refuses to touchdisconnect.xtctx/config.yamlfrom{}, losingstorePathoverridesThe ones worth explaining
Antigravity.
steps.tsreturnsnew Date(0)when neither the step metadata nor the summary carries a parseable time.fullSyncreads withsince = new Date(0), and the filter wastimestamp <= sincewith no guard — so0 <= 0dropped exactly those steps. Every other scraper guards this; claude-code even spells out why: "A zerosincemeans full sync: emit even epoch-sentinel timestamps."TOML comments. The write path checks
tomlHasCommentsand declines, telling the user to add the entry by hand and keep their comments. The remove path re-serialised and deleted them all — which made following that advice pointless.config.yaml. Degrading the read is right and deliberate: disconnect is what someone reaches for when things are already wrong. Using
{}as the base for the write is not. Everything else still disconnects; only this rewrite is skipped, with a warning naming the file.Cursor. Two skips did not advance
messageIndexwhile the two below them did, andscan.tshashes that index into the row id.The fifth is not a fix
The smoke suite failed four of five cases once under load and has passed every attempt since, including under deliberate contention. I could not reproduce it, so I have not claimed to fix it.
What made it unfixable is worth fixing on its own: nothing drained the spawned server's stderr, so the only evidence was
initialize did not answer within 60s— a symptom, with the cause piped into a buffer nobody read. That text is now attached to the failure, and an exited child fails its pending call immediately rather than waiting out the timeout. The next occurrence will say why.Verification
Each test fails against the previous behaviour — confirmed by stashing each fix and re-running:
expected [] to have a length of 1expected [ 0, 1 ] to deeply equal [ 0, 2 ]Full suite 695 passed across 95 files, drift 28, integration 5, smoke 20. Lint and typecheck clean.