Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
35 commits
Select commit Hold shift + click to select a range
bc0f391
fix: close the five findings left open by the audit
fstubner Sep 21, 2026
172e640
fix: close what the second audit found in the index and MCP layers
fstubner Sep 21, 2026
9a3bf33
fix(gates): make three release gates able to fail, and two tests able…
fstubner Sep 21, 2026
729d473
fix(scrapers): report where drift was seen, and make six tests able t…
fstubner Sep 21, 2026
6090fef
test: close the nine remaining findings from the third and fourth audits
fstubner Sep 21, 2026
a2fa63a
fix: stop the product claiming retrieval works without setup
fstubner Sep 21, 2026
f613ced
test(review): make "did we look everywhere" a command, not a memory
fstubner Sep 21, 2026
08678ba
fix: correct what the docs and the skill tell people, and detect the …
fstubner Sep 21, 2026
ae9ee93
feat(scripts): probe embedding devices on three operating systems
fstubner Sep 21, 2026
1436bc8
docs(perf): the GPU fallback chain does not survive three operating s…
fstubner Sep 21, 2026
b4bf8a8
feat(embeddings): optional OpenAI-compatible endpoint, per docs/embed…
fstubner Sep 21, 2026
1f56336
feat(embeddings): pick the execution provider by timing it on this ma…
fstubner Sep 21, 2026
03983ae
feat(calibrate): measure properly, and run it automatically before a …
fstubner Sep 21, 2026
4ddc4e5
fix: finish embedding on its own, instead of needing a command nobody…
fstubner Sep 21, 2026
d82239e
feat(embeddings): bge-small at 0.62/0.64, and a bake-off that reproduces
fstubner Sep 21, 2026
08b6862
fix(status): say when a backlog is too large to drain on its own
fstubner Sep 21, 2026
a8f604f
fix(search): stop a model that cannot load reading as one still loading
fstubner Sep 21, 2026
6dfc816
fix(mcp): reject an unknown tool_filter id instead of matching nothing
fstubner Sep 21, 2026
7e063fb
docs(readme): say that disconnect --global-mcp is not symmetric with …
fstubner Sep 21, 2026
a4bbada
fix(claims): the local-only promise, the thresholds, and a false stat…
fstubner Sep 21, 2026
3853585
fix(ux): tell people what happens next, and stop advice that cannot a…
fstubner Sep 21, 2026
ab22dc2
fix(ux): a broken config reaches agents, and an unset-up project stop…
fstubner Sep 21, 2026
92ad5e1
docs: catch the prose up with a day that changed four things undernea…
fstubner Sep 21, 2026
d9d21f3
fix(cli): make --help answer a stranger's questions, not a maintainer's
fstubner Sep 21, 2026
82b3d27
fix(tests): type-check the tests, which is where five errors were hiding
fstubner Sep 21, 2026
508adbc
perf(embeddings): batch 16, measured on the real path after the bench…
fstubner Sep 21, 2026
1117ea8
fix(calibrate): respect XTCTX_DISABLE_EMBEDDINGS, which CI caught and…
fstubner Sep 22, 2026
fd16a0b
fix: calibrate from the server too, because the uncalibrated state wa…
fstubner Sep 22, 2026
f28ec43
feat(calibrate): automate it properly, so nobody has to know it exists
fstubner Sep 22, 2026
eba3b93
fix: two merge blockers from the audit — an invalid release.yml, and …
fstubner Sep 23, 2026
b219f76
fix(calibrate): apply the verdict in the session that measured it, an…
fstubner Sep 23, 2026
ded22aa
fix: the remaining audit findings — claims, lows, and one tested lock
fstubner Sep 23, 2026
3318089
Merge remote-tracking branch 'origin/main' into fix/audit-findings
fstubner Sep 23, 2026
5857093
fix: an urgent embed stops waiting for calibration, and lock tests sp…
fstubner Sep 24, 2026
71feb0a
fix(demo): retry the temp-dir removal, and give the server a temp home
fstubner Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/copilot-instructions.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ Do not rely on this block for a generated summary; raw local transcripts are aut
- Transport: stdio

## Notes
- Indexing is on-demand from MCP recent, detail, and search calls.
- Indexing runs when the MCP server starts and on recent, detail, and search calls; `xtctx scan` does it on demand.
- There is no xtctx daemon, API server, dashboard, durable memory, or generated brief.
- Content outside this managed block is preserved.
<!-- xtctx:end -->
9 changes: 9 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -139,6 +139,15 @@ jobs:
npm run security:checklist
npm run audit:production

# Both used to run only inside `verify:release`, which only the release
# workflow invokes — so a broken workflow file or an unmapped file was
# found at release time, if at all. An invalid release.yml sat on a
# branch for two days with this job green.
- name: Workflow files and review coverage
run: |
npm run check:workflows
npm run check:review-coverage

- name: Landing dependency audit
# The landing site has only devDependencies (Astro is a build-time
# dependency of a static site), so --omit=dev would audit nothing.
Expand Down
17 changes: 12 additions & 5 deletions .github/workflows/drift-canary.yml
Original file line number Diff line number Diff line change
Expand Up @@ -99,14 +99,21 @@ jobs:
# the run, and `upstream-watch` is what files issues now.

# 78 means the canary could not run at all — no API key — which is a
# configuration gap, not upstream drift. Surfaced as a warning so it is
# visible in the run summary without turning the nightly job red and
# without filing an issue that says the scraper is broken.
- name: Note that the canary was skipped
# configuration gap, not upstream drift. It FAILS the job, with a message
# saying which of the two it is.
#
# It used to pass with a warning, on the reasoning that a missing key
# should not turn the nightly job red. There is no nightly job any more:
# this runs when a person dispatches it, to find out whether a tool's
# format drifted, and a green result that checked nothing answers that
# question wrongly. A check that could not run must not look like one
# that passed.
- name: Fail because the canary could not run
if: steps.canary.outputs.exit_code == '78'
run: |
echo "::warning title=drift canary skipped::${{ matrix.tool }} did not run — its API key is not configured, so this run produced no drift signal for it."
echo "::error title=drift canary did not run::${{ matrix.tool }} was not checked — its API key is not configured, so this run says nothing about whether its format drifted."
cat canary-stderr.log || true
exit 1

- name: Fail job on canary failure
# exit_code is empty when the canary step was skipped by a scoped dispatch.
Expand Down
73 changes: 73 additions & 0 deletions .github/workflows/embedding-device-probe.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,73 @@
# Which ONNX execution providers work, on operating systems nobody here owns.
#
# `docs/embedding-performance.md` has DirectML at ~6x the CPU path with
# numerically identical vectors, and stops there, because DirectML is
# Windows-only and WebGPU — the portable candidate — had only ever been run on
# the one Windows desktop this project is developed on. WSL on that host has
# `/dev/dxg` but no `/dev/dri`, so it exercises the fallback rather than a
# Linux GPU.
#
# This workflow is the second and third machine. `macos-latest` is real Apple
# Silicon with a Metal-backed WebGPU, and `ubuntu-latest` has no GPU at all,
# which makes it the more important of the two: it is the machine that has to
# fall back cleanly, and the one this project currently has no evidence about.
#
# Not on every push — it downloads a model and runs four processes per OS for
# a report a person reads, and nothing merges on its result. It runs when the
# probe itself changes, which is when its answer can have moved, and otherwise
# on request.
#
# The path filter is also what makes the report reachable before this file is
# on `main`: a `workflow_dispatch` workflow is only dispatchable from the
# default branch, so a probe that was manual-only could not be run on the pull
# request that introduces it.
name: embedding-device-probe

on:
workflow_dispatch:
pull_request:
paths:
- scripts/probe-embedding-device.mjs
- .github/workflows/embedding-device-probe.yml

permissions:
contents: read

jobs:
probe:
name: probe (${{ matrix.os }})
runs-on: ${{ matrix.os }}
# A runner where every device fails is a finding, not a broken workflow —
# the other two still have to report.
continue-on-error: true
strategy:
fail-fast: false
matrix:
os: [ubuntu-latest, macos-latest, windows-latest]

steps:
- name: Checkout
uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1

- name: Setup Node
uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
with:
node-version: 24
cache: npm
cache-dependency-path: package-lock.json

- name: Install dependencies
run: npm ci

# Same cache and the same reason as ci.yml: four processes per OS each
# want the 86MB MiniLM, and HuggingFace answers 429 to a burst of them.
- name: Cache the embedding model
uses: actions/cache@0057852bfaa89a56745cba8c7296529d2fc39830 # v4.3.0
with:
path: node_modules/@huggingface/transformers/.cache
key: hf-model-${{ runner.os }}-${{ hashFiles('src/handoff/embeddings.ts') }}

# One run, both formats. Probing twice would re-embed 48 segments per
# device for a second copy of the same numbers.
- name: Probe
run: node scripts/probe-embedding-device.mjs --json-also
18 changes: 16 additions & 2 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -95,6 +95,12 @@ jobs:
# one manually dispatched workflow, while `GITHUB_TOKEN` is present
# in every run of every workflow in the repository.
token: ${{ secrets.RELEASE_PLEASE_TOKEN }}
# Not left in `.git/config`. With the default, the admin PAT sits on
# disk while `npm ci`, `npm --prefix landing ci` and the whole
# `verify:release` suite run — third-party install scripts and every
# test in the repository, any one of which could read it. Only the
# push at the end needs it, and it is supplied there.
persist-credentials: false

- name: Setup Node
uses: actions/setup-node@820762786026740c76f36085b0efc47a31fe5020 # v7.0.0
Expand Down Expand Up @@ -185,6 +191,7 @@ jobs:
env:
VERSION: ${{ steps.bump.outputs.version }}
TAG: ${{ steps.bump.outputs.tag }}
RELEASE_TOKEN: ${{ secrets.RELEASE_PLEASE_TOKEN }}
run: |
set -euo pipefail
git config user.name "github-actions[bot]"
Expand All @@ -195,8 +202,15 @@ jobs:
git add package.json package-lock.json CHANGELOG.md plugin .claude-plugin landing/src/data/site.ts
git commit -m "chore(release): ${VERSION}"
git tag "$TAG"
git push origin "HEAD:${GITHUB_REF_NAME}"
git push origin "$TAG"
# The token reaches git here and nowhere else; see
# `persist-credentials: false` on the checkout. Through the
# environment rather than an Actions expression, so it is never
# substituted into the script text. (Do not write the expression
# syntax in this comment: Actions evaluates it inside `run:` blocks
# even in shell comments, and an empty one made this whole file
# invalid — see scripts/check-workflows.mjs.)
git push "https://x-access-token:${RELEASE_TOKEN}@github.com/${GITHUB_REPOSITORY}.git" "HEAD:${GITHUB_REF_NAME}"
git push "https://x-access-token:${RELEASE_TOKEN}@github.com/${GITHUB_REPOSITORY}.git" "$TAG"

- name: Create the GitHub release
env:
Expand Down
36 changes: 29 additions & 7 deletions .github/workflows/upstream-watch.yml
Original file line number Diff line number Diff line change
Expand Up @@ -60,25 +60,47 @@ jobs:
echo "**Run:** ${{ github.server_url }}/${{ github.repository }}/actions/runs/${{ github.run_id }}"
} > issue-body.md

# Dedupe on the exact version pair in the title, and across *all* states:
# a closed issue means the human already looked at that release, so
# reopening it nightly would recreate the noise this replaces.
# The version pair has to be IN the title for the dedupe below to mean
# anything. It was a fixed string with no versions in it, searched across
# all states — so one issue, ever, silenced this watcher permanently:
# every later upstream release matched that title and filed nothing.
- name: Build the title for this release set
id: title
if: steps.check.outputs.exit_code == '10'
run: |
set -euo pipefail
# Each MOVED line reads "MOVED <pkg>: <from> -> <to>"; the joined
# "to" versions identify this particular set of releases.
versions=$(grep '^MOVED' upstream-report.txt | sed 's/.*-> //' | paste -sd, -)
echo "value=upstream: tracked coding tools released (${versions})" >> "$GITHUB_OUTPUT"

# Across *all* states: a closed issue means the human already looked at
# THAT release set, so reopening it nightly would recreate the noise this
# replaces.
- name: Check whether this release was already reported
id: existing
if: steps.check.outputs.exit_code == '10'
env:
GH_TOKEN: ${{ secrets.GITHUB_TOKEN }}
TITLE: ${{ steps.title.outputs.value }}
run: |
title="upstream: tracked coding tool released — check transcript formats"
found=$(gh issue list --state all --search "\"$title\" in:title" --json number,title,state \
| jq -r --arg t "$title" '[.[] | select(.title == $t)] | length')
# `pipefail`, because a failing `gh` left `count` empty — which is
# not '0', which reads as "already reported". An API error silenced
# the watcher instead of failing the run.
set -euo pipefail
found=$(gh issue list --state all --search "\"$TITLE\" in:title" --json number,title,state \
| jq -r --arg t "$TITLE" '[.[] | select(.title == $t)] | length')
if [ -z "$found" ]; then
echo "could not determine whether this release set was already reported" >&2
exit 1
fi
echo "count=${found}" >> "$GITHUB_OUTPUT"

- name: Open an issue
if: steps.check.outputs.exit_code == '10' && steps.existing.outputs.count == '0'
uses: peter-evans/create-issue-from-file@fca9117c27cdc29c6c4db3b86c48e4115a786710 # v6.0.0
with:
title: "upstream: tracked coding tool released — check transcript formats"
title: ${{ steps.title.outputs.value }}
content-filepath: issue-body.md
labels: |
drift
Expand Down
8 changes: 5 additions & 3 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -43,8 +43,10 @@ npx -y xtctx status

## Release Notes

- Release Please manages versions and changelog.
- Release PRs are auto-merge enabled after required checks pass.
- Releases are cut by one manually dispatched workflow; nothing is released by
merging. See RELEASE.md.
- Commit prefixes are for readers, not for tooling: release notes come from
GitHub's own generator over the commit range.

<!-- xtctx:begin -->
Generated by xtctx setup. Do not edit inside this block.
Expand Down Expand Up @@ -75,7 +77,7 @@ Do not rely on this block for a generated summary; raw local transcripts are aut
- Transport: stdio

## Notes
- Indexing is on-demand from MCP recent, detail, and search calls.
- Indexing runs when the MCP server starts and on recent, detail, and search calls; `xtctx scan` does it on demand.
- There is no xtctx daemon, API server, dashboard, durable memory, or generated brief.
- Content outside this managed block is preserved.
<!-- xtctx:end -->
26 changes: 15 additions & 11 deletions ARCHITECTURE.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,8 @@ boundaries between them, and what each part is allowed to trust.

## Parts

- **CLI** (`src/cli/`) — `setup`, `status`, `disconnect`, and the internal
`--hook session-start` entry point. Bare `xtctx` on a non-TTY stdio pair
- **CLI** (`src/cli/`) — `setup`, `status`, `scan`, `calibrate`,
`disconnect`, and the internal `--hook session-start` entry point. Bare `xtctx` on a non-TTY stdio pair
starts the MCP server.
- **MCP server** (`src/mcp/`) — stdio JSON-RPC server exposing exactly five
read-only tools. Spawned by coding agents via `npx -y xtctx`.
Expand All @@ -19,7 +19,7 @@ boundaries between them, and what each part is allowed to trust.
- **Handoff index** (`src/handoff/`) — per-project SQLite database
(`.xtctx/state/xtctx.db`, WAL, schema-versioned) holding sessions,
messages, retrieval windows, FTS index, and embedding vectors. Refreshed
on demand from the scrapers; fully derived, rebuilt from scratch on
on demand from the scrapers and at MCP server start; fully derived, rebuilt from scratch on
corruption or schema mismatch.
- **Drift log** (`src/scrapers/drift-log.ts`) — per-tool record of the
places another tool's transcripts did not match what the scraper expected,
Expand Down Expand Up @@ -57,9 +57,10 @@ there is no daemon to leave behind.

**Serving a call.** A coding agent spawns `npx -y xtctx` over stdio, gets the
five read-only tools, and the process exits when the agent is done with it.
On the first call that needs data, the index refreshes: every scraper reads
its own tool's store, yields only chunks attributable to this project, and the
results land in `.xtctx/state/xtctx.db`.
The server starts a scan as it starts, and refreshes again on a call whose
indexed view has gone stale: every scraper reads its own tool's store, yields
only chunks attributable to this project, and the results land in
`.xtctx/state/xtctx.db`.

**How a conversation becomes searchable.** Messages are grouped into
overlapping retrieval windows — eight messages, stride four — so a hit carries
Expand All @@ -71,7 +72,9 @@ twice: into FTS5 for keyword search, and as one embedding vector.
comes from bm25 ordering but is rescored as a linear decay, because bm25
favours short documents and a one-line mention was outranking the paragraph
that decided something. Semantic matches are gated twice: a per-window floor
(0.15) and a per-query confidence floor (0.4). When nothing clears the second
(0.62) and a per-query confidence floor (0.64). Those numbers belong to the
model — they are swept per model, not carried between them, and MiniLM's were
0.15 and 0.36. When nothing clears the second
one, semantic results are dropped wholesale and only keyword hits remain —
whether a query found anything is a property of the query, not of each window,
and no answer beats a confident wrong one.
Expand All @@ -90,10 +93,11 @@ queue.

**Bounded, so a tool call always returns.** Scanning gets four seconds,
vectorizing six, and an indexed view is treated as current for thirty. Work
left over resumes on the next call. The embedding model loads lazily and only
for semantic search; `hybrid` deliberately answers from keyword while it is
still loading, so the first call after a cold start is fast rather than
blocked.
left over resumes on the next call. A scan also warms the embedding model and
builds vectors under the same cap, because a process spawned per agent session
would otherwise never vectorize anything; `hybrid` deliberately answers from
keyword while the model is still loading, so the first call after a cold start
is fast rather than blocked.

**What comes back is raw.** Sessions, message text, and pointers — never a
generated summary. A recap is the lossy artefact this exists to replace, and
Expand Down
9 changes: 7 additions & 2 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,11 @@ Entries are written by the `release` workflow when a release is cut by hand.

## [0.21.8](https://github.com/fstubner/xtctx/compare/xtctx-v0.21.7...xtctx-v0.21.8) (2026-08-31)

> **Not on npm.** 0.20.0 through 0.21.8 were tagged and given GitHub
> Releases, but none of them was ever published to npm — the last version
> published before them is 0.19.0. Everything listed from here down to 0.20.0
> reaches `npx -y xtctx` users only in the next version that is published.


### Bug Fixes

Expand Down Expand Up @@ -133,13 +138,13 @@ Entries are written by the `release` workflow when a release is cut by hand.

### Miscellaneous

* An automated release pipeline cut 112 versions between 0.19.1 and 0.74.0 in
* An automated release pipeline cut 76 release commits between 0.18.7 and 0.74.0 in
a few hours on 2026-08-29, none of which anyone asked for and none of which
were published to npm. The version was reset to 0.20.0 the same day and the
pipeline was replaced by a manually-triggered release; see the comments in
`.github/workflows/release.yml`. The individual "release xtctx X.Y.Z"
entries those runs generated are summarised here rather than listed, because
a reader scanning this file for what shipped was being shown 112 versions
a reader scanning this file for what shipped was being shown versions
that never existed.


Expand Down
2 changes: 1 addition & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ Do not rely on this block for a generated summary; raw local transcripts are aut
- Transport: stdio

## Notes
- Indexing is on-demand from MCP recent, detail, and search calls.
- Indexing runs when the MCP server starts and on recent, detail, and search calls; `xtctx scan` does it on demand.
- There is no xtctx daemon, API server, dashboard, durable memory, or generated brief.
- Content outside this managed block is preserved.
<!-- xtctx:end -->
4 changes: 3 additions & 1 deletion CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -42,7 +42,9 @@ Or run everything with:
1. Keep changes focused and atomic.
2. Add tests for behavior changes.
3. Update docs when CLI or MCP behavior changes.
4. Use conventional commits for release automation:
4. Use conventional commits. Nothing reads the prefix — release notes come
from GitHub's own generator over the commit range — but a reader scanning
the log does:
- `feat: ...`
- `fix: ...`
- `docs: ...`
Expand Down
2 changes: 1 addition & 1 deletion GEMINI.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ Do not rely on this block for a generated summary; raw local transcripts are aut
- Transport: stdio

## Notes
- Indexing is on-demand from MCP recent, detail, and search calls.
- Indexing runs when the MCP server starts and on recent, detail, and search calls; `xtctx scan` does it on demand.
- There is no xtctx daemon, API server, dashboard, durable memory, or generated brief.
- Content outside this managed block is preserved.
<!-- xtctx:end -->
Loading
Loading