Measurements taken on 2026-09-20 and 2026-09-21 while looking for a way to make indexing faster. They exist so the next person asking "can we speed this up" starts from evidence rather than from the same four guesses.
Three of the four guesses were wrong, which is the main reason this file is worth keeping.
Read this first, because the file is written in two voices. Everything up to and including "Model bake-off" was written on 2026-09-20, when none of it had changed the code and the conclusions were all "blocked on". On 2026-09-21 three of them shipped: device calibration, the model change to bge-small, and a bake-off harness that makes the model rows reproducible. Sections written before that day have been annotated where they are now false rather than rewritten, because how a wrong conclusion was reached is the part worth keeping. Where an annotation and a table disagree, the annotation is current.
tests/eval/results/ranking-baseline.json is regenerated by
npm run test:eval and now reads bge-small's numbers — hybrid
0.417 / 0.583 / 0.267. The MiniLM figures quoted throughout this file
(0.333 / 0.533 / 0.183) were that file's contents until 2026-09-21 and are
kept as the comparison the model change was made against.
Two more are now reproducible. The three-OS device table, which
scripts/probe-embedding-device.mjs regenerates and
.github/workflows/embedding-device-probe.yml runs on the same three runners.
And every model row below: scripts/embedding-bakeoff.ts takes --model and
--sweep, and its MiniLM row at 0.15/0.36 reproduces the old committed
baseline exactly, which is what establishes it measures the same thing.
Everything else — the single-machine device table, the batch sweep, the dtype
comparison, and the segment-duplication count — was measured in one session on
one machine with scratch scripts that are not in this repository. The method
is described precisely enough to redo, and the ratios are the durable part, but
nothing here re-runs them and nobody should treat the absolute numbers as
checkable. (The bge and gte rows said the same until 2026-09-21, when
scripts/embedding-bakeoff.ts was written and reproduced them.)
A number that reads as evidence while being untraceable is how the "18ms per embed" figure below survived long enough to drive a model change that had to be reverted.
Every number below comes from real segments pulled out of this project's own
index — the unvectorized windows, split with splitTextForEmbedding and
capped with capSegments, exactly as the indexer does it. Mean segment
length 883 characters.
An earlier round of this work measured
"18ms per embed" on strings like warm query number 5, concluded mpnet was
affordable, and shipped a model change that had to be reverted the next day.
A segment is ~1024 characters, and per-embed cost on short strings says
nothing about it. See the comment on DEFAULT_EMBEDDING_MODEL.
Each model and dtype ran in its own process: two providers in one process
fail intermittently with bad allocation.
Machine: Windows, 24 logical cores, discrete GPU. Absolute figures move with what else the machine is doing — several runs below differ by 20% for that reason. Ratios within a run are the durable part.
| device | ms/segment (separate runs) | cores used |
|---|---|---|
| CPU | 60.2, 81.7 | 10.7, 8.9 |
| DirectML | 10.1, 10.4, 11.8 | 0.9 |
| WebGPU | 15.1 | 1.0 |
cores used is process.cpuUsage() divided by wall time.
The vectors are identical. Embedding the same 128 segments on CPU and on DirectML and comparing the pairs: mean cosine 1.000000, worst pair 0.999999. It is the same computation on different silicon. No re-index, no threshold re-sweep, none of the vector-space mixing hazard that rules out quantization below.
For a 21,349-segment backlog that is roughly 21–29 minutes on CPU against about 4 minutes on the GPU.
onnxruntime-node already ships DirectML.dll in the package this project
installs. Nothing needed adding to package.json; the device option was
simply not passed at the time. It is now, chosen by xtctx calibrate — see
"Calibration, as implemented".
The section above was measured on one machine and left one question open —
DirectML is Windows-only, so is the portable WebGPU path safe to fall back
through? scripts/probe-embedding-device.mjs and
.github/workflows/embedding-device-probe.yml were written to answer it on
hardware nobody here owns. Run 2026-09-21, 48 segments of 1000 characters,
each device in its own process:
| runner | device | ms/segment | vs cpu | cores | worst cosine vs cpu |
|---|---|---|---|---|---|
| macos-latest (arm64) | cpu | 125.5 | — | 1.0 | baseline |
| macos-latest | webgpu | 21.4 | 5.9x | 0.3 | 0.999999 |
| macos-latest | dml | — | unsupported: coreml, webgpu, cpu |
||
| macos-latest | auto | 99.7 | 1.3x | 1.0 | 0.999999 |
| windows-latest | cpu | 35.4 | — | 2.0 | baseline |
| windows-latest | webgpu | 1784.6 | 0.02x | 3.9 | 1.000000 |
| windows-latest | dml | — | Specified display adapter handle is invalid |
||
| windows-latest | auto | 35.9 | 1.0x | 2.0 | 0.999999 |
| ubuntu-latest | cpu | 32.8 | — | 3.8 | baseline |
| ubuntu-latest | webgpu | — | Failed to get a WebGPU adapter: No supported adapters |
||
| ubuntu-latest | dml | — | unsupported: cuda, webgpu, cpu |
||
| ubuntu-latest | auto | 32.8 | 1.0x | 3.8 | 0.999999 |
Two results are worth more than the speed column.
auto is cpu, on all three. That was what the product passed at the
time, so the GPU was not merely underused but unused. It now passes the device
xtctx calibrate measured.
WebGPU on a GPU-less Windows machine is fifty times SLOWER, and does not fail. It found a software adapter, initialised cleanly, returned correct vectors, and took 1.78 seconds per segment against the CPU's 35ms. Linux with no GPU threw instead, which is the behaviour a fallback chain assumes. So "try WebGPU, fall back on error" is not implementable: on the one configuration where it is catastrophic, there is no error to fall back from, and the symptom is a backlog that never drains rather than anything that looks like a failure.
That rules out a silent chain, not the feature. What it leaves is a device
chosen on evidence: a short timed run against the real CPU path on this
machine. That is what xtctx calibrate does — see "Calibration, as
implemented" below.
Vectors agree everywhere. Worst pair across every device that ran on every runner is 0.999999. Whatever selects the device, it does not become part of vector identity, and a machine that ends up on CPU shares an index with one that does not.
DirectML is Windows-with-a-real-GPU only, and that is a feature here: the Windows runner has no display adapter and DirectML said so plainly rather than degrading, which is exactly the check WebGPU failed to perform.
Every ms/segment figure in this table and the one above it is warmup- contaminated, including the 5.1 quoted for DirectML — see "Warmup" below. The relative story survives; the absolute numbers are roughly 1.5–2x pessimistic on the GPU rows.
The CPU path already uses 9–11 of 24 cores, so ONNX Runtime is parallelising internally. Forking worker processes would mostly have them compete for cores already in use. There is some headroom to 24, but it costs a process pool, duplicated model memory and crash handling — against a 6x win from passing an option.
The GPU path uses 0.9 cores. The second prize after speed is that indexing stops eating the machine.
Three paired runs on CPU, fp32, same segments each time:
| batch | run 1 | run 2 | run 3 |
|---|---|---|---|
| 8 | 101.7 | ||
| 16 | 82.2 | 80.9 | 66.3 |
| 32 | 100.5 | 88.7 | 81.4 |
| 64 | 99.9 |
16 wins every pairing against 32, by roughly 10–20%.
Settled 2026-09-21, and MAX_BATCH_SIZE is 16. The direction above turned
out to be right, but nothing above it is why. Measured by running
xtctx scan --embed over this project's own index from an empty vector table,
150 seconds per size, on DirectML:
| batch | ms/window |
|---|---|
| 8 | 123.4 |
| 16 | 107.8, 107.9 |
| 32 | 123.4 |
| 128 | 542.9 |
scripts/probe-batch-size.mjs was written first and said the opposite for the
GPU — 4.0ms/segment at 128 against 5.3 at 32, a 1.37x win reproducing across
runs. Acting on that would have been a 5x regression. Its segments are all
exactly 1000 characters; real ones are not, and a batch is padded to its
longest member, so a wide batch of mixed lengths spends most of its work on
padding. Uniform inputs hide the dominant cost of the real workload.
That is the same failure as the "18ms per embed" figure at the top of this file. A benchmark that does not reproduce the shape of the real input is evidence for the wrong question, and it reproduced cleanly three times while pointing the wrong way.
| dtype | ms/segment |
|---|---|
| fp32 | 100.5 |
| q8 | 84.7 |
16% faster, well short of 2x.
And not free: comparing q8 vectors against fp32 vectors for the same 256 segments gives mean cosine 0.9889, worst pair 0.9789. Every vector moves slightly, which would shift ranking around the confidence threshold and require the eval to re-price it.
There is also a trap. dropVectorsFromOtherModels keys on the model name,
and dtype is not part of that key — so switching dtype would leave existing
fp32 vectors in place beside new q8 ones, silently mixing two slightly
different vector spaces. A dtype switch needs the key widened first.
16% does not buy any of that.
Windows are 8 messages with a stride of 4, so each message appears in about two windows. The obvious inference is that the same text is embedded twice and a content-hash cache would halve the work.
Measured across the real backlog:
unvectorized windows: 3505
segments to embed : 21349
distinct segments : 20155
duplicate segments : 1194 (5.6%)
Segments are packed by filling to a character budget, so the pack boundaries shift with each window's start and the text rarely repeats exactly. A cache would save about 5%. Not worth building.
Run against the 60-query eval corpus. The MiniLM row reproduced the committed
baseline exactly (hybrid 0.333 / 0.533 / 0.183), which is what establishes
that this harness and ranking.eval.test.ts are measuring the same thing.
First, every model at the thresholds tuned for MiniLM:
| model | hybrid mrr | recall@5 | top1 | false positives |
|---|---|---|---|---|
| MiniLM (0.15/0.36) | 0.333 | 0.533 | 0.183 | 0 |
| bge-small | 0.395 | 0.550 | 0.267 | 1.00 |
| gte-small | 0.385 | 0.600 | 0.250 | 1.00 |
A false-positive rate of 1.00 means every deliberately unanswerable query, gibberish included, returned something. The models are not worse — they place their cosine values higher, and a floor tuned to MiniLM's distribution stops excluding anything.
Swept to their own thresholds:
| model | min semantic / confident | hybrid mrr | recall@5 | top1 | vector mrr | FP |
|---|---|---|---|---|---|---|
| MiniLM | 0.15 / 0.36 | 0.333 | 0.533 | 0.183 | 0.246 | 0 |
| bge-small | 0.55 / 0.65 | 0.404 | 0.567 | 0.250 | 0.325 | 0 |
| bge-small | 0.65 / 0.75 | 0.433 | 0.600 | 0.283 | 0.000 | 0 |
| gte-small | 0.75 / 0.85 | 0.418 | 0.583 | 0.283 | 0.315 | 0 |
bge-small at 0.55/0.65 beats MiniLM on every metric at a false-positive rate of zero: hybrid ranking +21%, top-1 +37%, vector mode +32%.
Ignore the 0.65/0.75 row despite its better hybrid score. Vector mode there is
0.000 — nothing clears the semantic floor, so mode: "vector" returns
nothing at all, and the hybrid number is keyword hits being re-ranked. A
public option that silently returns empty is worse than a slightly lower
score.
The cost: indexing the eval corpus took 25.9s under bge-small against MiniLM's 14.2s, about 1.8x slower — the constraint the GPU result above loosens.
Caveat. The eval corpus is synthetic and 60 queries. Differences this size
are suggestive, not settled. Changing model means re-embedding everything,
which dropVectorsFromOtherModels already handles correctly by name.
Superseded 2026-09-21. bge-small is the default, at 0.62 / 0.64 rather
than the 0.55 / 0.65 recommended above. scripts/embedding-bakeoff.ts
reproduced every row in this section and then swept finer, which found that
the confidence floor was what destroyed vector mode: at 0.55/0.65 vector
scores 0.325, at 0.55/0.64 it scores 0.385, both at a false-positive rate of
zero. The shipped pair scores hybrid 0.417 / 0.583 / 0.267 and vector
0.398 / 0.533 / 0.317, and is what
tests/eval/results/ranking-baseline.json now holds.
The caveat above was not discharged and is not claimed to be: the corpus is still synthetic and still sixty queries. What changed is that the margin holds across three modes and every threshold pair swept rather than resting on one row, and that the 1.8x indexing cost stopped being decisive once calibration made embedding ~6x faster on a machine with a GPU.
Ranked by value against effort, on the evidence above:
- GPU chosen by measurement. Done —
xtctx calibrate, and automatically insidextctx scan --embed. See below. - Batch size 16. Done — measured on the real path, ~13% on the GPU.
- bge-small with its own thresholds. Done, at 0.62/0.64.
Closed, with reasons above: quantization, segment caching, multi-process
embedding, and (from DEFAULT_EMBEDDING_MODEL) mpnet and static models.
Every per-segment figure in this file above the three-OS table was taken after a two-segment warmup batch, which is enough for the CPU and not remotely enough for a GPU. Graph compilation and buffer allocation are a one-off cost, so a short warmup leaves them inside the timed window, and they are large enough on the GPU paths to change which device looks fastest.
Same desktop, same 16 segments, only the warmup differing:
| warmup | cpu | dml | webgpu | winner |
|---|---|---|---|---|
| 2 segments, 1 timed pass | 20.1 | 6.0 | 10.1 | dml |
| 2 segments, best of 3 | 21.0 | 4.6 | 3.4 | webgpu |
| full pass, best of 3 | 18.4–21.3 | 3.1 | 3.3 | dml |
The last row is three consecutive runs and varies by 0.0 on DirectML and 0.1
on WebGPU, against a CPU arm that still moves by 3ms with machine load. That
is the measurement xtctx calibrate takes.
Two things follow. The earlier DirectML figure of 5.1ms per segment was also warmup-contaminated and the real number here is about 3.1, so the GPU win on this machine is ~6x rather than ~4x. And a single timed pass is not a measurement: it ranked WebGPU as the slower GPU when it is within 7% of the faster one.
xtctx calibrate times the model on every execution provider the platform
offers, each in its own process, takes the fastest of three passes per device,
and writes the verdict to ~/.xtctx/device.json keyed by platform, arch, core
count, model and dtype. xtctx scan --embed runs it automatically when the
machine has no verdict yet — a minute against the hours that command is about
to spend — and --no-calibrate skips it.
The MCP server calibrates too, in the background when it starts on a machine
with no verdict — never behind a tool call, which has a four-second budget.
It starts calibration before its first scan and makes the model's first load
wait for the result, so the session that paid for the measurement is the one
that uses it. (An earlier version calibrated after the scan and tried to
retarget the provider, which every scan had already started loading; the
verdict only ever reached the next session, and this paragraph said the server
did not calibrate at all.) One server calibrates at a time, under a
machine-wide lock. xtctx status prints the device, read off the provider
rather than off the cache file, so the line is evidence that the indexer is
using it rather than evidence that a file exists.
The decision rule is "fastest measured device", with a 1.1x margin over the CPU, and every case where no comparison exists resolves to the CPU: a device that would not initialise, and a run where the CPU arm itself failed.
Measured end to end on this desktop, xtctx scan --embed before and after:
551.9ms per window on the CPU against 50.7ms on the GPU, taking this
project's remaining backlog from about 24 minutes to about 1.5.