Skip to content

Latest commit

 

History

History
360 lines (281 loc) · 16.7 KB

File metadata and controls

360 lines (281 loc) · 16.7 KB

Embedding performance

Measurements taken on 2026-09-20 and 2026-09-21 while looking for a way to make indexing faster. They exist so the next person asking "can we speed this up" starts from evidence rather than from the same four guesses.

Three of the four guesses were wrong, which is the main reason this file is worth keeping.

Read this first, because the file is written in two voices. Everything up to and including "Model bake-off" was written on 2026-09-20, when none of it had changed the code and the conclusions were all "blocked on". On 2026-09-21 three of them shipped: device calibration, the model change to bge-small, and a bake-off harness that makes the model rows reproducible. Sections written before that day have been annotated where they are now false rather than rewritten, because how a wrong conclusion was reached is the part worth keeping. Where an annotation and a table disagree, the annotation is current.

What is reproducible here, and what is not

tests/eval/results/ranking-baseline.json is regenerated by npm run test:eval and now reads bge-small's numbers — hybrid 0.417 / 0.583 / 0.267. The MiniLM figures quoted throughout this file (0.333 / 0.533 / 0.183) were that file's contents until 2026-09-21 and are kept as the comparison the model change was made against.

Two more are now reproducible. The three-OS device table, which scripts/probe-embedding-device.mjs regenerates and .github/workflows/embedding-device-probe.yml runs on the same three runners. And every model row below: scripts/embedding-bakeoff.ts takes --model and --sweep, and its MiniLM row at 0.15/0.36 reproduces the old committed baseline exactly, which is what establishes it measures the same thing.

Everything else — the single-machine device table, the batch sweep, the dtype comparison, and the segment-duplication count — was measured in one session on one machine with scratch scripts that are not in this repository. The method is described precisely enough to redo, and the ratios are the durable part, but nothing here re-runs them and nobody should treat the absolute numbers as checkable. (The bge and gte rows said the same until 2026-09-21, when scripts/embedding-bakeoff.ts was written and reproduced them.)

A number that reads as evidence while being untraceable is how the "18ms per embed" figure below survived long enough to drive a model change that had to be reverted.

How these were measured

Every number below comes from real segments pulled out of this project's own index — the unvectorized windows, split with splitTextForEmbedding and capped with capSegments, exactly as the indexer does it. Mean segment length 883 characters.

An earlier round of this work measured "18ms per embed" on strings like warm query number 5, concluded mpnet was affordable, and shipped a model change that had to be reverted the next day. A segment is ~1024 characters, and per-embed cost on short strings says nothing about it. See the comment on DEFAULT_EMBEDDING_MODEL.

Each model and dtype ran in its own process: two providers in one process fail intermittently with bad allocation.

Machine: Windows, 24 logical cores, discrete GPU. Absolute figures move with what else the machine is doing — several runs below differ by 20% for that reason. Ratios within a run are the durable part.

Device: the GPU is ~6x, for free

device ms/segment (separate runs) cores used
CPU 60.2, 81.7 10.7, 8.9
DirectML 10.1, 10.4, 11.8 0.9
WebGPU 15.1 1.0

cores used is process.cpuUsage() divided by wall time.

The vectors are identical. Embedding the same 128 segments on CPU and on DirectML and comparing the pairs: mean cosine 1.000000, worst pair 0.999999. It is the same computation on different silicon. No re-index, no threshold re-sweep, none of the vector-space mixing hazard that rules out quantization below.

For a 21,349-segment backlog that is roughly 21–29 minutes on CPU against about 4 minutes on the GPU.

onnxruntime-node already ships DirectML.dll in the package this project installs. Nothing needed adding to package.json; the device option was simply not passed at the time. It is now, chosen by xtctx calibrate — see "Calibration, as implemented".

Device, on three operating systems: WebGPU is not a fallback chain

The section above was measured on one machine and left one question open — DirectML is Windows-only, so is the portable WebGPU path safe to fall back through? scripts/probe-embedding-device.mjs and .github/workflows/embedding-device-probe.yml were written to answer it on hardware nobody here owns. Run 2026-09-21, 48 segments of 1000 characters, each device in its own process:

runner device ms/segment vs cpu cores worst cosine vs cpu
macos-latest (arm64) cpu 125.5 — 1.0 baseline
macos-latest webgpu 21.4 5.9x 0.3 0.999999
macos-latest dml — unsupported: coreml, webgpu, cpu
macos-latest auto 99.7 1.3x 1.0 0.999999
windows-latest cpu 35.4 — 2.0 baseline
windows-latest webgpu 1784.6 0.02x 3.9 1.000000
windows-latest dml — Specified display adapter handle is invalid
windows-latest auto 35.9 1.0x 2.0 0.999999
ubuntu-latest cpu 32.8 — 3.8 baseline
ubuntu-latest webgpu — Failed to get a WebGPU adapter: No supported adapters
ubuntu-latest dml — unsupported: cuda, webgpu, cpu
ubuntu-latest auto 32.8 1.0x 3.8 0.999999

Two results are worth more than the speed column.

auto is cpu, on all three. That was what the product passed at the time, so the GPU was not merely underused but unused. It now passes the device xtctx calibrate measured.

WebGPU on a GPU-less Windows machine is fifty times SLOWER, and does not fail. It found a software adapter, initialised cleanly, returned correct vectors, and took 1.78 seconds per segment against the CPU's 35ms. Linux with no GPU threw instead, which is the behaviour a fallback chain assumes. So "try WebGPU, fall back on error" is not implementable: on the one configuration where it is catastrophic, there is no error to fall back from, and the symptom is a backlog that never drains rather than anything that looks like a failure.

That rules out a silent chain, not the feature. What it leaves is a device chosen on evidence: a short timed run against the real CPU path on this machine. That is what xtctx calibrate does — see "Calibration, as implemented" below.

Vectors agree everywhere. Worst pair across every device that ran on every runner is 0.999999. Whatever selects the device, it does not become part of vector identity, and a machine that ends up on CPU shares an index with one that does not.

DirectML is Windows-with-a-real-GPU only, and that is a feature here: the Windows runner has no display adapter and DirectML said so plainly rather than degrading, which is exactly the check WebGPU failed to perform.

Every ms/segment figure in this table and the one above it is warmup- contaminated, including the 5.1 quoted for DirectML — see "Warmup" below. The relative story survives; the absolute numbers are roughly 1.5–2x pessimistic on the GPU rows.

Multi-process embedding: no headroom worth taking

The CPU path already uses 9–11 of 24 cores, so ONNX Runtime is parallelising internally. Forking worker processes would mostly have them compete for cores already in use. There is some headroom to 24, but it costs a process pool, duplicated model memory and crash handling — against a 6x win from passing an option.

The GPU path uses 0.9 cores. The second prize after speed is that indexing stops eating the machine.

Batch size: 16 beats 32

Three paired runs on CPU, fp32, same segments each time:

batch run 1 run 2 run 3
8 101.7
16 82.2 80.9 66.3
32 100.5 88.7 81.4
64 99.9

16 wins every pairing against 32, by roughly 10–20%.

Settled 2026-09-21, and MAX_BATCH_SIZE is 16. The direction above turned out to be right, but nothing above it is why. Measured by running xtctx scan --embed over this project's own index from an empty vector table, 150 seconds per size, on DirectML:

batch ms/window
8 123.4
16 107.8, 107.9
32 123.4
128 542.9

scripts/probe-batch-size.mjs was written first and said the opposite for the GPU — 4.0ms/segment at 128 against 5.3 at 32, a 1.37x win reproducing across runs. Acting on that would have been a 5x regression. Its segments are all exactly 1000 characters; real ones are not, and a batch is padded to its longest member, so a wide batch of mixed lengths spends most of its work on padding. Uniform inputs hide the dominant cost of the real workload.

That is the same failure as the "18ms per embed" figure at the top of this file. A benchmark that does not reproduce the shape of the real input is evidence for the wrong question, and it reproduced cleanly three times while pointing the wrong way.

Quantized weights: rejected

dtype ms/segment
fp32 100.5
q8 84.7

16% faster, well short of 2x.

And not free: comparing q8 vectors against fp32 vectors for the same 256 segments gives mean cosine 0.9889, worst pair 0.9789. Every vector moves slightly, which would shift ranking around the confidence threshold and require the eval to re-price it.

There is also a trap. dropVectorsFromOtherModels keys on the model name, and dtype is not part of that key — so switching dtype would leave existing fp32 vectors in place beside new q8 ones, silently mixing two slightly different vector spaces. A dtype switch needs the key widened first.

16% does not buy any of that.

Caching duplicate segments: rejected

Windows are 8 messages with a stride of 4, so each message appears in about two windows. The obvious inference is that the same text is embedded twice and a content-hash cache would halve the work.

Measured across the real backlog:

unvectorized windows: 3505
segments to embed   : 21349
distinct segments   : 20155
duplicate segments  : 1194 (5.6%)

Segments are packed by filling to a character budget, so the pack boundaries shift with each window's start and the text rarely repeats exactly. A cache would save about 5%. Not worth building.

Model bake-off: bge-small wins, once thresholds are per-model

Run against the 60-query eval corpus. The MiniLM row reproduced the committed baseline exactly (hybrid 0.333 / 0.533 / 0.183), which is what establishes that this harness and ranking.eval.test.ts are measuring the same thing.

First, every model at the thresholds tuned for MiniLM:

model hybrid mrr recall@5 top1 false positives
MiniLM (0.15/0.36) 0.333 0.533 0.183 0
bge-small 0.395 0.550 0.267 1.00
gte-small 0.385 0.600 0.250 1.00

A false-positive rate of 1.00 means every deliberately unanswerable query, gibberish included, returned something. The models are not worse — they place their cosine values higher, and a floor tuned to MiniLM's distribution stops excluding anything.

Swept to their own thresholds:

model min semantic / confident hybrid mrr recall@5 top1 vector mrr FP
MiniLM 0.15 / 0.36 0.333 0.533 0.183 0.246 0
bge-small 0.55 / 0.65 0.404 0.567 0.250 0.325 0
bge-small 0.65 / 0.75 0.433 0.600 0.283 0.000 0
gte-small 0.75 / 0.85 0.418 0.583 0.283 0.315 0

bge-small at 0.55/0.65 beats MiniLM on every metric at a false-positive rate of zero: hybrid ranking +21%, top-1 +37%, vector mode +32%.

Ignore the 0.65/0.75 row despite its better hybrid score. Vector mode there is 0.000 — nothing clears the semantic floor, so mode: "vector" returns nothing at all, and the hybrid number is keyword hits being re-ranked. A public option that silently returns empty is worse than a slightly lower score.

The cost: indexing the eval corpus took 25.9s under bge-small against MiniLM's 14.2s, about 1.8x slower — the constraint the GPU result above loosens.

Caveat. The eval corpus is synthetic and 60 queries. Differences this size are suggestive, not settled. Changing model means re-embedding everything, which dropVectorsFromOtherModels already handles correctly by name.

Superseded 2026-09-21. bge-small is the default, at 0.62 / 0.64 rather than the 0.55 / 0.65 recommended above. scripts/embedding-bakeoff.ts reproduced every row in this section and then swept finer, which found that the confidence floor was what destroyed vector mode: at 0.55/0.65 vector scores 0.325, at 0.55/0.64 it scores 0.385, both at a false-positive rate of zero. The shipped pair scores hybrid 0.417 / 0.583 / 0.267 and vector 0.398 / 0.533 / 0.317, and is what tests/eval/results/ranking-baseline.json now holds.

The caveat above was not discharged and is not claimed to be: the corpus is still synthetic and still sixty queries. What changed is that the margin holds across three modes and every threshold pair swept rather than resting on one row, and that the 1.8x indexing cost stopped being decisive once calibration made embedding ~6x faster on a machine with a GPU.

What this leaves

Ranked by value against effort, on the evidence above:

  1. GPU chosen by measurement. Done — xtctx calibrate, and automatically inside xtctx scan --embed. See below.
  2. Batch size 16. Done — measured on the real path, ~13% on the GPU.
  3. bge-small with its own thresholds. Done, at 0.62/0.64.

Closed, with reasons above: quantization, segment caching, multi-process embedding, and (from DEFAULT_EMBEDDING_MODEL) mpnet and static models.

Warmup: a short one inverts the ranking

Every per-segment figure in this file above the three-OS table was taken after a two-segment warmup batch, which is enough for the CPU and not remotely enough for a GPU. Graph compilation and buffer allocation are a one-off cost, so a short warmup leaves them inside the timed window, and they are large enough on the GPU paths to change which device looks fastest.

Same desktop, same 16 segments, only the warmup differing:

warmup cpu dml webgpu winner
2 segments, 1 timed pass 20.1 6.0 10.1 dml
2 segments, best of 3 21.0 4.6 3.4 webgpu
full pass, best of 3 18.4–21.3 3.1 3.3 dml

The last row is three consecutive runs and varies by 0.0 on DirectML and 0.1 on WebGPU, against a CPU arm that still moves by 3ms with machine load. That is the measurement xtctx calibrate takes.

Two things follow. The earlier DirectML figure of 5.1ms per segment was also warmup-contaminated and the real number here is about 3.1, so the GPU win on this machine is ~6x rather than ~4x. And a single timed pass is not a measurement: it ranked WebGPU as the slower GPU when it is within 7% of the faster one.

Calibration, as implemented

xtctx calibrate times the model on every execution provider the platform offers, each in its own process, takes the fastest of three passes per device, and writes the verdict to ~/.xtctx/device.json keyed by platform, arch, core count, model and dtype. xtctx scan --embed runs it automatically when the machine has no verdict yet — a minute against the hours that command is about to spend — and --no-calibrate skips it.

The MCP server calibrates too, in the background when it starts on a machine with no verdict — never behind a tool call, which has a four-second budget. It starts calibration before its first scan and makes the model's first load wait for the result, so the session that paid for the measurement is the one that uses it. (An earlier version calibrated after the scan and tried to retarget the provider, which every scan had already started loading; the verdict only ever reached the next session, and this paragraph said the server did not calibrate at all.) One server calibrates at a time, under a machine-wide lock. xtctx status prints the device, read off the provider rather than off the cache file, so the line is evidence that the indexer is using it rather than evidence that a file exists.

The decision rule is "fastest measured device", with a 1.1x margin over the CPU, and every case where no comparison exists resolves to the CPU: a device that would not initialise, and a run where the CPU arm itself failed.

Measured end to end on this desktop, xtctx scan --embed before and after: 551.9ms per window on the CPU against 50.7ms on the GPU, taking this project's remaining backlog from about 24 minutes to about 1.5.