feat(eval): dense score diagnostics, and the answer is not a threshold - #206
Merged
mrsibe merged 2 commits intoSep 30, 2026
Merged
Conversation
Adds `npm run eval:scores`, which measures whether a single dense similarity threshold is capable of separating "this passage answers the question" from "this one does not" — before any threshold grid is chosen. Child 2 of #192, and it changes what the next step should be. ## First, the score is not a cosine `SQLiteVectorStore` computes `score = 1 - distance / 2`, and sqlite-vec's cosine distance is `1 - cosine`, so: ``` score = (1 + cosine) / 2 => threshold 0.5 == cosine 0.0 ``` The shipped `threshold = 0.5` is therefore **cosine ≥ 0**, not "cosine ≥ 0.5". Every guard that mattered here — the sweep grid, the threshold report, the source comment — was describing an affine map, not the cosine. Fixed in the store's comment, in `eval/README.md`, and reported as two columns throughout the new diagnostics. ## The measurement Dense only (hybrid's `score` is an RRF value, `1 / (60 + rank)`, and not on the same scale), `validation` split only (the distribution is used to choose where to sweep, so it must not see `test`), `threshold = 0` and `candidateK = 500` (above the 53-chunk index, so every chunk is scored for every query). ``` best relevant p10 0.8906 (cosine 0.781) p50 0.9359 (0.872) best non-relevant p50 0.9243 (cosine 0.849) p90 0.9386 (0.877) margin p10 -0.0272 p50 +0.0026 unanswerable max p50 0.9259 (cosine 0.852) p90 0.9424 (0.885) ``` Three things fall out of that: 1. **The best non-relevant passage outranks the best relevant one about half the time.** The margin's p50 is +0.0026 and its p10 is −0.0272. Dense similarity is barely informative about relevance on this corpus — the hard-negative clusters are doing exactly what they were built for. 2. **The unanswerable max sits above the worst required relevant passage** (p90 0.9424 vs p10 0.8906), so the two distributions are not separable. 3. **`cross-lingual` is the lowest of every type** (best relevant p10 0.8808) against `semantic` at 0.9208 — a threshold tuned on the aggregate would cut cross-lingual first. ## The threshold curve is the answer | threshold | raw cosine | answerable hit | answerable full recall | unanswerable abstain | | --- | --- | --- | --- | --- | | 0.875 | 0.75 | 1.0000 | 1.0000 | 0.0000 | | 0.900 | 0.80 | 0.8947 | 0.8947 | 0.2000 | | 0.925 | 0.85 | 0.6316 | 0.6316 | 0.4000 | | 0.950 | 0.900 | 0.1579 | 0.1579 | 1.0000 | **Abstention never rises without full recall falling.** Every threshold that refuses an unanswerable question refuses required relevant passages at the same rate. There is no operating point that buys the first without paying the second. So the conclusion is not "sweep 0.7–0.8 instead of 0–0.6". It is: **a single dense similarity threshold cannot carry both recall and abstention on this corpus**, and the next mechanism to evaluate is a different signal — reranker score, top1−top2 margin, per-query thresholds, or claim-level answerability. That is a much more useful finding than a best threshold would have been, and it is why this PR was worth doing before more clusters. ## What this PR does not do - It does not choose a threshold. That is now blocked on a signal that can separate the distributions. - The `eval:threshold` sweep grid is left as it is. Re-gridding it would move a knob whose curve is flat where it matters and precipitous where it does not — the diagnostic is the replacement for re-gridding, not a preamble to it. ## Testing - `npm run typecheck` — clean - `npm test` — 504 pass - `npm run eval:scores` — writes `docs/eval/scores-v1.6.{json,md}` - The committed baseline is **unchanged**: `--eval-scores` is off by default, so the CI-diffed JSON does not grow a score series it has no use for Part of #192 (child 2: corpus diagnostics).
…ne transform
`score = (1 + cosine) / 2` is an affine map, so an **absolute** score converts as
`cosine = 2·score − 1`. A **difference** of scores does not: the `+1` cancels and
`Δcosine = 2·Δscore`.
The margin row applied the absolute transform, which reported `2Δscore − 1` — for a
margin near zero that is a cosine margin near **−1**, a sign flip on top of a scale error:
```
before after
margin p10 -0.0272 (-1.054) -0.0272 (-0.054)
margin p50 +0.0026 (-0.995) +0.0026 (+0.005)
```
`quantileRow` now takes the cosine transform as an argument and the margin row passes
`toCosineMargin`, so the two cannot be confused at the call site again. The report also
says which mapping applies to which row.
This does not change the separability conclusion — the p90 for the best non-relevant
passage is 0.9386 against a worst-required-relevant p10 of 0.8906, and that comparison
uses absolute scores, which were always converted correctly.
The margin itself is also now labelled in the report as an **oracle** quantity: at runtime
nothing knows which result is relevant, so it measures how much the score separates the
two, and it is not a signal the product could use.
Part of #192 (correction to the score diagnostics).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #205 (
feat/eval-corpus-hard-negatives). Retarget tomainafter the chain merges.What does this PR do?
Adds
npm run eval:scores, which answers one question before any threshold grid is chosen: can a single dense similarity threshold separate "this passage answers the question" from "this one does not"?It gives a more useful answer than a best threshold would have.
First, the score is not a cosine
SQLiteVectorStorecomputesscore = 1 - distance / 2, and sqlite-vec's cosine distance is1 - cosine, so:The shipped
threshold = 0.5is cosine ≥ 0, not "cosine ≥ 0.5". Everything downstream — the sweep grid0 / 0.3 / 0.4 / 0.5 / 0.6, the threshold report, the source comment — was describing an affine map of the cosine rather than the cosine. That is a large part of why the sweep looked so flat.Fixed in the store's comment, in
eval/README.md, and reported as two columns throughout the new diagnostics so the two never get conflated again.Method
scoreis an RRF value (1 / (60 + rank)) and not on the same scale as a normalised cosine; mixing them would manufacture a fresh misleading number.validationonly — the distribution is used to choose where to sweep, so looking attestfirst would be tuning on the reporting side.threshold = 0,candidateK = 500(above the 53-chunk index) so the distribution is not already truncated by the parameter under investigation.What it measured
cross-lingualis the lowest of every type — best relevant p10 0.8808, againstsemanticat 0.9208 andhard-negativeat 0.9359. A threshold tuned on the aggregate would cut cross-lingual first, which is the concern you raised when asking for the by-type breakdown.The threshold curve is the answer
Abstention never rises without full recall falling. Every threshold that refuses an unanswerable question refuses required relevant passages at the same rate — they are the same events. There is no operating point that buys the first without paying the second.
So the conclusion is not "sweep 0.7–0.8 instead of 0–0.6". It is:
That is why this was worth doing before more clusters. Re-gridding the sweep would have moved a knob whose curve is flat where it matters and precipitous where it does not.
What this PR does not do
eval:thresholdgrid. The diagnostic is the replacement for re-gridding, not a preamble to it.Testing
npm run typecheck— cleannpm test— 504 passnpm run eval:scores— writesdocs/eval/scores-v1.6.{json,md}--eval-scoresis off by default, so the CI-diffed JSON does not grow a score series it has no use forAlso in this push
#200 is now genuinely in the chain. Retargeting a base does not inject commits — you were right, and the previous "it cannot be missed" claim did not hold. I rebased it properly:
#203 → #200 → #204 → #205, verified withgit merge-base --is-ancestor, and force-pushed only those three branches. Every downstream PR re-ran its checks after the rebase.Related
Part of #192. Child 2 (corpus diagnostics).