Skip to content

feat(eval): dense score diagnostics, and the answer is not a threshold - #206

Merged
mrsibe merged 2 commits into
feat/eval-corpus-hard-negativesfrom
feat/eval-score-diagnostics
Sep 30, 2026
Merged

mrsibe merged 2 commits into
feat/eval-corpus-hard-negativesfrom
feat/eval-score-diagnostics

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #205 (feat/eval-corpus-hard-negatives). Retarget to main after the chain merges.

What does this PR do?

Adds npm run eval:scores, which answers one question before any threshold grid is chosen: can a single dense similarity threshold separate "this passage answers the question" from "this one does not"?

It gives a more useful answer than a best threshold would have.

First, the score is not a cosine

SQLiteVectorStore computes score = 1 - distance / 2, and sqlite-vec's cosine distance is 1 - cosine, so:

score = (1 + cosine) / 2      =>      threshold 0.5 == cosine 0.0

The shipped threshold = 0.5 is cosine ≥ 0, not "cosine ≥ 0.5". Everything downstream — the sweep grid 0 / 0.3 / 0.4 / 0.5 / 0.6, the threshold report, the source comment — was describing an affine map of the cosine rather than the cosine. That is a large part of why the sweep looked so flat.

Fixed in the store's comment, in eval/README.md, and reported as two columns throughout the new diagnostics so the two never get conflated again.

Method

  • dense only — a hybrid score is an RRF value (1 / (60 + rank)) and not on the same scale as a normalised cosine; mixing them would manufacture a fresh misleading number.
  • validation only — the distribution is used to choose where to sweep, so looking at test first would be tuning on the reporting side.
  • threshold = 0, candidateK = 500 (above the 53-chunk index) so the distribution is not already truncated by the parameter under investigation.

What it measured

best relevant      p10 0.8906 (cosine 0.781)   p50 0.9359 (0.872)
best non-relevant  p50 0.9243 (cosine 0.849)   p90 0.9386 (0.877)
margin             p10 -0.0272                 p50 +0.0026
unanswerable max   p50 0.9259 (cosine 0.852)   p90 0.9424 (0.885)
  1. The best non-relevant passage outranks the best relevant passage about half the time. The margin's p50 is +0.0026 and its p10 is −0.0272. Dense similarity is barely informative about relevance here — the hard-negative clusters are doing exactly what they were built for.
  2. The unanswerable max sits above the worst required relevant passage (p90 0.9424 vs p10 0.8906). The distributions are not separable.
  3. cross-lingual is the lowest of every type — best relevant p10 0.8808, against semantic at 0.9208 and hard-negative at 0.9359. A threshold tuned on the aggregate would cut cross-lingual first, which is the concern you raised when asking for the by-type breakdown.

The threshold curve is the answer

threshold raw cosine answerable hit answerable full recall unanswerable abstain
0.500 0.000 1.0000 1.0000 0.0000
0.875 0.750 1.0000 1.0000 0.0000
0.900 0.800 0.8947 0.8947 0.2000
0.925 0.850 0.6316 0.6316 0.4000
0.950 0.900 0.1579 0.1579 1.0000

Abstention never rises without full recall falling. Every threshold that refuses an unanswerable question refuses required relevant passages at the same rate — they are the same events. There is no operating point that buys the first without paying the second.

So the conclusion is not "sweep 0.7–0.8 instead of 0–0.6". It is:

A single dense similarity threshold cannot carry both recall and abstention on this corpus. The next mechanism to evaluate is a different signal — reranker score, top1−top2 margin, per-query thresholds, or claim-level answerability — not a finer grid.

That is why this was worth doing before more clusters. Re-gridding the sweep would have moved a knob whose curve is flat where it matters and precipitous where it does not.

What this PR does not do

  • It does not choose a threshold. That is now blocked on finding a signal that separates the distributions.
  • It does not change the eval:threshold grid. The diagnostic is the replacement for re-gridding, not a preamble to it.

Testing

  • npm run typecheck — clean
  • npm test — 504 pass
  • npm run eval:scores — writes docs/eval/scores-v1.6.{json,md}
  • The committed baseline is unchanged: --eval-scores is off by default, so the CI-diffed JSON does not grow a score series it has no use for

Also in this push

#200 is now genuinely in the chain. Retargeting a base does not inject commits — you were right, and the previous "it cannot be missed" claim did not hold. I rebased it properly: #203 → #200 → #204 → #205, verified with git merge-base --is-ancestor, and force-pushed only those three branches. Every downstream PR re-ran its checks after the rebase.

Related

Part of #192. Child 2 (corpus diagnostics).

Adds `npm run eval:scores`, which measures whether a single dense similarity threshold is
capable of separating "this passage answers the question" from "this one does not" — before
any threshold grid is chosen. Child 2 of #192, and it changes what the next step should be.

## First, the score is not a cosine

`SQLiteVectorStore` computes `score = 1 - distance / 2`, and sqlite-vec's cosine distance is
`1 - cosine`, so:

```
score = (1 + cosine) / 2      =>      threshold 0.5 == cosine 0.0
```

The shipped `threshold = 0.5` is therefore **cosine ≥ 0**, not "cosine ≥ 0.5". Every guard
that mattered here — the sweep grid, the threshold report, the source comment — was
describing an affine map, not the cosine. Fixed in the store's comment, in `eval/README.md`,
and reported as two columns throughout the new diagnostics.

## The measurement

Dense only (hybrid's `score` is an RRF value, `1 / (60 + rank)`, and not on the same scale),
`validation` split only (the distribution is used to choose where to sweep, so it must not
see `test`), `threshold = 0` and `candidateK = 500` (above the 53-chunk index, so every chunk
is scored for every query).

```
best relevant     p10 0.8906 (cosine 0.781)   p50 0.9359 (0.872)
best non-relevant p50 0.9243 (cosine 0.849)   p90 0.9386 (0.877)
margin            p10 -0.0272                 p50 +0.0026
unanswerable max  p50 0.9259 (cosine 0.852)   p90 0.9424 (0.885)
```

Three things fall out of that:

1. **The best non-relevant passage outranks the best relevant one about half the time.**
   The margin's p50 is +0.0026 and its p10 is −0.0272. Dense similarity is barely
   informative about relevance on this corpus — the hard-negative clusters are doing exactly
   what they were built for.
2. **The unanswerable max sits above the worst required relevant passage** (p90 0.9424 vs
   p10 0.8906), so the two distributions are not separable.
3. **`cross-lingual` is the lowest of every type** (best relevant p10 0.8808) against
   `semantic` at 0.9208 — a threshold tuned on the aggregate would cut cross-lingual first.

## The threshold curve is the answer

| threshold | raw cosine | answerable hit | answerable full recall | unanswerable abstain |
| --- | --- | --- | --- | --- |
| 0.875 | 0.75 | 1.0000 | 1.0000 | 0.0000 |
| 0.900 | 0.80 | 0.8947 | 0.8947 | 0.2000 |
| 0.925 | 0.85 | 0.6316 | 0.6316 | 0.4000 |
| 0.950 | 0.900 | 0.1579 | 0.1579 | 1.0000 |

**Abstention never rises without full recall falling.** Every threshold that refuses an
unanswerable question refuses required relevant passages at the same rate. There is no
operating point that buys the first without paying the second.

So the conclusion is not "sweep 0.7–0.8 instead of 0–0.6". It is: **a single dense
similarity threshold cannot carry both recall and abstention on this corpus**, and the next
mechanism to evaluate is a different signal — reranker score, top1−top2 margin, per-query
thresholds, or claim-level answerability. That is a much more useful finding than a best
threshold would have been, and it is why this PR was worth doing before more clusters.

## What this PR does not do

- It does not choose a threshold. That is now blocked on a signal that can separate the
  distributions.
- The `eval:threshold` sweep grid is left as it is. Re-gridding it would move a knob whose
  curve is flat where it matters and precipitous where it does not — the diagnostic is the
  replacement for re-gridding, not a preamble to it.

## Testing

- `npm run typecheck` — clean
- `npm test` — 504 pass
- `npm run eval:scores` — writes `docs/eval/scores-v1.6.{json,md}`
- The committed baseline is **unchanged**: `--eval-scores` is off by default, so the
  CI-diffed JSON does not grow a score series it has no use for

Part of #192 (child 2: corpus diagnostics).
@github-actions github-actions Bot added the enhancement New feature or request label Sep 30, 2026
…ne transform

`score = (1 + cosine) / 2` is an affine map, so an **absolute** score converts as
`cosine = 2·score − 1`. A **difference** of scores does not: the `+1` cancels and
`Δcosine = 2·Δscore`.

The margin row applied the absolute transform, which reported `2Δscore − 1` — for a
margin near zero that is a cosine margin near **−1**, a sign flip on top of a scale error:

```
                    before            after
margin p10   -0.0272 (-1.054)   -0.0272 (-0.054)
margin p50   +0.0026 (-0.995)   +0.0026 (+0.005)
```

`quantileRow` now takes the cosine transform as an argument and the margin row passes
`toCosineMargin`, so the two cannot be confused at the call site again. The report also
says which mapping applies to which row.

This does not change the separability conclusion — the p90 for the best non-relevant
passage is 0.9386 against a worst-required-relevant p10 of 0.8906, and that comparison
uses absolute scores, which were always converted correctly.

The margin itself is also now labelled in the report as an **oracle** quantity: at runtime
nothing knows which result is relevant, so it measures how much the score separates the
two, and it is not a signal the product could use.

Part of #192 (correction to the score diagnostics).
@mrsibe
mrsibe merged commit fd91d5b into feat/eval-corpus-hard-negatives Sep 30, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant