Skip to content

feat(eval): derive the similarity threshold on validation, report on test - #196

Merged
mrsibe merged 1 commit into
feat/eval-metrics-v2from
feat/eval-threshold-sweep
Sep 30, 2026
Merged

mrsibe merged 1 commit into
feat/eval-metrics-v2from
feat/eval-threshold-sweep

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #195 (feat/eval-metrics-v2). Retarget to main after that merges.

What does this PR do?

Derives the similarity threshold on a held-out split instead of hand-picking it, and adds a no-result-rate metric that a threshold sweep needs. Child 4 of #192.

The mechanism

A deterministic split, owned by the harness. --eval-split=validation|test partitions questions by a hash of the id, so the same questions.jsonl cuts the same way on every machine and both arms go through the same code path that produces the frozen baseline. A parameter chosen on the questions it is scored on is fitted, not measured.

npm run eval:threshold sweeps 0 / 0.3 / 0.4 / 0.5 / 0.6, selects on validation, and reports the winner on test.

The no-result rate. A higher threshold can look better on a ranking metric while quietly making the product answer "not in your sources" more often. That trade is invisible unless it is counted, so the sweep counts it on both splits.

The result here is a non-result, and it is reported as one

Threshold Recall@5 (val) nDCG@10 (val) No-result (val) nDCG@10 (test) No-result (test)
0 1.0000 0.9500 0.0000 0.9406 0.0000
0.3 1.0000 0.9500 0.0000 0.9406 0.0000
0.4 1.0000 0.9500 0.0000 0.9406 0.0000
0.5 1.0000 0.9500 0.0000 0.9406 0.0000
0.6 1.0000 0.9500 0.0000 0.9406 0.0000

The sweep is flat: every threshold produces identical metrics and never filters a passage. E5 does not score these query/chunk pairs below 0.6, so on this corpus the threshold is non-binding.

The report says "no evidence to change threshold = 0.5" rather than nominating the tie-break winner. The rule nominates 0 only because it prefers the widest threshold among equals; that is a tie-break, not a finding. Moving a product parameter on a flat sweep would be noise dressed as a result. The selection rule is stated in the output so it can be argued with.

This is the honest outcome, and it is another instance of the corpus limitation #192 describes. The mechanism is what ships here; the number is what the corpus cannot yet support.

Testing

  • npm run typecheck — clean
  • npm test — 489 pass, 4 new: split bounds, split determinism, validation/test partition without overlap, all preserves order
  • npm run eval — docs/eval/baseline-v1.6.json regenerated with config.split
  • npm run eval:threshold — re-runs cleanly and rewrites the same tables

Related

Part of #192. Child 4 (re-derive threshold).

…test

`threshold: 0.5` was hand-picked, and a cosine score has no universal meaning — its
distribution depends on the embedding model, the language, the query type and the
chunk length. Child 4 of #192.

**A deterministic split, owned by the harness.** `--eval-split=validation|test`
partitions the questions by a hash of the id, so the same `questions.jsonl` cuts the
same way on every machine and both arms go through the same code path that produces
the frozen baseline. A parameter chosen on the questions it is scored on is fitted,
not measured.

**`npm run eval:threshold`** sweeps `0 / 0.3 / 0.4 / 0.5 / 0.6`, selects on
`validation`, and reports the winner on `test`. It also counts a metric the baseline
does not carry: the **no-result rate**. A higher threshold can look better on a
ranking metric while quietly making the product answer "not in your sources" more
often, and that trade is invisible unless it is counted.

## The result here is a non-result, and it is reported as one

| Threshold | Recall@5 (val) | nDCG@10 (val) | No-result (val) | nDCG@10 (test) | No-result (test) |
| --- | --- | --- | --- | --- | --- |
| 0 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 |
| 0.3 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 |
| 0.4 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 |
| 0.5 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 |
| 0.6 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 |

The sweep is **flat**: every threshold produces identical metrics and never filters a
passage. E5 does not score these query/chunk pairs below 0.6, so on this corpus the
threshold is non-binding. The report says **"no evidence to change
`threshold = 0.5`"** rather than nominating the tie-break winner, because moving a
product parameter on a flat sweep would be noise dressed as a result. The tie-break
rule (widest threshold among equals) is stated so it can be argued with.

That is the honest outcome, and it is another instance of the corpus limitation
#192 describes. The machinery is what ships here; the number is what the corpus
cannot yet support.

## Testing

- `npm run typecheck` — clean
- `npm test` — 489 pass, 4 new: split bounds, split determinism, validation/test
  partition without overlap, `all` preserves order
- `npm run eval` — `docs/eval/baseline-v1.6.json` regenerated with `config.split`
- `npm run eval:threshold` — reruns cleanly and rewrites the same tables

Part of #192 (child 4).
@github-actions github-actions Bot added the enhancement New feature or request label Sep 30, 2026
@mrsibe
mrsibe merged commit 003d467 into feat/eval-metrics-v2 Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant