feat(eval): derive the similarity threshold on validation, report on test - #196
Merged
Merged
Conversation
…test `threshold: 0.5` was hand-picked, and a cosine score has no universal meaning — its distribution depends on the embedding model, the language, the query type and the chunk length. Child 4 of #192. **A deterministic split, owned by the harness.** `--eval-split=validation|test` partitions the questions by a hash of the id, so the same `questions.jsonl` cuts the same way on every machine and both arms go through the same code path that produces the frozen baseline. A parameter chosen on the questions it is scored on is fitted, not measured. **`npm run eval:threshold`** sweeps `0 / 0.3 / 0.4 / 0.5 / 0.6`, selects on `validation`, and reports the winner on `test`. It also counts a metric the baseline does not carry: the **no-result rate**. A higher threshold can look better on a ranking metric while quietly making the product answer "not in your sources" more often, and that trade is invisible unless it is counted. ## The result here is a non-result, and it is reported as one | Threshold | Recall@5 (val) | nDCG@10 (val) | No-result (val) | nDCG@10 (test) | No-result (test) | | --- | --- | --- | --- | --- | --- | | 0 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 | | 0.3 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 | | 0.4 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 | | 0.5 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 | | 0.6 | 1.0000 | 0.9500 | 0.0000 | 0.9406 | 0.0000 | The sweep is **flat**: every threshold produces identical metrics and never filters a passage. E5 does not score these query/chunk pairs below 0.6, so on this corpus the threshold is non-binding. The report says **"no evidence to change `threshold = 0.5`"** rather than nominating the tie-break winner, because moving a product parameter on a flat sweep would be noise dressed as a result. The tie-break rule (widest threshold among equals) is stated so it can be argued with. That is the honest outcome, and it is another instance of the corpus limitation #192 describes. The machinery is what ships here; the number is what the corpus cannot yet support. ## Testing - `npm run typecheck` — clean - `npm test` — 489 pass, 4 new: split bounds, split determinism, validation/test partition without overlap, `all` preserves order - `npm run eval` — `docs/eval/baseline-v1.6.json` regenerated with `config.split` - `npm run eval:threshold` — reruns cleanly and rewrites the same tables Part of #192 (child 4).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #195 (
feat/eval-metrics-v2). Retarget tomainafter that merges.What does this PR do?
Derives the similarity threshold on a held-out split instead of hand-picking it, and adds a no-result-rate metric that a threshold sweep needs. Child 4 of #192.
The mechanism
A deterministic split, owned by the harness.
--eval-split=validation|testpartitions questions by a hash of the id, so the samequestions.jsonlcuts the same way on every machine and both arms go through the same code path that produces the frozen baseline. A parameter chosen on the questions it is scored on is fitted, not measured.npm run eval:thresholdsweeps0 / 0.3 / 0.4 / 0.5 / 0.6, selects onvalidation, and reports the winner ontest.The no-result rate. A higher threshold can look better on a ranking metric while quietly making the product answer "not in your sources" more often. That trade is invisible unless it is counted, so the sweep counts it on both splits.
The result here is a non-result, and it is reported as one
The sweep is flat: every threshold produces identical metrics and never filters a passage. E5 does not score these query/chunk pairs below 0.6, so on this corpus the threshold is non-binding.
The report says "no evidence to change
threshold = 0.5" rather than nominating the tie-break winner. The rule nominates0only because it prefers the widest threshold among equals; that is a tie-break, not a finding. Moving a product parameter on a flat sweep would be noise dressed as a result. The selection rule is stated in the output so it can be argued with.This is the honest outcome, and it is another instance of the corpus limitation #192 describes. The mechanism is what ships here; the number is what the corpus cannot yet support.
Testing
npm run typecheck— cleannpm test— 489 pass, 4 new: split bounds, split determinism, validation/test partition without overlap,allpreserves ordernpm run eval—docs/eval/baseline-v1.6.jsonregenerated withconfig.splitnpm run eval:threshold— re-runs cleanly and rewrites the same tablesRelated
Part of #192. Child 4 (re-derive threshold).