feat(eval): unanswerable questions, and an explicit split manifest - #202
Merged
mrsibe merged 1 commit intoSep 30, 2026
Merged
Conversation
The corpus contract corpus v2 needs. Without it, two of the four things #192 asks the dataset to cover cannot be represented at all: a question whose right answer is "your sources do not say", and a validation/test assignment that is chosen rather than computed. **Unanswerable questions.** `answerable: false` with `relevant: []`. They are excluded from `metrics` and `byType` — `recallAtK` treats no ground truth as 0/0, and averaging "correctly refused" into "missed" would corrupt the table — and reported as their own group: `noResultCount`, `noResultRate` and `meanRetrieved`. On this row **higher no-result is better**, which is the opposite of how it reads everywhere else. `assertQuestionShape` refuses the reverse combination in either direction, because both are silent: an answerable question with no ground truth reads as a permanent miss, and an unanswerable one carrying ground truth reads as a normal hit. **An explicit split manifest.** `eval/splits.json` replaces the hash of the question id. A hash is reproducible but not *stable*: adding a question moved others between the sides, and a rare query type could end up entirely on one side without anyone choosing that — which is what had happened (`multi-hop` and `cross-lingual` were both entirely on `test`). A question with no manifest entry is now **refused** rather than defaulted, so a new question cannot leak into the reporting side. **Six unanswerable questions** are added (three topic-adjacent, three plainly out-of-scope). Their specifics were checked absent from the corpus before being written, so they are genuinely unanswerable rather than believed to be. ## What the first measurement shows ``` unanswerable: { questions: 6, noResultRate: 0, meanRetrieved: 19 } ``` Every unanswerable question returns **the entire corpus** — 19 of 19 chunks — at `threshold: 0.5`, and the threshold sweep is flat all the way to `0.6`: refusal stays `0/6` and `meanRetrieved` stays at the cap. So on this corpus the threshold cannot separate a relevant passage from an unrelated one, and "Who won the 2018 FIFA World Cup?" is answered with 19 passages of river monitoring, tidal energy and tea storage. That is the evidence the threshold decision did not have before, and it is also the corpus limitation #192 describes: it does not say `0.5` is wrong, it says this corpus cannot tell. The reports say so rather than recommending a change. Retrieval metrics are unchanged (`recallAt5` 1.0000), because the six new questions are correctly kept out of them. ## Testing - `npm run typecheck` — clean - `npm test` — 504 pass, including the manifest contract (unknown side rejected, missing entry refused, `all` needs no manifest), the answerable/unanswerable tie, and the harness's window invariant - `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json` - `npm run eval:retrieval`, `eval:threshold`, `eval:sweep` — all regenerated; the threshold rule now gates on answerable quality and then maximises unanswerable refusal ## Scope This is the **contract**, plus the smallest honest corpus increment that exercises it. Hard negatives and near-duplicate documents — the thing that would bring `Recall@5` off 1.0000 and give `candidateK` something to do — are the next step, not this PR. Part of #192 (child 2, stage 1 of the corpus).
mrsibe
merged commit Sep 30, 2026
1d5df53
into
fix/eval-sweep-context-window-invariant
4 checks passed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #201 (
fix/eval-sweep-context-window-invariant). Retarget tomainafter the chain merges.What does this PR do?
Adds the corpus contract that corpus v2 needs. Without it, two of the four things #192 asks the dataset to cover cannot be represented at all: a question whose right answer is "your sources do not say", and a validation/test assignment that is chosen rather than computed.
Unanswerable questions
answerable: falsewithrelevant: [].metricsandbyType.recallAtKtreats no ground truth as 0/0 rather than 0, and averaging "correctly refused" into "missed" would corrupt the table.noResultCount,noResultRate,meanRetrieved. On this row higher no-result is better, which is the opposite of how it reads everywhere else.assertQuestionShaperefuses the reverse combination in either direction, because both are silent: an answerable question with no ground truth reads as a permanent miss, and an unanswerable one carrying ground truth reads as a normal hit.An explicit split manifest
eval/splits.jsonreplaces the hash of the question id.A hash is reproducible but not stable: adding a question moved others between the sides, and a rare query type could end up entirely on one side without anyone choosing that — which is what had happened. Under the hash,
multi-hop(n=2) andcross-lingual(n=1) were both entirely ontest. The manifest places one of the twomulti-hopquestions and the singlecross-lingualquestion onvalidation, and says so in a file a reviewer can read.A question with no manifest entry is now refused rather than defaulted, so a new question cannot leak into the reporting side. That matters because
testis the side a choice must not be fitted to.Current distribution:
cross-lingualstill cannot be split at n=1. That is what the corpus growth is for; the manifest at least makes the imbalance explicit instead of an accident.Six unanswerable questions
Three topic-adjacent ("Which laboratory published the river monitoring protocol?"), three plainly out-of-scope ("Who won the 2018 FIFA World Cup?"). Their specifics were checked absent from the corpus before being written (
capacity,laboratory,published,price,megawatt,FIFA,mercuryall return nothing), so they are genuinely unanswerable rather than believed to be.What the first measurement shows
Every unanswerable question returns the entire corpus — 19 of 19 chunks — at
threshold: 0.5. The threshold sweep is flat all the way to0.6: refusal stays0/6andmeanRetrievedstays at the cap. In the sweep,meanRetrievedfor unanswerable questions trackscandidateKexactly (5 → 5.0, 20 → 19.0), i.e. the full pool at every setting.So on this corpus the threshold cannot separate a relevant passage from an unrelated one, and "Who won the 2018 FIFA World Cup?" is answered with 19 passages of river monitoring, tidal energy and tea storage.
This is evidence, not a product conclusion. It does not say
0.5is wrong; it says this corpus cannot tell, and the reports say exactly that rather than recommending a change. It is the same corpus limitation #192 describes, now visible on the axis that was previously unmeasurable.Retrieval metrics are unchanged (
recallAt51.0000) because the six new questions are correctly kept out of them.The threshold rule is updated accordingly: hold answerable quality at the
threshold = 0level (validation nDCG@10 and context recall must not regress), then take the threshold that refuses the most unanswerable questions. A refusal gain can never be bought with a retrieval loss.Testing
npm run typecheck— cleannpm test— 504 pass, including the manifest contract (unknown side rejected, missing entry refused,allneeds no manifest), the answerable/unanswerable tie in both directions, andanswerabledefaulting to true for the pre-existing datasetnpm run evaltwice — byte-identicaldocs/eval/baseline-v1.6.jsonnpm run eval:retrieval,eval:threshold,eval:sweep— all regeneratedScope
This is the contract, plus the smallest honest corpus increment that exercises it. Hard negatives and near-duplicate documents — the thing that would bring
Recall@5off 1.0000 and givecandidateKsomething to do — are the next step, not this PR.Related
Part of #192. Child 2, stage 1 of the corpus.