Skip to content

feat(eval): unanswerable questions, and an explicit split manifest - #202

Merged
mrsibe merged 1 commit into
fix/eval-sweep-context-window-invariantfrom
feat/eval-corpus-contract
Sep 30, 2026
Merged

mrsibe merged 1 commit into
fix/eval-sweep-context-window-invariantfrom
feat/eval-corpus-contract

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #201 (fix/eval-sweep-context-window-invariant). Retarget to main after the chain merges.

What does this PR do?

Adds the corpus contract that corpus v2 needs. Without it, two of the four things #192 asks the dataset to cover cannot be represented at all: a question whose right answer is "your sources do not say", and a validation/test assignment that is chosen rather than computed.

Unanswerable questions

answerable: false with relevant: [].

  • Excluded from metrics and byType. recallAtK treats no ground truth as 0/0 rather than 0, and averaging "correctly refused" into "missed" would corrupt the table.
  • Reported as their own group: noResultCount, noResultRate, meanRetrieved. On this row higher no-result is better, which is the opposite of how it reads everywhere else.
  • assertQuestionShape refuses the reverse combination in either direction, because both are silent: an answerable question with no ground truth reads as a permanent miss, and an unanswerable one carrying ground truth reads as a normal hit.

An explicit split manifest

eval/splits.json replaces the hash of the question id.

A hash is reproducible but not stable: adding a question moved others between the sides, and a rare query type could end up entirely on one side without anyone choosing that — which is what had happened. Under the hash, multi-hop (n=2) and cross-lingual (n=1) were both entirely on test. The manifest places one of the two multi-hop questions and the single cross-lingual question on validation, and says so in a file a reviewer can read.

A question with no manifest entry is now refused rather than defaulted, so a new question cannot leak into the reporting side. That matters because test is the side a choice must not be fitted to.

Current distribution:

semantic       validation=5  test=13
exact          validation=2  test=4
multi-hop      validation=1  test=1
zh             validation=1  test=2
cross-lingual  validation=1  test=0
unanswerable   validation=3  test=3

cross-lingual still cannot be split at n=1. That is what the corpus growth is for; the manifest at least makes the imbalance explicit instead of an accident.

Six unanswerable questions

Three topic-adjacent ("Which laboratory published the river monitoring protocol?"), three plainly out-of-scope ("Who won the 2018 FIFA World Cup?"). Their specifics were checked absent from the corpus before being written (capacity, laboratory, published, price, megawatt, FIFA, mercury all return nothing), so they are genuinely unanswerable rather than believed to be.

What the first measurement shows

unanswerable: { questions: 6, noResultRate: 0, meanRetrieved: 19 }

Every unanswerable question returns the entire corpus — 19 of 19 chunks — at threshold: 0.5. The threshold sweep is flat all the way to 0.6: refusal stays 0/6 and meanRetrieved stays at the cap. In the sweep, meanRetrieved for unanswerable questions tracks candidateK exactly (5 → 5.0, 20 → 19.0), i.e. the full pool at every setting.

So on this corpus the threshold cannot separate a relevant passage from an unrelated one, and "Who won the 2018 FIFA World Cup?" is answered with 19 passages of river monitoring, tidal energy and tea storage.

This is evidence, not a product conclusion. It does not say 0.5 is wrong; it says this corpus cannot tell, and the reports say exactly that rather than recommending a change. It is the same corpus limitation #192 describes, now visible on the axis that was previously unmeasurable.

Retrieval metrics are unchanged (recallAt5 1.0000) because the six new questions are correctly kept out of them.

The threshold rule is updated accordingly: hold answerable quality at the threshold = 0 level (validation nDCG@10 and context recall must not regress), then take the threshold that refuses the most unanswerable questions. A refusal gain can never be bought with a retrieval loss.

Testing

  • npm run typecheck — clean
  • npm test — 504 pass, including the manifest contract (unknown side rejected, missing entry refused, all needs no manifest), the answerable/unanswerable tie in both directions, and answerable defaulting to true for the pre-existing dataset
  • npm run eval twice — byte-identical docs/eval/baseline-v1.6.json
  • npm run eval:retrieval, eval:threshold, eval:sweep — all regenerated

Scope

This is the contract, plus the smallest honest corpus increment that exercises it. Hard negatives and near-duplicate documents — the thing that would bring Recall@5 off 1.0000 and give candidateK something to do — are the next step, not this PR.

Related

Part of #192. Child 2, stage 1 of the corpus.

The corpus contract corpus v2 needs. Without it, two of the four things #192 asks the
dataset to cover cannot be represented at all: a question whose right answer is "your
sources do not say", and a validation/test assignment that is chosen rather than
computed.

**Unanswerable questions.** `answerable: false` with `relevant: []`. They are excluded
from `metrics` and `byType` — `recallAtK` treats no ground truth as 0/0, and averaging
"correctly refused" into "missed" would corrupt the table — and reported as their own
group: `noResultCount`, `noResultRate` and `meanRetrieved`. On this row **higher
no-result is better**, which is the opposite of how it reads everywhere else.
`assertQuestionShape` refuses the reverse combination in either direction, because both
are silent: an answerable question with no ground truth reads as a permanent miss, and
an unanswerable one carrying ground truth reads as a normal hit.

**An explicit split manifest.** `eval/splits.json` replaces the hash of the question id.
A hash is reproducible but not *stable*: adding a question moved others between the
sides, and a rare query type could end up entirely on one side without anyone choosing
that — which is what had happened (`multi-hop` and `cross-lingual` were both entirely on
`test`). A question with no manifest entry is now **refused** rather than defaulted, so a
new question cannot leak into the reporting side.

**Six unanswerable questions** are added (three topic-adjacent, three plainly
out-of-scope). Their specifics were checked absent from the corpus before being written,
so they are genuinely unanswerable rather than believed to be.

## What the first measurement shows

```
unanswerable: { questions: 6, noResultRate: 0, meanRetrieved: 19 }
```

Every unanswerable question returns **the entire corpus** — 19 of 19 chunks — at
`threshold: 0.5`, and the threshold sweep is flat all the way to `0.6`: refusal stays
`0/6` and `meanRetrieved` stays at the cap. So on this corpus the threshold cannot
separate a relevant passage from an unrelated one, and "Who won the 2018 FIFA World Cup?"
is answered with 19 passages of river monitoring, tidal energy and tea storage.

That is the evidence the threshold decision did not have before, and it is also the
corpus limitation #192 describes: it does not say `0.5` is wrong, it says this corpus
cannot tell. The reports say so rather than recommending a change.

Retrieval metrics are unchanged (`recallAt5` 1.0000), because the six new questions are
correctly kept out of them.

## Testing

- `npm run typecheck` — clean
- `npm test` — 504 pass, including the manifest contract (unknown side rejected, missing
  entry refused, `all` needs no manifest), the answerable/unanswerable tie, and the
  harness's window invariant
- `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json`
- `npm run eval:retrieval`, `eval:threshold`, `eval:sweep` — all regenerated; the
  threshold rule now gates on answerable quality and then maximises unanswerable refusal

## Scope

This is the **contract**, plus the smallest honest corpus increment that exercises it.
Hard negatives and near-duplicate documents — the thing that would bring `Recall@5` off
1.0000 and give `candidateK` something to do — are the next step, not this PR.

Part of #192 (child 2, stage 1 of the corpus).
@github-actions github-actions Bot added the enhancement New feature or request label Sep 30, 2026
@mrsibe
mrsibe merged commit 1d5df53 into fix/eval-sweep-context-window-invariant Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant