feat(eval): confusion clusters, and the corpus reaches 53 chunks - #205
Merged
mrsibe merged 2 commits intoSep 30, 2026
Merged
Conversation
…egatives
The first of the confusion clusters corpus v2 is built from. Not "near-duplicate documents
that lower Recall@5" — deliberately confusable material, so the benchmark can tell retrieval
strategies apart at all.
## The cluster
Four documents that share almost all their vocabulary and differ in every number:
| | v1 | v2 | v3 |
| --- | --- | --- | --- |
| Path | `/v1/complete` | `/v2/generate` | `/v3/chat` |
| Default timeout | 30000 ms | 60000 ms | 45000 ms |
| Recommended retries | 2 | 5 | 3 |
| Backoff base | 500 ms | 2000 ms | 1000 ms |
| Rate limit | 600 rpm | 3000 rpm | 1200 rpm |
| Context window | 8192 tokens | 32768 tokens | 65536 tokens |
All four documents discuss *timeout, retry, backoff, rate limit and context window* in that
order, so a query naming one field matches four passages and only one is right. A migration
guide restates every number in comparison tables, which is the hardest negative of the set:
it contains v1's timeout **and** v2's **and** v3's in one block.
## The tool that made it possible
`npm run eval:blocks <file.md>` prints the real `document_blocks.order` values through the
same loader and block builder the harness uses. Ground truth `block` is a pipeline ordinal,
not a line number a human counted; authoring a corpus by guessing it is how a dataset
quietly drifts. `eval/README.md` now says so and points at the command.
## What it measured
```
before after
chunks 19 35
Recall@5 0.9211 0.8913
nDCG@10 0.8476 0.7915
MAP@10 0.7965 0.7361
```
By type, the new cluster behaves as designed:
| type | n | Recall@5 | nDCG@10 |
| --- | --- | --- | --- |
| hard-negative | 4 | **1.0000** | **0.7827** |
| multi-hop | 4 | 0.7500 | 0.7289 |
| semantic | 19 | 0.9474 | 0.8974 |
| cross-lingual | 9 | 0.6667 | 0.4311 |
`hard-negative` has perfect Recall@5 and a poor nDCG, which is the exact signature of this
kind of material: the correct passage **is** in the top 5, but a confusable sibling outranks
it. That is a ranking problem the strategy comparison can now see, and it is invisible on a
corpus where the right answer is always at rank 1.
`multi-hop` fell to 0.75 because the two comparison questions need four locations each and
do not get all four.
Also added: 8 questions in total (2 exact, 4 hard-negative, 1 semantic, 2 multi-hop, and two
**near-miss unanswerable** — "How much GPU memory does Gateway API v2 require?" and "What is
the monthly subscription price of Gateway API v3?", both about documents that discuss the
subject at length and never mention the answer). The FIFA questions were a sanity check for
a totally out-of-scope query; these are the ones that actually test abstention, because the
subject *is* in the library.
Unanswerable abstention is still 0/8, with `meanCandidatesRetrieved` at the `candidateK`
cap of 20 and `meanContextPassages` at 3.
## Scope, honestly
**35 chunks, not 150–300.** One cluster is not the target. The mechanism is proven and
verified end to end, and the remaining clusters are the same exercise repeated — but this PR
does not pretend the index is large enough for `candidateK ∈ {5,10,20,40}` to be a real
sweep. That arrives with the remaining clusters.
Part of #192 (child 2, stage 3: first confusion cluster).
Four more deliberately confusable documents, this time in the monitoring family the corpus
already had two members of, so the cluster is confusable with the *existing* documents as
well as internally.
| | reservoir | coastal | estuary | groundwater |
| --- | --- | --- | --- | --- |
| Sampling interval | fortnightly | hourly | daily | monthly |
| Replicates | 4 | 3 | 5 | 2 |
| Sensor depth | 5 m below surface | 1 m below surface | 2 m below surface | 15 m below water table |
Every document discusses *sampling interval, replicate samples and sensor depth* in the same
order, and each states the network's lowest or highest value ("the least frequent cadence in
the network", "the highest of any programme"), so a query lands on several plausible passages
and only one is right.
Corpus: 19 → 53 chunks, 13 → 21 documents, 54 → 63 questions.
## What it measured
```
after cluster 1 after cluster 2
chunks 35 53
Recall@5 0.8913 0.8774
nDCG@10 0.7915 0.7804
MAP@10 0.7361 0.7280
```
| type | n | Recall@5 | nDCG@10 |
| --- | --- | --- | --- |
| cross-lingual | 9 | 0.6667 | 0.4173 |
| multi-hop | 5 | 0.7000 | 0.7218 |
| semantic | 21 | 0.9048 | 0.8542 |
| exact | 8 | 1.0000 | 0.8663 |
| hard-negative | 7 | **1.0000** | **0.8758** |
| zh | 3 | 1.0000 | 1.0000 |
`hard-negative` keeps the signature the cluster was built for: the correct passage is always
in the top 5 and a confusable sibling keeps outranking it.
## A result worth pausing on
The sweep (all 63 questions) now shows **hybrid ahead of dense on the deciding metric**:
| | Recall@5 | nDCG@10 |
| --- | --- | --- |
| dense, candidateK=20 | 0.8774 | 0.7804 |
| hybrid, candidateK=20 | **0.8962** | **0.8098** |
| hybrid, candidateK=5 | **0.9104** | 0.8101 |
But on the **validation** split the same comparison is a dead heat (both 0.8684), so the
held-out rule still declines to adopt.
That discrepancy is the honest state of things, and it is a property of the split rather
than of hybrid: 24 validation questions is not enough to see a gain that the full set shows.
The conclusion is to grow the validation split, **not** to go back to selecting on
everything — which is what produced the earlier, over-confident "hybrid clears the rule".
## Unanswerable is still 0/10
`meanCandidatesRetrieved` sits at the `candidateK` cap and `meanContextPassages` at
`contextK` for every unanswerable question, so abstention remains unmeasurable here. Two of
the ten are near-miss questions ("How much GPU memory does Gateway API v2 require?", "How
many litres per second does the reservoir release downstream?") about subjects the corpus
discusses at length — the realistic hallucination shape rather than the FIFA sanity check.
## Scope, honestly
**53 chunks, not 150–300.** Two clusters is not the target either. `candidateK ∈
{5,10,20,40}` now has more room than it did at 19 chunks, but `candidateK ≥ 10` still
behaves identically, so the axis is only partly unlocked. The remaining clusters are the same
exercise and the same verification loop; this PR stops where the verified work stops.
Part of #192 (child 2, stage 3: confusion clusters).
mrsibe
force-pushed
the
fix/eval-review-followups
branch
from
September 30, 2026 10:07
22f070c to
84e9e0a
Compare
mrsibe
force-pushed
the
feat/eval-corpus-hard-negatives
branch
from
September 30, 2026 10:07
d171f8f to
3ffee33
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #204 (
fix/eval-review-followups). Retarget tomainafter the chain merges.What does this PR do?
The first two confusion clusters of corpus v2 — deliberately confusable material, not "near-duplicate documents that lower Recall@5". The goal is a benchmark that can tell retrieval strategies apart at all, and lowering Recall is a consequence of that rather than the objective.
The clusters
1. Gateway API v1 / v2 / v3 + a migration guide. Four documents sharing almost all their vocabulary and differing in every number:
All four discuss timeout, retry, backoff, rate limit, context window in that order. The migration guide restates every number in comparison tables, so it contains v1's timeout and v2's and v3's in a single block — the hardest negative in the set.
2. Monitoring protocols. Four documents that are confusable with each other and with the
river-monitoring.md/lake-monitoring.mdalready in the corpus:Each states the network's extreme value ("the least frequent cadence in the network", "the highest of any programme"), so several passages read as plausible and only one is right.
Corpus: 19 → 53 chunks, 13 → 21 documents, 44 → 63 questions.
The tool that made it possible
npm run eval:blocks <file.md>prints the realdocument_blocks.ordervalues through the same loader and block builder the harness uses. Ground truthblockis a pipeline ordinal, not a line number a human counted; authoring a corpus by guessing it is how a dataset quietly drifts.eval/README.mdnow says so and points at the command.What it measured
hard-negativehas perfect Recall@5 and a much worse nDCG — the signature the material was built for: the correct passage is in the top 5, and a confusable sibling keeps outranking it. That is a ranking problem the strategy comparison can now see, and it is invisible on a corpus where the right answer is always at rank 1.multi-hopfell to 0.7000: the comparison questions need four locations each and do not get all four.A result worth pausing on
The sweep (all 63 questions) now shows hybrid ahead of dense on the deciding metric:
But on the validation split the same comparison is a dead heat (both 0.8684), so the held-out rule still declines to adopt and
testreports dense only.That discrepancy is the honest state of things, and it is a property of the split, not of hybrid: 24 validation questions is not enough to see a gain the full set shows. The conclusion is to grow the validation split, not to go back to selecting on everything — which is exactly what produced the earlier, over-confident "hybrid clears the rule".
Unanswerable is still 0/10
meanCandidatesRetrievedsits at thecandidateKcap andmeanContextPassagesatcontextKfor every unanswerable question, so abstention remains unmeasurable here. Two of the ten are near-miss questions — "How much GPU memory does Gateway API v2 require?" and "How many litres per second does the reservoir release downstream?" — about subjects the corpus discusses at length and never answers. Those are the realistic hallucination shape; the FIFA questions were only ever a sanity check.Scope, honestly
53 chunks, not 150–300. Two clusters is not the target either.
candidateK ∈ {5,10,20,40}has more room than it did at 19 chunks, butcandidateK ≥ 10still behaves identically, so the axis is only partly unlocked andthresholdremains unmeasurable.The remaining clusters are the same exercise and the same verification loop. This PR stops where the verified work stops.
Testing
npm test— 504 passnpm run evaltwice — byte-identical baseline (63 questions, 53 chunks)npm run eval:retrieval,eval:threshold,eval:sweep— all regeneratedeval/splits.jsonextended: 24 validation / 39 test, stratified by typeRelated
Part of #192. Child 2, stage 3 (confusion clusters).