feat(eval): nine cross-lingual questions, and the corpus stops being saturated - #203
Merged
Merged
Conversation
…saturated The corpus had exactly **one** cross-lingual question (`q030`, Chinese over an English source), which is not enough to call anything a weakness — correctly flagged in the #192 review as a signal rather than a finding. This adds eight more Chinese questions over **English** sources, reusing the already-verified ground truth of their English counterparts: the fact is the same, only the query language changes, which is precisely what the multilingual model has to bridge. Cross-lingual goes from n=1 to n=9. ## Recall@5 is no longer 1.0000 ``` before after Recall@5 1.0000 0.9211 nDCG@10 0.9437 0.8476 MAP@10 0.9222 0.7965 hitRate@5 1.0000 0.9211 ``` The benchmark was saturated on this corpus, which is why no retrieval change could ever clear the adoption rule and why `candidateK` appeared to do nothing. Eight questions were enough to bring Recall@5 off the ceiling. ## The finding, at a size that supports it | type | n | Recall@5 | nDCG@10 | MAP@10 | | --- | --- | --- | --- | --- | | **cross-lingual** | **9** | **0.6667** | **0.5032** | **0.3446** | | exact | 6 | 1.0000 | 1.0000 | 1.0000 | | multi-hop | 2 | 1.0000 | 0.9599 | 0.9167 | | semantic | 18 | 1.0000 | 0.9312 | 0.9074 | | zh | 3 | 1.0000 | 1.0000 | 1.0000 | Chinese questions over a **Chinese** source (`zh`) still score 1.0000. Chinese questions over an **English** source score 0.6667. So the gap is specifically **cross-lingual retrieval**, not Chinese, and it is now measured on nine questions rather than asserted from one. This is the kind of finding that was not defensible before: `q030` alone gave nDCG 0.63 and the honest conclusion was "a signal worth testing". With n=9 the same direction holds at the same magnitude, so it is a real weakness of the current pipeline on this corpus. ## What it does to the other reports - **Strategy comparison**: now that `recallAt5` has headroom it becomes the deciding metric again, and **hybrid does not clear the rule** — it matches dense on Recall@5 (0.9211) while improving nDCG@10 (0.8574 vs 0.8476), MRR and MAP@10. So the amended rule gives the conservative answer in the unsaturated regime and the permissive one in the saturated regime; both are reported rather than one being quietly preferred. - **Sweep**: `candidateK` finally moves something — dense nDCG@10 0.8212 at `candidateK=5` versus 0.8476 at 10 and above. `contextK=8` now buys context recall (1.0000 versus 0.9211) at a precision cost (0.3246 → 0.1316), which is the trade the column exists for. - **Threshold**: still flat from 0 to 0.6. Even an out-of-scope question scores above 0.6 against every chunk, so the threshold remains unmeasurable on a 19-chunk corpus. ## Testing - `npm test` — 504 pass - `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json` - `npm run eval:retrieval`, `eval:threshold`, `eval:sweep` — all regenerated - `eval/splits.json` extended: 3 of the 8 new questions on `validation`, 5 on `test` ## Scope Hard negatives and near-duplicate *documents* are still not added; this increment is about question coverage over the existing corpus. The corpus is still 19 chunks, which is why `threshold` cannot be measured and `candidateK` saturates at 10. Part of #192 (child 2, stage 2: cross-lingual coverage).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #202 (
feat/eval-corpus-contract). Retarget tomainafter the chain merges.What does this PR do?
Adds eight Chinese questions over English sources, taking cross-lingual from n=1 to n=9. The ground truth is reused verbatim from the already-verified English counterparts: the fact is the same, only the query language changes, which is exactly what the multilingual model has to bridge.
This directly answers the point you raised — one question cannot support "cross-lingual is a real weakness".
Recall@5 is no longer 1.0000
The benchmark was saturated on this corpus. That is why no retrieval change could clear the adoption rule, and why
candidateKappeared to do nothing. Eight questions were enough to bring Recall@5 off the ceiling.The finding, at a size that supports it
Chinese questions over a Chinese source (
zh) still score 1.0000. Chinese questions over an English source score 0.6667.So the gap is specifically cross-lingual retrieval, not Chinese.
q030alone gave nDCG 0.63 and the honest reading was "a signal worth testing"; at n=9 the same direction holds at the same magnitude, so it is now a defensible finding about the current pipeline on this corpus.What it does to the other reports
Strategy comparison — now that
recallAt5has headroom, it becomes the deciding metric again, and hybrid does not clear the rule: it matches dense on Recall@5 (0.9211) while improving nDCG@10 (0.8574 vs 0.8476), MRR (0.8141 vs 0.8009) and MAP@10 (0.8097 vs 0.7965).Worth noting: the amended rule now gives the conservative answer in the unsaturated regime and the permissive one in the saturated regime. Both are reported rather than one being quietly preferred. Whether "better ranking, equal Recall@5, +3 ms p95" should be adopted is a product judgement, and the rule currently says no because a retrieval miss cannot be repaired downstream.
Sweep —
candidateKfinally moves something: dense nDCG@10 is 0.8212 atcandidateK=5versus 0.8476 at 10 and above.contextK=8now buys context recall (1.0000 vs 0.9211) at a precision cost (0.3246 → 0.1316), which is the trade the column exists to show.Threshold — still flat from 0 to 0.6. Even an out-of-scope question scores above 0.6 against every chunk, so the threshold remains unmeasurable on a 19-chunk corpus. This is the clearest remaining argument for the document-level expansion.
Testing
npm test— 504 passnpm run evaltwice — byte-identicaldocs/eval/baseline-v1.6.jsonnpm run eval:retrieval,eval:threshold,eval:sweep— all regeneratedeval/splits.jsonextended: 3 of the 8 new questions onvalidation, 5 ontestScope
Hard negatives and near-duplicate documents are still not added; this increment is about question coverage over the existing corpus. The corpus is still 19 chunks, which is why
thresholdcannot be measured andcandidateKsaturates at 10.Related
Part of #192. Child 2, stage 2 (cross-lingual coverage).