Skip to content

feat(eval): nine cross-lingual questions, and the corpus stops being saturated - #203

Merged
mrsibe merged 1 commit into
feat/eval-corpus-contractfrom
feat/eval-cross-lingual-corpus
Sep 30, 2026
Merged

mrsibe merged 1 commit into
feat/eval-corpus-contractfrom
feat/eval-cross-lingual-corpus

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #202 (feat/eval-corpus-contract). Retarget to main after the chain merges.

What does this PR do?

Adds eight Chinese questions over English sources, taking cross-lingual from n=1 to n=9. The ground truth is reused verbatim from the already-verified English counterparts: the fact is the same, only the query language changes, which is exactly what the multilingual model has to bridge.

This directly answers the point you raised — one question cannot support "cross-lingual is a real weakness".

Recall@5 is no longer 1.0000

                 before    after
Recall@5         1.0000    0.9211
nDCG@10          0.9437    0.8476
MAP@10           0.9222    0.7965
hitRate@5        1.0000    0.9211

The benchmark was saturated on this corpus. That is why no retrieval change could clear the adoption rule, and why candidateK appeared to do nothing. Eight questions were enough to bring Recall@5 off the ceiling.

The finding, at a size that supports it

type n Recall@5 nDCG@10 MAP@10
cross-lingual 9 0.6667 0.5032 0.3446
exact 6 1.0000 1.0000 1.0000
multi-hop 2 1.0000 0.9599 0.9167
semantic 18 1.0000 0.9312 0.9074
zh 3 1.0000 1.0000 1.0000

Chinese questions over a Chinese source (zh) still score 1.0000. Chinese questions over an English source score 0.6667.

So the gap is specifically cross-lingual retrieval, not Chinese. q030 alone gave nDCG 0.63 and the honest reading was "a signal worth testing"; at n=9 the same direction holds at the same magnitude, so it is now a defensible finding about the current pipeline on this corpus.

What it does to the other reports

Strategy comparison — now that recallAt5 has headroom, it becomes the deciding metric again, and hybrid does not clear the rule: it matches dense on Recall@5 (0.9211) while improving nDCG@10 (0.8574 vs 0.8476), MRR (0.8141 vs 0.8009) and MAP@10 (0.8097 vs 0.7965).

Worth noting: the amended rule now gives the conservative answer in the unsaturated regime and the permissive one in the saturated regime. Both are reported rather than one being quietly preferred. Whether "better ranking, equal Recall@5, +3 ms p95" should be adopted is a product judgement, and the rule currently says no because a retrieval miss cannot be repaired downstream.

Sweep — candidateK finally moves something: dense nDCG@10 is 0.8212 at candidateK=5 versus 0.8476 at 10 and above. contextK=8 now buys context recall (1.0000 vs 0.9211) at a precision cost (0.3246 → 0.1316), which is the trade the column exists to show.

Threshold — still flat from 0 to 0.6. Even an out-of-scope question scores above 0.6 against every chunk, so the threshold remains unmeasurable on a 19-chunk corpus. This is the clearest remaining argument for the document-level expansion.

Testing

  • npm test — 504 pass
  • npm run eval twice — byte-identical docs/eval/baseline-v1.6.json
  • npm run eval:retrieval, eval:threshold, eval:sweep — all regenerated
  • eval/splits.json extended: 3 of the 8 new questions on validation, 5 on test

Scope

Hard negatives and near-duplicate documents are still not added; this increment is about question coverage over the existing corpus. The corpus is still 19 chunks, which is why threshold cannot be measured and candidateK saturates at 10.

Related

Part of #192. Child 2, stage 2 (cross-lingual coverage).

…saturated

The corpus had exactly **one** cross-lingual question (`q030`, Chinese over an English
source), which is not enough to call anything a weakness — correctly flagged in the #192
review as a signal rather than a finding. This adds eight more Chinese questions over
**English** sources, reusing the already-verified ground truth of their English
counterparts: the fact is the same, only the query language changes, which is precisely
what the multilingual model has to bridge. Cross-lingual goes from n=1 to n=9.

## Recall@5 is no longer 1.0000

```
                 before    after
Recall@5         1.0000    0.9211
nDCG@10          0.9437    0.8476
MAP@10           0.9222    0.7965
hitRate@5        1.0000    0.9211
```

The benchmark was saturated on this corpus, which is why no retrieval change could ever
clear the adoption rule and why `candidateK` appeared to do nothing. Eight questions were
enough to bring Recall@5 off the ceiling.

## The finding, at a size that supports it

| type | n | Recall@5 | nDCG@10 | MAP@10 |
| --- | --- | --- | --- | --- |
| **cross-lingual** | **9** | **0.6667** | **0.5032** | **0.3446** |
| exact | 6 | 1.0000 | 1.0000 | 1.0000 |
| multi-hop | 2 | 1.0000 | 0.9599 | 0.9167 |
| semantic | 18 | 1.0000 | 0.9312 | 0.9074 |
| zh | 3 | 1.0000 | 1.0000 | 1.0000 |

Chinese questions over a **Chinese** source (`zh`) still score 1.0000. Chinese questions
over an **English** source score 0.6667. So the gap is specifically **cross-lingual
retrieval**, not Chinese, and it is now measured on nine questions rather than asserted
from one.

This is the kind of finding that was not defensible before: `q030` alone gave nDCG 0.63
and the honest conclusion was "a signal worth testing". With n=9 the same direction holds
at the same magnitude, so it is a real weakness of the current pipeline on this corpus.

## What it does to the other reports

- **Strategy comparison**: now that `recallAt5` has headroom it becomes the deciding
  metric again, and **hybrid does not clear the rule** — it matches dense on Recall@5
  (0.9211) while improving nDCG@10 (0.8574 vs 0.8476), MRR and MAP@10. So the amended rule
  gives the conservative answer in the unsaturated regime and the permissive one in the
  saturated regime; both are reported rather than one being quietly preferred.
- **Sweep**: `candidateK` finally moves something — dense nDCG@10 0.8212 at `candidateK=5`
  versus 0.8476 at 10 and above. `contextK=8` now buys context recall (1.0000 versus
  0.9211) at a precision cost (0.3246 → 0.1316), which is the trade the column exists for.
- **Threshold**: still flat from 0 to 0.6. Even an out-of-scope question scores above 0.6
  against every chunk, so the threshold remains unmeasurable on a 19-chunk corpus.

## Testing

- `npm test` — 504 pass
- `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json`
- `npm run eval:retrieval`, `eval:threshold`, `eval:sweep` — all regenerated
- `eval/splits.json` extended: 3 of the 8 new questions on `validation`, 5 on `test`

## Scope

Hard negatives and near-duplicate *documents* are still not added; this increment is about
question coverage over the existing corpus. The corpus is still 19 chunks, which is why
`threshold` cannot be measured and `candidateK` saturates at 10.

Part of #192 (child 2, stage 2: cross-lingual coverage).
@github-actions github-actions Bot added the enhancement New feature or request label Sep 30, 2026
@mrsibe
mrsibe merged commit 1fb1945 into feat/eval-corpus-contract Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant