feat(eval): per-question paired deltas, and hybrid advantage is five questions - #207
Merged
Merged
Conversation
…stions `npm run eval:paired` reports the dense ↔ hybrid pair per question and the classification per query type. It answers the one thing the aggregate comparison could not. The open question was: hybrid is ahead on the full set and level on `validation`, and an average cannot say whether that is a broad small gain or a handful of rescued cases. Those two readings imply different next steps. Child 2 of #192. ## What it found ``` validation 19 questions improved 2 / tied 16 / regressed 1 mean ΔnDCG@10 +0.0047 test 34 questions improved 5 / tied 29 / regressed 0 mean ΔnDCG@10 +0.0432 ``` **The gain is not broad — it is five questions on `test` and two on `validation`.** On the reporting side, 29 of 34 questions are untouched. The mean ΔnDCG@10 of +0.0432 is carried by `q013` (rank 11 → 4, ΔnDCG +0.43), `q059` (5 → 2), `q026` (2 → 1), `q049` (3 → 2) and `q052` (2 → 1). ## And it is not the story we had started telling ourselves The plausible story was "hybrid helps the hard negatives". The breakdown says otherwise: | type | n (test) | improved | tied | regressed | | --- | --- | --- | --- | --- | | cross-lingual | 5 | **0** | 5 | 0 | | hard-negative | 5 | **1** | 4 | 0 | | multi-hop | 3 | 1 | 2 | 0 | | semantic | 15 | **3** | 12 | 0 | | exact | 4 | 0 | 4 | 0 | Three of the five wins are `semantic`, one is `hard-negative` and one `multi-hop`. With n=1 of 5, "hybrid helps hard-negative questions" is **not** supported at this size — which is exactly the pattern that would have been asserted from the aggregate number alone. **`cross-lingual` is untouched on both splits: 0 improved, 0 regressed.** Adding a sparse channel does nothing for the category that is measurably weakest, which is itself a useful negative result and a reason not to reach for hybrid as the answer to it. No question was found by one strategy and missed by the other, on either split, so no regression hides behind the `0 = not found` sentinel. ## How Metrics come from `src/main/eval/metrics.ts` — imported, not reimplemented — so a delta cannot disagree with the metric it is a delta of. `firstRelevantRank` is `0` for "never found", so the miss cases are classified explicitly rather than folded into an arithmetic delta that cannot express them. `validation` is the selecting side; the `test` breakdown **explains** the observed difference and is labelled as not-for-selection in the report itself. Nothing here changes the shipped strategy. ## Testing - `npm run typecheck` — clean - `npm run eval:paired` — 4 harness runs, writes `docs/eval/paired-v1.6.{json,md}` - The committed baseline is unchanged Part of #192 (child 2: corpus diagnostics).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #206 (
feat/eval-score-diagnostics, which this also fixes). Retarget tomainafter the chain merges.What does this PR do?
Adds
npm run eval:paired: the dense ↔ hybrid pair per question, classified per query type. It answers the one thing the aggregate comparison could not.The open question was: hybrid is ahead on the full set and level on
validation— is that a broad small gain or a handful of rescued cases? Those two readings imply different next steps, and a mean over forty questions cannot distinguish them.What it found
The gain is not broad — it is five questions on
testand two onvalidation. On the reporting side, 29 of 34 questions are untouched. The +0.0432 comes fromq013(rank 11 → 4, ΔnDCG +0.43),q059(5 → 2),q026(2 → 1),q049(3 → 2) andq052(2 → 1).And it is not the story we had started telling ourselves
The plausible story was "hybrid helps the hard negatives". The breakdown says otherwise:
Three of the five wins are
semantic, onehard-negative, onemulti-hop. With 1 of 5, "hybrid helps hard-negative questions" is not supported at this size — and that is exactly the kind of claim the aggregate number would have let us assert.cross-lingualis untouched on both splits: 0 improved, 0 regressed. Adding a sparse channel does nothing for the category that is measurably weakest. That is a useful negative result, and a reason not to reach for hybrid as the answer to cross-lingual.No question was found by one strategy and missed by the other on either split, so no regression is hiding behind the
0 = not foundsentinel.On the margin bug you found
Fixed in #206 as a separate commit, since that PR introduced it:
quantileRownow takes the cosine transform as an argument and the margin row passestoCosineMargin(Δcosine = 2·Δscore, the+1cancels), so the two cannot be conflated at the call site again.You were also right that the margin is an oracle quantity — at runtime nothing knows which result is relevant — and the report now says so explicitly, so it cannot be read as a candidate mechanism.
Method
Metrics come from
src/main/eval/metrics.ts— imported, not reimplemented — so a delta cannot disagree with the metric it is a delta of.firstRelevantRankis0for "never found", so miss cases are classified explicitly rather than folded into an arithmetic delta that cannot express them.validationis the selecting side; thetestbreakdown explains the observed difference and is labelled not-for-selection in the report itself. Nothing here changes the shipped strategy.Testing
npm run typecheck— cleannpm test— 504 passnpm run eval:paired— 4 harness runs, writesdocs/eval/paired-v1.6.{json,md}Next
As you set out: not margin-as-abstention. The next step is more near-miss unanswerable questions (validation has 5, which makes abstention move in 20 % steps), then a formal Abstention Signal Evaluation comparing runtime-available signals — top1 score, top1−top2 gap, top1−top3 gap, dense/BM25 agreement, reranker score — with
top1−top2as one candidate rather than the presumed answer.Related
Part of #192. Child 2 (corpus diagnostics).