Skip to content

feat(eval): per-question paired deltas, and hybrid advantage is five questions - #207

Merged
mrsibe merged 1 commit into
feat/eval-score-diagnosticsfrom
feat/eval-paired-delta
Sep 30, 2026
Merged

mrsibe merged 1 commit into
feat/eval-score-diagnosticsfrom
feat/eval-paired-delta

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #206 (feat/eval-score-diagnostics, which this also fixes). Retarget to main after the chain merges.

What does this PR do?

Adds npm run eval:paired: the dense ↔ hybrid pair per question, classified per query type. It answers the one thing the aggregate comparison could not.

The open question was: hybrid is ahead on the full set and level on validation — is that a broad small gain or a handful of rescued cases? Those two readings imply different next steps, and a mean over forty questions cannot distinguish them.

What it found

validation  19 questions   improved 2  / tied 16 / regressed 1   mean ΔnDCG@10 +0.0047
test        34 questions   improved 5  / tied 29 / regressed 0   mean ΔnDCG@10 +0.0432

The gain is not broad — it is five questions on test and two on validation. On the reporting side, 29 of 34 questions are untouched. The +0.0432 comes from q013 (rank 11 → 4, ΔnDCG +0.43), q059 (5 → 2), q026 (2 → 1), q049 (3 → 2) and q052 (2 → 1).

And it is not the story we had started telling ourselves

The plausible story was "hybrid helps the hard negatives". The breakdown says otherwise:

type n (test) improved tied regressed
cross-lingual 5 0 5 0
hard-negative 5 1 4 0
multi-hop 3 1 2 0
semantic 15 3 12 0
exact 4 0 4 0

Three of the five wins are semantic, one hard-negative, one multi-hop. With 1 of 5, "hybrid helps hard-negative questions" is not supported at this size — and that is exactly the kind of claim the aggregate number would have let us assert.

cross-lingual is untouched on both splits: 0 improved, 0 regressed. Adding a sparse channel does nothing for the category that is measurably weakest. That is a useful negative result, and a reason not to reach for hybrid as the answer to cross-lingual.

No question was found by one strategy and missed by the other on either split, so no regression is hiding behind the 0 = not found sentinel.

On the margin bug you found

Fixed in #206 as a separate commit, since that PR introduced it:

                    before              after
margin p10   -0.0272 (-1.054)     -0.0272 (-0.054)
margin p50   +0.0026 (-0.995)     +0.0026 (+0.005)

quantileRow now takes the cosine transform as an argument and the margin row passes toCosineMargin (Δcosine = 2·Δscore, the +1 cancels), so the two cannot be conflated at the call site again.

You were also right that the margin is an oracle quantity — at runtime nothing knows which result is relevant — and the report now says so explicitly, so it cannot be read as a candidate mechanism.

Method

Metrics come from src/main/eval/metrics.ts — imported, not reimplemented — so a delta cannot disagree with the metric it is a delta of. firstRelevantRank is 0 for "never found", so miss cases are classified explicitly rather than folded into an arithmetic delta that cannot express them.

validation is the selecting side; the test breakdown explains the observed difference and is labelled not-for-selection in the report itself. Nothing here changes the shipped strategy.

Testing

  • npm run typecheck — clean
  • npm test — 504 pass
  • npm run eval:paired — 4 harness runs, writes docs/eval/paired-v1.6.{json,md}
  • Committed baseline unchanged

Next

As you set out: not margin-as-abstention. The next step is more near-miss unanswerable questions (validation has 5, which makes abstention move in 20 % steps), then a formal Abstention Signal Evaluation comparing runtime-available signals — top1 score, top1−top2 gap, top1−top3 gap, dense/BM25 agreement, reranker score — with top1−top2 as one candidate rather than the presumed answer.

Related

Part of #192. Child 2 (corpus diagnostics).

…stions

`npm run eval:paired` reports the dense ↔ hybrid pair per question and the classification
per query type. It answers the one thing the aggregate comparison could not.

The open question was: hybrid is ahead on the full set and level on `validation`, and an
average cannot say whether that is a broad small gain or a handful of rescued cases. Those
two readings imply different next steps. Child 2 of #192.

## What it found

```
validation  19 questions   improved 2  / tied 16 / regressed 1   mean ΔnDCG@10 +0.0047
test        34 questions   improved 5  / tied 29 / regressed 0   mean ΔnDCG@10 +0.0432
```

**The gain is not broad — it is five questions on `test` and two on `validation`.** On the
reporting side, 29 of 34 questions are untouched. The mean ΔnDCG@10 of +0.0432 is carried by
`q013` (rank 11 → 4, ΔnDCG +0.43), `q059` (5 → 2), `q026` (2 → 1), `q049` (3 → 2) and `q052`
(2 → 1).

## And it is not the story we had started telling ourselves

The plausible story was "hybrid helps the hard negatives". The breakdown says otherwise:

| type | n (test) | improved | tied | regressed |
| --- | --- | --- | --- | --- |
| cross-lingual | 5 | **0** | 5 | 0 |
| hard-negative | 5 | **1** | 4 | 0 |
| multi-hop | 3 | 1 | 2 | 0 |
| semantic | 15 | **3** | 12 | 0 |
| exact | 4 | 0 | 4 | 0 |

Three of the five wins are `semantic`, one is `hard-negative` and one `multi-hop`. With n=1
of 5, "hybrid helps hard-negative questions" is **not** supported at this size — which is
exactly the pattern that would have been asserted from the aggregate number alone.

**`cross-lingual` is untouched on both splits: 0 improved, 0 regressed.** Adding a sparse
channel does nothing for the category that is measurably weakest, which is itself a useful
negative result and a reason not to reach for hybrid as the answer to it.

No question was found by one strategy and missed by the other, on either split, so no
regression hides behind the `0 = not found` sentinel.

## How

Metrics come from `src/main/eval/metrics.ts` — imported, not reimplemented — so a delta
cannot disagree with the metric it is a delta of. `firstRelevantRank` is `0` for "never
found", so the miss cases are classified explicitly rather than folded into an arithmetic
delta that cannot express them.

`validation` is the selecting side; the `test` breakdown **explains** the observed
difference and is labelled as not-for-selection in the report itself. Nothing here changes
the shipped strategy.

## Testing

- `npm run typecheck` — clean
- `npm run eval:paired` — 4 harness runs, writes `docs/eval/paired-v1.6.{json,md}`
- The committed baseline is unchanged

Part of #192 (child 2: corpus diagnostics).
@github-actions github-actions Bot added the enhancement New feature or request label Sep 30, 2026
@mrsibe
mrsibe merged commit b2e0b9d into feat/eval-score-diagnostics Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant