Skip to content

feat(eval): hit rate, MAP, and per-query-type metrics - #195

Merged
mrsibe merged 1 commit into
feat/eval-config-parityfrom
feat/eval-metrics-v2
Sep 30, 2026
Merged

mrsibe merged 1 commit into
feat/eval-config-parityfrom
feat/eval-metrics-v2

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #194 (feat/eval-config-parity). Retarget to main after that merges.

What does this PR do?

Adds hit rate and MAP, and reports every metric per query type. Child 3 of #192.

The two metrics

  • hitRate@5 — whether any of the first five passages covers ground truth. Deliberately blunt next to Recall@5: a two-passage question that finds only one scores 1.0 here and 0.5 on Recall@5. Both facts matter — "the model had a chance" is not "the material was complete".
  • mapAt10 — mean average precision. The one metric in the harness that combines ranking position with coverage, so pulling a second relevant passage from rank 9 to rank 2 moves it while Recall@5 sits still.

Per-query-type reporting

questions.jsonl gained an optional type; the harness groups by it and the report renders a table. Untagged questions group under untagged rather than disappearing.

The aggregate was already hiding a real gap, on the existing 30 questions:

Type n Recall@5 nDCG@10 Hit rate@5 MAP@10
cross-lingual 1 1.0000 0.6309 1.0000 0.5000
exact 6 1.0000 1.0000 1.0000 1.0000
multi-hop 2 1.0000 0.9599 1.0000 0.9167
semantic 18 1.0000 0.9312 1.0000 0.9074
zh 3 1.0000 1.0000 1.0000 1.0000
all 30 1.0000 0.9437 1.0000 0.9222

The single cross-lingual question (q030, Chinese over an English source) ranks far worse than everything else. Recall@5 = 1.0000 reported that as a success; nDCG@10 and MAP@10 are what make the multilingual gap visible. That is the metric doing its job before the corpus even grows.

Types in use: exact (number/name/detail), semantic (why/how), multi-hop (two or more blocks), cross-lingual (question language ≠ source language), zh (Chinese over a Chinese source). Documented in eval/README.md.

Testing

  • npm run typecheck — clean
  • npm test — 485 pass, 4 new: hit rate vs recall, hit-rate@k bounds, AP position sensitivity, AP's repeat-counts-once rule
  • npm run eval twice — byte-identical docs/eval/baseline-v1.6.json

Related

Part of #192. Child 3 (retrieval metrics v2).

Recall@K, MRR and nDCG@10 could not express three things the epic needs, and a
single aggregate average was hiding a gap the corpus already contained. Child 3 of
#192.

**Two metrics added.** `hitRate@5` is whether *any* of the first five passages
covers ground truth; `MAP@10` combines ranking position with coverage, so pulling a
second relevant passage from rank 9 to rank 2 moves it while Recall@5 sits still.
Hit rate is deliberately blunt next to Recall@5: a two-passage question that finds
one scores 1.0 and 0.5 respectively, and both facts matter — "the model had a
chance" is not "the material was complete".

**Every metric is now reported per query type.** `questions.jsonl` gained an optional
`type`, the harness groups by it, and the report renders a table. The aggregate was
already concealing something:

| Type | n | Recall@5 | nDCG@10 | Hit rate@5 | MAP@10 |
| --- | --- | --- | --- | --- | --- |
| cross-lingual | 1 | 1.0000 | **0.6309** | 1.0000 | **0.5000** |
| exact | 6 | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| multi-hop | 2 | 1.0000 | 0.9599 | 1.0000 | 0.9167 |
| semantic | 18 | 1.0000 | 0.9312 | 1.0000 | 0.9074 |
| zh | 3 | 1.0000 | 1.0000 | 1.0000 | 1.0000 |
| **all** | 30 | 1.0000 | 0.9437 | 1.0000 | 0.9222 |

The one cross-lingual question (`q030`, Chinese over an English source) ranks far
worse than everything else. Recall@5 = 1.0000 reported that as a success; nDCG@10
and MAP@10 are what make the multilingual gap visible. That is the metric doing its
job on the existing 30 questions, before the corpus grows.

Untagged questions group under `untagged` rather than being dropped, and the types
in use are documented in `eval/README.md`.

## Testing

- `npm run typecheck` — clean
- `npm test` — 485 pass, 4 new: hit rate vs recall, hit rate@k bounds, AP position
  sensitivity, AP's repeat-counts-once rule
- `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json`

Part of #192 (child 3).
@github-actions github-actions Bot added the enhancement New feature or request label Sep 30, 2026
@mrsibe
mrsibe merged commit 2255096 into feat/eval-config-parity Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant