feat(eval): hit rate, MAP, and per-query-type metrics - #195
Merged
Merged
Conversation
Recall@K, MRR and nDCG@10 could not express three things the epic needs, and a single aggregate average was hiding a gap the corpus already contained. Child 3 of #192. **Two metrics added.** `hitRate@5` is whether *any* of the first five passages covers ground truth; `MAP@10` combines ranking position with coverage, so pulling a second relevant passage from rank 9 to rank 2 moves it while Recall@5 sits still. Hit rate is deliberately blunt next to Recall@5: a two-passage question that finds one scores 1.0 and 0.5 respectively, and both facts matter — "the model had a chance" is not "the material was complete". **Every metric is now reported per query type.** `questions.jsonl` gained an optional `type`, the harness groups by it, and the report renders a table. The aggregate was already concealing something: | Type | n | Recall@5 | nDCG@10 | Hit rate@5 | MAP@10 | | --- | --- | --- | --- | --- | --- | | cross-lingual | 1 | 1.0000 | **0.6309** | 1.0000 | **0.5000** | | exact | 6 | 1.0000 | 1.0000 | 1.0000 | 1.0000 | | multi-hop | 2 | 1.0000 | 0.9599 | 1.0000 | 0.9167 | | semantic | 18 | 1.0000 | 0.9312 | 1.0000 | 0.9074 | | zh | 3 | 1.0000 | 1.0000 | 1.0000 | 1.0000 | | **all** | 30 | 1.0000 | 0.9437 | 1.0000 | 0.9222 | The one cross-lingual question (`q030`, Chinese over an English source) ranks far worse than everything else. Recall@5 = 1.0000 reported that as a success; nDCG@10 and MAP@10 are what make the multilingual gap visible. That is the metric doing its job on the existing 30 questions, before the corpus grows. Untagged questions group under `untagged` rather than being dropped, and the types in use are documented in `eval/README.md`. ## Testing - `npm run typecheck` — clean - `npm test` — 485 pass, 4 new: hit rate vs recall, hit rate@k bounds, AP position sensitivity, AP's repeat-counts-once rule - `npm run eval` twice — byte-identical `docs/eval/baseline-v1.6.json` Part of #192 (child 3).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #194 (
feat/eval-config-parity). Retarget tomainafter that merges.What does this PR do?
Adds hit rate and MAP, and reports every metric per query type. Child 3 of #192.
The two metrics
hitRate@5— whether any of the first five passages covers ground truth. Deliberately blunt next to Recall@5: a two-passage question that finds only one scores 1.0 here and 0.5 on Recall@5. Both facts matter — "the model had a chance" is not "the material was complete".mapAt10— mean average precision. The one metric in the harness that combines ranking position with coverage, so pulling a second relevant passage from rank 9 to rank 2 moves it while Recall@5 sits still.Per-query-type reporting
questions.jsonlgained an optionaltype; the harness groups by it and the report renders a table. Untagged questions group underuntaggedrather than disappearing.The aggregate was already hiding a real gap, on the existing 30 questions:
The single cross-lingual question (
q030, Chinese over an English source) ranks far worse than everything else.Recall@5 = 1.0000reported that as a success; nDCG@10 and MAP@10 are what make the multilingual gap visible. That is the metric doing its job before the corpus even grows.Types in use:
exact(number/name/detail),semantic(why/how),multi-hop(two or more blocks),cross-lingual(question language ≠ source language),zh(Chinese over a Chinese source). Documented ineval/README.md.Testing
npm run typecheck— cleannpm test— 485 pass, 4 new: hit rate vs recall, hit-rate@k bounds, AP position sensitivity, AP's repeat-counts-once rulenpm run evaltwice — byte-identicaldocs/eval/baseline-v1.6.jsonRelated
Part of #192. Child 3 (retrieval metrics v2).