feat(eval): fifteen near-miss unanswerable questions — closes Phase 1 - #208
Merged
Merged
Conversation
…al PR Closes RAG Eval v2 Phase 1. This is the final expansion of the eval system — the conclusions below are the point of the phase, not more infrastructure. ## What was added Fifteen unanswerable questions whose **subject is discussed at length in the corpus and whose answer is not there** — "How much GPU memory does Gateway API v3 need?", "Which laboratory is accredited to analyse the estuary transect samples?", "How many staff work on the reservoir monitoring programme?". Every specific term was checked absent before the question was written (`gpu`, `vram`, `sla`, `uptime`, `expiry`, `sdk`, `self-hosted`, `accredit`, `vendor`, `supplier`, `funding`, `grant`, `staff`, `warranty`, `retention`, `encrypt`, `pricing`, `invoice`, `audit` all return nothing). That check is what makes them unanswerable rather than believed to be, and it is the part that cannot be automated away. `validation` unanswerable goes from **5 to 15**, so abstention resolution goes from 20 % per question to 6.7 %. The earlier "Who won the 2018 FIFA World Cup?" questions were a sanity check for a query that is obviously out of scope; these are the realistic shape — the user's actual hallucination risk is asking a question about a document that *is* in the library. ## The conclusion survives the larger sample ``` unanswerable max candidate min 0.8978 (cosine 0.796) p50 0.9279 (0.856) max 0.9435 (0.887) worst required relevant p10 0.8906 (cosine 0.781) ``` Even the **lowest** unanswerable top score is above the p10 of the worst required relevant passage. With n=15 instead of n=5 the distributions still do not separate, and the threshold curve is unchanged in shape: | threshold | raw cosine | answerable full recall | unanswerable abstain | | --- | --- | --- | --- | | 0.875 | 0.750 | 1.0000 | 0.0000 | | 0.900 | 0.800 | 0.8947 | 0.0667 | | 0.925 | 0.850 | 0.6316 | 0.4000 | | 0.950 | 0.900 | 0.1579 | 1.0000 | Abstention still never rises without full recall falling. The finer resolution makes the finding *more* solid, not different. Regenerated to stay consistent: baseline, scores, threshold and sweep. `paired-v1.6` is unchanged, because adding unanswerable questions does not touch any answerable one. ## Where Phase 1 leaves the product 1. The original 19-chunk benchmark could not guide a product decision; the current one (53 chunks, 78 questions, per-type metrics, held-out split) can. 2. **Cross-lingual retrieval is the largest measured gap**: Recall@5 0.6667 and nDCG@10 0.4173 against 0.93–1.00 for same-language questions. Chinese over a Chinese source scores 1.0000, so it is the cross-language matching, not Chinese. 3. **Hybrid has ranking value and no held-out mandate.** Its full-set advantage is five questions out of 34 (`paired-v1.6`), mostly `semantic`, and it does nothing for cross-lingual. Dense stays the default. 4. `contextK = 3` is justified against 5 and 8 on this corpus; `candidateK` above 10 is not distinguishable. 5. `threshold = 0.5` means **cosine ≥ 0**, not cosine ≥ 0.5. 6. **A single embedding similarity threshold cannot decide answerability** — the relevant and non-relevant and unanswerable score distributions overlap, and every threshold that abstains also drops required evidence. (6) is the phase's real result: it says stop tuning this knob and reach for a different mechanism. ## Stop line Phase 1 ends here. Not done, and deliberately not started: public benchmarks, reranker eval, generator faithfulness, citation entailment, growing the corpus to 300 chunks, an `Abstention Signal Evaluation`. Those build a more complete benchmark; they do not unblock KnowNote development, and (6) already says the next useful work is a different retrieval mechanism rather than a better measurement of this one.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #207. The last eval PR of RAG Eval v2 Phase 1.
What was added
Fifteen unanswerable questions whose subject is discussed at length in the corpus and whose answer is not there — "How much GPU memory does Gateway API v3 need?", "Which laboratory is accredited to analyse the estuary transect samples?", "How many staff work on the reservoir monitoring programme?".
Every specific term was checked absent before the question was written (
gpu,vram,sla,uptime,expiry,sdk,self-hosted,accredit,vendor,supplier,funding,grant,staff,warranty,retention,encrypt,pricing,invoice,auditall return nothing). That check is what makes them unanswerable rather than believed to be, and it is the part that cannot be automated away.validationunanswerable goes from 5 to 15, so abstention resolution goes from 20 % per question to 6.7 %. The FIFA questions were a sanity check for an obviously out-of-scope query; these are the realistic shape — the actual hallucination risk is a question about a document that is in the library.The conclusion survives the larger sample
Even the lowest unanswerable top score sits above the p10 of the worst required relevant passage. With n=15 instead of n=5 the distributions still do not separate, and the threshold curve is unchanged in shape:
Abstention still never rises without full recall falling. The finer resolution makes the finding more solid, not different.
Regenerated to stay consistent: baseline, scores, threshold and sweep.
paired-v1.6is unchanged, because adding unanswerable questions touches no answerable one.Where Phase 1 leaves the product
paired-v1.6), mostlysemantic, and it does nothing for cross-lingual. Dense stays the default.contextK = 3is justified against 5 and 8;candidateKabove 10 is not distinguishable.threshold = 0.5means cosine ≥ 0, not cosine ≥ 0.5.(6) is the phase result that matters: stop tuning this knob and reach for a different mechanism.
Stop line
Phase 1 ends here. Deliberately not started: public benchmarks, reranker eval, generator faithfulness, citation entailment, growing to 300 chunks, an
Abstention Signal Evaluation. Those build a more complete benchmark; they do not unblock KnowNote development, and (6) already says the next useful work is a different retrieval mechanism rather than a better measurement of this one.Testing
npm test— 504 passnpm run eval— baseline regenerated (78 questions, 53 chunks)npm run eval:scores,eval:threshold,eval:sweep— regenerated