Skip to content

feat(eval): fifteen near-miss unanswerable questions — closes Phase 1 - #208

Merged
mrsibe merged 1 commit into
feat/eval-paired-deltafrom
feat/eval-near-miss-unanswerable
Sep 30, 2026
Merged

mrsibe merged 1 commit into
feat/eval-paired-deltafrom
feat/eval-near-miss-unanswerable

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Stacked on #207. The last eval PR of RAG Eval v2 Phase 1.

What was added

Fifteen unanswerable questions whose subject is discussed at length in the corpus and whose answer is not there — "How much GPU memory does Gateway API v3 need?", "Which laboratory is accredited to analyse the estuary transect samples?", "How many staff work on the reservoir monitoring programme?".

Every specific term was checked absent before the question was written (gpu, vram, sla, uptime, expiry, sdk, self-hosted, accredit, vendor, supplier, funding, grant, staff, warranty, retention, encrypt, pricing, invoice, audit all return nothing). That check is what makes them unanswerable rather than believed to be, and it is the part that cannot be automated away.

validation unanswerable goes from 5 to 15, so abstention resolution goes from 20 % per question to 6.7 %. The FIFA questions were a sanity check for an obviously out-of-scope query; these are the realistic shape — the actual hallucination risk is a question about a document that is in the library.

The conclusion survives the larger sample

unanswerable max candidate    min 0.8978 (cosine 0.796)   p50 0.9279 (0.856)   max 0.9435 (0.887)
worst required relevant       p10 0.8906 (cosine 0.781)

Even the lowest unanswerable top score sits above the p10 of the worst required relevant passage. With n=15 instead of n=5 the distributions still do not separate, and the threshold curve is unchanged in shape:

threshold raw cosine answerable full recall unanswerable abstain
0.875 0.750 1.0000 0.0000
0.900 0.800 0.8947 0.0667
0.925 0.850 0.6316 0.4000
0.950 0.900 0.1579 1.0000

Abstention still never rises without full recall falling. The finer resolution makes the finding more solid, not different.

Regenerated to stay consistent: baseline, scores, threshold and sweep. paired-v1.6 is unchanged, because adding unanswerable questions touches no answerable one.

Where Phase 1 leaves the product

  1. The original 19-chunk benchmark could not guide a product decision; the current one (53 chunks, 78 questions, per-type metrics, held-out split) can.
  2. Cross-lingual retrieval is the largest measured gap: Recall@5 0.6667 and nDCG@10 0.4173 against 0.93–1.00 for same-language questions. Chinese over a Chinese source scores 1.0000, so it is cross-language matching, not Chinese.
  3. Hybrid has ranking value and no held-out mandate. Its full-set advantage is five questions out of 34 (paired-v1.6), mostly semantic, and it does nothing for cross-lingual. Dense stays the default.
  4. contextK = 3 is justified against 5 and 8; candidateK above 10 is not distinguishable.
  5. threshold = 0.5 means cosine ≥ 0, not cosine ≥ 0.5.
  6. A single embedding similarity threshold cannot decide answerability. Relevant, non-relevant and unanswerable score distributions overlap, and every threshold that abstains also drops required evidence.

(6) is the phase result that matters: stop tuning this knob and reach for a different mechanism.

Stop line

Phase 1 ends here. Deliberately not started: public benchmarks, reranker eval, generator faithfulness, citation entailment, growing to 300 chunks, an Abstention Signal Evaluation. Those build a more complete benchmark; they do not unblock KnowNote development, and (6) already says the next useful work is a different retrieval mechanism rather than a better measurement of this one.

Testing

  • npm test — 504 pass
  • npm run eval — baseline regenerated (78 questions, 53 chunks)
  • npm run eval:scores, eval:threshold, eval:sweep — regenerated

…al PR

Closes RAG Eval v2 Phase 1. This is the final expansion of the eval system — the
conclusions below are the point of the phase, not more infrastructure.

## What was added

Fifteen unanswerable questions whose **subject is discussed at length in the corpus and
whose answer is not there** — "How much GPU memory does Gateway API v3 need?", "Which
laboratory is accredited to analyse the estuary transect samples?", "How many staff work on
the reservoir monitoring programme?".

Every specific term was checked absent before the question was written (`gpu`, `vram`, `sla`,
`uptime`, `expiry`, `sdk`, `self-hosted`, `accredit`, `vendor`, `supplier`, `funding`,
`grant`, `staff`, `warranty`, `retention`, `encrypt`, `pricing`, `invoice`, `audit` all
return nothing). That check is what makes them unanswerable rather than believed to be, and
it is the part that cannot be automated away.

`validation` unanswerable goes from **5 to 15**, so abstention resolution goes from 20 % per
question to 6.7 %. The earlier "Who won the 2018 FIFA World Cup?" questions were a sanity
check for a query that is obviously out of scope; these are the realistic shape — the user's
actual hallucination risk is asking a question about a document that *is* in the library.

## The conclusion survives the larger sample

```
unanswerable max candidate    min 0.8978 (cosine 0.796)   p50 0.9279 (0.856)   max 0.9435 (0.887)
worst required relevant       p10 0.8906 (cosine 0.781)
```

Even the **lowest** unanswerable top score is above the p10 of the worst required relevant
passage. With n=15 instead of n=5 the distributions still do not separate, and the threshold
curve is unchanged in shape:

| threshold | raw cosine | answerable full recall | unanswerable abstain |
| --- | --- | --- | --- |
| 0.875 | 0.750 | 1.0000 | 0.0000 |
| 0.900 | 0.800 | 0.8947 | 0.0667 |
| 0.925 | 0.850 | 0.6316 | 0.4000 |
| 0.950 | 0.900 | 0.1579 | 1.0000 |

Abstention still never rises without full recall falling. The finer resolution makes the
finding *more* solid, not different.

Regenerated to stay consistent: baseline, scores, threshold and sweep. `paired-v1.6` is
unchanged, because adding unanswerable questions does not touch any answerable one.

## Where Phase 1 leaves the product

1. The original 19-chunk benchmark could not guide a product decision; the current one
   (53 chunks, 78 questions, per-type metrics, held-out split) can.
2. **Cross-lingual retrieval is the largest measured gap**: Recall@5 0.6667 and nDCG@10
   0.4173 against 0.93–1.00 for same-language questions. Chinese over a Chinese source
   scores 1.0000, so it is the cross-language matching, not Chinese.
3. **Hybrid has ranking value and no held-out mandate.** Its full-set advantage is five
   questions out of 34 (`paired-v1.6`), mostly `semantic`, and it does nothing for
   cross-lingual. Dense stays the default.
4. `contextK = 3` is justified against 5 and 8 on this corpus; `candidateK` above 10 is not
   distinguishable.
5. `threshold = 0.5` means **cosine ≥ 0**, not cosine ≥ 0.5.
6. **A single embedding similarity threshold cannot decide answerability** — the relevant
   and non-relevant and unanswerable score distributions overlap, and every threshold that
   abstains also drops required evidence.

(6) is the phase's real result: it says stop tuning this knob and reach for a different
mechanism.

## Stop line

Phase 1 ends here. Not done, and deliberately not started: public benchmarks, reranker eval,
generator faithfulness, citation entailment, growing the corpus to 300 chunks, an
`Abstention Signal Evaluation`. Those build a more complete benchmark; they do not unblock
KnowNote development, and (6) already says the next useful work is a different retrieval
mechanism rather than a better measurement of this one.
@github-actions github-actions Bot added the enhancement New feature or request label Sep 30, 2026
@mrsibe
mrsibe merged commit f423300 into feat/eval-paired-delta Sep 30, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant