Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
484 changes: 482 additions & 2 deletions docs/eval/baseline-v1.6.json

Large diffs are not rendered by default.

10 changes: 5 additions & 5 deletions docs/eval/baseline-v1.6.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ Generated by `npm run eval`. The numbers below are harness output — do not edi
| Ranks | `candidateK=20, threshold=0.5` |
| Context width | `contextK=3` |
| Corpus | `eval/corpus` (21 documents) |
| Split | `all` (63 questions, 53 answerable) |
| Split | `all` (78 questions, 53 answerable) |
| Index size | 53 chunks |

## Metrics
Expand Down Expand Up @@ -63,13 +63,13 @@ with a small second one means the threshold filters nothing and the window is al

| Metric | Value |
| --- | --- |
| Unanswerable questions | 10 |
| Retrieval abstained | 0.0000 (0/10) |
| Unanswerable questions | 25 |
| Retrieval abstained | 0.0000 (0/25) |
| Mean candidates passing the threshold | 20.00 |
| Mean passages in the context window | 3.00 |

Timing is informational only and is **not** frozen: indexing 2796 ms, query
p50 13.67 ms, p95 16.34 ms on the
Timing is informational only and is **not** frozen: indexing 5921 ms, query
p50 53.92 ms, p95 85.84 ms on the
machine that produced this file. Timing and index size depend on hardware and on the
corpus, so they must never be the reason two runs differ.

Expand Down
16 changes: 8 additions & 8 deletions docs/eval/scores-v1.6.json
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
"indexSize": 53,
"counts": {
"answerable": 19,
"unanswerable": 5
"unanswerable": 15
},
"distributions": {
"bestRelevant": {
Expand Down Expand Up @@ -47,12 +47,12 @@
},
"unanswerableMax": {
"p0": 0.897848,
"p10": 0.897848,
"p25": 0.912751,
"p50": 0.925936,
"p75": 0.934276,
"p10": 0.908616,
"p25": 0.914179,
"p50": 0.927896,
"p75": 0.936526,
"p90": 0.942441,
"p100": 0.942441
"p100": 0.943454
}
},
"byType": {
Expand Down Expand Up @@ -186,7 +186,7 @@
"overlap": {
"worstRelevantP10": 0.89056,
"bestNonRelevantP90": 0.9386,
"unanswerableMaxP50": 0.925936,
"unanswerableMaxP50": 0.927896,
"unanswerableMaxP90": 0.942441
},
"separable": false,
Expand Down Expand Up @@ -308,7 +308,7 @@
"cosine": 0.8,
"hitRate": 0.8947368421052632,
"fullRecallRate": 0.8947368421052632,
"abstentionRate": 0.2
"abstentionRate": 0.06666666666666667
},
{
"threshold": 0.925,
Expand Down
8 changes: 4 additions & 4 deletions docs/eval/scores-v1.6.md
Original file line number Diff line number Diff line change
Expand Up @@ -33,7 +33,7 @@ Dense only, `validation` split only, `threshold = 0`, `candidateK = 500` (above
index size, so every chunk is scored for every query), `contextK = 3`.

- answerable questions: 19
- unanswerable questions: 5
- unanswerable questions: 15
- index size: 53 chunks

Hybrid is deliberately excluded: its `score` is an RRF value (`1 / (60 + rank)`) and is not
Expand Down Expand Up @@ -61,7 +61,7 @@ Where a row shows a raw cosine in brackets: an **absolute** score maps as
| worst relevant | 19 | 0.8808 (0.762) | 0.8906 (0.781) | 0.9086 (0.817) | 0.9355 (0.871) | 0.9428 (0.886) | 0.9530 (0.906) | 0.9537 (0.907) |
| best non-relevant | 19 | 0.8904 (0.781) | 0.9081 (0.816) | 0.9178 (0.836) | 0.9243 (0.849) | 0.9336 (0.867) | 0.9386 (0.877) | 0.9459 (0.892) |
| margin (best rel − best non-rel) | 19 | -0.0294 (-0.059) | -0.0272 (-0.054) | -0.0092 (-0.018) | 0.0026 (0.005) | 0.0081 (0.016) | 0.0329 (0.066) | 0.0623 (0.125) |
| unanswerable max candidate | 5 | 0.8978 (0.796) | 0.8978 (0.796) | 0.9128 (0.826) | 0.9259 (0.852) | 0.9343 (0.869) | 0.9424 (0.885) | 0.9424 (0.885) |
| unanswerable max candidate | 15 | 0.8978 (0.796) | 0.9086 (0.817) | 0.9142 (0.828) | 0.9279 (0.856) | 0.9365 (0.873) | 0.9424 (0.885) | 0.9435 (0.887) |

## By query type

Expand Down Expand Up @@ -101,7 +101,7 @@ relevant passage; **full recall** = share whose every ground-truth block is stil
| 0.825 | 0.650 | 1.0000 | 1.0000 | 0.0000 |
| 0.850 | 0.700 | 1.0000 | 1.0000 | 0.0000 |
| 0.875 | 0.750 | 1.0000 | 1.0000 | 0.0000 |
| 0.900 | 0.800 | 0.8947 | 0.8947 | 0.2000 |
| 0.900 | 0.800 | 0.8947 | 0.8947 | 0.0667 |
| 0.925 | 0.850 | 0.6316 | 0.6316 | 0.4000 |
| 0.950 | 0.900 | 0.1579 | 0.1579 | 1.0000 |
| 0.975 | 0.950 | 0.0000 | 0.0000 | 1.0000 |
Expand All @@ -115,7 +115,7 @@ The threshold decision turns on whether these overlap:
| --- | --- | --- |
| worst relevant, p10 | 0.8906 | 0.781 |
| best non-relevant, p90 | 0.9386 | 0.877 |
| unanswerable max, p50 | 0.9259 | 0.852 |
| unanswerable max, p50 | 0.9279 | 0.856 |
| unanswerable max, p90 | 0.9424 | 0.885 |

**The distributions **overlap**, so a higher threshold buys abstention by giving up required relevant passages. If the curve above shows abstention rising only as full recall falls, then the honest conclusion is that **a single dense similarity threshold cannot carry both recall and abstention** — and the next mechanism to evaluate is not a finer threshold grid but a different signal (reranker score, top1−top2 margin, per-query thresholds, or claim-level answerability).
Expand Down
Loading
Loading