Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
224 changes: 223 additions & 1 deletion docs/eval/baseline-v1.6.json

Large diffs are not rendered by default.

20 changes: 17 additions & 3 deletions docs/eval/baseline-v1.6.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ Generated by `npm run eval`. The numbers below are harness output — do not edi
| Ranks | `candidateK=20, threshold=0.5` |
| Context width | `contextK=3` |
| Corpus | `eval/corpus` (13 documents) |
| Split | `all` (30 questions) |
| Split | `all` (36 questions, 30 answerable) |
| Index size | 19 chunks |

## Metrics
Expand Down Expand Up @@ -43,8 +43,22 @@ The type comes from `type` in `questions.jsonl`; untagged questions report as
| semantic | 18 | 1.0000 | 0.9312 | 1.0000 | 0.9074 |
| zh | 3 | 1.0000 | 1.0000 | 1.0000 | 1.0000 |

Timing is informational only and is **not** frozen: indexing 2015 ms, query
p50 71.63 ms, p95 94.77 ms on the
### Unanswerable questions

These carry no ground truth, so the correct outcome is that retrieval finds nothing. They
are excluded from every metric above — a missing ground truth is not a miss — and reported
here instead. A higher **no-results** rate is better on this row, which is the opposite of
how it reads everywhere else, and `mean passages retrieved` is how much irrelevant context
was pulled in anyway. This is the row a threshold decision should move.

| Metric | Value |
| --- | --- |
| Unanswerable questions | 6 |
| Returned no results | 0.0000 (0/6) |
| Mean passages retrieved | 19.00 |

Timing is informational only and is **not** frozen: indexing 1496 ms, query
p50 11.42 ms, p95 17.31 ms on the
machine that produced this file. Timing and index size depend on hardware and on the
corpus, so they must never be the reason two runs differ.

Expand Down
21 changes: 15 additions & 6 deletions docs/eval/retrieval-v1.6.json
Original file line number Diff line number Diff line change
Expand Up @@ -7,6 +7,9 @@
"label": "dense (vector)",
"chunking": "1000/100",
"chunkCount": 19,
"split": "all",
"questions": 36,
"answerableCount": 30,
"recallAt1": 0.833333,
"recallAt5": 1,
"recallAt10": 1,
Expand All @@ -16,14 +19,17 @@
"mapAt10": 0.922222,
"contextPrecision": 0.355556,
"contextRecall": 1,
"indexingMs": 1548,
"latencyP95Ms": 20.51
"indexingMs": 1518,
"latencyP95Ms": 14.2
},
{
"id": "sparse",
"label": "sparse (BM25)",
"chunking": "1000/100",
"chunkCount": 19,
"split": "all",
"questions": 36,
"answerableCount": 30,
"recallAt1": 0.733333,
"recallAt5": 0.866667,
"recallAt10": 0.866667,
Expand All @@ -33,14 +39,17 @@
"mapAt10": 0.8,
"contextPrecision": 0.311111,
"contextRecall": 0.866667,
"indexingMs": 1490,
"latencyP95Ms": 1.96
"indexingMs": 1513,
"latencyP95Ms": 2.46
},
{
"id": "hybrid",
"label": "hybrid (RRF of dense + BM25)",
"chunking": "1000/100",
"chunkCount": 19,
"split": "all",
"questions": 36,
"answerableCount": 30,
"recallAt1": 0.866667,
"recallAt5": 1,
"recallAt10": 1,
Expand All @@ -50,8 +59,8 @@
"mapAt10": 0.938889,
"contextPrecision": 0.355556,
"contextRecall": 1,
"indexingMs": 1507,
"latencyP95Ms": 20.66
"indexingMs": 1557,
"latencyP95Ms": 19.55
}
]
}
11 changes: 6 additions & 5 deletions docs/eval/retrieval-v1.6.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,14 +4,15 @@ Generated by `node scripts/eval-retrieval.mjs`. Numbers are harness output; do n

## What was measured

Every strategy runs the real RAG eval harness against the same corpus and the same 30
questions as `baseline-v1.6.json`, with chunking held fixed at 1000/100. Only the retrieval strategy changes.
Every strategy runs the real RAG eval harness against the same corpus and the same questions
as `baseline-v1.6.json` (split `all`, 36 questions of which
30 are answerable), with chunking held fixed at 1000/100. Only the retrieval strategy changes.

| Strategy | Recall@1 | Recall@5 | MRR | nDCG@10 | MAP@10 | Context P | Query p95 |
| --- | --- | --- | --- | --- | --- | --- | --- |
| dense (vector) | 0.8333 | 1.0000 | 0.9278 | 0.9437 | 0.9222 | 0.3556 | 20.51 ms |
| sparse (BM25) | 0.7333 | 0.8667 | 0.8056 | 0.8184 | 0.8000 | 0.3111 | 1.96 ms |
| hybrid (RRF of dense + BM25) | 0.8667 | 1.0000 | 0.9444 | 0.9561 | 0.9389 | 0.3556 | 20.66 ms |
| dense (vector) | 0.8333 | 1.0000 | 0.9278 | 0.9437 | 0.9222 | 0.3556 | 14.20 ms |
| sparse (BM25) | 0.7333 | 0.8667 | 0.8056 | 0.8184 | 0.8000 | 0.3111 | 2.46 ms |
| hybrid (RRF of dense + BM25) | 0.8667 | 1.0000 | 0.9444 | 0.9561 | 0.9389 | 0.3556 | 19.55 ms |

## Not evaluated

Expand Down
Loading
Loading