Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions .prettierignore
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,11 @@ docs/eval/baseline-*.md
docs/eval/chunking-*.json
docs/eval/chunking-*.md

# Same reason again: generated by `scripts/eval-threshold.mjs`. Regenerate with
# `npm run eval:threshold`.
docs/eval/threshold-*.json
docs/eval/threshold-*.md

# And the same again for `scripts/eval-retrieval.mjs` (#77).
docs/eval/retrieval-*.json
docs/eval/retrieval-*.md
1 change: 1 addition & 0 deletions docs/eval/baseline-v1.6.json
Original file line number Diff line number Diff line change
Expand Up @@ -11,6 +11,7 @@
"respectHeadings": false
},
"retrieval": "dense",
"split": "all",
"candidateK": 20,
"contextK": 3,
"threshold": 0.5,
Expand Down
7 changes: 4 additions & 3 deletions docs/eval/baseline-v1.6.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,8 @@ Generated by `npm run eval`. The numbers below are harness output — do not edi
| Ranks | `candidateK=20, threshold=0.5` |
| Context width | `contextK=3` |
| Evidence per query | `evidenceK=5` |
| Corpus | `eval/corpus` (13 documents, 30 questions) |
| Corpus | `eval/corpus` (13 documents) |
| Split | `all` (30 questions) |
| Index size | 19 chunks |

## Metrics
Expand Down Expand Up @@ -42,8 +43,8 @@ The type comes from `type` in `questions.jsonl`; untagged questions report as
| semantic | 18 | 1.0000 | 0.9312 | 1.0000 | 0.9074 |
| zh | 3 | 1.0000 | 1.0000 | 1.0000 | 1.0000 |

Timing is informational only and is **not** frozen: indexing 1545 ms, query
p50 11.04 ms, p95 14.14 ms on the
Timing is informational only and is **not** frozen: indexing 1542 ms, query
p50 15.01 ms, p95 69.16 ms on the
machine that produced this file. Timing and index size depend on hardware and on the
corpus, so they must never be the reason two runs differ.

Expand Down
183 changes: 183 additions & 0 deletions docs/eval/threshold-v1.6.json
Original file line number Diff line number Diff line change
@@ -0,0 +1,183 @@
{
"baseline": "v1.6",
"productionThreshold": 0.5,
"flat": true,
"recommended": 0.5,
"rows": [
{
"threshold": 0,
"validation": {
"questions": 10,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19,
"metrics": {
"recallAt1": 0.9,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.933333,
"ndcgAt10": 0.95,
"hitRateAt5": 1,
"mapAt10": 0.933333,
"evidencePrecisionAt5": 0.2
}
},
"test": {
"questions": 20,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19,
"metrics": {
"recallAt1": 0.8,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.925,
"ndcgAt10": 0.940626,
"hitRateAt5": 1,
"mapAt10": 0.916667,
"evidencePrecisionAt5": 0.22
}
}
},
{
"threshold": 0.3,
"validation": {
"questions": 10,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19,
"metrics": {
"recallAt1": 0.9,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.933333,
"ndcgAt10": 0.95,
"hitRateAt5": 1,
"mapAt10": 0.933333,
"evidencePrecisionAt5": 0.2
}
},
"test": {
"questions": 20,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19,
"metrics": {
"recallAt1": 0.8,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.925,
"ndcgAt10": 0.940626,
"hitRateAt5": 1,
"mapAt10": 0.916667,
"evidencePrecisionAt5": 0.22
}
}
},
{
"threshold": 0.4,
"validation": {
"questions": 10,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19,
"metrics": {
"recallAt1": 0.9,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.933333,
"ndcgAt10": 0.95,
"hitRateAt5": 1,
"mapAt10": 0.933333,
"evidencePrecisionAt5": 0.2
}
},
"test": {
"questions": 20,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19,
"metrics": {
"recallAt1": 0.8,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.925,
"ndcgAt10": 0.940626,
"hitRateAt5": 1,
"mapAt10": 0.916667,
"evidencePrecisionAt5": 0.22
}
}
},
{
"threshold": 0.5,
"validation": {
"questions": 10,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19,
"metrics": {
"recallAt1": 0.9,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.933333,
"ndcgAt10": 0.95,
"hitRateAt5": 1,
"mapAt10": 0.933333,
"evidencePrecisionAt5": 0.2
}
},
"test": {
"questions": 20,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19,
"metrics": {
"recallAt1": 0.8,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.925,
"ndcgAt10": 0.940626,
"hitRateAt5": 1,
"mapAt10": 0.916667,
"evidencePrecisionAt5": 0.22
}
}
},
{
"threshold": 0.6,
"validation": {
"questions": 10,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19,
"metrics": {
"recallAt1": 0.9,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.933333,
"ndcgAt10": 0.95,
"hitRateAt5": 1,
"mapAt10": 0.933333,
"evidencePrecisionAt5": 0.2
}
},
"test": {
"questions": 20,
"noResultCount": 0,
"noResultRate": 0,
"meanRetrieved": 19,
"metrics": {
"recallAt1": 0.8,
"recallAt5": 1,
"recallAt10": 1,
"mrr": 0.925,
"ndcgAt10": 0.940626,
"hitRateAt5": 1,
"mapAt10": 0.916667,
"evidencePrecisionAt5": 0.22
}
}
}
]
}
51 changes: 51 additions & 0 deletions docs/eval/threshold-v1.6.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# Threshold derivation — v1.6 (#192)

Generated by `node scripts/eval-threshold.mjs`. Numbers are harness output; do not edit them by hand.

## What was measured

The real harness, the same corpus and the production retrieval config
(`candidateK=20, contextK=3`), once per candidate threshold. The **validation**
split selects; the **test** split reports. Both come from the same
`--eval-split=` code path, and the split is a deterministic function of the
question id, so this is reproducible.

| Threshold | n (val) | Recall@5 (val) | nDCG@10 (val) | MAP@10 (val) | No-result (val) | nDCG@10 (test) | No-result (test) |
| --- | --- | --- | --- | --- | --- | --- | --- |
| 0 | 10 | 1.0000 | 0.9500 | 0.9333 | 0.0000 | 0.9406 | 0.0000 |
| 0.3 | 10 | 1.0000 | 0.9500 | 0.9333 | 0.0000 | 0.9406 | 0.0000 |
| 0.4 | 10 | 1.0000 | 0.9500 | 0.9333 | 0.0000 | 0.9406 | 0.0000 |
| 0.5 | 10 | 1.0000 | 0.9500 | 0.9333 | 0.0000 | 0.9406 | 0.0000 |
| 0.6 | 10 | 1.0000 | 0.9500 | 0.9333 | 0.0000 | 0.9406 | 0.0000 |

## Selection rule

Best validation nDCG@10, then fewest validation no-results, then the **lowest**
threshold — the first stage is supposed to favour recall and let a later stage
filter, so among equals the wider one is the safer default.

## Outcome

The sweep is **flat**: every threshold from 0 to 0.6 produces the same validation nDCG@10 (0.9500), the same Recall@5 (1.0000) and a no-result rate of 0.0000. On this corpus the threshold is simply **non-binding** — E5 never scores these query/chunk pairs below the top of the swept range, so no passage is ever filtered out.

**No evidence to change `threshold = 0.5`.** The tie-break rule nominates `0` only because it prefers the widest threshold among equals; that is a tie-break, not a finding. What this run establishes is that the current value cannot be validated *or* falsified here, which is a property of the corpus, not of the threshold. Re-run after #192 child 2 grows it.

The app currently ships `threshold = 0.5`: validation nDCG@10
0.9500, test nDCG@10
0.9406, test no-result rate
0.0000.

## Caveat on this corpus

The split removes the most obvious form of overfitting, but 10
validation questions is a thin basis for a decision, and the corpus is still small. A
threshold is a product decision with a **no-result-rate** cost attached, so a
recommendation here is only as good as the corpus behind it. Re-run this after the
corpus grows (#192 child 2).

## Reproduce

```bash
npm run eval:prepare # one-time, networked model bootstrap
npm run eval:threshold # offline; rewrites this file
```
14 changes: 12 additions & 2 deletions eval/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,8 +7,10 @@ every experiment (#77, #78) is reported as a delta against that file.
## Commands

```bash
npm run eval:prepare # one-time, networked: download the pinned embedding model
npm run eval # offline and deterministic: run the harness, rewrite the baseline
npm run eval:prepare # one-time, networked: download the pinned embedding model
npm run eval # offline and deterministic: run the harness, rewrite the baseline
npm run eval:retrieval # strategy comparison (#77)
npm run eval:threshold # derive the similarity threshold on validation, report on test
```

### The harness runs the production configuration
Expand All @@ -22,13 +24,21 @@ the product rather than a research setup. Two Ks, because they answer different
| `--eval-context-k=` | `3` | passages the chat prompt actually takes (`chatHandlers.ts`) |
| `--eval-threshold=` | `0.5` | the similarity floor the app ships |
| `--eval-retrieval=` | `dense` | `dense`, `sparse`, or `hybrid` |
| `--eval-split=` | `all` | `all`, `validation`, or `test` — a deterministic id-based split |
| `--eval-baseline=` | `v1.6` | name written into `docs/eval/baseline-<name>.{json,md}` |

Ranking metrics are computed at `candidateK` depth, not at `contextK`: `Recall@10` needs
at least ten results, and truncation only takes a prefix of the candidate list, so the
truncation cannot change the ranking it is measured on. `contextK` is recorded so the
report describes the whole online path.

### Swept parameters are chosen on `validation`, reported on `test`

A parameter picked on the same questions it is scored on is a fitted number, not a
result. `--eval-split=validation` selects roughly a third of the questions by a
deterministic hash of the id; `test` is the rest. `npm run eval:threshold` uses this
to pick a similarity threshold on `validation` and report it on `test`.

`eval:prepare` downloads the pinned `multilingual-e5-small` revision into the app's model
cache and verifies it. `eval` never touches the network: if the model is missing it stops
with
Expand Down
3 changes: 2 additions & 1 deletion package.json
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,8 @@
"db:migrate": "drizzle-kit migrate",
"db:push": "drizzle-kit push",
"db:studio": "drizzle-kit studio",
"eval:retrieval": "npm run build && node scripts/eval-retrieval.mjs"
"eval:retrieval": "npm run build && node scripts/eval-retrieval.mjs",
"eval:threshold": "npm run build && node scripts/eval-threshold.mjs"
},
"//test": [
"`node --test` strips TypeScript types rather than compiling them, and strip-only",
Expand Down
Loading
Loading