Skip to content

feat(eval): freeze a reproducible v1.4 retrieval baseline - #115

Merged
mrsibe merged 2 commits into
mainfrom
feat/eval-harness
Sep 25, 2026
Merged

mrsibe merged 2 commits into
mainfrom
feat/eval-harness

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 25, 2026 •

Copy link
Copy Markdown
Owner

What does this PR do?

Adds the RAG eval harness and commits the frozen v1.4 retrieval baseline. It indexes a first-party corpus through the normal ingestion path, runs the real Retriever, and writes a deterministic metrics JSON plus a markdown summary. npm run eval is offline; npm run eval:prepare is the one-time networked model bootstrap.

Why?

There was no way to answer "did this change make retrieval better?", so every future v1.5 claim about BM25, fusion, rerankers or chunking would be folklore. #75 is the reference point #77 and #78 must be measured against (and absorbs #76's "freeze the baseline" DoD).

Related issue

Fixes #75

What changed?

  • eval/corpus/ (13 first-party markdown documents) + eval/questions.jsonl (30 questions) + eval/README.md.
  • src/main/eval/metrics.ts — pure Recall@1/5/10, MRR, nDCG@10, evidence precision, percentile; unit-tested.
  • src/main/eval/harness.ts / run.ts / report.ts / types.ts — ingestion, ground-truth resolution, retrieval, report.
  • npm run eval:prepare and npm run eval; scripts/eval.mjs launcher.
  • docs/eval/baseline-v1.4.{json,md} — the frozen report.
  • CI: cache the pinned model by revision, run the harness twice, diff the output.
  • docs/architecture.md / CONTRIBUTING.md — document the measurement seam and commands.

Design notes

  • Ground truth is corpus identity, not database identity: { document: "river-monitoring.md", page: null, block: 4, quote: "..." }. Runtime documentIds are random and blockIds embed them, so a dataset keyed on them would break instead of measure a chunking change. The runner maps corpus → runtime ids after ingestion; production id generation is untouched. The optional quote fails the run if a parser change moves the block ordinal.
  • The dataset is built not to saturate. The corpus pairs near-duplicate documents (lake vs river monitoring, supercapacitors vs lithium cells, wild vs hive pollinators, green roofs vs street canopy, vinegar vs lactic fermentation, wave vs tidal energy) and includes a Chinese document. Questions cover similar-document discrimination, multi-location ground truth, paraphrases with little keyword overlap, Chinese questions, and a Chinese question over an English source. Every question is answerable; difficulty comes from discrimination, not unanswerable queries.
  • evidencePrecisionAt5, not citation recall. The harness runs no model and produces no answer, so a metric named "answer citation recall" would be a false claim. It is the share of the first 5 retrieved passages that cover ground truth — retrieval precision. Answer-level citation correctness stays with the resolver ([Feat] Resolve and validate citations; never render an ungrounded [n] #70); a model-driven answer eval is a separate deliverable.
  • eval:prepare vs eval split: a clean checkout means eval:prepare && eval — no user model/provider configuration, but not "no download". eval refuses to download and forces the local backend, so a developer's remote embedding config cannot leak into the baseline.
  • Timing is not frozen: the committed JSON excludes latency so two runs are byte-identical; p50/p95 appears in the markdown as informational only.
  • Main-process entry (--eval-harness): the DB layer, vector store and loaders do not exist outside Electron, and a plain Node runner would have to reimplement the ingestion path it is supposed to measure. It uses a throwaway profile and never opens the developer's database.

Frozen numbers

Metric Value
Recall@1 0.7667
Recall@5 0.9333
Recall@10 1.0000
MRR 0.8492
nDCG@10 0.8899
Evidence precision@5 0.2000

Recall@5 has room to move, so the adoption rule ("adopt only if Recall@5 improves and nDCG@10 does not regress") is now satisfiable.

How was this tested?

  • npm run typecheck — passes.
  • npm test — passes (metrics cases updated for the rename).
  • npm run build — passes.
  • node scripts/eval.mjs twice: byte-identical baseline-v1.4.json.
  • npm run eval:prepare — no-op with a warm cache; eval with a cold cache exits 1 with the documented message.
  • npm run build:unpack and npm run smoke:packaged — pass (18 checks).
  • Lint: no new warnings.

Not verified: cross-machine reproducibility. The committed baseline was produced on the author's Linux machine; CI asserts two consecutive runs on ubuntu agree, not that they equal this file. If the first CI run disagrees, the committed snapshot should be regenerated on CI's platform.

Screenshots / recordings

Not applicable (no UI).

Checklist

  • I have reviewed my own changes.
  • npm run typecheck passes.
  • npm run build passes.
  • I have tested the affected user workflow.
  • I have not included unrelated changes.
  • I have updated documentation when necessary.

Desktop / build changes

  • Not applicable
  • npm run build:unpack passes.
  • npm run smoke:packaged passes.

There was no way to answer "did this change make retrieval better?", which makes
every v1.5 hybrid/reranker/chunking claim folklore. This adds the harness and the
reference point they must be measured against.

- `eval/corpus` + `eval/questions.jsonl`: a first-party corpus with ground truth.
  Relevance is expressed in **corpus identity** (`document` relative path + `block`
  ordinal + optional `quote`), never a runtime `documentId`/`blockId`. Runtime ids
  are random and `blockId` embeds them, so a dataset keyed on them would break —
  rather than measure — the chunking change #78 is about.
- `src/main/eval/`: metrics (pure, unit-tested), the harness, and the report
  renderer. It indexes the corpus through the normal ingestion path and runs the
  real `Retriever`, so it measures the shipped pipeline, not a shortcut.
- Runs as a main-process entry (`--eval-harness`) like the packaged smoke test,
  because the DB layer, vector store and loaders do not exist outside Electron. It
  uses a throwaway profile and never opens the developer's database.
- `npm run eval:prepare` is the only networked step: it bootstraps the pinned
  `multilingual-e5-small` revision. `npm run eval` is offline, refuses to download,
  and forces the local backend regardless of the developer's embedding config.
- The committed JSON excludes wall-clock timing, so two runs are byte-identical and
  a PR can diff them. Latency p50/p95 is reported in the markdown as informational.
- CI caches the pinned model and asserts two consecutive runs produce identical
  output.

Baseline on this machine: Recall@1 0.9231, Recall@5 1.0000, MRR 0.9487,
nDCG@10 0.9615.

Refs #75 (frozen baseline, absorbed from #76)
@github-actions github-actions Bot added the enhancement New feature or request label Sep 25, 2026
…he dataset

Two corrections to the benchmark before it is frozen.

1. `citationRecall` was not recall and not about answers: it was the share of
   the top-k retrieved passages covering ground truth, i.e. retrieval precision.
   It also never ran a model, so the old name misdescribed what was measured.
   Rename it to `evidencePrecisionAt5` (`evidencePrecisionAtK`), document that
   each retrieved passage counts once however many blocks it covers, and leave
   answer-level citation correctness to the resolver (#70). A model-driven answer
   eval is a separate deliverable.

2. The dataset saturated: 12 of 13 questions were rank 1 and Recall@5 was 1.0,
   which makes the frozen adoption rule ("Recall@5 improves and nDCG@10 does not
   regress") unsatisfiable. Add six near-duplicate distractor documents (lake vs
   river monitoring, supercapacitors vs lithium cells, wild vs hive pollinators,
   green roofs vs street canopy, vinegar vs lactic fermentation, wave vs tidal
   energy) and a Chinese document, and grow to 30 questions: similar-document
   discrimination, multi-location ground truth, paraphrases, Chinese questions
   and a Chinese question over an English source.

New baseline: Recall@1 0.7667, Recall@5 0.9333, Recall@10 1.0000, MRR 0.8492,
nDCG@10 0.8899, evidence precision@5 0.2000. Two runs are byte-identical.

Refs #75
@mrsibe
mrsibe merged commit f1aac4f into main Sep 25, 2026
4 checks passed
@mrsibe
mrsibe deleted the feat/eval-harness branch September 25, 2026 17:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feat] RAG eval harness: dataset format, runner, metrics

1 participant