Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
21 changes: 21 additions & 0 deletions .github/workflows/verify.yml
Original file line number Diff line number Diff line change
Expand Up @@ -78,3 +78,24 @@ jobs:
- name: Smoke test the packaged app (macOS / Windows)
if: runner.os != 'Linux'
run: npm run smoke:packaged

# The eval baseline (#75) is the reference point every v1.5 experiment is
# measured against, so CI proves the committed numbers still reproduce. The
# model cache is keyed on the pinned model file, so the 134 MB download
# happens once per pin, not once per run. A hit is required for the run to
# be offline: `eval:prepare` only downloads when the cache is cold.
- name: Cache eval embedding model
if: runner.os == 'Linux'
uses: actions/cache@v4
with:
path: ~/.config/knownote/models
key: knownote-eval-model-${{ hashFiles('src/main/embedding/localModel.ts') }}

- name: Eval harness is deterministic
if: runner.os == 'Linux'
run: |
xvfb-run -a npm run eval:prepare
xvfb-run -a node scripts/eval.mjs
cp docs/eval/baseline-v1.4.json /tmp/eval-a.json
xvfb-run -a node scripts/eval.mjs
diff /tmp/eval-a.json docs/eval/baseline-v1.4.json
26 changes: 14 additions & 12 deletions CONTRIBUTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,18 +39,20 @@ npm run dev

## Commands

| Command | What it does |
| ------------------------ | ------------------------------------------------------------- |
| `npm run dev` | Start the app in development with HMR. |
| `npm run typecheck` | Typecheck the main/preload, renderer and test projects. |
| `npm test` | Run the Node test suite (`test/**/*.test.ts`). |
| `npm run lint` | ESLint (see the note on the current baseline below). |
| `npm run format` | Prettier over the whole repository. |
| `npm run build` | Typecheck, then bundle with electron-vite. |
| `npm run build:unpack` | `build`, then produce an unpacked app in `dist/`. |
| `npm run smoke:packaged` | Launch the packaged app's `--smoke-test` and check it starts. |
| `npm run db:generate` | Generate a Drizzle migration from `src/main/db/schema.ts`. |
| `npm run db:studio` | Inspect the development database. |
| Command | What it does |
| ------------------------ | -------------------------------------------------------------- |
| `npm run dev` | Start the app in development with HMR. |
| `npm run typecheck` | Typecheck the main/preload, renderer and test projects. |
| `npm test` | Run the Node test suite (`test/**/*.test.ts`). |
| `npm run lint` | ESLint (see the note on the current baseline below). |
| `npm run format` | Prettier over the whole repository. |
| `npm run build` | Typecheck, then bundle with electron-vite. |
| `npm run build:unpack` | `build`, then produce an unpacked app in `dist/`. |
| `npm run smoke:packaged` | Launch the packaged app's `--smoke-test` and check it starts. |
| `npm run eval:prepare` | One-time, networked: download the pinned eval embedding model. |
| `npm run eval` | Run the RAG eval harness offline; rewrites `docs/eval/`. |
| `npm run db:generate` | Generate a Drizzle migration from `src/main/db/schema.ts`. |
| `npm run db:studio` | Inspect the development database. |

## Before you open a pull request

Expand Down
25 changes: 25 additions & 0 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -78,3 +78,28 @@ Rules:
compatibility and delegates to the default retriever. `SearchResult` gained a
`locator` field, which is populated from `RetrievedEvidence`; ranking, scores and
the existing fields are unchanged.

## Retrieval measurement

Retrieval quality is a number before it is an opinion. `eval/` holds a committed
corpus and a ground-truth dataset; `src/main/eval/` runs them through the normal
ingestion path and the real `Retriever`, and writes
`docs/eval/baseline-<version>.json` (deterministic) plus a markdown summary.

Rules:

- **Ground truth is corpus identity, not database identity.** A relevant location
is `{ document: <corpus-relative path>, page, block: <ordinal>, quote? }`. Runtime
`documentId`s are random and `blockId`s embed them, so a dataset keyed on them
would break — instead of measuring — a chunking or parser change.
- **The baseline is frozen and the delta is explicit.** v1.5 experiments (#77,
#78) are reported as a delta against the committed baseline, with an
adopted-change threshold. A change that is not measured against it is not
adopted.
- **`npm run eval` is offline.** Only `npm run eval:prepare` may download the
pinned model. The pinned revision is part of the embedding space identity, so a
model change is a baseline change.

The harness is a main-process entry (`--eval-harness`), like the packaged smoke
test, because the DB layer, vector store and loaders do not exist outside
Electron. See `eval/README.md` for the dataset format and commands.
Loading
Loading