Skip to content

feat(provenance): persist chunk↔block mapping and page range - #108

Merged
mrsibe merged 1 commit into
mainfrom
feat/chunk-block-mapping
Sep 25, 2026
Merged

mrsibe merged 1 commit into
mainfrom
feat/chunk-block-mapping

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 25, 2026

Copy link
Copy Markdown
Owner

What does this PR do?

Persists the chunk → block mapping and the denormalized page range, so a retrieval result can be resolved back to its ordered blocks and its page without re-parsing the document.

This is the #68 half of #67, now rebased directly onto main after #106 (the chunking half) merged. It supersedes #107, which GitHub closed when its stacked base branch was deleted.

Why?

A retrieval result only carries a chunk id. Even with correct canonical offsets, provenance cannot survive in the database without the mapping — "chunking without a persisted block mapping is not a partial feature, it is an unverifiable one" (#68, epic #82).

Related issue

Fixes #68. Completes the persistence half of #67.

What changed?

  • Schema + additive migration 0015_colossal_solo.sql
    • chunk_blocks(chunk_id, block_id, start_in_block, end_in_block) with a composite primary key and an index on block_id.
    • Denormalized chunks.page_start / chunks.page_end (both NULL for non-paged sources), so a citation does not need a MIN/MAX join on the hot path.
    • Existing documents are untouched and keep working at document-level citations.
  • src/main/services/chunkProvenance.ts
    • insertChunkBlocks() writes the mapping.
    • resolveChunkProvenance() resolves a chunk id with a single join (chunks → chunk_blocks → document_blocks) and returns the ordered block spans plus the page range.
    • projectChunkProvenance() is the pure ordering/projection step and is unit-tested.
  • KnowledgeService
    • saveChunks() writes each chunk and its mapping in one transaction, so a partially indexed document never has a chunk without its mapping, and stores page_start/page_end.
    • getChunkProvenance() exposes the resolver.
    • deleteDocument() and reindexDocument() delete mappings explicitly, because the connection does not enable foreign-key cascades.
  • Packaged smoke test now round-trips the real database: insert a document + its blocks + a chunk + the mapping, then resolve the chunk and check the block spans and page range.

How was this tested?

  • npm test — 110 tests pass. test/chunkBlocks.test.ts gains the ordered-span projection and the page-range invariant (page_start/page_end equal the min/max page of the mapped blocks).
  • The packaged smoke test covers the database round trip on all three CI platforms, which is the only environment with a working better-sqlite3 (the Node unit-test runner has an Electron-ABI build).
  • npm run typecheck — clean. npm run build — clean. npx eslint — clean.

Screenshots / recordings

Not applicable — no UI change.

Checklist

  • I have reviewed my own changes.
  • npm run typecheck passes.
  • npm run build passes.
  • I have tested the affected user workflow. (DB path covered by the packaged smoke test in CI)
  • I have not included unrelated changes.
  • I have updated documentation when necessary.

Desktop / build changes

  • Not applicable
  • npm run build:unpack passes. (delegated to CI)
  • npm run smoke:packaged passes. (delegated to CI; includes the provenance round trip)

A retrieval result only carries a chunk id. Without a persisted mapping it
cannot be resolved back to a page, a heading or a character range without
re-parsing the document — which is what #68 (and the second half of #67)
requires.

- `chunk_blocks` (chunk_id, block_id, start_in_block, end_in_block, PK on
  chunk+block) plus denormalized `chunks.page_start`/`page_end`. Additive
  migration 0015; existing documents keep working at document-level.
- `chunkProvenance.ts`: `insertChunkBlocks()` writes the mapping, and
  `resolveChunkProvenance()` resolves a chunk id with a single join through
  `chunk_blocks` into `document_blocks`, returning the ordered block spans
  and the page range. `projectChunkProvenance()` (pure) does the ordering and
  projection.
- `KnowledgeService.saveChunks()` writes each chunk and its mapping in one
  transaction, so a partially indexed document never has chunks without
  their mapping. `getChunkProvenance()` exposes the resolver. `deleteDocument`
  and `reindexDocument` delete the mapping explicitly rather than relying on
  a foreign-key cascade the connection does not enable.
- The packaged smoke test now round-trips a fixture through the real
  database: document + blocks + chunk + mapping, then resolves the chunk and
  checks the block spans and page range.

Verification: `test/chunkBlocks.test.ts` gains two tests — the ordered-span
projection and the page-range invariant (page_start/page_end equal the
min/max page of the mapped blocks) — 110 tests pass. `npm run typecheck`,
`npm run build` and `npx eslint` are clean. The packaged smoke check runs in
CI (no xvfb locally).

Fixes #68
Refs #67
@github-actions github-actions Bot added the enhancement New feature or request label Sep 25, 2026
@mrsibe
mrsibe merged commit 0967b23 into main Sep 25, 2026
4 checks passed
@mrsibe
mrsibe deleted the feat/chunk-block-mapping branch September 25, 2026 06:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feat] Persist chunk↔block mapping and denormalized page range on chunks

1 participant