Repository navigation
feat(chat): record and show what each answer was built from - #93
Merged
Merged
Conversation
The differentiator the product record names — a citation — had nowhere to live
because the retrieval step threw its own identity away. `buildRAGContext` accepted
`{ documentTitle, content, score }` and nothing else, so `documentId`, `chunkId`
and `chunkIndex` were discarded between search and prompt. The answer arrived
looking identical whether it came from the reader's documents or from nowhere.
**Retrieval now leaves a record.** `SearchResult` already carried all of it; the
prompt builder returns the sources alongside the context string (prompt text
unchanged, so answer behaviour is unchanged), and they are persisted on the
assistant message's `metadata` — a JSON column that already existed and was unused,
so no migration. The retrieval outcome is recorded too: `used`, `none` or `failed`.
Today a failed search is a log line and the model is called with no context at all,
which is indistinguishable from a sourced answer.
**Each answer shows its evidence.** `AnswerSources` renders a disclosure —
"Based on N source passages" → the passages **quoted verbatim**, each with its
document title and an in-app "Show in library" action. Three states, deliberately
not collapsed into one: `used` (the passages), `none` (a T3 line saying nothing
from your sources was used) and `failed` (a T3 line saying the search broke). A
message with no recorded status renders nothing — an older message is *unknown*,
and calling it ungrounded would accuse a grounded answer.
The click travels through `uiStore` to `SourcePanel`, which derives the open
document **during render** rather than consuming the request in an effect. Two
shapes were tried and rejected: `setState` in an effect fails
`react-hooks/set-state-in-effect` and cascades renders, and clearing the request
from render writes to a store mid-render. The clear belongs in the event handlers.
**Typed and parsed defensively.** `ChatMessageMetadata` replaces a
`Record<string, any>`, and `shared/utils/answerSources.ts` is the only reader.
Malformed entries are dropped individually — two usable passages out of three still
show two — and "not recorded" is kept distinct from "none".
**One conversation per notebook.** The maintainer confirmed a notebook holds one
growing conversation. Two things were needed to make that true in the interface:
the composer no longer says "select a session first" (that state was a load
transient pointing at a picker that does not exist) and the textarea is no longer
disabled while `currentSession` is loading.
Also: three locale files carried dead English-only copy for surfaces that do not
exist (an editor menubar, a notes/trash/tags feature) — 84 keys total, unreachable
and making the app read as half-translated. Removed, so all eight namespaces are in
parity. Plural forms are the one intentional difference and are kept explicitly:
`t('ankiCards', { count })` never renders the literal `ankiCards_one`, so a
string-matching prune would have deleted a live translation.
Deliberately not here, with reasons in PRODUCT.md and DESIGN.md: page/span-accurate
citations need #82 (chunk offsets are computed against a preprocessed string, not
`documents.content`, so a span citation would be a false claim), and
model-emitted inline `[cite:id]` markers — the contract both Cherry Studio and Open
Notebook use — need a system-prompt change and real model testing that this
environment cannot provide. What landed is the half both of those implementations
also have: the retrieved set is persisted, visible and inspectable.
mrsibe
force-pushed
the
feat/answer-sources
branch
from
September 24, 2026 19:33
c80cb58 to
3039b24
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Makes each answer record and show what it was built from, and makes "one growing conversation per notebook" true in the interface.
Why: the retrieval step threw its own identity away
buildRAGContextaccepted{ documentTitle, content, score }and nothing else.SearchResultalready carrieddocumentId,chunkId,chunkIndexanddocumentType— all discarded between search and prompt. So the answer arrived looking identical whether it came from the reader's own documents or from nowhere, which is the one distinction this product cannot afford. A failed search was a log line; the model was called with no context and the answer looked sourced.How the two closest products do it
I read both, at the maintainer's suggestion.
Cherry Studio mints a citation id per lookup call and puts the contract in a dedicated system-prompt block:
The id prefix carries a full 32 bits specifically so a collision cannot silently mis-attribute a marker to an earlier result. Retrieval is a tool call, so the tool result is already the provenance record and the renderer resolves markers against tool results from this message and earlier turns.
Open Notebook uses
[source:<id>]/[note:<id>]/[source_insight:<id>]markers in the model's text, converted to compact numbered[1](#ref-source-…)links plus an appended reference list, resolved through a customReactMarkdownlink component. Their parser is defensive to the point of telling: single brackets, double brackets, bold-wrapped, comma-separated, and aninsight:alias forsource_insightbecause "some models emit the short form". Their references are document-level, not passage-level — a useful data point that document-level references ship in a NotebookLM-style product.Both share one prerequisite: a persisted set that markers resolve against. KnowNote had neither the set nor the markers. This PR adds the set and renders it. The markers are deliberately not here.
What changed
Retrieval leaves a record. The prompt builder now returns the sources alongside the context string — the prompt text is byte-identical, so answer behaviour is unchanged — and they are persisted on the assistant message's
metadata, a JSON column that already existed and was unused, so no migration. The outcome is recorded too:used,noneorfailed.Each answer shows its evidence.
AnswerSourcesis a disclosure: "Based on N source passages" → the passages quoted verbatim, each with its document title and an in-app "Show in library" action.Three states, deliberately not collapsed into one:
usednonefailedClicking a source opens it in place. The transcript (centre) and library (left) are siblings, so the request travels through
uiStoreandSourcePanelderives the open document during render. Two shapes were tried and rejected, and the comments say why:setStateinside an effect failsreact-hooks/set-state-in-effect(it caught me) and cascades renders; clearing the store request from render writes to a store mid-render. The clear belongs in the event handlers, which is why the derivation needs no cleanup step.Typed and parsed defensively.
ChatMessageMetadatareplaces aRecord<string, any>, andshared/utils/answerSources.tsis the only reader. Malformed entries are dropped individually — two usable passages out of three still show two — and "not recorded" stays distinct from "none".One conversation per notebook. The composer no longer says "select a session first"; that state was a load transient pointing at a picker that does not exist, and the textarea is no longer disabled while the session loads.
Three locale files carried dead English-only copy for surfaces that do not exist — an editor menubar, a notes/trash/tags feature — 84 keys, unreachable, and making the app read as half-translated. Removed: all eight namespaces are now in parity. Plural forms are the one intentional difference and are kept explicitly, because
t('ankiCards', { count })never renders the literalankiCards_one, so a string-matching prune would delete a live translation.What is deliberately not here
ChunkingServicecomputes offsets against a preprocessed string rather thandocuments.content, andparseResult.structureis still discarded. A span-accurate citation built on those offsets would be a false claim, which the product record forbids. Owned by [Epic] v1.4 — Trusted Research Loop: source provenance end to end #82.[cite:id]markers. Both reference products do this, and it is the next step — but it changes the system prompt, and answer quality cannot be verified in this environment. It should be done with the maintainer watching real output, not landed blind.<pre>, so there is nothing to highlight per chunk yet. That is the reader work in [Feat] Source Reader architecture — PDF first (replace the chunk-concatenation <pre> viewer) #71, and the "Show in library" action deliberately stops at opening the document rather than pretending to target a span.One correction to the previous critique
Assessment A reported that
session-auto-switchedhas no renderer subscriber. That was wrong —chatStore.ts:350subscribes and switches the current session. I re-checked before writing this because the claim was in a review I delivered.What is actually true, and is the remaining gap for "one growing conversation":
getActiveSessionByNotebookfiltersstatus = 'active', and the rollover archives the previous session. So after a rollover, reloading the notebook shows only the messages that came after it — the earlier part of the thread is not loaded, and there is no visible boundary saying so. That needs a decision (load archived ancestors into the transcript, or mark the summarised boundary), so it is reported rather than guessed. The rollover triggers at 100k tokens, so it is rare but real in a long research thread.How was this tested?
npm test— 84/84 (was 73; 11 added for the metadata parsers).npm run typecheck— passes.npm run lint— 0 errors, 111 warnings (was 112).npm run check:design— no violations.npx electron-vite build— builds.npx prettier --check— clean.impeccable detect --json src—[].Not verified: the rendered evidence region, the "Show in library" round trip, and the three states as a reader sees them. No browser or Electron capture is available here and I will not publish the local database. The parsers are covered by tests; the UI is not. The prompt is unchanged, so answer behaviour is unchanged — that much is by construction rather than by observation.
Checklist
npm run typecheckpasses.npm run buildpasses.Desktop / build changes