Skip to content

feat(chat): record and show what each answer was built from - #93

Merged
mrsibe merged 1 commit into
mainfrom
feat/answer-sources
Sep 25, 2026
Merged

mrsibe merged 1 commit into
mainfrom
feat/answer-sources

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 24, 2026

Copy link
Copy Markdown
Owner

What does this PR do?

Makes each answer record and show what it was built from, and makes "one growing conversation per notebook" true in the interface.

Why: the retrieval step threw its own identity away

buildRAGContext accepted { documentTitle, content, score } and nothing else. SearchResult already carried documentId, chunkId, chunkIndex and documentType — all discarded between search and prompt. So the answer arrived looking identical whether it came from the reader's own documents or from nowhere, which is the one distinction this product cannot afford. A failed search was a log line; the model was called with no context and the answer looked sourced.

How the two closest products do it

I read both, at the maintainer's suggestion.

Cherry Studio mints a citation id per lookup call and puts the contract in a dedicated system-prompt block:

When a statement in your answer is based on a web_search, web_fetch, kb_search, or kb_read
result, append a citation marker immediately after that statement: [cite:ID] …
- Do not add a "References" or "Sources" section at the end — the app renders citations
  from the inline markers.
- Statements from your own knowledge take no marker.

The id prefix carries a full 32 bits specifically so a collision cannot silently mis-attribute a marker to an earlier result. Retrieval is a tool call, so the tool result is already the provenance record and the renderer resolves markers against tool results from this message and earlier turns.

Open Notebook uses [source:<id>] / [note:<id>] / [source_insight:<id>] markers in the model's text, converted to compact numbered [1](#ref-source-…) links plus an appended reference list, resolved through a custom ReactMarkdown link component. Their parser is defensive to the point of telling: single brackets, double brackets, bold-wrapped, comma-separated, and an insight: alias for source_insight because "some models emit the short form". Their references are document-level, not passage-level — a useful data point that document-level references ship in a NotebookLM-style product.

Both share one prerequisite: a persisted set that markers resolve against. KnowNote had neither the set nor the markers. This PR adds the set and renders it. The markers are deliberately not here.

What changed

Retrieval leaves a record. The prompt builder now returns the sources alongside the context string — the prompt text is byte-identical, so answer behaviour is unchanged — and they are persisted on the assistant message's metadata, a JSON column that already existed and was unused, so no migration. The outcome is recorded too: used, none or failed.

Each answer shows its evidence. AnswerSources is a disclosure: "Based on N source passages" → the passages quoted verbatim, each with its document title and an in-app "Show in library" action.

Three states, deliberately not collapsed into one:

State Shown
used the disclosure and the passages
none one T3 line: nothing from your sources was used
failed one T3 line: searching your sources failed
no status recorded nothing — an older message is unknown, and calling it ungrounded would accuse a grounded answer

Clicking a source opens it in place. The transcript (centre) and library (left) are siblings, so the request travels through uiStore and SourcePanel derives the open document during render. Two shapes were tried and rejected, and the comments say why: setState inside an effect fails react-hooks/set-state-in-effect (it caught me) and cascades renders; clearing the store request from render writes to a store mid-render. The clear belongs in the event handlers, which is why the derivation needs no cleanup step.

Typed and parsed defensively. ChatMessageMetadata replaces a Record<string, any>, and shared/utils/answerSources.ts is the only reader. Malformed entries are dropped individually — two usable passages out of three still show two — and "not recorded" stays distinct from "none".

One conversation per notebook. The composer no longer says "select a session first"; that state was a load transient pointing at a picker that does not exist, and the textarea is no longer disabled while the session loads.

Three locale files carried dead English-only copy for surfaces that do not exist — an editor menubar, a notes/trash/tags feature — 84 keys, unreachable, and making the app read as half-translated. Removed: all eight namespaces are now in parity. Plural forms are the one intentional difference and are kept explicitly, because t('ankiCards', { count }) never renders the literal ankiCards_one, so a string-matching prune would delete a live translation.

What is deliberately not here

  • Page and span precision. ChunkingService computes offsets against a preprocessed string rather than documents.content, and parseResult.structure is still discarded. A span-accurate citation built on those offsets would be a false claim, which the product record forbids. Owned by [Epic] v1.4 — Trusted Research Loop: source provenance end to end #82.
  • Model-emitted inline [cite:id] markers. Both reference products do this, and it is the next step — but it changes the system prompt, and answer quality cannot be verified in this environment. It should be done with the maintainer watching real output, not landed blind.
  • Jumping to the passage inside the document. The viewer renders a document's chunks as one joined <pre>, so there is nothing to highlight per chunk yet. That is the reader work in [Feat] Source Reader architecture — PDF first (replace the chunk-concatenation <pre> viewer) #71, and the "Show in library" action deliberately stops at opening the document rather than pretending to target a span.

One correction to the previous critique

Assessment A reported that session-auto-switched has no renderer subscriber. That was wrong — chatStore.ts:350 subscribes and switches the current session. I re-checked before writing this because the claim was in a review I delivered.

What is actually true, and is the remaining gap for "one growing conversation": getActiveSessionByNotebook filters status = 'active', and the rollover archives the previous session. So after a rollover, reloading the notebook shows only the messages that came after it — the earlier part of the thread is not loaded, and there is no visible boundary saying so. That needs a decision (load archived ancestors into the transcript, or mark the summarised boundary), so it is reported rather than guessed. The rollover triggers at 100k tokens, so it is rare but real in a long research thread.

How was this tested?

  • npm test — 84/84 (was 73; 11 added for the metadata parsers).
  • npm run typecheck — passes.
  • npm run lint — 0 errors, 111 warnings (was 112).
  • npm run check:design — no violations.
  • npx electron-vite build — builds.
  • npx prettier --check — clean.
  • impeccable detect --json src — [].

Not verified: the rendered evidence region, the "Show in library" round trip, and the three states as a reader sees them. No browser or Electron capture is available here and I will not publish the local database. The parsers are covered by tests; the UI is not. The prompt is unchanged, so answer behaviour is unchanged — that much is by construction rather than by observation.

Checklist

  • I have reviewed my own changes.
  • npm run typecheck passes.
  • npm run build passes.
  • I have tested the affected user workflow. (not run on screen)
  • I have not included unrelated changes.
  • I have updated documentation when necessary.

Desktop / build changes

  • Not applicable

@github-actions github-actions Bot added the enhancement New feature or request label Sep 24, 2026
The differentiator the product record names — a citation — had nowhere to live
because the retrieval step threw its own identity away. `buildRAGContext` accepted
`{ documentTitle, content, score }` and nothing else, so `documentId`, `chunkId`
and `chunkIndex` were discarded between search and prompt. The answer arrived
looking identical whether it came from the reader's documents or from nowhere.

**Retrieval now leaves a record.** `SearchResult` already carried all of it; the
prompt builder returns the sources alongside the context string (prompt text
unchanged, so answer behaviour is unchanged), and they are persisted on the
assistant message's `metadata` — a JSON column that already existed and was unused,
so no migration. The retrieval outcome is recorded too: `used`, `none` or `failed`.
Today a failed search is a log line and the model is called with no context at all,
which is indistinguishable from a sourced answer.

**Each answer shows its evidence.** `AnswerSources` renders a disclosure —
"Based on N source passages" → the passages **quoted verbatim**, each with its
document title and an in-app "Show in library" action. Three states, deliberately
not collapsed into one: `used` (the passages), `none` (a T3 line saying nothing
from your sources was used) and `failed` (a T3 line saying the search broke). A
message with no recorded status renders nothing — an older message is *unknown*,
and calling it ungrounded would accuse a grounded answer.

The click travels through `uiStore` to `SourcePanel`, which derives the open
document **during render** rather than consuming the request in an effect. Two
shapes were tried and rejected: `setState` in an effect fails
`react-hooks/set-state-in-effect` and cascades renders, and clearing the request
from render writes to a store mid-render. The clear belongs in the event handlers.

**Typed and parsed defensively.** `ChatMessageMetadata` replaces a
`Record<string, any>`, and `shared/utils/answerSources.ts` is the only reader.
Malformed entries are dropped individually — two usable passages out of three still
show two — and "not recorded" is kept distinct from "none".

**One conversation per notebook.** The maintainer confirmed a notebook holds one
growing conversation. Two things were needed to make that true in the interface:
the composer no longer says "select a session first" (that state was a load
transient pointing at a picker that does not exist) and the textarea is no longer
disabled while `currentSession` is loading.

Also: three locale files carried dead English-only copy for surfaces that do not
exist (an editor menubar, a notes/trash/tags feature) — 84 keys total, unreachable
and making the app read as half-translated. Removed, so all eight namespaces are in
parity. Plural forms are the one intentional difference and are kept explicitly:
`t('ankiCards', { count })` never renders the literal `ankiCards_one`, so a
string-matching prune would have deleted a live translation.

Deliberately not here, with reasons in PRODUCT.md and DESIGN.md: page/span-accurate
citations need #82 (chunk offsets are computed against a preprocessed string, not
`documents.content`, so a span citation would be a false claim), and
model-emitted inline `[cite:id]` markers — the contract both Cherry Studio and Open
Notebook use — need a system-prompt change and real model testing that this
environment cannot provide. What landed is the half both of those implementations
also have: the retrieved set is persisted, visible and inspectable.
@mrsibe
mrsibe force-pushed the feat/answer-sources branch from c80cb58 to 3039b24 Compare September 24, 2026 19:33
@mrsibe
mrsibe merged commit 29771b6 into main Sep 25, 2026
3 checks passed
@mrsibe
mrsibe deleted the feat/answer-sources branch September 25, 2026 03:48
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant