Skip to content

fix: index in the background, coalesce streaming, fix toasts, recover truncation (#176, #177, #178, #179) - #180

Merged
mrsibe merged 4 commits into
mainfrom
fix/176-179-responsiveness-and-truncation
Sep 29, 2026
Merged

mrsibe merged 4 commits into
mainfrom
fix/176-179-responsiveness-and-truncation

Conversation

@mrsibe

@mrsibe mrsibe commented Sep 29, 2026

Copy link
Copy Markdown
Owner

What does this PR do?

Four commits, one per issue. They keep the app responsive while it works and give a truncated answer a way out:

  1. #176 — local ONNX inference moves off the Electron main process, and imports are registered immediately and indexed in the background.
  2. #177 — streaming snapshots are coalesced and each message renders from its own turn instead of a per-chunk remap of the whole transcript.
  3. #178 — one <Toaster /> at the app root, with stable ids for repeatable toasts.
  4. #179 — a truncated answer can continue in place, bounded by a setting, and connections can declare reasoning effort.

Why?

  • #176: importing a book-sized document froze the window; ingestion was also one long IPC call.
  • #177: every streamed chunk wrote to the store and re-rendered every historical MessageItem (Markdown + highlight + KaTeX).
  • #178: sonner replays active toasts to a late subscriber, so toasts fired with no note open were held and released in a burst.
  • #179: the truncation notice was correct but there was no recovery, and reasoning models spend the output budget on thinking first.

Related issue

Fixes #176
Fixes #177
Fixes #178
Fixes #179

What changed?

How was this tested?

  • npm run typecheck — passes (node, web, test).
  • npm test — 469 pass / 0 fail.
  • npm run check:design — no violations.
  • npx eslint src test — 0 errors.
  • npx electron-vite build — succeeds; out/main/embeddingWorker.js is emitted.

Not tested here: packaged worker_threads loading from app.asar (needs npm run build:unpack && npm run smoke:packaged), real ONNX inference in the worker, the React Profiler claim, and a live provider for the reasoning-effort effect. No Fixes claim on those.

Checklist

  • I have reviewed my own changes.
  • npm run typecheck passes.
  • npm run build passes.
  • I have tested the affected user workflow.
  • I have not included unrelated changes.
  • I have updated documentation when necessary.

Desktop / build changes

  • Not applicable
  • npm run build:unpack passes.
  • npm run smoke:packaged passes.

…the background (#176)

Local ONNX inference ran in the Electron main process, so a book-sized import
held the event loop for the whole tokenize + forward pass and the window
stalled. Ingestion was also a single long IPC call that resolved only after
parse, chunk, embed and persist.

- add embeddingWorker (worker_threads) with batching, per-batch progress and
  per-batch retry inside the worker
- add WorkerEmbeddingBackend; EmbeddingService selects it in production and
  keeps the in-process backend for the injected test loader
- build the worker as a second main-process entry (electron.vite.config.ts)
- add IngestionQueue and route knowledge:add-files / add-folder through it:
  documents are registered as `pending` and returned immediately, indexing
  continues over knowledge:index-progress
- broadcast index progress to all windows instead of the original sender
- reload the document list when a background import completes or fails
… its own turn (#177)

A streaming answer wrote to the store on every snapshot, remapped the whole
message array, and re-rendered every historical MessageItem — each with a
Markdown parse, syntax highlight and KaTeX pass.

- hold the latest snapshot per message and commit on a 40ms timer; flush
  synchronously before an outcome so terminal state is never delayed
- keep the per-message sequence counter out of the store and write it only when
  a gap is actually detected
- guard the pending→streaming transition so a chunk that changes nothing does
  not commit
- render live content and reasoning from the message's own turn, and memoize
  MessageItem so unchanged answers do not re-render
<Toaster /> lived inside NoteEditor, and sonner replays every still-active
toast to a late subscriber. Import, save and excerpt toasts fired while no note
was open were therefore held and released together the next time a note opened.

- mount a single <Toaster /> at the app root and remove it from NoteEditor
- give repeatable toasts a stable id so a later call updates the existing toast
  instead of stacking a duplicate
- add a guard test that pins the single mount and the keyed call sites
A reasoning model spends the output budget on thinking first, so the visible
answer can end with finishReason 'length'. The notice was correct; there was no
way out of it.

- add the 截断时自动继续 setting and a continuation bound (default 2); on a
  truncated attempt the manager continues in place, seeding the next call with
  the accumulated answer, and stops at the bound
- apply the same bound to the manual 继续生成 chain, recorded in the message
  metadata so it survives a reload
- show the truncation notice and the recovery actions as one unit under the
  answer
- extend the model connection with contextWindow, reasoning, reasoningEffort
  and reasoningBudget, and translate them per protocol; a connection that
  declares none sends exactly the request it did before
@github-actions github-actions Bot added the bug Something isn't working label Sep 29, 2026
@mrsibe
mrsibe merged commit 2d7d5c5 into main Sep 29, 2026
4 checks passed
@mrsibe
mrsibe deleted the fix/176-179-responsiveness-and-truncation branch September 29, 2026 09:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

1 participant