Skip to content

feat: 重构向量存储架构和优化思维导图生成 - #21

Closed
jimmyken wants to merge 1 commit into
mrsibe:mainfrom
jimmyken:feature/vector-store-improvements
Closed

jimmyken wants to merge 1 commit into
mrsibe:mainfrom
jimmyken:feature/vector-store-improvements

Conversation

@jimmyken

@jimmyken jimmyken commented Jan 5, 2026

Copy link
Copy Markdown
Contributor

主要改进:

向量存储架构优化

  • 实现每个笔记本独立的向量表(vec_{notebookId}),替代全局 vec_embeddings
  • 添加 vec_metadata 表追踪每个笔记本的向量维度
  • 支持动态向量维度(768, 1024, 1536等),自动检测嵌入模型输出维度
  • 修复维度不匹配错误(Expected 1024 dimensions but received 768)

AI 提供商兼容性

  • 修复 Qwen 提供商与 AI SDK v5 的兼容性问题
  • 将 Qwen 从 qwen-ai-provider 迁移到 @ai-sdk/openai-compatible
  • 解决 UnsupportedModelVersionError 错误

思维导图生成改进

  • 优化提示词结构,添加明确的格式要求和示例
  • 修复 schema 验证错误:chunkIds 和 keywords 字段支持 null 值
  • 增强内容聚合:MAX_CHUNKS_PER_DOC 从 10 提升到 30
  • 添加详细的调试日志(输入提示词、模型输出、错误信息)

技术细节

  • 修改文件:9个核心文件
  • 新增功能:动态向量表管理、维度自动检测
  • 性能优化:提升内容聚合效率
  • 调试改进:完整的输入输出日志追踪

此次重构解决了多个关键问题,提升了系统的灵活性和稳定性。

主要改进:

**向量存储架构优化**
- 实现每个笔记本独立的向量表(vec_{notebookId}),替代全局 vec_embeddings
- 添加 vec_metadata 表追踪每个笔记本的向量维度
- 支持动态向量维度(768, 1024, 1536等),自动检测嵌入模型输出维度
- 修复维度不匹配错误(Expected 1024 dimensions but received 768)

**AI 提供商兼容性**
- 修复 Qwen 提供商与 AI SDK v5 的兼容性问题
- 将 Qwen 从 qwen-ai-provider 迁移到 @ai-sdk/openai-compatible
- 解决 UnsupportedModelVersionError 错误

**思维导图生成改进**
- 优化提示词结构,添加明确的格式要求和示例
- 修复 schema 验证错误:chunkIds 和 keywords 字段支持 null 值
- 增强内容聚合:MAX_CHUNKS_PER_DOC 从 10 提升到 30
- 添加详细的调试日志(输入提示词、模型输出、错误信息)

**技术细节**
- 修改文件:9个核心文件
- 新增功能:动态向量表管理、维度自动检测
- 性能优化:提升内容聚合效率
- 调试改进:完整的输入输出日志追踪

此次重构解决了多个关键问题,提升了系统的灵活性和稳定性。
@mrsibe

mrsibe commented Jan 15, 2026

Copy link
Copy Markdown
Owner

Code review

Found 1 issue:

  1. Missing data migration - PR changes vector storage architecture from single vec_embeddings table to per-notebook tables (vec_{notebookId}), but provides no migration strategy. Existing users will lose all vector indexes on upgrade.

https://github.com/MrSibe/KnowNote/blob/d37edfa5d0577f2f0f87e07249a33fa1d4fd8e6e/src/main/db/index.ts#L225-L246

The PR introduces a new createNotebookVectorTable() function that creates separate vector tables for each notebook, completely changing from the original architecture that used a single vec_embeddings table with notebook_id column (established in commit a785075). Without migration code, existing users' vector data will be inaccessible after upgrading.

🤖 Generated with Claude Code

- If this code review was useful, please react with 👍. Otherwise, react with 👎.

@mrsibe

mrsibe commented Jan 15, 2026

Copy link
Copy Markdown
Owner

Code review

Found 1 issue:

  1. Missing data migration - PR changes vector storage architecture from single vec_embeddings table to per-notebook tables (vec_{notebookId}), but provides no migration strategy. Existing users will lose all vector indexes on upgrade.

https://github.com/MrSibe/KnowNote/blob/d37edfa5d0577f2f0f87e07249a33fa1d4fd8e6e/src/main/db/index.ts#L225-L246

The PR introduces a new createNotebookVectorTable() function that creates separate vector tables for each notebook, completely changing from the original architecture that used a single vec_embeddings table with notebook_id column (established in commit a785075). Without migration code, existing users' vector data will be inaccessible after upgrading.

🤖 Generated with Claude Code

  • If this code review was useful, please react with 👍. Otherwise, react with 👎.

@jimmyken 感谢贡献!可以运行一下数据库迁移的指令吗?自动生成一下SQLite的迁移SQL

@mrsibe

mrsibe commented Sep 24, 2026

Copy link
Copy Markdown
Owner

谢谢这份实现 —— 每 notebook 一张向量表 + 把宽度记在 vec_metadata 里的方向是对的,这次就是按你的思路落地的。

但没有按原样合并,原因是其中的维度处理会丢用户数据:

initialize() 里只要调用点给的维度与既有表不一致,就会 DROP TABLE 重建。而好几个调用点(KnowledgeService.search、重建单篇文档索引等)根本不传维度,会回落到 VectorStoreManager 里那个全局可变的 defaultDimensions。于是应用重启后第一次搜索、或者索引完另一个笔记本之后,就可能拿一个错的默认值去初始化,把那个笔记本的向量整张删掉——文档状态还是 indexed,谁也不会再重嵌,检索就此永久为空。另外 review 里提到的 vec_embeddings 迁移也确实缺了。

这两点改起来比一个 PR 里塞更多东西要省事,所以这里直接接手重做了,提交在 #47:

  • 宽度以 vec_metadata 里存的为准,任何调用点都不再拿默认值建表/删表;
  • 宽度不一致直接报错,只有索引链路(刚量到真实维度)能通过 rebuildNotebookVectorTable() 显式重建,并把该笔记本的文档标回 pending;
  • 补上旧全局表的迁移(vec0 → vec0 INSERT..SELECT,逐笔记本搬完再删旧表),并在打包 smoke test 里加了这条升级路径的断言;
  • 查询、删除、计数在没有向量表时按空处理,不再顺手建表。

#47 的提交里保留了你的署名(Co-authored-by)。你原本那版里的 Qwen provider 修复和思维导图 prompt 调整没有一起带过来:provider 那部分已被 #44 的 protocol 重构取代(AISDKProvider.ts 已删除,Qwen 现在走 openai-completions),prompt 那部分是独立的东西,想推进的话可以单独开一个 PR。

这个 PR 我先关掉,如果你想自己基于 #47 继续调整,欢迎在原分支或新分支上提。

@mrsibe mrsibe closed this Sep 24, 2026
mrsibe added a commit that referenced this pull request Sep 24, 2026
…th (#47)

Embedding width was hardcoded to 1024 in three places at once: the vec0 DDL
(FLOAT[1024]), the `dimensions: 1024` passed to embedBatch(), and the
VectorStoreManager default. A vec0 table's width is fixed at creation, so any
model that does not return 1024 dimensions could not index at all - the insert
fails after the embedding call has already been paid for. That is #33: every
embedding model except a 1024-dimensional one fails to index.

Vectors now live in one vec0 table per notebook (`vec_<notebookId>`) with the
width recorded in `vec_metadata`. That table is the single authority on the
width:

- The embedding width comes from the model (`embeddings[0].dimensions`) instead
  of being requested as 1024.
- Search, delete and single-document reindex read the width from `vec_metadata`.
  They no longer pass a guessed default, so they can never create or drop a table
  based on a wrong number.
- A width change is only possible through `rebuildNotebookVectorTable()`, called
  from the indexing path, which also marks the notebook's documents `pending` so
  they get re-indexed. A mismatch anywhere else throws.
- Existing installs are migrated in `initVectorStore()`: the old global
  `vec_embeddings` is copied notebook by notebook (`INSERT..SELECT` between vec0
  tables) and then dropped. Without this, upgrading users would keep a table the
  new code never reads.
- Deleting a notebook drops its table and metadata row; cascade deletes do not
  reach vec0 shadow tables.

`vectorTableSql.ts` collects every statement that has to interpolate a table
name - a SQL identifier cannot be a bind parameter - behind an allowlist, and is
covered by `test/vectorTableSql.test.ts`.

The packaged smoke test now seeds a legacy 1024-dimensional `vec_embeddings`
table, asserts it is migrated and still searchable, and asserts the data-loss
case explicitly: a width mismatch must be refused, must leave existing vectors
untouched, and only an explicit rebuild may replace them.

Design follows #21 by @jimmyken (d37edfa), which introduced per-notebook tables
and dynamic widths. That version dropped and recreated a notebook's table
whenever a caller asked for a different width, and several call sites (search,
reindex) had no width to give, so a single wrong default silently destroyed a
notebook's vectors; the migration was also missing. This reworks it so the
stored width, not a default, decides.

Co-authored-by: jimmyken <four498@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants