Repository navigation
Conversation
主要改进:
**向量存储架构优化**
- 实现每个笔记本独立的向量表(vec_{notebookId}),替代全局 vec_embeddings
- 添加 vec_metadata 表追踪每个笔记本的向量维度
- 支持动态向量维度(768, 1024, 1536等),自动检测嵌入模型输出维度
- 修复维度不匹配错误(Expected 1024 dimensions but received 768)
**AI 提供商兼容性**
- 修复 Qwen 提供商与 AI SDK v5 的兼容性问题
- 将 Qwen 从 qwen-ai-provider 迁移到 @ai-sdk/openai-compatible
- 解决 UnsupportedModelVersionError 错误
**思维导图生成改进**
- 优化提示词结构,添加明确的格式要求和示例
- 修复 schema 验证错误:chunkIds 和 keywords 字段支持 null 值
- 增强内容聚合:MAX_CHUNKS_PER_DOC 从 10 提升到 30
- 添加详细的调试日志(输入提示词、模型输出、错误信息)
**技术细节**
- 修改文件:9个核心文件
- 新增功能:动态向量表管理、维度自动检测
- 性能优化:提升内容聚合效率
- 调试改进:完整的输入输出日志追踪
此次重构解决了多个关键问题,提升了系统的灵活性和稳定性。
Code reviewFound 1 issue:
The PR introduces a new 🤖 Generated with Claude Code - If this code review was useful, please react with 👍. Otherwise, react with 👎. |
@jimmyken 感谢贡献!可以运行一下数据库迁移的指令吗?自动生成一下SQLite的迁移SQL |
|
谢谢这份实现 —— 每 notebook 一张向量表 + 把宽度记在 但没有按原样合并,原因是其中的维度处理会丢用户数据:
这两点改起来比一个 PR 里塞更多东西要省事,所以这里直接接手重做了,提交在 #47:
#47 的提交里保留了你的署名( 这个 PR 我先关掉,如果你想自己基于 #47 继续调整,欢迎在原分支或新分支上提。 |
…th (#47) Embedding width was hardcoded to 1024 in three places at once: the vec0 DDL (FLOAT[1024]), the `dimensions: 1024` passed to embedBatch(), and the VectorStoreManager default. A vec0 table's width is fixed at creation, so any model that does not return 1024 dimensions could not index at all - the insert fails after the embedding call has already been paid for. That is #33: every embedding model except a 1024-dimensional one fails to index. Vectors now live in one vec0 table per notebook (`vec_<notebookId>`) with the width recorded in `vec_metadata`. That table is the single authority on the width: - The embedding width comes from the model (`embeddings[0].dimensions`) instead of being requested as 1024. - Search, delete and single-document reindex read the width from `vec_metadata`. They no longer pass a guessed default, so they can never create or drop a table based on a wrong number. - A width change is only possible through `rebuildNotebookVectorTable()`, called from the indexing path, which also marks the notebook's documents `pending` so they get re-indexed. A mismatch anywhere else throws. - Existing installs are migrated in `initVectorStore()`: the old global `vec_embeddings` is copied notebook by notebook (`INSERT..SELECT` between vec0 tables) and then dropped. Without this, upgrading users would keep a table the new code never reads. - Deleting a notebook drops its table and metadata row; cascade deletes do not reach vec0 shadow tables. `vectorTableSql.ts` collects every statement that has to interpolate a table name - a SQL identifier cannot be a bind parameter - behind an allowlist, and is covered by `test/vectorTableSql.test.ts`. The packaged smoke test now seeds a legacy 1024-dimensional `vec_embeddings` table, asserts it is migrated and still searchable, and asserts the data-loss case explicitly: a width mismatch must be refused, must leave existing vectors untouched, and only an explicit rebuild may replace them. Design follows #21 by @jimmyken (d37edfa), which introduced per-notebook tables and dynamic widths. That version dropped and recreated a notebook's table whenever a caller asked for a different width, and several call sites (search, reindex) had no width to give, so a single wrong default silently destroyed a notebook's vectors; the migration was also missing. This reworks it so the stored width, not a default, decides. Co-authored-by: jimmyken <four498@gmail.com>
主要改进:
向量存储架构优化
AI 提供商兼容性
思维导图生成改进
技术细节
此次重构解决了多个关键问题,提升了系统的灵活性和稳定性。