From a5b8aca418ff35e3ed3cc95cb26f13d34c87aa92 Mon Sep 17 00:00:00 2001 From: Winshare Date: Fri, 14 Aug 2026 04:07:57 +0800 Subject: [PATCH 1/5] docs: define documentation information architecture --- ...14-docs-information-architecture-design.md | 114 ++++++++++++++++++ 1 file changed, 114 insertions(+) create mode 100644 docs/archive/plans/2026-08-14-docs-information-architecture-design.md diff --git a/docs/archive/plans/2026-08-14-docs-information-architecture-design.md b/docs/archive/plans/2026-08-14-docs-information-architecture-design.md new file mode 100644 index 0000000..cd95926 --- /dev/null +++ b/docs/archive/plans/2026-08-14-docs-information-architecture-design.md @@ -0,0 +1,114 @@ +# Lumin 文档信息架构重整设计 + +> **设计日期**:2026-08-14 +> **适用基线**:`9f1bf56` +> **状态**:已确认,待实施 + +## 1. 目标 + +将 `docs/` 重整为清晰的当前文档入口,同时确保所有历史分析、完成计划、阶段测量和旧素材统一位于 `docs/archive/`。重整后,读者应能从唯一入口快速区分外部接口、内部架构、工程运维状态和历史记录。 + +## 2. 设计原则 + +1. `docs/` 顶层只保留 `README.md`,作为唯一导航和文档生命周期规则入口。 +2. 当前文档按职责组织,不按语言或运行时拆分,避免 TS/Rust/Embedded 交叉内容重复。 +3. 所有不再承担当前规范作用的文档和素材均位于 `docs/archive/`。 +4. 文件移动使用 Git rename,保留历史;不因归档而改写旧结论、旧命令和测量环境。 +5. 当前导航和真实 Markdown 链接必须有效;归档正文中的历史字面路径可以保留。 +6. 根目录或源码中的当前引用若因迁移失效,允许做最小路径修复,不改变运行时行为。 + +## 3. 目标目录 + +```text +docs/ +├── README.md +├── reference/ +│ └── API.md +├── architecture/ +│ ├── DUAL_LOOP_ARCHITECTURE.md +│ ├── MEMORY.md +│ └── TOOL_COMPLETION_INVARIANTS.md +├── operations/ +│ ├── PROJECT_STATUS.md +│ └── TEST_COVERAGE.md +└── archive/ + ├── README.md + ├── analysis/ + ├── plans/ + └── assets/ +``` + +目录职责: + +| 目录 | 内容 | 生命周期 | +|------|------|----------| +| `reference/` | 对外协议、调用方式和公共 API | 接口变化时同步更新 | +| `architecture/` | 运行时结构、关键机制和必须维持的不变量 | 实现边界变化时同步更新 | +| `operations/` | 工程状态、测试矩阵、安全与发布准备信息 | 每次审计、发布或验证结果变化时更新 | +| `archive/analysis/` | 已被实现或后续设计取代的分析 | 只读历史 | +| `archive/plans/` | 已完成设计、实施计划、阶段测量和审计记录 | 只读历史 | +| `archive/assets/` | 当前文档不再引用的旧素材 | 只读历史 | + +## 4. 文件映射 + +| 当前路径 | 目标路径 | +|----------|----------| +| `docs/API.md` | `docs/reference/API.md` | +| `docs/DUAL_LOOP_ARCHITECTURE.md` | `docs/architecture/DUAL_LOOP_ARCHITECTURE.md` | +| `docs/MEMORY.md` | `docs/architecture/MEMORY.md` | +| `docs/TOOL_COMPLETION_INVARIANTS.md` | `docs/architecture/TOOL_COMPLETION_INVARIANTS.md` | +| `docs/PROJECT_STATUS.md` | `docs/operations/PROJECT_STATUS.md` | +| `docs/TEST_COVERAGE.md` | `docs/operations/TEST_COVERAGE.md` | +| `docs/archive/superpowers/plans/*` | `docs/archive/plans/*` | +| `docs/superpowers/plans/2026-08-14-docs-refresh-and-archive.md` | `docs/archive/plans/2026-08-14-docs-refresh-and-archive.md` | + +本设计与随后生成的实施计划在工作完成后同样保留于 `docs/archive/plans/`。迁移完成后不保留空的 `docs/superpowers/` 或 `docs/archive/superpowers/` 层级。 + +## 5. 导航与引用更新 + +`docs/README.md` 按以下顺序导航: + +1. 当前工程状态; +2. API 参考; +3. 架构与关键不变量; +4. 测试与验证; +5. 历史归档。 + +链接更新范围: + +- 六份当前文档之间的相对链接; +- `docs/README.md` 和 `docs/archive/README.md`; +- 根 `README.md` 中指向 API 文档的链接; +- `src/task/message-queue.ts` 和 capability test 注释中指向历史设计/审计的路径; +- 本次设计和实施计划中的最终路径。 + +不批量改写归档正文中的旧路径。它们描述的是原执行环境,除非某个路径是当前导航所依赖的 Markdown 链接。 + +## 6. 内容边界 + +本次只改变文档组织和链接: + +- 不修改 TypeScript/Rust 运行时行为; +- 不重新定义 API、架构结论或测试数字; +- 不重命名现有参考文档文件名,只改变目录; +- 不删除历史文件或素材; +- 不引入文档站点生成器、侧边栏框架或额外依赖。 + +## 7. 验证标准 + +完成迁移后必须满足: + +1. `docs/` 顶层只有 `README.md` 和职责目录。 +2. 所有历史文档和素材均位于 `docs/archive/`。 +3. `archive/analysis`、`archive/plans`、`archive/assets` 的文件数量与迁移前后清单一致。 +4. 根 README、ROADMAP 和全部 Markdown 文档不存在缺失的相对链接目标。 +5. 当前源码/测试注释不再引用已移除的 `docs/superpowers/` 路径。 +6. `git diff --check` 通过,Git 对历史文件识别为 rename。 +7. `npm run typecheck` 通过,确认最小注释路径调整没有影响 TypeScript。 + +## 8. 失败处理 + +- 若文件名冲突,停止迁移并保留两个来源,不覆盖历史文件。 +- 若相对链接检查失败,按引用方所在目录重新计算路径后再验证。 +- 若 Git 未识别为 rename,核对源文件与目标文件内容,避免无意改写归档正文。 +- 若发现未分类文档,先根据“当前契约”或“历史记录”确定生命周期,再放入对应职责目录;不留在顶层临时堆放。 From 082ccad6a686077e36bcaf477c55a55ef0a83f94 Mon Sep 17 00:00:00 2001 From: Winshare Date: Fri, 14 Aug 2026 04:10:52 +0800 Subject: [PATCH 2/5] docs: add agent guide scope to docs design --- ...14-docs-information-architecture-design.md | 44 ++++++++++++++++--- 1 file changed, 38 insertions(+), 6 deletions(-) diff --git a/docs/archive/plans/2026-08-14-docs-information-architecture-design.md b/docs/archive/plans/2026-08-14-docs-information-architecture-design.md index cd95926..a7de976 100644 --- a/docs/archive/plans/2026-08-14-docs-information-architecture-design.md +++ b/docs/archive/plans/2026-08-14-docs-information-architecture-design.md @@ -6,7 +6,7 @@ ## 1. 目标 -将 `docs/` 重整为清晰的当前文档入口,同时确保所有历史分析、完成计划、阶段测量和旧素材统一位于 `docs/archive/`。重整后,读者应能从唯一入口快速区分外部接口、内部架构、工程运维状态和历史记录。 +将 `docs/` 重整为清晰的当前文档入口,同时确保所有历史分析、完成计划、阶段测量和旧素材统一位于 `docs/archive/`。同步更新 `CLAUDE.md`,新增仓库级 `AGENTS.md`,并刷新运行时模板 `templates/base/AGENTS.md`。重整后,人类开发者和编码代理都应从一致入口快速区分外部接口、内部架构、工程运维状态和历史记录。 ## 2. 设计原则 @@ -16,6 +16,7 @@ 4. 文件移动使用 Git rename,保留历史;不因归档而改写旧结论、旧命令和测量环境。 5. 当前导航和真实 Markdown 链接必须有效;归档正文中的历史字面路径可以保留。 6. 根目录或源码中的当前引用若因迁移失效,允许做最小路径修复,不改变运行时行为。 +7. 仓库级代理说明与运行时工作区模板职责分离,不能把开发流程规则注入最终用户的 agent prompt。 ## 3. 目标目录 @@ -84,29 +85,60 @@ docs/ 不批量改写归档正文中的旧路径。它们描述的是原执行环境,除非某个路径是当前导航所依赖的 Markdown 链接。 -## 6. 内容边界 +## 6. 开发者与代理入口 -本次只改变文档组织和链接: +### `CLAUDE.md` + +`CLAUDE.md` 作为开发者快速入口进行结构化重写: + +- 删除约 4,900 行 TypeScript 等失效静态描述,改为链接 `docs/operations/PROJECT_STATUS.md` 的日期化审计数据; +- 补充 `src/loop/`、`src/task/`、`src/memory-file-backend.ts`、`src/embedded.ts` 和 Rust workspace; +- 区分 single、dual 和 embedded 的执行边界; +- 使用当前可复现的安装、类型检查、测试和构建命令; +- 不再把尚未配置 provider 的 coverage 命令写成默认可用能力; +- 链接 `docs/README.md`,避免复制 API、风险和测试数字的完整正文。 + +### 根 `AGENTS.md` + +新增根 `AGENTS.md`,作为仓库级编码代理协作指南。内容包括: + +- 工程交付面和权威文档入口; +- 当前目录职责及 single/dual/embedded 重要边界; +- 保留用户改动、避免编辑生成目录、使用 `rg` 和按比例验证等工作规则; +- 文档生命周期和归档约定; +- 常用验证命令及真实 LLM/capability 测试的显式前置条件。 + +它不复制个人工具说明,也不覆盖宿主平台的更高优先级规则。 + +### `templates/base/AGENTS.md` + +现有模板继续作为生成到用户 workspace 的运行时 agent prompt,不承担仓库开发指南职责。只更新与当前运行时行为相关的默认说明:工具结果确认、来源引用、破坏性操作确认、模糊需求处理和能力边界;不加入 Git、测试、归档等仓库维护规则。 + +## 7. 内容边界 + +本次改变文档组织、导航和三份协作入口: - 不修改 TypeScript/Rust 运行时行为; - 不重新定义 API、架构结论或测试数字; - 不重命名现有参考文档文件名,只改变目录; - 不删除历史文件或素材; - 不引入文档站点生成器、侧边栏框架或额外依赖。 +- 除 `CLAUDE.md`、根 `AGENTS.md` 和 `templates/base/AGENTS.md` 外,不扩展根目录文档改写范围。 -## 7. 验证标准 +## 8. 验证标准 完成迁移后必须满足: 1. `docs/` 顶层只有 `README.md` 和职责目录。 2. 所有历史文档和素材均位于 `docs/archive/`。 3. `archive/analysis`、`archive/plans`、`archive/assets` 的文件数量与迁移前后清单一致。 -4. 根 README、ROADMAP 和全部 Markdown 文档不存在缺失的相对链接目标。 +4. 根 README、CLAUDE、AGENTS、ROADMAP 和全部 Markdown 文档不存在缺失的相对链接目标。 5. 当前源码/测试注释不再引用已移除的 `docs/superpowers/` 路径。 6. `git diff --check` 通过,Git 对历史文件识别为 rename。 7. `npm run typecheck` 通过,确认最小注释路径调整没有影响 TypeScript。 +8. `CLAUDE.md` 和根 `AGENTS.md` 不复制易漂移的完整测试数量或 API wire schema;模板 `AGENTS.md` 不包含仓库开发指令。 -## 8. 失败处理 +## 9. 失败处理 - 若文件名冲突,停止迁移并保留两个来源,不覆盖历史文件。 - 若相对链接检查失败,按引用方所在目录重新计算路径后再验证。 From f79c739ab7df166ef14935d6f4b04dea2ba8576d Mon Sep 17 00:00:00 2001 From: Winshare Date: Fri, 14 Aug 2026 04:20:09 +0800 Subject: [PATCH 3/5] docs: reorganize documentation and agent guides --- AGENTS.md | 59 ++ CLAUDE.md | 243 +++---- README.md | 20 +- docs/API.md | 571 --------------- docs/DUAL_LOOP_ARCHITECTURE.md | 661 ------------------ docs/MEMORY.md | 313 --------- docs/README.md | 50 ++ docs/TEST_COVERAGE.md | 140 ---- docs/TOOL_COMPLETION_INVARIANTS.md | 209 ------ docs/architecture/DUAL_LOOP_ARCHITECTURE.md | 330 +++++++++ docs/architecture/MEMORY.md | 245 +++++++ .../TOOL_COMPLETION_INVARIANTS.md | 125 ++++ docs/archive/README.md | 32 + .../analysis}/01-agent-loop-analysis.md | 0 .../02-tool-orchestration-analysis.md | 0 .../03-context-management-analysis.md | 0 .../analysis}/04-multi-agent-analysis.md | 0 .../05-hooks-permissions-analysis.md | 0 .../analysis}/06-improvement-roadmap.md | 0 .../2026-03-26-dual-loop-v2-design.md | 0 docs/{ => archive/assets}/luminlogo.png | Bin docs/{ => archive/assets}/luminpulse logo.svg | 0 .../plans/2026-03-31-rust-ts-parity.md | 0 .../2026-04-12-builtin-tools-test-plan.md | 0 .../2026-04-12-rust-builtin-tools-mvp.md | 0 .../plans/2026-04-13-c1-after-phase-a.md | 0 .../plans/2026-04-13-c1-baseline.md | 0 .../plans/2026-04-13-c3-after-phase-b.md | 0 .../plans/2026-04-13-c4-after-phase-c.md | 0 .../plans/2026-04-13-c7-after-phase-e.md | 0 ...026-04-13-dual-loop-architecture-design.md | 0 .../2026-04-13-dual-loop-audit-and-roadmap.md | 0 .../2026-04-13-phase-a-message-queue-impl.md | 0 ...026-04-13-phase-b-disk-persistence-impl.md | 0 ...026-04-13-phase-c-structured-abort-impl.md | 0 ...4-13-phase-d-permissions-plan-mode-impl.md | 0 ...6-04-13-phase-e-knowledge-eviction-impl.md | 0 .../2026-04-13-plan-mode-after-phase-d.md | 0 ...2026-04-15-phase-h-embedded-bundle-impl.md | 0 ...14-docs-information-architecture-design.md | 4 +- .../2026-08-14-docs-refresh-and-archive.md | 108 +++ ...ture-reorganization-implementation-plan.md | 157 +++++ docs/operations/PROJECT_STATUS.md | 127 ++++ docs/operations/TEST_COVERAGE.md | 269 +++++++ docs/reference/API.md | 590 ++++++++++++++++ src/task/message-queue.ts | 2 +- templates/base/AGENTS.md | 27 +- .../capability/dual-loop-capabilities.test.ts | 2 +- 48 files changed, 2226 insertions(+), 2058 deletions(-) create mode 100644 AGENTS.md delete mode 100644 docs/API.md delete mode 100644 docs/DUAL_LOOP_ARCHITECTURE.md delete mode 100644 docs/MEMORY.md create mode 100644 docs/README.md delete mode 100644 docs/TEST_COVERAGE.md delete mode 100644 docs/TOOL_COMPLETION_INVARIANTS.md create mode 100644 docs/architecture/DUAL_LOOP_ARCHITECTURE.md create mode 100644 docs/architecture/MEMORY.md create mode 100644 docs/architecture/TOOL_COMPLETION_INVARIANTS.md create mode 100644 docs/archive/README.md rename docs/{ => archive/analysis}/01-agent-loop-analysis.md (100%) rename docs/{ => archive/analysis}/02-tool-orchestration-analysis.md (100%) rename docs/{ => archive/analysis}/03-context-management-analysis.md (100%) rename docs/{ => archive/analysis}/04-multi-agent-analysis.md (100%) rename docs/{ => archive/analysis}/05-hooks-permissions-analysis.md (100%) rename docs/{ => archive/analysis}/06-improvement-roadmap.md (100%) rename docs/{ => archive/analysis}/2026-03-26-dual-loop-v2-design.md (100%) rename docs/{ => archive/assets}/luminlogo.png (100%) rename docs/{ => archive/assets}/luminpulse logo.svg (100%) rename docs/{superpowers => archive}/plans/2026-03-31-rust-ts-parity.md (100%) rename docs/{superpowers => archive}/plans/2026-04-12-builtin-tools-test-plan.md (100%) rename docs/{superpowers => archive}/plans/2026-04-12-rust-builtin-tools-mvp.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-c1-after-phase-a.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-c1-baseline.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-c3-after-phase-b.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-c4-after-phase-c.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-c7-after-phase-e.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-dual-loop-architecture-design.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-dual-loop-audit-and-roadmap.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-phase-a-message-queue-impl.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-phase-b-disk-persistence-impl.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-phase-c-structured-abort-impl.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-phase-d-permissions-plan-mode-impl.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-phase-e-knowledge-eviction-impl.md (100%) rename docs/{superpowers => archive}/plans/2026-04-13-plan-mode-after-phase-d.md (100%) rename docs/{superpowers => archive}/plans/2026-04-15-phase-h-embedded-bundle-impl.md (100%) create mode 100644 docs/archive/plans/2026-08-14-docs-refresh-and-archive.md create mode 100644 docs/archive/plans/2026-08-14-docs-structure-reorganization-implementation-plan.md create mode 100644 docs/operations/PROJECT_STATUS.md create mode 100644 docs/operations/TEST_COVERAGE.md create mode 100644 docs/reference/API.md diff --git a/AGENTS.md b/AGENTS.md new file mode 100644 index 0000000..bb6380a --- /dev/null +++ b/AGENTS.md @@ -0,0 +1,59 @@ +# Repository Agent Guide + +## Scope + +These instructions apply to repository maintenance and development. They are not runtime prompt content. The workspace prompt template used by Lumin agents is `templates/base/AGENTS.md`. + +## Start Here + +- `docs/README.md` is the documentation index and lifecycle policy. +- `docs/reference/` contains current external contracts. +- `docs/architecture/` contains implementation structure and invariants. +- `docs/operations/` contains dated project and verification status. +- `docs/archive/` contains historical evidence only. + +When prose conflicts with implementation, prefer source code, runtime schemas, and reproducible tests. Use archived files only to understand earlier decisions. + +## Repository Map + +- `src/`: primary TypeScript runtime. +- `src/loop/` and `src/task/`: single/dual execution and task lifecycle. +- `src/embedded.ts`: Node-free embedded entry; keep its dependency graph platform-neutral. +- `rust/crates/`: partially aligned Rust runtime and server. +- `tests/`: default, real-LLM, capability, parity, and benchmark tests. +- `templates/`: files copied into user workspaces. +- `scripts/`: build and verification helpers. + +## Working Rules + +- Inspect the relevant source and tests before changing behavior or documentation. +- Preserve unrelated and pre-existing worktree changes. +- Search with `rg`/`rg --files`; make focused, reviewable edits. +- Do not manually edit generated or local-state directories: `dist/`, `node_modules/`, `rust/target/`, or `workspace/`. +- Do not claim TypeScript/Rust parity without checking the exact capability on both sides. +- Do not treat dual-loop `processMessage()` as a final-result call; it returns task creation/queueing state. +- Keep Node APIs out of the embedded entry and its transitive imports. +- Keep root/developer instructions out of `templates/base/AGENTS.md`. + +## Documentation Rules + +- Keep `docs/README.md` as the only top-level document under `docs/`. +- Put API contracts in `docs/reference/`, architecture in `docs/architecture/`, and dated engineering evidence in `docs/operations/`. +- Put completed designs, plans, analyses, and measurements under `docs/archive/`. +- Repair inbound and cross-document links whenever a document moves. +- Do not silently modernize historical commands or conclusions inside archived records. +- Avoid copying volatile test totals or protocol schemas into multiple entry documents; link to the authoritative current document. + +## Verification + +For TypeScript or documentation changes, start with: + +```bash +npm run typecheck +npm test -- --reporter=dot +git diff --check +``` + +Also run `npm run build:embedded` when changing the embedded dependency graph or exports. Run the relevant Cargo command when changing Rust. Real-LLM and capability suites require explicit credentials/flags; skipped tests are not successful capability verification. + +Before handing off, report the commands actually run, skipped or environment-sensitive cases, and whether changes remain uncommitted. diff --git a/CLAUDE.md b/CLAUDE.md index e3b1639..18ecdea 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -1,166 +1,137 @@ -# CLAUDE.md — @prismer/agent-core +# CLAUDE.md — Lumin Agent Core ## Project Overview -`@prismer/agent-core` is a lightweight, standalone agent runtime (~4,900 LOC TypeScript). -OpenAI-compatible, zero heavy dependencies (only Zod). Designed for the -[Prismer.AI](https://prismer.ai) academic research platform but works independently -as a general-purpose agent framework. +`@prismer/agent-core` is a standalone agent runtime with three delivery surfaces: -**npm**: `@prismer/agent-core` v0.3.1 -**Repository**: https://github.com/prismer-ai/agent-core -**Main project**: gitlab.app:prismer/library (this repo is at `docker/agent/`) -**License**: MIT +- Node.js/TypeScript, the primary implementation and public npm package; +- a Rust workspace containing `lumin-core` and `lumin-server` with partial feature parity; +- a Node-free embedded TypeScript entry for JavaScriptCore, Hermes, and Electron hosts. + +The package version is defined by `package.json` and `src/version.ts`. Use the dated [project status audit](docs/operations/PROJECT_STATUS.md) for current source size, test results, security findings, and known gaps instead of copying those values here. + +## Documentation + +Start at [docs/README.md](docs/README.md): + +- [API reference](docs/reference/API.md) +- [single/dual-loop architecture](docs/architecture/DUAL_LOOP_ARCHITECTURE.md) +- [memory architecture](docs/architecture/MEMORY.md) +- [tool completion invariants](docs/architecture/TOOL_COMPLETION_INVARIANTS.md) +- [project status](docs/operations/PROJECT_STATUS.md) +- [test and verification status](docs/operations/TEST_COVERAGE.md) + +Historical analyses, completed plans, measurements, and unused assets live under `docs/archive/` and are not current contracts. ## Quick Start +Requirements: Node.js 20 or newer and npm. + ```bash -npm install -npx tsc # Compile TypeScript -npm test # Run all tests (vitest) +npm ci +npm run typecheck +npm test +npm run build +``` -# Run agent CLI +Run the CLI after building: + +```bash node dist/cli.js agent --message "Hello" -node dist/cli.js serve --port 3001 # Start HTTP + WebSocket server +node dist/cli.js serve --port 3001 +node dist/cli.js health --url http://localhost:3001 ``` -## Architecture - -### Source Layout +Development mode: -``` -src/ -├── agent.ts # Core agent loop (LLM → tool → response cycle) -├── server.ts # HTTP + WebSocket gateway (SSE streaming) -├── index.ts # runAgent() entry + PromptBuilder + module exports -├── provider.ts # OpenAI-compatible LLM client + FallbackProvider -├── prompt.ts # Dynamic system prompt builder (SOUL.md / TOOLS.md / Skills) -├── cli.ts # CLI entry point -├── sse.ts # EventBus + SSE writer (Zod schemas) -├── version.ts # Centralized version string (single source of truth) -├── agents.ts # Sub-agent registry (6 built-in agents) -├── skills.ts # SKILL.md loader + YAML frontmatter -├── compaction.ts # Context overflow → memory flush → LLM summarize -├── session.ts # Session management + directive accumulation -├── workspace.ts # Workspace config (AGENTS.md/USER.md) -├── memory.ts # File-based persistent memory (keyword recall) -├── hooks.ts # Lifecycle hooks (before_prompt, before_tool, after_tool, agent_end) -├── config.ts # Runtime configuration from env vars -├── tools.ts # Tool registry interface -├── directives.ts # UI directive types -├── observer.ts # Event type definitions -├── ipc.ts # Inter-process communication -├── log.ts # Structured logging -├── schemas.ts # Zod schema exports -├── channels/ # Communication adapters -│ ├── types.ts # ChannelAdapter interface -│ ├── manager.ts # Auto-detect from env vars -│ ├── cloud-im.ts # Prismer Cloud IM (SSE) -│ └── telegram.ts # Telegram Bot (long-polling) -└── tools/ - ├── index.ts # Tool registry - ├── loader.ts # prismer-workspace plugin adapter - └── clawhub.ts # Pure JS skill installer (git clone) +```bash +npm run dev -- agent --message "Hello" +npm run dev -- serve --port 3001 ``` -### Templates +Agent calls require an OpenAI-compatible provider configuration. Server startup itself may succeed without a key, but the health endpoint currently reports missing provider/plugin configuration as degraded; see the API reference. + +## Source Layout + +```text +src/ +├── agent.ts core AsyncGenerator agent loop +├── index.ts Node runtime composition and public exports +├── embedded.ts Node-free embedded runtime factory +├── server.ts / cli.ts / ipc.ts HTTP, WebSocket, CLI, and IPC entrypoints +├── provider.ts OpenAI-compatible and fallback providers +├── tools.ts / tools/ registry, built-ins, plan mode, plugin loader +├── loop/ single/dual loop adapters and task UI routing +├── task/ task state, store, queue, and disk metadata +├── memory.ts cross-platform memory facade/interfaces +├── memory-file-backend.ts Node-only Markdown file backend +├── world-model/ dual-loop world model and handoff builder +├── channels/ Telegram and Prismer Cloud IM adapters +└── prompt.ts / workspace.ts dynamic workspace prompt composition + +rust/ +├── Cargo.toml workspace definition +└── crates/ + ├── lumin-core/ Rust runtime library + └── lumin-server/ Rust HTTP/WebSocket server -``` templates/ -├── base/ # Default workspace templates -│ ├── AGENTS.md # Agent priority/routing config -│ ├── USER.md # User preferences -│ ├── SOUL.md # Agent personality/identity -│ ├── TOOLS.md # Available tools reference -│ └── HEARTBEAT.md # Health check format -├── lite/ # Minimal template -├── researcher/ # Academic researcher template -├── mathematician/ # Theorem-proving template -└── financial-analyst/ # Quantitative analysis template +└── base/AGENTS.md runtime workspace prompt template ``` -### Docker Integration +The root `AGENTS.md` is a repository collaboration guide. It is intentionally different from `templates/base/AGENTS.md`, which is copied into user workspaces and becomes model prompt content. -``` -Dockerfile.lumin # Container image build (base: prismer-academic + lumin + plugins) -lumin-entrypoint.sh # Container entrypoint (starts lumin serve + container gateway) -``` +## Runtime Modes -## Key Concepts - -### Agent Loop -`agent.ts` implements the core loop: prompt LLM → execute tool calls → accumulate response → repeat until done. Supports sub-agent delegation, doom-loop detection, and context guard. - -### Compaction -When context exceeds `MAX_CONTEXT_CHARS` (default 600K), the compaction system: -1. Flushes extractable facts to memory (`{workspace}/.prismer/memory/YYYY-MM-DD.md`) -2. LLM-summarizes the conversation -3. Injects the summary as a compact message pair - -### Thinking Control -`/think` and `/nothink` directives control provider-level thinking: -- Kimi: `enable_thinking` parameter -- Claude: `thinking.budget_tokens` - -### Workspace Config -- `AGENTS.md` (priority 9) + `USER.md` (priority 3.5) are loaded into the system prompt -- Modifying `AGENTS.md` changes agent behavior immediately - -### Channels -ChannelAdapter interface + auto-detection from env vars: -- `TELEGRAM_BOT_TOKEN` → TelegramAdapter (long-polling) -- `PRISMER_IM_*` → CloudIMAdapter (SSE) - -### Skills -SKILL.md files with YAML frontmatter define installable skills. -ClawHub integration: `lumin skill install ` (pure JS git clone, no external CLI). - -## Environment Variables - -| Variable | Description | Default | -|----------|-------------|---------| -| `OPENAI_API_KEY` | LLM provider API key | required | -| `OPENAI_API_BASE_URL` | LLM provider base URL | `https://api.openai.com/v1` | -| `AGENT_DEFAULT_MODEL` | Default model ID | `gpt-4o` | -| `WORKSPACE_DIR` | Working directory | `./workspace` | -| `LUMIN_PORT` | HTTP/WS server port | `3001` | -| `MAX_CONTEXT_CHARS` | Compaction threshold | `600000` | -| `PRISMER_PLUGIN_PATH` | Path to workspace plugin | — | -| `TELEGRAM_BOT_TOKEN` | Telegram bot token (optional) | — | -| `PRISMER_IM_BASE_URL` | Cloud IM base URL (optional) | — | -| `PRISMER_IM_CONVERSATION_ID` | Cloud IM conversation (optional) | — | -| `PRISMER_IM_TOKEN` | Cloud IM auth token (optional) | — | - -## Testing +| Mode | Selection | Completion model | +|------|-----------|------------------| +| `single` | default | request waits for the final agent result | +| `dual` | `LUMIN_LOOP_MODE=dual` | request returns a task ID; clients poll task state/result | +| embedded | `@prismer/agent-core/embedded` | injected provider/tools/memory, single-loop only | -```bash -npm test # All tests -npx vitest run tests/agent.test.ts # Single test -npx vitest --coverage # Coverage report -``` +Do not assume TypeScript/Rust feature parity or that internal EventBus events are all forwarded over WebSocket. Consult the architecture and API documents before changing protocol behavior. + +## Configuration + +Configuration is validated in `src/config.ts`. Common environment variables: -## Submodule Usage (Main Project) +| Variable | Purpose | +|----------|---------| +| `OPENAI_API_KEY` | provider credential | +| `OPENAI_API_BASE_URL` | OpenAI-compatible endpoint | +| `AGENT_DEFAULT_MODEL` | default model identifier | +| `MODEL_FALLBACK_CHAIN` | comma-separated fallback models | +| `WORKSPACE_DIR` | workspace root | +| `PRISMER_PLUGIN_PATH` | optional workspace plugin entry | +| `LUMIN_PORT` | server port | +| `LUMIN_LOOP_MODE` | `single` or `dual` | +| `MAX_CONTEXT_CHARS` | context character budget | +| `APPROVAL_TIMEOUT_MS` | sensitive-tool approval timeout | +| `PRISMER_ENABLED_MODULES` | comma-separated tool module filter | +| `LOG_LEVEL` / `DEBUG` | logging controls | -In the main Prismer project, this repo is mounted as a submodule: +Some schema fields are not fully wired into runtime selection. Check `src/config.ts`, its consumers, and the project-status audit before documenting a setting as operational. + +## Verification + +Use checks proportional to the change: ```bash -# In the main project root -git submodule update --init docker/agent - -# Update to latest luminclaw -cd docker/agent -git pull origin main -cd ../.. -git add docker/agent -git commit -m "chore: update luminclaw submodule" +npm run typecheck +npm test -- --reporter=dot +npm run build:embedded + +cd rust +cargo test --workspace ``` -The Dockerfile.lumin in the main project references `docker/agent/` which maps to this repo. +Real-LLM and capability suites require explicit credentials/flags and may otherwise skip. The Rust suite also contains environment-sensitive network cases. Follow [TEST_COVERAGE.md](docs/operations/TEST_COVERAGE.md) for reproducible commands and interpretation; do not infer coverage percentages from test counts. -## Development Practices +## Development Rules -- **TypeScript strict mode** — All code is strictly typed -- **Zod validation** — Runtime type safety for configs, events, schemas -- **Zero heavy deps** — Only Zod in production; TypeScript/vitest in dev -- **File-based memory** — No vector DB dependency, keyword-based recall -- **OpenAI-compatible** — Works with any OpenAI-compatible LLM provider +- Preserve existing user changes and keep runtime edits separate from documentation-only work. +- Do not edit generated `dist/`, `node_modules/`, `rust/target/`, or workspace state. +- Treat synchronous bash cancellation, dual-loop persistence, WebSocket task delivery, and approval wiring as known boundaries, not completed guarantees. +- Keep Node-only imports out of `src/embedded.ts` and its transitive dependency graph. +- Update current docs when API, task state, memory behavior, runtime parity, or verification results change. +- Move completed designs and plans to `docs/archive/plans/`; do not use archived documents as current specifications. diff --git a/README.md b/README.md index 1918303..9fb4f6f 100644 --- a/README.md +++ b/README.md @@ -20,19 +20,19 @@ ## Features - **Agent loop** — tool calling, sub-agent delegation, doom-loop detection, context guard -- **Dual-loop mode** (`LUMIN_LOOP_MODE=dual`) — async task execution decoupled from the dialogue loop. User messages enqueue into running tasks, `task.progress` streams per iteration, tasks survive server restart via disk-backed transcripts, `POST /v1/tasks/:id/resume` picks up from the last persisted turn +- **Dual-loop mode** (`LUMIN_LOOP_MODE=dual`) — background task execution, message queueing, task polling, cancellation, and basic restart recovery. See the [architecture guide](docs/architecture/DUAL_LOOP_ARCHITECTURE.md) for current delivery, persistence, and concurrency limits - **Structured cancel** — `POST /v1/tasks/:id/cancel` with `AbortReason` propagation; in-flight LLM fetch and tool executions honor the signal; dangling tool calls are synthesized as `[Aborted: ]` so history stays well-formed - **Permission modes** — `default` / `plan` / `auto` / `bypass`. Dual-loop runs in `auto` (auto-denies `requiresUserInteraction` tools). `enter_plan_mode` / `exit_plan_mode` tools flip the mode mid-conversation - **OpenAI-compatible** — works with any `/chat/completions` endpoint (OpenAI, Anthropic, Ollama, etc.) -- **File-based memory** — keyword recall, zero vector DB dependency. Beats Letta/MemGPT on LoCoMo (86% vs 74%). Cross-task knowledge: `WorldModel.knowledgeBase` facts persisted to MemoryStore on task completion, recalled on next task start +- **File-based memory** — keyword recall with zero vector DB dependency, plus cross-task WorldModel fact persistence. Current behavior and historical benchmark scope are documented in the [memory guide](docs/architecture/MEMORY.md) - **Context compaction** — automatic fact extraction + LLM summarization when context overflows - **Lifecycle hooks** — `before_prompt`, `before_tool`, `after_tool`, `agent_end` - **Skills** — installable SKILL.md extensions with ClawHub (pure JS git clone) - **Channels** — Telegram, Cloud IM adapters (auto-detected from env) -- **HTTP + WebSocket gateway** — zero external dependencies, real-time streaming +- **HTTP + WebSocket gateway** — zero external server-framework dependencies; protocol and dual-loop streaming limits are documented in the [API reference](docs/reference/API.md) - **CLI** — `lumin agent`, `lumin serve`, `lumin health` -- **~4,900 LOC** — single production dependency (Zod) -- **Capability-validated** — end-to-end dual-loop tests C1 (dialogue decoupling), C3 (polling), C4 (reliable cancel), C5 (per-iteration progress), C6 (concurrent isolation), C7 (cross-task knowledge) all pass against a real LLM. Run with `RUN_CAPABILITY_TESTS=1 npx vitest run tests/capability/` +- **Small production dependency surface** — one direct production dependency (Zod) +- **Capability suites included** — real-LLM checks are opt-in and require explicit credentials/flags; see the [verification status](docs/operations/TEST_COVERAGE.md) ## Quick Start @@ -173,7 +173,7 @@ import { PrismerAgent } from '@prismer/agent-core/agent'; import { OpenAICompatibleProvider } from '@prismer/agent-core/provider'; import { ToolRegistry } from '@prismer/agent-core/tools'; import { SessionStore } from '@prismer/agent-core/session'; -import { MemoryStore, FileMemoryBackend } from '@prismer/agent-core/memory'; +import { FileMemoryBackend, MemoryStore } from '@prismer/agent-core'; import { HookRegistry } from '@prismer/agent-core/hooks'; import { EventBus } from '@prismer/agent-core/sse'; import { loadConfig } from '@prismer/agent-core/config'; @@ -194,12 +194,12 @@ Zero-dependency file-based memory with keyword recall. Tested on the [LoCoMo](ht Zero-dependency keyword search + strong LLM outperforms Letta's embedding+rerank pipeline. ```typescript -import { MemoryStore } from '@prismer/agent-core/memory'; +import { FileMemoryBackend, MemoryStore } from '@prismer/agent-core'; -const memory = new MemoryStore('./workspace'); +const memory = new MemoryStore(new FileMemoryBackend('./workspace')); await memory.store('The calibration coefficient is 0.03847', ['numeric']); const results = await memory.search('calibration coefficient'); -console.log(results); // [{ content: '...', score: 1.0, tags: ['numeric'] }] +console.log(results); // [{ text: '...', score: 1.0, tags: ['numeric'] }] ``` ## Workspace Templates @@ -229,7 +229,7 @@ Each skill is a directory with a `SKILL.md` file (YAML frontmatter + markdown bo ## API Reference -See [docs/API.md](docs/API.md) for the complete HTTP, WebSocket, and IPC protocol documentation. +See [docs/reference/API.md](docs/reference/API.md) for the complete HTTP, WebSocket, and IPC protocol documentation. ## Contributing diff --git a/docs/API.md b/docs/API.md deleted file mode 100644 index 70f7203..0000000 --- a/docs/API.md +++ /dev/null @@ -1,571 +0,0 @@ -# API Reference - -Lumin exposes three interfaces: **HTTP API**, **WebSocket**, and **IPC** (stdin/stdout JSON). - ---- - -## HTTP API - -Start the server with `lumin serve --port 3001` or programmatically via `startServer()`. - -### `GET /health` - -Health check endpoint. - -**Response:** -```json -{ - "status": "ok", - "version": "0.3.1", - "uptime": 12345 -} -``` - -### `GET /v1/tools` - -List all registered tools. - -**Response:** -```json -{ - "tools": [ - { - "name": "bash", - "description": "Execute a bash command in the container.", - "parameters": { ... } - } - ], - "count": 42 -} -``` - -### `POST /v1/chat` - -Send a message and receive the complete response (synchronous). - -**Request:** -```json -{ - "content": "Write a LaTeX survey on attention mechanisms", - "sessionId": "optional-session-id", - "config": { - "model": "gpt-4o", - "agentId": "researcher", - "maxIterations": 20, - "tools": ["latex", "arxiv"] - } -} -``` - -**Routing behavior (dual-loop mode, Phase A):** - -Before a new task is created, the server checks whether the session already has -an active task (via `getActiveForSession(sessionId)`): - -- If **no active task** exists → a new task is created and dispatched in the - background. The response contains `taskId` for the newly-created task. -- If an **active task** exists → the message is enqueued to the task's process- - global message queue (delivered at the next inner-loop iteration boundary), - and the response returns immediately with `queued: true` and the existing - task's `taskId`. No new task is created. - -A `task.message.enqueued` WebSocket event is emitted for each enqueued message. - -**Response:** -```json -{ - "status": "success", - "response": "I'll help you write...", - "thinking": "Let me plan the survey structure...", - "directives": [ - { "type": "SWITCH_COMPONENT", "payload": { "component": "latex-editor" } } - ], - "toolsUsed": ["latex_project", "arxiv_search"], - "usage": { "promptTokens": 1500, "completionTokens": 800, "totalTokens": 2300 }, - "sessionId": "session-1234567890", - "iterations": 3, - "taskId": "task-abc123", - "queued": false, - "loopMode": "dual" -} -``` - -| Field | Type | Description | -|-------|------|-------------| -| `taskId` | `string?` | Present in dual-loop mode. ID of the newly-created or existing active task. | -| `queued` | `boolean?` | `true` if the message was enqueued to an already-running task rather than starting a new one. | -| `loopMode` | `"single"` \| `"dual"` | Current server loop mode. | - -**Error Response:** -```json -{ - "status": "error", - "error": "LLM request failed: timeout", - "sessionId": "session-1234567890" -} -``` - ---- - -### `GET /v1/tasks` - -List all tasks known to the dual-loop agent (active, completed, interrupted, -failed, cancelled). Available only when the server runs in dual-loop mode; the -array will be empty in single-loop mode. - -**Response:** -```json -{ - "tasks": [ - { - "id": "task-abc123", - "sessionId": "session-1234567890", - "instruction": "Write a survey on attention", - "status": "executing", - "checkpoints": [], - "progress": { "iterations": 3, "toolsUsed": ["bash"], "lastActivity": 1712700000000 }, - "createdAt": 1712699000000, - "updatedAt": 1712700000000 - } - ], - "count": 1 -} -``` - -### `GET /v1/tasks/:id` - -Return a single task by ID. - -**Response:** same shape as an element of `GET /v1/tasks`: - -```json -{ - "id": "task-abc123", - "sessionId": "session-1234567890", - "instruction": "Write a survey on attention", - "status": "completed", - "checkpoints": [...], - "progress": { "iterations": 5, "toolsUsed": ["bash","arxiv_search"], "lastActivity": 1712700500000 }, - "plan": { "steps": ["...", "..."] }, - "result": "Survey complete. See /workspace/survey.pdf.", - "error": null, - "createdAt": 1712699000000, - "updatedAt": 1712700500000 -} -``` - -**404 Response** if the task ID is unknown: -```json -{ "error": "Task task-abc123 not found" } -``` - -### `POST /v1/tasks/:id/cancel` - -Cancel the active task identified by `:id` with a structured `AbortReason` -(Phase C7). The cancellation propagates into the inner loop via a per-task -`AbortController`, aborts any in-flight LLM fetch + tool execution, and injects -a synthetic `[Aborted: ]` tool_result into the transcript for any -unresolved tool calls. - -**Request body (optional):** -```json -{ "reason": "user_explicit_cancel" } -``` - -`reason` must be one of: - -| Value | Meaning | -|-------|---------| -| `user_interrupted` | User pressed Ctrl-C / closed WS connection | -| `user_explicit_cancel` | User pressed "Cancel" in UI (default when omitted) | -| `timeout` | Task exceeded its deadline | -| `sibling_error` | A sibling sub-agent failed and cancellation cascaded | -| `server_shutdown` | Server is shutting down (SIGTERM / SIGINT) | - -**Response 200:** -```json -{ "status": "cancelled", "taskId": "task-abc123", "reason": "user_explicit_cancel" } -``` - -**Response 400** — invalid `reason` string. -**Response 404** — task not found. -**Response 409** — task is not in an active status (`executing` / `planning` / `paused`); nothing to cancel. - -> **v1 limitation**: see `handleCancelTask` JSDoc in `src/server.ts` for the -> current per-task cancellation scope and semantics. - -### `POST /v1/tasks/:id/resume` - -Resume an `interrupted` task (Phase B5). On server startup `loadPersistedTasks()` -scans `{workspaceDir}/.lumin/sessions/*/tasks/*.meta.json`; non-terminal tasks -from a prior run are re-registered with status `interrupted`. This endpoint -replays the persisted transcript (`*.jsonl`) and re-dispatches the inner loop. - -**Request body:** none. - -**Response 200:** -```json -{ "status": "resumed", "taskId": "task-abc123", "sessionId": "session-1234567890" } -``` - -**Response 404** — task not found. -**Response 405** — current loop mode does not support resume (single-loop). -**Response 409** — task is not in a resumable state (e.g. already completed). - ---- - -## WebSocket API - -Connect to `ws://:/v1/stream` for real-time streaming. - -### Client → Server - -#### `chat.send` - -Send a message to the agent. - -```json -{ - "type": "chat.send", - "content": "Write a LaTeX paper on transformers", - "sessionId": "optional-session-id", - "config": { - "model": "gpt-4o", - "agentId": "researcher" - } -} -``` - -#### `ping` - -Keep-alive ping. - -```json -{ "type": "ping" } -``` - -### Server → Client - -#### `connected` - -Sent immediately after WebSocket connection. - -```json -{ - "type": "connected", - "sessionId": "session-1234567890", - "version": "0.3.1" -} -``` - -#### `lifecycle.start` - -Agent loop has started processing. - -```json -{ - "type": "lifecycle.start", - "sessionId": "session-1234567890" -} -``` - -#### `text.delta` - -Streaming text token from the LLM. - -```json -{ - "type": "text.delta", - "delta": "I'll help you " -} -``` - -#### `tool.start` - -Tool execution has begun. `toolId` uniquely identifies this invocation (format: `toolName:index`). - -```json -{ - "type": "tool.start", - "tool": "latex_project", - "toolId": "latex_project:0", - "args": { "action": "compile", "file": "main.tex" } -} -``` - -#### `tool.end` - -Tool execution completed. - -```json -{ - "type": "tool.end", - "tool": "latex_project", - "toolId": "latex_project:0", - "result": "Compilation successful. PDF: output/main.pdf" -} -``` - -#### `directive` - -UI directive from a plugin tool. - -```json -{ - "type": "directive", - "directive": { - "type": "SWITCH_COMPONENT", - "payload": { "component": "latex-editor" } - } -} -``` - -#### `chat.final` - -Agent loop completed. Contains the full response. - -```json -{ - "type": "chat.final", - "content": "I've compiled your LaTeX paper...", - "thinking": "...", - "directives": [...], - "toolsUsed": ["latex_project"], - "sessionId": "session-1234567890", - "iterations": 2, - "usage": { "promptTokens": 1500, "completionTokens": 800, "totalTokens": 2300 } -} -``` - -#### `task.created` - -A dual-loop task was created. Emitted immediately after the HTTP `/v1/chat` -response (before the inner loop starts). - -```json -{ - "type": "task.created", - "data": { "taskId": "task-abc123", "sessionId": "session-1234567890", "instruction": "Write a survey..." } -} -``` - -#### `task.planning` - -Emitted when the task enters the planning phase. - -```json -{ - "type": "task.planning", - "data": { "taskId": "task-abc123", "goal": "Write a survey..." } -} -``` - -#### `task.planned` - -Emitted after planning, with the ordered step list. - -```json -{ - "type": "task.planned", - "data": { "taskId": "task-abc123", "steps": ["Search arxiv", "Outline", "Draft", "Compile"] } -} -``` - -#### `task.progress` (Phase A5) - -Emitted once per inner-loop iteration. - -```json -{ - "type": "task.progress", - "data": { - "taskId": "task-abc123", - "iteration": 3, - "toolsUsed": ["bash", "arxiv_search"], - "lastActivity": 1712700000000 - } -} -``` - -#### `task.message.enqueued` (Phase A3) - -Emitted when `POST /v1/chat` enqueues a message into an already-running task -rather than spawning a new one. - -```json -{ - "type": "task.message.enqueued", - "data": { - "taskId": "task-abc123", - "messageId": "msg-xyz", - "content": "also include figure-level results (truncated to 500 chars)" - } -} -``` - -#### `task.message.orphaned` (Phase C5) - -Emitted when a task terminates (completion, failure, or cancellation) while -queued messages remain undrained. One event is emitted per orphaned message. - -```json -{ - "type": "task.message.orphaned", - "data": { - "taskId": "task-abc123", - "messageId": "msg-xyz", - "content": "also include figure-level results", - "reason": "task_completed" - } -} -``` - -`reason` is one of `task_completed` | `task_aborted`. - -#### `task.completed` - -Emitted when the inner loop finishes successfully. - -```json -{ - "type": "task.completed", - "data": { - "taskId": "task-abc123", - "sessionId": "session-1234567890", - "result": "Survey complete. See /workspace/survey.pdf.", - "toolsUsed": ["bash", "arxiv_search"] - } -} -``` - -#### `error` - -An error occurred during processing. - -```json -{ - "type": "error", - "error": "LLM request failed: timeout" -} -``` - -#### `pong` - -Response to a `ping`. - -```json -{ "type": "pong" } -``` - -### Event Flow - -A typical agent interaction produces events in this order: - -``` -Client: chat.send -Server: lifecycle.start -Server: text.delta (0..N tokens) ← LLM streaming -Server: tool.start ← Tool begins -Server: directive (0..N) ← UI directives from tool -Server: tool.end ← Tool completes -Server: text.delta (0..N tokens) ← LLM continues -Server: chat.final ← Done -``` - -Multiple `tool.start → tool.end` cycles may occur in a single turn. Pair them by `toolId`. - ---- - -## IPC Protocol (stdin/stdout) - -Used in embedded mode (`docker exec lumin agent --message "..."`) and CLI piping. - -### Input Format - -JSON messages on stdin, one per line: - -```json -{"type":"message","content":"hello","sessionId":"sess-1","config":{"model":"us-kimi-k2.5"}} -``` - -### Output Format - -Responses are wrapped with markers for reliable parsing: - -``` ----LUMIN_OUTPUT_START--- -{"status":"success","response":"Hello! How can I help?","sessionId":"sess-1","iterations":1} ----LUMIN_OUTPUT_END--- -``` - -### Output Fields - -| Field | Type | Description | -|-------|------|-------------| -| `status` | `"success"` \| `"error"` | Result status | -| `response` | `string` | Agent's text response | -| `thinking` | `string?` | Reasoning content (if thinking model used) | -| `directives` | `Directive[]` | UI directives emitted during execution | -| `toolsUsed` | `string[]` | Tool names invoked during the turn | -| `usage` | `object?` | Token usage stats | -| `sessionId` | `string` | Session identifier | -| `iterations` | `number` | Agent loop iterations | -| `error` | `string?` | Error message (when `status: "error"`) | - ---- - -## Programmatic API - -```typescript -import { runAgent, EventBus } from '@prismer/agent-core'; - -// Simple usage -await runAgent({ - type: 'message', - content: 'Hello', - sessionId: 'my-session', - config: { model: 'us-kimi-k2.5' }, -}); - -// Server mode with custom EventBus -const bus = new EventBus(); -bus.subscribe('*', (event) => console.log(event)); - -await runAgent( - { type: 'message', content: 'Hello' }, - { - bus, - onResult: (result, sessionId) => { - console.log('Agent response:', result.text); - console.log('Tools used:', result.toolsUsed); - }, - }, -); -``` - -### Key Exports - -```typescript -// Core -export { runAgent, RunAgentOptions } from './index.js'; -export { PrismerAgent, AgentResult, AgentOptions } from './agent.js'; -export { loadConfig, resetConfig, LuminConfigSchema, LuminConfig } from './config.js'; -export { createLogger, Logger, LogLevel } from './log.js'; - -// Infrastructure -export { OpenAICompatibleProvider, FallbackProvider, Provider } from './provider.js'; -export { ToolRegistry, Tool, ToolContext } from './tools.js'; -export { EventBus, StdoutSSEWriter } from './sse.js'; -export { SessionStore, Session } from './session.js'; -export { PromptBuilder, PromptSection } from './prompt.js'; -export { SkillLoader, LoadedSkill, SkillMeta } from './skills.js'; -export { MemoryStore } from './memory.js'; -export { HookRegistry, Hook, HookType, HookContext } from './hooks.js'; -export { ChannelManager } from './channels/manager.js'; - -// Tools -export { loadWorkspaceToolsFromPlugin, createTool, createClawHubTool } from './tools/index.js'; - -// Protocol -export { InputMessage, OutputMessage, writeOutput, parseOutput } from './ipc.js'; -``` diff --git a/docs/DUAL_LOOP_ARCHITECTURE.md b/docs/DUAL_LOOP_ARCHITECTURE.md deleted file mode 100644 index 51e30fc..0000000 --- a/docs/DUAL_LOOP_ARCHITECTURE.md +++ /dev/null @@ -1,661 +0,0 @@ -# Agent Architecture — Dual-Loop Execution Mode - -## Lumin — Runtime Mode Switching: Single-Loop / Dual-Loop - -> **Status**: Capability-validated against real LLM. Phase F tests C1, C3, C4, C5, C6, C7 all pass (6/6) — see [`tests/capability/dual-loop-capabilities.test.ts`](../tests/capability/dual-loop-capabilities.test.ts) and [`docs/superpowers/plans/2026-04-13-dual-loop-audit-and-roadmap.md`](./superpowers/plans/2026-04-13-dual-loop-audit-and-roadmap.md). C2 (mid-flight steering) not yet translated to an automated test. Run with `RUN_CAPABILITY_TESTS=1 npx vitest run tests/capability/`. -> **Runtimes**: TypeScript (primary), Rust (core-parity — see §10 for divergences) -> **Mode**: `LUMIN_LOOP_MODE=single` (default) | `dual` -> **Tests**: TS ~700, Rust 619, total ~1,319 + stress 16 -> **Rev**: 6 — Incorporates Phase A/B/C/D/E/F/G (2026-04-15) - ---- - -## 1. Architecture Overview - -``` -┌──────────────────────────────────────────────┐ -│ Single-Loop Mode (default) │ -│ chat.send → PrismerAgent loop → chat.final │ -│ Synchronous: caller blocked until complete │ -└──────────────────────────────────────────────┘ - -┌──────────────────────────────────────────────────────────────┐ -│ Dual-Loop Mode (LUMIN_LOOP_MODE=dual) │ -│ │ -│ Outer Loop (HIL): │ -│ POST /v1/chat → check active task for session │ -│ ├── no active → create Task, spawn inner loop │ -│ └── active → enqueue on MessageQueue, return queued │ -│ artifact store, task state machine (+ `interrupted`) │ -│ │ -│ ┌──────────────────┐ │ -│ │ MessageQueue │ (per DualLoopAgent) │ -│ │ keyed by taskId │ │ -│ └────────┬─────────┘ │ -│ │ drained at iteration boundary │ -│ ▼ │ -│ Inner Loop (EL): background execution │ -│ onIterationStart: drain queue → insert user messages │ -│ PrismerAgent → tools → checkpoints │ -│ per iteration: emit task.progress │ -│ results via EventBus (chat.final / task.completed) │ -│ │ -│ Per-task context: │ -│ taskContexts: Map │ -│ │ -│ Disk persistence (Phase B, TS only): │ -│ {workspaceDir}/.lumin/sessions/{sessionId}/tasks/ │ -│ {taskId}.meta.json — task metadata │ -│ {taskId}.jsonl — transcript (append-only) │ -│ Re-registered as `interrupted` at server startup. │ -└──────────────────────────────────────────────────────────────┘ -``` - -Both modes conform to the same `IAgentLoop` interface. Server code (`server.ts` / `http.rs`) is mode-agnostic. - ---- - -## 2. Implementation Status - -### What's Implemented (TS + Rust) - -| Component | TS | Rust | Status / Notes | -|-----------|:--:|:----:|----------------| -| `IAgentLoop` interface | ✓ | ✓ | Both conform | -| `SingleLoopAgent` → `PrismerAgent` | ✓ | ✓ | Full agent loop | -| `DualLoopAgent` with background execution | ✓ | ✓ | `tokio::spawn` / fire-and-forget | -| Task state machine (including `interrupted`) | ✓ | ✓ | 7 states in TS (adds `interrupted`, Phase B2) | -| `InMemoryTaskStore` (CRUD, active detection) | ✓ | ✓ | — | -| `InMemoryArtifactStore` (add, assign, filter) | ✓ | ✓ | — | -| `WorldModel` (handoff context, fact extraction) | ✓ | ✓ | Regex path + measurement extraction | -| `DirectiveRouter` (realtime/checkpoint/HIL routing) | ✓ | — | TS only | -| `AgentViewStack` (multi-agent UI state) | ✓ | — | TS only | -| `FallbackProvider` (retry + model chain) | ✓ | ✓ | 429/5xx retry, exponential backoff | -| Session persistence (user input in history) | ✓ | ✓ | User messages persisted | -| Memory tools (`memory_store`, `memory_recall`) | ✓ | ✓ | File-based, keyword search | -| HTTP `/v1/chat` (single + dual mode) | ✓ | ✓ | Same JSON schema (camelCase) | -| WebSocket `/v1/stream` | ✓ | ✓ | Same event protocol | -| `chat.final` emission from dual-loop | ✓ | ✓ | Background → EventBus → client | -| Approval gates (sensitive tool confirmation) | ✓ | — | TS only | -| **MessageQueue (Phase A1)** | ✓ | — | Process-global per `DualLoopAgent`, keyed by taskId | -| **`onIterationStart` callback in PrismerAgent (Phase A4)** | ✓ | — | Drains queue at iteration boundary | -| **Per-task `AbortController` (Phase A / C-review)** | ✓ | partial | Rust: iteration-boundary check only | -| **Per-iteration `task.progress` event (Phase A5)** | ✓ | — | Data: `{taskId, iteration, toolsUsed, lastActivity}` | -| **Disk persistence (JSONL + meta.json) (Phase B1)** | ✓ | — | `{workspaceDir}/.lumin/sessions/{sessionId}/tasks/` | -| **`interrupted` task status (Phase B2)** | ✓ | — | Set on non-terminal tasks at restart | -| **Server startup re-register persisted tasks (Phase B4)** | ✓ | — | `loadPersistedTasks()` in `startServer` | -| **Resume endpoint + `resumeTask` (Phase B5)** | ✓ | — | `POST /v1/tasks/:id/resume` | -| **`AbortReason` structured enum (Phase C1/C6)** | ✓ | ✓ | Wire-parity: snake_case serde | -| **Abort propagation into LLM fetch (Phase C2)** | ✓ | — | `fetch(…, { signal })` | -| **Abort-aware tool context (Phase C3)** | ✓ | — | `ToolContext.signal` | -| **Synthetic `[Aborted: ]` tool_result (Phase C4)** | ✓ | — | Filled for unresolved tool_calls | -| **Termination drain of queued messages (Phase C5, Gap 3)** | ✓ | — | Emits `task.message.orphaned` | -| **`POST /v1/tasks/:id/cancel` endpoint (Phase C7)** | ✓ | — | See §4.3 | -| **`POST /v1/tasks/:id/resume` endpoint (Phase B5)** | ✓ | — | See §4.3 / B-series | -| **PermissionMode + ToolPermissionContext (Phase D1/D3)** | ✓ | — | `src/permissions.ts` | -| **`Tool.requiresUserInteraction` + `checkPermissions` (Phase D2)** | ✓ | — | Per-tool override | -| **`enter_plan_mode` / `exit_plan_mode` tools (Phase D4)** | ✓ | — | Flip mode at runtime | -| **Auto-deny in headless mode (Phase D3)** | ✓ | — | Default for dual-loop | -| **WorldModel `knowledgeBase` persisted to MemoryStore (Phase E1/E2)** | ✓ | — | On task completion | -| **TTL-based task eviction (Phase E3)** | ✓ | — | Replaces unbounded in-memory store | -| **Capability test suite C1–C7 (Phase F)** | ✓ | — | `tests/capability/dual-loop-capabilities.test.ts` | - -### What's Not Yet Implemented - -| Feature | Notes | -|---------|-------| -| Clarification-gate mid-iteration pause/resume | Different from Phase B resume (interrupted tasks). State machine supports `paused`, not wired. | -| Sub-agent delegation in dual-loop | TS `@mention` works in single-loop only | -| Multimodal `ContentBlock[]` in Rust | Rust uses `Option`; image/file blocks downgraded | -| `thinkingLevel` control in Rust | TS only | -| Channel adapters in Rust (Telegram, CloudIM) | TS implemented, Rust stubs | -| Directive file scanning | TS only | -| ComponentSpec serialization (Level 1/2/3) | Design only | -| PEP (Prismer Extension Protocol) | Design only | - ---- - -## 3. Core Abstractions - -### 3.1 IAgentLoop Interface - -```typescript -interface IAgentLoop { - readonly mode: 'single' | 'dual'; - processMessage(input: AgentLoopInput, opts?: AgentLoopCallOpts): Promise; - addArtifact(artifact: Artifact): void; - resume(clarification: string): void; - cancel(): void; - shutdown(): Promise; -} -``` - -**Behavioral contract:** -- **Single-loop**: `processMessage()` resolves when agent finishes. Result has full text. -- **Dual-loop**: `processMessage()` resolves immediately (< 100ms). Result has task ID. Actual result arrives via `chat.final` event on EventBus. - -### 3.2 TaskStatus State Machine - -``` -pending → planning → executing → paused → executing (resume) - → executing → completed (terminal) - → executing → failed (terminal) -pending → failed (direct) -``` - -All transitions validated by `TaskStateMachine`. Invalid transitions throw. - -> **Note**: `planning` is a reserved state for a future lightweight planning LLM call before execution dispatch. Current implementation transitions directly from `pending` to `executing`. - -### 3.3 Artifact Store - -```typescript -interface Artifact { - id: string; // UUID, auto-generated - url: string; - mimeType: string; // JSON: "mimeType" - type: ArtifactType; // "image" | "file" | "url", JSON: "type" - addedBy: 'user' | 'agent'; - taskId: string | null; // null = unassigned, assigned on task creation - addedAt: number; -} -``` - -Unassigned artifacts are automatically assigned to new tasks in dual-loop mode. - -### 3.4 WorldModel - -```typescript -interface WorldModel { - taskId: string; - goal: string; - completedWork: AgentCompletionRecord[]; - knowledgeBase: KnowledgeFact[]; // regex-extracted facts - activeComponent: string; - openFiles: string[]; - recentArtifacts: string[]; - componentSummaries: Map; - handoffNotes: Map; -} -``` - -`buildHandoffContext(model, targetAgentId)` produces a compact string injected into the inner loop's system prompt. Budget default is 3,000 chars (~750 tokens) — sized to leave >99% of the context window for the sub-agent's actual work. Configurable via `HANDOFF_BUDGET` constant. - -`extractStructuredFacts(text, agentId)` extracts file paths (`/workspace/...`) and measurements (`42 citations`, `3 figures`) via regex — zero LLM cost. - ---- - -## 4. Execution Flow - -### 4.1 Single-Loop (Default) - -``` -POST /v1/chat { content, sessionId } - │ - ├── SessionStore.getOrCreate(sessionId) - ├── session.addMessage(user input) ← persisted for multi-turn - ├── session.buildMessages(systemPrompt) - │ - └── PrismerAgent.processMessage(input, session) - │ - └── for iteration 1..maxIterations: - ├── LLM call (streaming via EventBus) - ├── if no tool_calls → break, return text - └── execute tools → push results to session - ├── doom-loop detection (configurable, default: 3 consecutive errors) - ├── repetition detection (configurable, default: 5 identical calls) - └── tool result compaction (>140K truncated) - │ - └── Return { status, response, toolsUsed, sessionId, iterations } -``` - -Doom-loop and repetition thresholds are configurable via `AgentOptions` (`doomLoopThreshold`, `repetitionThreshold`). Defaults are empirical values that balance between premature termination and runaway loops. - -### 4.2 Dual-Loop (Phase A–C) - -``` -POST /v1/chat { content, sessionId } - │ - ├── active = getActiveForSession(sessionId) - │ - ├── if active exists: - │ ├── messageQueue.enqueue(active.taskId, { content, messageId }) - │ ├── bus.publish(task.message.enqueued) - │ └── Return: { status: "success", queued: true, taskId: active.taskId } - │ - ├── else (create new task): - │ ├── Assign unassigned artifacts to new task - │ ├── Create Task (pending → executing) - │ ├── Create WorldModel for task - │ ├── taskContexts.set(taskId, { abortController, bus }) - │ ├── Persist initial {taskId}.meta.json + first transcript entry (Phase B) - │ ├── bus.publish(task.created) - │ ├── Return: { status: "success", taskId, loopMode: "dual" } - │ │ - │ └── Background (tokio::spawn / fire-and-forget): - │ ├── bus.publish(task.planning) - │ ├── (optional) planning LLM call → bus.publish(task.planned) - │ ├── Build system prompt with handoff context - │ ├── Create fresh PrismerAgent with: - │ │ - signal: taskContexts.get(taskId).abortController.signal - │ │ - onIterationStart: drain messageQueue → insert user msgs - │ │ + bus.publish(task.progress) - │ ├── Run inner loop: - │ │ for each iteration: - │ │ ├── onIterationStart() — drain queue - │ │ ├── LLM call (signal-aware, Phase C2) - │ │ ├── tool execution (ToolContext.signal, Phase C3) - │ │ └── append to {taskId}.jsonl (Phase B1) - │ ├── On success: - │ │ ├── Complete task (→ completed) - │ │ ├── Persist knowledgeBase → MemoryStore (Phase E1/E2) - │ │ ├── drainQueueOnTermination(taskId, 'task_completed') - │ │ ├── bus.publish(task.completed) - │ │ └── bus.publish(chat.final) - │ └── On error / abort: - │ ├── Fail task (→ failed, or keep `interrupted`) - │ ├── Fill synthetic [Aborted: ] tool_result (Phase C4) - │ ├── drainQueueOnTermination(taskId, 'task_aborted') - │ └── bus.publish(error) -``` - -**Startup (Phase B4):** - -``` -startServer() - └── sharedLoop.loadPersistedTasks() - └── enumerate {workspaceDir}/.lumin/sessions/*/tasks/*.meta.json - ├── terminal tasks (completed/failed/cancelled) → restored as-is - └── non-terminal tasks → re-registered with status: interrupted - (requires explicit POST /v1/tasks/:id/resume to continue) -``` - -**Client-side integration for dual-loop:** - -1. Read `taskId` + `queued` from the HTTP response. -2. Listen for `task.progress`, `task.completed`, `chat.final` on WebSocket, - keyed by `taskId`. -3. If the WebSocket disconnects and reconnects, poll `GET /v1/tasks/:id` to - recover task state and result. -4. Post additional user messages to the same session — the server auto-routes - them to the active task's queue and emits `task.message.enqueued`. - -### 4.3 Cancellation (Phase C) - -``` -Client sends POST /v1/tasks/:id/cancel { reason: "user_explicit_cancel" } - (or chat.cancel WebSocket message, legacy path) - │ - ├── TS (full mid-execution abort): - │ ├── loop.cancel(taskId, reason) - │ ├── taskContexts.get(taskId).abortController.abort(reason) - │ ├── AbortSignal propagates into: - │ │ - LLM fetch (Phase C2) - │ │ - ToolContext.signal → tool implementations (Phase C3) - │ ├── Inner loop unwinds; unresolved tool_calls get [Aborted: ] - │ │ tool_result (Phase C4) - │ ├── drainQueueOnTermination(taskId, 'task_aborted') (Phase C5) - │ └── Task status → cancelled (terminal); bus emits chat.cancelled - │ - └── Rust (iteration-boundary only): - ├── DualLoopAgent.cancel_with_reason(reason) sets Mutex> - ├── Inner loop checks the flag at the top of each iteration - │ (agent.rs:488-493) and returns Err(...) early. (Phase C6 wire-parity) - └── In-flight LLM call + current tool execution are NOT interruptible. -``` - -**Mid-execution abort availability:** - -| | LLM fetch | Tool execution (normal) | Tool execution (`execFileSync` bash) | -|-|:---------:|:-----------------------:|:------------------------------------:| -| TS | ✓ (AbortSignal) | ✓ (ToolContext.signal) | ✗ (synchronous, non-interruptible) | -| Rust | ✗ (boundary only) | ✗ (boundary only) | ✗ | - -> **Per Gate 1 = c**: the Rust-parity policy is "wire-schema parity only" for -> abort — the structured `AbortReason` enum matches between TS and Rust via -> snake_case serde, but runtime PARA (pause / abort / resume / ack) semantics -> are TS-first and deferred to v2.0 in Rust. - ---- - -## 5. JSON Schema (camelCase, unified across TS/Rust) - -### POST /v1/chat Request - -```json -{ "content": "string", "sessionId": "string?" } -``` - -> **Alias deprecation**: Rust currently accepts both `sessionId` (canonical) and `session_id` (legacy alias). The `session_id` alias will be removed in a future version. Clients should use `sessionId`. - -### POST /v1/chat Response - -**Single-loop mode:** -```json -{ - "status": "success", - "response": "The answer is 42.", - "thinking": "string?", - "sessionId": "session-abc", - "toolsUsed": ["bash", "memory_recall"], - "iterations": 3, - "durationMs": 4521, - "usage": { "promptTokens": 1200, "completionTokens": 350 } -} -``` - -**Dual-loop mode** (same endpoint, different semantics): -```json -{ - "status": "success", - "response": "Task a1b2c3 created and executing.", - "sessionId": "session-abc", - "toolsUsed": [], - "iterations": 0, - "durationMs": 8 -} -``` - -Callers can distinguish modes by: `iterations === 0` and `response` contains "Task". A future version will add an explicit `taskId` field and `mode` field to the response. - -### GET /health Response - -```json -{ - "status": "ok", - "version": "0.3.1", - "runtime": "lumin", - "loopMode": "single", - "uptime": 42.5 -} -``` - -### WebSocket /v1/stream Protocol - -``` -Server → { type: "connected", sessionId, version, runtime } -Client → { type: "chat.send", content: "..." } -Server → { type: "text.delta", delta: "..." } (0..N) -Server → { type: "tool.start", tool, toolId } (0..N) -Server → { type: "tool.end", tool, toolId, result } -Server → { type: "chat.final", content, thinking, toolsUsed, sessionId } - -# Dual-loop additions (Phase A/B/C) -Server → { type: "task.created", data: { taskId, sessionId, instruction } } -Server → { type: "task.planning", data: { taskId, goal } } -Server → { type: "task.planned", data: { taskId, steps: string[] } } -Server → { type: "task.progress", data: { taskId, iteration, toolsUsed, lastActivity } } -Server → { type: "task.message.enqueued", - data: { taskId, messageId, content } } # content truncated to 500 chars -Server → { type: "task.message.orphaned", - data: { taskId, messageId, content, - reason: "task_completed" | "task_aborted" } } -Server → { type: "task.completed", - data: { taskId, sessionId, result?, toolsUsed? } } -``` - ---- - -## 6. Factory & Configuration - -### Mode Resolution (4-level priority) - -``` -1. Explicit argument to createAgentLoop(mode) — highest -2. LUMIN_LOOP_MODE environment variable -3. DB field dbLoopMode (per-container) -4. Default: 'single' — lowest -``` - -### Environment Variables - -| Variable | Default | Description | -|----------|---------|-------------| -| `LUMIN_LOOP_MODE` | `single` | `single` or `dual` | -| `OPENAI_API_KEY` | (required) | LLM provider API key | -| `OPENAI_API_BASE_URL` | `https://api.openai.com/v1` | LLM endpoint | -| `AGENT_DEFAULT_MODEL` | `gpt-4o` | Model ID (prefix `prismer-gateway/` stripped) | -| `WORKSPACE_DIR` | `./workspace` (TS) / `/workspace` (Rust) | Working directory | -| `MAX_CONTEXT_CHARS` | `600000` | Compaction threshold | -| `LUMIN_PORT` | `3001` | Server port | - ---- - -## 7. Session Persistence - -User input is persisted to `session.messages` **before** calling `buildMessages()`. This ensures multi-turn conversations work correctly — the LLM sees all prior user messages and assistant responses when building the next response. - -``` -Request 1: "Remember code FALCON99" - → session.messages: [user("Remember..."), assistant("I'll remember...")] - -Request 2: "What was the code?" - → buildMessages(): [system, user("Remember..."), assistant("I'll remember..."), user("What was...")] - → LLM sees full history → recalls FALCON99 -``` - -Previously, user input was only added to the temporary `messages` array for the LLM call but not persisted to `session.messages`, causing recall failure across requests. - ---- - -## 8. Directive System - -### 8.1 Directive Types (21 total) - -| Delivery | Types | -|----------|-------| -| **Realtime** (14) | SWITCH_COMPONENT, TIMELINE_EVENT, THINKING_UPDATE, OPERATION_STATUS, UPDATE_CONTENT, UPDATE_LATEX, UPDATE_CODE, UPDATE_DATA_GRID, UPDATE_GALLERY, JUPYTER_ADD_CELL, JUPYTER_CELL_OUTPUT, EXTENSION_UPDATE, AGENT_CURSOR, HUMAN_CURSOR | -| **Checkpoint** (3) | COMPILE_COMPLETE, NOTIFICATION, COMPONENT_STATE_SYNC | -| **HIL-only** (4) | TASK_UPDATE, UPDATE_TASKS, ACTION_REQUEST, REQUEST_CONFIRMATION | - -### 8.2 DirectiveRouter (TS only) - -Routes directives by delivery mode. Realtime → publish immediately to EventBus. Checkpoint → buffer until next checkpoint event. HIL-only → consumed by outer loop, not forwarded. - -### 8.3 AgentViewStack (TS only) - -Tracks component ownership across multi-agent delegation. Push on delegate, pop on completion, restore parent's active component. - ---- - -## 9. Built-in Agents - -6 agents, identical in TS and Rust: - -| ID | Mode | Tools | Purpose | -|----|------|-------|---------| -| `researcher` | Primary | `null` (no filter — all registered tools available) | Orchestrate, delegate to sub-agents | -| `latex-expert` | Subagent | latex_compile, latex_project, switch_component, update_content, bash | LaTeX writing/compilation | -| `data-analyst` | Subagent | jupyter_execute, jupyter_notebook, switch_component, update_content, bash | Data analysis | -| `literature-scout` | Subagent | arxiv_search, load_pdf, context_search, switch_component, bash | Paper discovery | -| `compaction` | Hidden | `[]` (no tools) | Conversation summarization | -| `summarizer` | Hidden | `[]` (no tools) | Title generation | - -`tools: null` means no tool filtering — the agent sees all tools registered in the `ToolRegistry`. `tools: [...]` restricts to the listed names via `getSpecs(allowedTools)`. - ---- - -## 10. Cross-Runtime Parity - -### Verified by Sync Parity Test (11 tests, real LLM) - -Both TS and Rust servers are started, identical requests are sent, structural equivalence is verified: - -| Test | Verification | -|------|-------------| -| P1 | Health endpoint: same fields, same `loopMode` | -| P2 | Chat response: same `status`, `response`, `sessionId` | -| P3 | Tool calling: both execute `bash`, report in `toolsUsed` | -| P4 | Multi-step: both handle 2+ tool iterations | -| P5 | Session: both recall context across requests | -| P6 | Memory: both `memory_store` + `memory_recall` | -| P7 | Dual-loop: both return quickly with task info (validates immediate `taskId` return only; does NOT validate end-to-end task execution with client subscription — see audit doc §1.3) | -| P8 | Dual-loop health: both report `loopMode: "dual"` | -| P9 | WebSocket: both produce `open → connected → chat.final` | -| P10 | Errors: both handle empty content gracefully | -| P11 | Concurrency: both handle 3 parallel requests | - -### Rust Divergences - -Rust implements **core agent loop parity** (single-loop, dual-loop, tools, sessions, memory, config, provider with fallback). Per the **TS-first / Rust-parity** policy (Gate 1 = c), advanced runtime features land TS-first and Rust tracks wire-schema parity only until v2.0. - -| Capability | TS | Rust | Notes | -|------------|:--:|:----:|-------| -| `AbortReason` enum (wire format) | ✓ | ✓ | snake_case serde, parity verified | -| `POST /v1/tasks/:id/cancel` with `reason` | ✓ | partial | Rust: iteration-boundary abort only | -| Mid-execution abort into LLM fetch | ✓ | ✗ | TS: AbortSignal on `fetch` | -| Mid-execution abort into tool context | ✓ | ✗ | TS: `ToolContext.signal` | -| Synthetic `[Aborted: ]` tool_result | ✓ | ✗ | Phase C4, TS only | -| `PermissionMode` + per-tool policy | ✓ | ✗ | Phase D, TS only | -| Plan-mode tools (`enter_plan_mode` / `exit_plan_mode`) | ✓ | ✗ | Phase D4 | -| Disk persistence (`.lumin/sessions/.../tasks/*`) | ✓ | ✗ | Phase B1 | -| Task resume (`POST /v1/tasks/:id/resume`) | ✓ | ✗ | Phase B5 | -| MessageQueue routing mid-task | ✓ | ✗ | Phase A | -| Termination drain / `task.message.orphaned` | ✓ | ✗ | Phase C5 | -| Approval gates (`needsApproval`, `waitForApproval`) | ✓ | — | **Security: Rust tools execute without confirmation** | -| `DirectiveRouter`, `AgentViewStack` | ✓ | — | No directive routing by delivery mode | -| `@mention` delegation, `delegate` tool | ✓ | — | Single-loop only | -| Multimodal `ContentBlock[]` in Message | ✓ | — | Rust uses `Option` | -| `thinkingLevel` / `temperature` control | ✓ | — | No per-request LLM parameter tuning | -| Channel adapters (Telegram, CloudIM) | ✓ | — | Rust stubs | - -> **Security warning**: Rust runtime should not be used for untrusted tool execution in dual-loop mode until approval gates are implemented. In dual-loop mode, the inner loop executes tools autonomously without human confirmation. - ---- - -## 11. Test Coverage - -See [docs/TEST_COVERAGE.md](./TEST_COVERAGE.md) for full breakdown. - -| Metric | Count | -|--------|------:| -| TS unit tests | 510 | -| TS integration + sync parity | 21 | -| Rust unit tests | 483 | -| Rust integration (LLM) | 8 | -| Stress test scenarios | 16 (8 × 2 runtimes) | -| **Total** | **1,038** | - ---- - -## 12. Future Work - -### Phase Next: Functional Dual-Loop - -1. **Clarification gates**: Inner loop pauses on `REQUEST_CONFIRMATION` directive, outer loop resumes with user response -2. **Dynamic artifact reassignment**: Artifacts uploaded during task execution injected into inner loop context -3. **Task result polling**: `GET /v1/tasks/:id` endpoint for clients that miss the `chat.final` event -4. **Dual-mode response field**: Add explicit `taskId` and `mode` fields to `/v1/chat` response - -### Phase Later: Multi-Agent Orchestration - -5. **SubAgentManager**: `spawn_agent`, `agent_status`, `await_agent` tools for primary agent -6. **Parallel sub-agents**: Primary spawns N sub-agents, awaits all in parallel -7. **File-level write locks**: Per-path async mutex for concurrent sub-agent file access -8. **Rust approval gates**: Port TS approval mechanism to Rust - -### Phase Future: ComponentSpec & PEP - -9. **ComponentSpec serialization**: Level 1 (brief) / Level 2 (structured) / Level 3 (full) per component -10. **PEP (Prismer Extension Protocol)**: Agent-built runtime extensions with hot reload -11. **OT/CRDT for concurrent editing**: Agent + human co-editing same document - ---- - -## 13. Known Limitations - -### Closed since Rev 5 - -- ~~**WorldModel is per-task, not persisted**~~ — CLOSED by Phase E1/E2. `knowledgeBase` facts are written to MemoryStore on task completion. -- ~~**In-memory stores have no eviction**~~ — CLOSED by Phase E3. TTL-based eviction lands for completed/failed/cancelled tasks. -- ~~**Dual-loop result delivery is fire-and-forget**~~ — CLOSED. `GET /v1/tasks/:id` returns status + result; reconnecting clients can poll. -- ~~**Dialogue cannot steer running tasks**~~ — CLOSED by Phase A. MessageQueue delivers dialogue-layer messages at the next iteration boundary; `POST /v1/chat` against an active session auto-routes to the existing task. - -### Still open - -**Rust has no PermissionMode / approval gates.** Per Gate 1 = c, runtime PARA semantics are deferred to v2.0 in Rust. Rust tools execute autonomously in dual-loop mode without human confirmation — do not use Rust runtime for untrusted tool execution. - -**Rust has no disk persistence.** Phase B is TS-only. A Rust server restart loses all in-flight task state; no `interrupted` status, no resume. - -**Rust cancellation is boundary-only.** `DualLoopAgent.cancel_with_reason` sets a `Mutex>` that the inner loop checks at the top of each iteration (agent.rs:488-493). In-flight LLM calls and synchronous tool execution continue until the current iteration yields. Per Gate 1 = c, mid-execution abort is TS-only. - -**bash `execFileSync` cannot be aborted mid-execution.** Even in TS, the synchronous bash tool implementation cannot observe an `AbortSignal` once `execFileSync` has started — the process runs to completion (or kills on timeout) before the signal is checked. Flagged multiple times in review, open. - -**Clarification-gate mid-iteration pause is not wired.** The task state machine has a `paused` state, but no runtime path transitions into it mid-iteration awaiting a user clarification. This is separate from Phase B resume (`interrupted` → rerun from disk); clarification pause means the inner loop halts between a tool call and its tool_result while waiting for a dialogue answer. Future work. - -**`IAgentLoop.cancel(taskId?)` has a v1 simplification.** See the JSDoc on `handleCancelTask` in `src/server.ts` for the current scope — broadly, the cancel endpoint targets the specified `taskId`, but the underlying loop's per-task cancellation is still being refined and may not distinguish siblings when multiple tasks are concurrent in future multi-task-per-session scenarios. - -**Sub-agent delegation is single-loop only.** `@mention` works in the single-loop agent. Dual-loop sub-agent orchestration (`spawn_agent`, `agent_status`, `await_agent`) is future work (§12). - ---- - -## 14. Plan Mode & Permissions (Phase D) - -Phase D introduces a **permission mode** system that gates sensitive tool -execution per-task. It is the TS-side foundation for PARA (pause / abort / -resume / ack) semantics. - -### 14.1 `PermissionMode` - -Defined in `src/permissions.ts`: - -| Mode | Behavior | -|------|----------| -| `default` | Interactive: user is prompted on `requiresUserInteraction` tools | -| `plan` | Plan mode: `requiresUserInteraction` tools are auto-denied; the agent must produce a plan rather than mutate state | -| `auto` | Headless: `requiresUserInteraction` tools are auto-denied (this is dual-loop's default) | -| `bypass` | All gates disabled (dangerous; local dev only) | - -Dual-loop mode defaults to `auto` so background tasks never block awaiting a user confirmation that never arrives. - -### 14.2 `Tool.requiresUserInteraction` + `checkPermissions` - -Every `Tool` may declare two permission hooks: - -```ts -interface Tool { - // Coarse flag — if true, this tool is auto-denied in plan/auto modes unless - // checkPermissions overrides. - requiresUserInteraction?: boolean; - - // Fine-grained override. Called with the current ToolPermissionContext; - // returns allow / deny with an optional reason string injected into the - // tool_result when denied. - checkPermissions?(ctx: ToolPermissionContext): { allow: boolean; reason?: string }; -} -``` - -The `ToolPermissionContext` carries the current `mode`, the `prePlanMode` -(restored on `exit_plan_mode`), tool args, and session info. - -### 14.3 Plan-mode tools - -Two built-in tools flip the mode at runtime: - -- **`enter_plan_mode`** — saves `prePlanMode = mode`, sets `mode = 'plan'`. The - agent is expected to use this before producing a multi-step plan so it can - research read-only without accidentally mutating the workspace. -- **`exit_plan_mode`** — restores `mode = prePlanMode` (or `default`). Normally - called after the agent has drafted the plan and is ready to execute. - -### 14.4 Transitional note - -The `PermissionMode` type currently lives in `src/permissions.ts`. Per -`ReleasePlan-1.9.0.md` item D12, it will be replaced by a re-export from -`@prismer/sandbox-runtime` once that package publishes. Consumers should import -from `@prismer/agent-core` rather than directly from `./permissions.js` so the -eventual swap is transparent. - ---- - -## Appendix: Design Principles - -**Zero regression**: Dual-loop mode is opt-in via `LUMIN_LOOP_MODE=dual`. Single-loop path is completely unchanged. - -**In-process first**: Task store, artifact store, world model are all in-memory. Can be replaced with Redis/DB without changing the `IAgentLoop` interface. - -**Context budget per agent**: Each sub-agent gets a fresh session with configurable handoff context (default ≤ 3K chars, ~750 tokens). Context budget is per-agent, not per-task — scales to arbitrarily complex tasks. - -**Additive events**: New SSE/WS event types (`task.completed`, `chat.final`) are additive. Old clients ignore unknown types. - -**camelCase JSON**: All HTTP/WS/IPC interfaces use camelCase field names. Rust structs use `#[serde(rename_all = "camelCase")]`. Legacy `session_id` alias accepted but deprecated. diff --git a/docs/MEMORY.md b/docs/MEMORY.md deleted file mode 100644 index 8607d10..0000000 --- a/docs/MEMORY.md +++ /dev/null @@ -1,313 +0,0 @@ -# Memory System - -Lumin 的持久化记忆系统,基于 pluggable backend 抽象 + 文件默认实现。 - ---- - -## 架构 - -``` -MemoryStore (facade) - │ - ├── store(text, tags?) → 写入记忆 - ├── recall(query, maxChars?) → 关键词召回(返回格式化字符串) - ├── search(query, opts?) → 结构化搜索(返回 MemorySearchResult[]) - ├── loadRecentContext(max?) → 加载最近记忆(注入 system prompt) - └── close() - │ - └── MemoryBackend (interface) - ├── FileMemoryBackend ← 默认,零依赖 - ├── CloudMemoryBackend ← 预留(Prismer Cloud) - └── VectorMemoryBackend ← 预留(embedding search) -``` - -### FileMemoryBackend (TS + Rust 完全对齐) - -- 存储路径: `/workspace/.prismer/memory/YYYY-MM-DD.md` -- 每条记忆以 `---` 分隔,带时间戳和可选 tags -- 格式: `## HH:MM — [tag1, tag2]\ncontent\n\n---\n\n` -- 搜索: 关键词匹配,score = 匹配关键词数 / 总关键词数(0–1 归一化) -- 关键词过滤: 长度 < 3 的词被忽略 -- Turn-level chunking: 对话格式自动检测 → 3-turn 滑窗(1-turn 重叠) -- Multi-query fallback: 5+ 关键词时自动生成 3 关键词子查询 -- Tag filtering: search 支持按 tag 过滤 -- `recent()` 只加载 today + yesterday 的文件 - -**跨运行时对齐状态** (TS `src/memory.ts` ↔ Rust `rust/crates/lumin-core/src/memory.rs`): - -| 方法 | TS | Rust | 一致 | -|------|:--:|:----:|:----:| -| `store(content, tags)` | ✓ | ✓ | ✓ | -| `recall(query, maxChars)` | ✓ | ✓ | ✓ | -| `search(query, opts)` → `MemorySearchResult[]` | ✓ | ✓ | ✓ | -| `loadRecentContext(maxChars)` | ✓ | ✓ | ✓ | -| `close()` | ✓ | ✓ | ✓ | -| Turn-level chunking | ✓ | ✓ | ✓ | -| Multi-query fallback (5+ keywords) | ✓ | ✓ | ✓ | -| Tag filtering | ✓ | ✓ | ✓ | -| Score normalization (0-1) | ✓ | ✓ | ✓ | -| `capabilities()` | ✓ | ✓ | ✓ | - -### Compaction → Memory Flush 流程 - -当上下文超过 `MAX_CONTEXT_CHARS`(默认 600K)时触发: - -``` -context overflow → memoryFlushBeforeCompaction() - │ - ├── serialize dropped messages (≤8K chars) - ├── LLM 提取关键事实(maxTokens: 500) - └── memoryStore.store(facts, ['auto-flush', 'compaction']) - → compactConversation() - │ - ├── LLM 摘要(maxTokens: 2000) - └── 注入 session.compactionSummary -``` - -注意: `agent.ts:303` 的 `!session.compactionSummary` 守卫使得**每个 Session 只触发一次** memory flush。 - -### 记忆注入 - -`session.buildMessages()` 将记忆上下文追加到 system prompt: - -``` -[system prompt] -## Memory from Previous Sessions -[loadRecentContext() 输出,最多 3000 chars] -``` - ---- - -## 配置 - -| 环境变量 | Config 路径 | 默认值 | 说明 | -|----------|------------|--------|------| -| `MEMORY_BACKEND` | `memory.backend` | `file` | 后端类型: `file` / `cloud` / `vector` | -| — | `memory.recentContextMaxChars` | `3000` | system prompt 中记忆上下文的最大字符数 | - ---- - -## 关键文件 - -| 文件 | 说明 | -|------|------| -| `src/memory.ts` | TS MemoryBackend 接口 + FileMemoryBackend + MemoryStore facade | -| `rust/crates/lumin-core/src/memory.rs` | Rust MemoryStore (完全对齐 TS) | -| `src/compaction.ts` | TS `memoryFlushBeforeCompaction()` + `compactConversation()` | -| `rust/crates/lumin-core/src/compaction.rs` | Rust compaction (含 LLM-based memory flush) | -| `src/agent.ts:296-320` | TS context guard + compaction 触发逻辑 | -| `src/session.ts:53-80` | TS `buildMessages()` — compaction summary + 记忆注入 | -| `src/config.ts:122-128` | TS memory 配置 schema | -| `rust/crates/lumin-core/src/config.rs` | Rust MemoryConfig | - ---- - -## 评测结果 - -### Benchmark 1: 自定义事实召回(Fact Recall) - -测试 Lumin compaction pipeline 对精确事实的记忆保持能力。 - -**方法**: 4 轮 compaction cycle,每轮种入 3 条独立事实(共 12 条),每轮通过真实 LLM 调用种入 → `memoryFlushBeforeCompaction()` 提取 → `compactConversation()` 压缩。最终用 LLM 验证每条事实的召回。 - -**模型**: `us-kimi-k2.5` | **耗时**: 334s | **LLM 调用**: ~36 次 - -``` -Memory Store Recall by Compaction Cycle -─────────────────────────────────────────────────────── -Cycle 1 │████████████████████ 100% 3/3 -Cycle 2 │████████████████████ 100% 6/6 -Cycle 3 │████████████████████ 100% 9/9 -Cycle 4 │████████████████████ 100% 12/12 -─────────────────────────────────────────────────────── -Final LLM Recall (with memory context): 92% (11/12) -``` - -| 事实类别 | 示例 | Store 召回 | LLM 召回 | -|---------|------|-----------|---------| -| numeric | 校准系数 0.03847 | ✓ | ✓ | -| person | Prof. Yolanda Marchetti, Bologna | ✓ | ✓ | -| path | /data/.../run-47b.parquet | ✓ | ✓ | -| decision | Fourier-Bessel 替代 wavelet | ✓ | ✓ | -| config | Redis 7200s, 24 workers | ✓ | ✓ | -| architecture | Apache Pulsar, 5 partitions | ✓ | ✗ (空响应) | -| deadline | ICML 2027, Jan 23 | ✓ | ✓ | -| formula | D = kT/(6πηr) | ✓ | ✓ | -| version | PyTorch 2.4.1, CUDA 12.6 | ✓ | ✓ | -| credential | GCP prismer-research-42 | ✓ | ✓ | - -**结论**: FileMemoryBackend 的关键词搜索对精确事实(数字、人名、路径)保持 100% 召回。LLM 端 92% 丢失主要来自模型偶发空响应,非记忆系统问题。 - ---- - -### Benchmark 2: LoCoMo 公开基准(Long-term Conversation Memory) - -使用 Snap Research 的 [LoCoMo](https://github.com/snap-research/locomo) 数据集评测。LoCoMo 是学术界广泛引用的长期对话记忆基准,包含 10 组多会话对话(~300 轮/组)和标注的 QA 对。 - -**方法**: 将 19 个会话(369 轮,43K chars)存入 FileMemoryBackend → 对每个 QA 问题进行关键词搜索(`maxChars: 6000`)→ 用 LLM 基于检索的记忆片段回答 → LLM-as-judge 评分。 - -**样本**: conv-30 (最小样本) | **QA**: 56 题 - -#### 多模型 × 多运行时对比 - -``` -LoCoMo Benchmark — Model × Runtime Comparison (FileMemoryBackend, keyword search) -═══════════════════════════════════════════════════════════════════════════════════════════ - Opus 4.6 (Rust) Opus 4.6 (TS) kimi-k2.5 (TS) glm-4.6 (TS) -─────────────────────────────────────────────────────────────────────────────────────────── -Single-hop (11) ██████████ 100% ██████████ 100% ██░░░░░░░░ 18% ██░░░░░░░░ 27% -Temporal (15) ██████████ 100% █████████░ 93% ███████░░░ 73% █░░░░░░░░░ 13% -Open-domain(15) ██████████ 100% █████████░ 93% ██████░░░░ 67% ███░░░░░░░ 33% -Adversarial(15) ████░░░░░░ 40% ██████░░░░ 60% ████████░░ 80% █████░░░░░ 53% -─────────────────────────────────────────────────────────────────────────────────────────── -Overall 84% 86% 63% ~48% -No adversarial 100% 95% 56% ~35% -═══════════════════════════════════════════════════════════════════════════════════════════ -Baseline: Letta/MemGPT filesystem ≈ 74% on full LoCoMo -``` - -| 模型 | 运行时 | Overall | No Adv. | Single-hop | Temporal | Open-domain | Adversarial | -|------|--------|---------|---------|------------|----------|-------------|-------------| -| **Claude Opus 4.6** | **Rust** | **84%** | **100%** | **100%** (11/11) | **100%** (15/15) | **100%** (15/15) | 40% (6/15) | -| **Claude Opus 4.6** | **TS** | **86%** | **95%** | **100%** (11/11) | **93%** (14/15) | **93%** (14/15) | 60% (9/15) | -| us-kimi-k2.5 | TS | 63% | 56% | 18% (2/11) | 73% (11/15) | 67% (10/15) | **80%** (12/15) | -| glm-4.6 | TS | ~48%* | ~35%* | ~27% (3/11) | ~13% (2/15) | ~33% (5/15) | ~53% (8/15) | -| *Letta/MemGPT* | — | *~74%* | — | — | — | — | — | - -\* glm-4.6 在 Q50/56 超时(20 分钟限制),数据为部分结果。 - -#### Rust vs TS Memory Pipeline 对比 - -Claude Opus 4.6 在两个运行时上的 LoCoMo 结果验证了 **Rust memory pipeline 与 TS 完全对齐**: - -| 维度 | Rust pipeline | TS pipeline | 差异 | -|------|:----------:|:--------:|:----:| -| Single-hop | 100% | 100% | = | -| Temporal | **100%** | 93% | Rust +7pp | -| Open-domain | **100%** | 93% | Rust +7pp | -| Non-adversarial | **100%** | 95% | Rust +5pp | -| Adversarial | 40% | **60%** | TS +20pp | -| Overall | 84% | **86%** | TS +2pp | - -Rust 在事实类问题上达到**完美 100%(41/41 non-adversarial)**,Temporal 和 Open-domain 均超过 TS。TS overall 略高(86% vs 84%)是因为 Adversarial 类别 TS 更好(60% vs 40%)——这是 **LLM 判定差异**(不同运行的模型 calibration 随机性),非 memory pipeline 差异。两个运行时使用相同的搜索算法(keyword matching + turn-level chunking + multi-query fallback),结果一致性证明了代码对齐的正确性。 - -#### 各模型分析 - -**Claude Opus 4.6 — Rust (84%, non-adv 100%)**: -- **Non-adversarial 完美得分**: Single-hop 100%, Temporal 100%, Open-domain 100% -- Rust memory pipeline 的 turn-level chunking 和 multi-query fallback 完全生效 -- Adversarial 40% 低于 TS 运行(60%)——这是 LLM 不同运行间的随机性,非代码差异 -- **验证了 Rust ↔ TS memory pipeline 的功能对齐** - -**Claude Opus 4.6 — TS (86%, non-adv 95%)**: -- **超越 Letta/MemGPT baseline (+12pp)**,即使使用零依赖关键词搜索 -- Single-hop 100%: 能从噪声较多的检索结果中精准定位事实 -- Temporal 93%: 准确解析会话头部的日期时间标记 -- Open-domain 93%: 叙事理解和细节提取能力极强 -- Adversarial 60%: 唯一短板 — 倾向于尝试回答陷阱题而非拒绝(过度自信) -- **结论**: 模型能力是记忆系统效果的决定性因素,强 LLM 可以弥补搜索质量不足 - -**Kimi K2.5 (63%)**: -- Temporal 73%: 日期关键词匹配有效 -- Adversarial 80%: 识别陷阱题的能力优于 Claude -- Single-hop 18%: 弱 — 从长上下文中提取分散事实的能力不足 -- 性价比适中,适合非关键场景 - -**GLM-4.6 (~48%, partial)**: -- 推理速度最慢(20 分钟未完成 56 题) -- 经常返回空响应或超时错误 -- Temporal 最差(~13%)— 无法有效解析日期上下文 -- 不推荐用于记忆召回场景 - -#### P0+P1 搜索优化实验(回归 — 已弃用) - -额外测试了两项搜索优化(使用 kimi-k2.5): -- **P0 Turn-level chunking**: `splitIntoChunks()` 对 >500 chars 条目进行滑动窗口切分 -- **P1 Multi-query**: 对 ≥5 关键词查询生成 3 关键词子窗口 - -**结果**: Overall 从 63% 降至 46%(↓17pp),Temporal 从 73% 暴跌至 20%(↓53pp)。 - -| 原因 | 影响 | -|------|------| -| Turn-level chunks 丢失会话头部时间戳 | Temporal ↓53pp | -| 小 chunks 关键词密度高,挤掉完整会话上下文 | Open-domain ↓14pp | -| LLM 收到碎片化片段,无法拼出连贯叙事 | 大量 "I don't have that information" | - -**结论**: P0+P1 对对话记忆场景有害。代码保留在 `src/memory.ts` 中但不推荐启用。真正的瓶颈不是搜索粒度,而是语义理解 — 如 Claude Opus 所证明,强 LLM + 粗粒度搜索 > 弱 LLM + 细粒度搜索。 - -#### 与 Letta/MemGPT Baseline 对比 - -| 维度 | Lumin + Claude Opus | Lumin + Kimi K2.5 | Letta (filesystem) | -|------|---------------------|-------------------|-------------------| -| 整体准确率 | **86%** | 63% | ~74% | -| 检索方式 | 关键词匹配(score = hits/keywords) | 同左 | 全文搜索 + embedding rerank | -| 存储粒度 | 整个会话 (~2K chars/条) | 同左 | 可配置 chunk | -| 依赖 | 零依赖(纯 Node.js fs) | 同左 | Python + OpenAI embedding API | - ---- - -### 瓶颈分析与改进方向 - -#### 核心发现 - -多模型对比揭示了记忆系统的真正瓶颈: - -1. **模型能力 >> 搜索优化**: Claude Opus 86% vs Kimi K2.5 63%(相同搜索引擎,+23pp),而 P0+P1 搜索优化在 Kimi 上反而 -17pp。LLM 从噪声中提取信号的能力是决定性因素。 -2. **关键词搜索已够用**: 配合强 LLM,零依赖的关键词搜索(FileMemoryBackend)已超越 Letta/MemGPT 的 embedding + rerank 方案(86% vs 74%)。 -3. **搜索粒度不宜过细**: Turn-level chunking 丢失上下文(日期、叙事弧),对对话记忆有害。整会话存储(~2K chars)是当前最优粒度。 -4. **Adversarial 与模型个性相关**: Claude 倾向回答(60%),Kimi 倾向拒绝(80%),GLM 居中(53%)。这是模型 calibration 差异,非记忆系统问题。 - -#### 改进路径(修订) - -| 优先级 | 方向 | 预期提升 | 实现复杂度 | 状态 | -|--------|------|---------|-----------|------| -| ~~P0~~ | ~~Turn-level chunking~~ | ~~+10-15%~~ **实际: -17%** | 低 | ✗ 已验证有害 | -| ~~P1~~ | ~~Multi-query search~~ | ~~+5-8%~~ **贡献不明** | 低 | ✗ 已验证无效 | -| P2 | **升级默认模型** — 使用 Claude Opus 或同级别模型 | +23pp (已验证) | 零 (仅改配置) | ✓ 已验证 | -| P3 | **Adversarial calibration** — 调整 system prompt 使强模型适度拒答 | +5-10% overall | 低 | 待实现 | -| P4 | **Semantic search backend** — embedding 向量检索 | +5-10% (边际) | 中 (需向量 DB 或 API) | 待评估 | -| P5 | **Multi-sample evaluation** — 扩展到 LoCoMo 全部 10 samples | 验证鲁棒性 | 低 (仅增加测试时间) | 待实现 | - ---- - -## 测试 - -### TypeScript -```bash -# 单元测试 (27 tests, 0 LLM calls) -npx vitest run tests/memory.test.ts - -# 自定义事实召回 benchmark (需要 LLM gateway, ~5 min) -npx vitest run tests/memory-recall-benchmark.test.ts - -# LoCoMo 公开基准 (需要 LLM gateway, ~14 min) -npx vitest run tests/locomo-benchmark.test.ts -``` - -### Rust -```bash -# 单元测试 (49 tests, 0 LLM calls) -cargo test --workspace -- memory::tests - -# LoCoMo benchmark via Rust memory pipeline + Claude -node tests/benchmark/locomo-rust-claude.mjs -# 需要 local Claude at http://localhost:3456/v1/messages -``` - -### 查看结果 -```bash -cat tests/output/locomo-benchmark.json | jq '.categoryScores' -cat tests/output/locomo-benchmark-claude-opus-4-6.json | jq '.categoryScores' -``` - -**测试文件**: - -| 文件 | 运行时 | 类型 | 说明 | -|------|--------|------|------| -| `tests/memory.test.ts` | TS | 单元测试 | 27 tests: FileMemoryBackend + MemoryStore facade | -| `rust/crates/lumin-core/src/memory.rs` (tests) | Rust | 单元测试 | 49 tests: search, chunking, multi-query, tags | -| `tests/memory-recall-benchmark.test.ts` | TS | 集成 benchmark | 自定义 12 事实 × 4 compaction cycles | -| `tests/locomo-benchmark.test.ts` | TS | 公开 benchmark | LoCoMo 数据集, 56 QA, LLM-as-judge | -| `tests/benchmark/locomo-rust-claude.mjs` | Rust pipeline | 公开 benchmark | LoCoMo + Rust memory + Claude Opus 4.6 | -| `tests/fixtures/locomo10.json` | — | 数据集 | LoCoMo 10 samples (2.8MB) | -| `tests/output/*.json` | — | 结果 | Benchmark 输出 (gitignored) | diff --git a/docs/README.md b/docs/README.md new file mode 100644 index 0000000..b487738 --- /dev/null +++ b/docs/README.md @@ -0,0 +1,50 @@ +# Lumin 文档中心 + +> **工程事实基线**:2026-08-14,源码提交 `9f1bf56`,包版本 `0.3.1`。 + +`docs/README.md` 是当前文档的唯一入口。参考文档按职责维护;历史分析、已完成计划、阶段测量和旧素材统一位于 [`archive/`](./archive/README.md)。当文档与实现冲突时,以源码、类型定义和可执行测试为准。 + +## 当前工程状态 + +- [工程状态审计](./operations/PROJECT_STATUS.md):版本、交付面、验证快照、风险和建议。 +- [测试与验证状态](./operations/TEST_COVERAGE.md):可复现命令、跳过策略、Rust 和安全审计结果。 + +## API 参考 + +- [API Reference](./reference/API.md):HTTP、WebSocket、IPC、程序化 API 和嵌入式入口。 + +## 架构与约束 + +- [单循环 / 双循环架构](./architecture/DUAL_LOOP_ARCHITECTURE.md):任务状态、持久化、取消和跨运行时边界。 +- [Memory System](./architecture/MEMORY.md):后端接口、压缩集成、TS/Rust 差异和历史评测。 +- [工具完成与事件顺序不变量](./architecture/TOOL_COMPLETION_INVARIANTS.md):工具调度、事件因果关系和终止条件。 + +## 开发者与代理入口 + +- [`../CLAUDE.md`](../CLAUDE.md):开发者快速开始和工程概览。 +- [`../AGENTS.md`](../AGENTS.md):仓库级编码代理协作规则。 +- [`../templates/base/AGENTS.md`](../templates/base/AGENTS.md):生成到用户 workspace 的运行时 agent 模板,不是仓库开发指南。 + +## 历史归档 + +- [归档索引](./archive/README.md) +- [历史分析](./archive/analysis/) +- [设计、实施计划与阶段测量](./archive/plans/) +- [旧素材](./archive/assets/) + +归档内容只代表原记录时点,不作为当前 API 或能力承诺。 + +## 事实优先级 + +1. 源码、类型定义和运行时 schema。 +2. 在当前提交上可复现的测试、构建和审计输出。 +3. 本目录中的当前参考文档。 +4. `archive/` 中的历史材料。 + +## 维护约定 + +- 当前文档应写明复核日期或适用提交。 +- API 只记录网关实际透传的协议;内部 EventBus 事件单独标注。 +- 测试结果区分发现、通过、跳过和失败,不用测试数量代替 coverage。 +- TS/Rust 对齐按能力逐项记录,不使用“完全一致”概括存在差异的实现。 +- 完成、被替代或超过一个发布周期未维护的设计与计划移入 `archive/plans/`,不留在当前目录。 diff --git a/docs/TEST_COVERAGE.md b/docs/TEST_COVERAGE.md deleted file mode 100644 index 1bc3321..0000000 --- a/docs/TEST_COVERAGE.md +++ /dev/null @@ -1,140 +0,0 @@ -# Test Coverage — @prismer/agent-core - -## Summary - -| Runtime | Unit Tests | Integration Tests | Sync Parity | Stress Test | -|---------|-----------|------------------|-------------|-------------| -| **TypeScript** | 510 | 10 (LLM) + 11 (sync) | 11 | 8/8 | -| **Rust** | 483 | 8 (LLM) | (covered by sync) | 8/8 | - -## TypeScript Tests (521 total) - -Run: `npx vitest run` - -LLM integration tests require env vars: -```bash -OPENAI_API_KEY=... OPENAI_API_BASE_URL=... AGENT_DEFAULT_MODEL=... npx vitest run -``` - -### By Module - -| Module | Tests | Description | -|--------|------:|-------------| -| locomo-benchmark | 60 | Memory recall accuracy benchmark | -| directives | 42 | Directive types, serialization, payloads | -| task | 40 | Task state machine, store CRUD, transitions | -| directive-router | 37 | Routing by delivery mode (realtime/checkpoint/HIL) | -| builtins | 32 | Built-in tool implementations | -| loop | 29 | Loop factory, mode resolution, SingleLoopAgent | -| memory-recall-bench | 29 | Memory system benchmark | -| sse | 28 | EventBus pub/sub, backpressure, schemas | -| memory | 27 | File-based memory store/recall | -| ipc | 22 | IPC protocol, markers, serialization | -| session | 20 | Session history, buildMessages, child sessions | -| server | 19 | HTTP/WS server endpoints | -| config | 17 | Config loading, env vars, defaults | -| agents | 17 | Agent registry, built-in agents, mention resolution | -| skills | 17 | SKILL.md parsing, frontmatter, scanning | -| dual-loop | 16 | DualLoopAgent, WorldModel, task lifecycle | -| log | 16 | Structured logging | -| approval | 16 | Sensitive tool approval flow | -| provider | 15 | LLM provider, streaming, thinking models | -| agent | 13 | Core agent loop, doom detection, compaction | -| sync-parity | 11 | TS ↔ Rust behavioral equivalence | -| prompt | 11 | Prompt builder, priority ordering | -| observer | 11 | Event observation | -| hooks | 11 | Lifecycle hooks | -| compaction | 10 | Context overflow compaction | -| integration | 10 | LLM integration (real API calls) | -| workspace | 2 | Workspace middleware | - -## Rust Tests (491 total) - -Run: `cargo test --workspace` - -LLM integration tests require the same env vars: -```bash -OPENAI_API_KEY=... OPENAI_API_BASE_URL=... AGENT_DEFAULT_MODEL=... cargo test --workspace -- --test-threads=1 -``` - -### By Module - -| Module | Tests | Description | -|--------|------:|-------------| -| directives | 72 | Directive types, serialization, camelCase JSON | -| provider | 35 | Message constructors, parse_response, FallbackProvider | -| task | 33 | State machine, all transitions, store CRUD | -| sse | 31 | EventBus pub/sub, clone, event types | -| tools | 30 | ToolRegistry CRUD, bash tool, get_specs | -| ipc | 27 | IPC protocol, InputMessage/OutputMessage, roundtrip | -| memory | 25 | File-based store/recall, keywords, scoring | -| agents | 22 | Built-in agents, registry, mention resolution | -| config | 21 | Config defaults, from_env, serde roundtrip | -| session | 20 | Session CRUD, buildMessages, child sessions | -| world_model | 19 | KnowledgeFact, handoff context, fact extraction | -| loop_factory | 18 | Mode resolution, factory creation | -| loop_types | 18 | LoopMode, AgentLoopInput/Result, ImageRef | -| loop_dual | 18 | DualLoopAgent, task creation, cancel, shutdown | -| skills | 17 | SKILL.md parsing, directory scanning | -| artifacts | 15 | Artifact store, assignment, filtering | -| prompt | 13 | PromptBuilder, priority ordering, sections | -| agent | 13 | MockProvider, doom loop, tool compaction, usage | -| loop_single | 11 | SingleLoopAgent, no-op methods | -| compaction | 10 | Truncation, orphan repair | -| hooks | 8 | Hook registry, before_prompt/tool, after_tool | -| workspace | 7 | WorkspaceConfig loading | -| integration | 8 | Real LLM: provider, agent, loops, thinking | - -## Sync Parity Test (11 tests) - -Run: `npx vitest run tests/sync-parity.test.ts` (requires LLM env vars) - -Starts both TS and Rust servers, sends identical requests, verifies structural equivalence: - -| Test | Verification | -|------|-------------| -| P1 | Health endpoint: same fields, same `loopMode` value | -| P2 | Basic text response: same `status`, `response`, `sessionId` fields | -| P3 | Tool calling: both execute `bash`, report in `toolsUsed` | -| P4 | Multi-step tools: both handle 2+ tool iterations | -| P5 | Session persistence: both recall context across requests | -| P6 | Memory tools: both `memory_store` + `memory_recall` via LLM | -| P7 | Dual-loop: both return quickly with task info | -| P8 | Dual-loop health: both report `loopMode: "dual"` | -| P9 | WebSocket: both produce `open → connected → chat.final` sequence | -| P10 | Error handling: both handle empty content gracefully | -| P11 | Concurrency: both handle 3 parallel requests | - -## Stress Test (8 scenarios × 2 runtimes) - -Run: `node tests/benchmark/dual-loop-stress.mjs` - -| Scenario | Description | -|----------|-------------| -| S1 | Single-loop basic text response | -| S2 | Single-loop tool calling (bash) | -| S3 | Multi-turn tool usage | -| S4 | Dual-loop quick return | -| S5 | Dual-loop mode verification | -| S6 | Concurrent 3× requests | -| S7 | Session persistence (with retry) | -| S8 | WebSocket lifecycle events | - -## JSON Schema Alignment - -All HTTP/WS/IPC interfaces use **camelCase** field names: - -| Field | JSON Key | -|-------|----------| -| Session ID | `sessionId` | -| Tools used | `toolsUsed` | -| Loop mode | `loopMode` | -| Prompt tokens | `promptTokens` | -| Completion tokens | `completionTokens` | -| Duration | `durationMs` | -| MIME type | `mimeType` | -| Artifact ID | `artifactId` | -| Task ID | `taskId` | -| Created at | `createdAt` | - -Rust structs use `#[serde(rename_all = "camelCase")]` with `#[serde(alias = "snake_case")]` for backward compatibility. diff --git a/docs/TOOL_COMPLETION_INVARIANTS.md b/docs/TOOL_COMPLETION_INVARIANTS.md deleted file mode 100644 index 55546c8..0000000 --- a/docs/TOOL_COMPLETION_INVARIANTS.md +++ /dev/null @@ -1,209 +0,0 @@ -# Tool Completion Ordering Invariants - -This document records critical architectural invariants for tool execution -ordering and stream finalization. These constraints are derived from a -real-world bug class (OpenClaw "dead state") where the agent runtime sends -`agent.end` / `chat.final` before all tool results have been delivered to -clients. Violating these invariants causes the frontend to show an -incomplete response with no further progress — a silent failure that is -extremely hard to diagnose in production. - ---- - -## The OpenClaw Bug (Root Cause Analysis) - -OpenClaw v2026.2.26 has a three-layer failure: - -### Layer 1 — Pi Framework (fire-and-forget tool handlers) - -```typescript -// pi-embedded-subscribe.handlers.ts (v2026.2.26) -case "tool_execution_start": - handleToolExecutionStart(ctx, evt).catch(err => { ... }); - return; // Non-blocking — does not await -case "tool_execution_end": - handleToolExecutionEnd(ctx, evt).catch(err => { ... }); - return; // Non-blocking — does not await -``` - -Tool handlers run as fire-and-forget promises. The `handleAgentEnd` lifecycle -event fires independently, with **no coordination** against pending tool -handlers. When a tool takes >100ms to complete, `agent_end` arrives first. - -### Layer 2 — Gateway (immediate cleanup on lifecycle:end) - -```typescript -// server-chat.ts (v2026.2.26) -// On lifecycle:end: -emitChatFinal(...); // Sends chat.final + cleans up buffers -// No flush of 150ms-throttled deltas before final -// No check for pending tool_result events -``` - -The gateway has a 150ms delta throttle — assistant text deltas are batched -and sent at 150ms intervals. When `lifecycle:end` fires, `emitChatFinal` -does **not** call `flushBufferedChatDeltaIfNeeded()` first (fixed in later -versions), so the last batch of text can be lost. More critically, it does -not wait for any pending `tool_result` events. - -### Layer 3 — Client (immediate WS close) - -The WebSocket client (our `openclawGatewayClient.ts`) originally closed the -connection immediately on `lifecycle:end` or `chat.final`, discarding any -late-arriving tool results. - -**Result:** Frontend shows 10 `tool.start` events, 0 `tool.end` events, -then silence. The agent appears frozen ("dead state"). - ---- - -## Why Lumin Is Safe (Current Architecture) - -Lumin's agent loop in `agent.ts` is **synchronous at the iteration level**: - -```typescript -// agent.ts — The critical section -const toolResults = await Promise.all( - response.toolCalls.map(async (call) => { - bus.publish({ type: 'tool.start', ... }); // ← Client sees tool start - const result = await tools.execute(...); // ← Await completion - bus.publish({ type: 'tool.end', ... }); // ← Client sees tool end - return result; - }) -); -// Only AFTER all tools complete: -// → Push tool results into messages[] -// → Loop back to LLM, or break if no more tool calls - -// agent.end is published AFTER the loop exits: -bus.publish({ type: 'agent.end', ... }); -``` - -The guarantee chain: - -1. `tool.start` → `await execute()` → `tool.end` — each tool completes - before its end event fires. -2. `Promise.all` — all parallel tools must resolve before the iteration - continues. -3. `agent.end` — only published after the `while` loop exits, which - requires `!response.toolCalls?.length` (no more tools to call). -4. **Single process** — EventBus is in-memory; `publish()` is synchronous - delivery to all subscribers. No network hop, no message reordering. - -**Therefore:** A client that receives `agent.end` is guaranteed to have -already received every `tool.start` / `tool.end` pair. - ---- - -## Invariants That MUST Be Maintained - -### INV-1: tool.end before agent.end - -`agent.end` MUST NOT be published until every `tool.end` for the current -iteration has been published. This is currently guaranteed by `await -Promise.all(...)` gating the loop iteration. - -**If violated:** Frontend shows tools "in progress" that never complete. - -### INV-2: No fire-and-forget tool execution - -Tool execution MUST be awaited. Never use `.catch(() => {})` / `.then()` -patterns that detach tool execution from the agent loop control flow. - -**If violated:** Same as OpenClaw Layer 1 — agent loop proceeds while tools -are still running. - -### INV-3: Text flush before final - -If text deltas are throttled or buffered (e.g., for streaming optimization), -all buffered text MUST be flushed before `agent.end` is published. - -**If violated:** Last few tokens of the assistant response are silently -dropped. User sees truncated text. - -### INV-4: EventBus ordering preserves causality - -Events published in sequence on the EventBus MUST be delivered to -subscribers in the same order. The current synchronous `publish()` loop -guarantees this. If EventBus ever becomes async (batched, networked), this -invariant needs explicit enforcement (sequence numbers, ordered delivery). - -**If violated:** Client may receive `tool.end` before `tool.start`, or -`agent.end` before `tool.end`, even if the agent published them in order. - -### INV-5: Directive scan after tool batch - -`scanDirectiveFiles()` MUST run after `Promise.all(toolResults)` resolves, -not concurrently. Directive files are written by tools during execution — -scanning before all tools finish will miss files from slower tools. - -**If violated:** UI directives (component switches, content updates) are -lost for slow-completing tools. - ---- - -## Future Risk Scenarios - -### Scenario A: Parallel tool execution with timeouts - -If Lumin adds per-tool timeouts (e.g., kill a tool after 60s), the timeout -handler must still emit `tool.end` with an error result. The agent loop -must not proceed to `agent.end` until all timed-out tools have their -`tool.end` events emitted. - -```typescript -// WRONG — breaks INV-1: -const timeout = setTimeout(() => { /* silently cancel */ }, 60_000); - -// CORRECT: -const timeout = setTimeout(() => { - bus.publish({ type: 'tool.end', data: { tool, toolId, result: 'Error: timeout' } }); - // resolve the Promise.all entry with error -}, 60_000); -``` - -### Scenario B: Streaming tool results - -If tools emit incremental results (e.g., streaming shell output), the -`tool.end` event must only fire after the stream is fully consumed. Do -not emit `tool.end` when the stream starts — emit it when it ends. - -### Scenario C: Sub-agent tool delegation - -When a tool delegates to a sub-agent (`handleDelegateCall`), the outer -agent's `Promise.all` already awaits the sub-agent's full completion. -If sub-agents ever run in a separate process or over a network boundary, -ensure the delegation call blocks until the sub-agent emits its own -`agent.end`. - -### Scenario D: WebSocket gateway layer - -If Lumin adds a gateway process (separate from the agent process) that -proxies events over WebSocket, the gateway MUST NOT close the connection -or emit `chat.final` until it has forwarded all events through `agent.end`. -Use the same pattern as the client-side fix: track pending `tool.start` -events and only finalize when all `tool.end` events have been forwarded. - -### Scenario E: Multi-agent orchestration - -If multiple agents share an EventBus or event transport, ensure that -`agent.end` from Agent A does not cause the transport layer to drop -pending events from Agent B. Each agent's event stream should have -independent lifecycle management. - ---- - -## Testing Checklist - -When modifying tool execution, the agent loop, or the EventBus: - -- [ ] Emit a message with 5+ parallel tool calls. Verify every `tool.start` - has a matching `tool.end` before `agent.end` arrives. -- [ ] Add an artificial 2s delay to one tool. Verify `agent.end` waits for - the slow tool. -- [ ] Simulate a tool that throws. Verify `tool.end` still fires (with - error result) and `agent.end` still fires after. -- [ ] If using streaming: verify text deltas are fully flushed before - `agent.end`. -- [ ] If using WebSocket proxy: verify WS stays open until `agent.end` is - forwarded to the client. diff --git a/docs/architecture/DUAL_LOOP_ARCHITECTURE.md b/docs/architecture/DUAL_LOOP_ARCHITECTURE.md new file mode 100644 index 0000000..378af58 --- /dev/null +++ b/docs/architecture/DUAL_LOOP_ARCHITECTURE.md @@ -0,0 +1,330 @@ +# Lumin 单循环 / 双循环架构 + +> **状态**:TypeScript 双循环已具备任务化执行骨架和基础恢复能力,但仍有协议透传、完整转录和并发隔离缺口。 +> **复核日期**:2026-08-14,提交 `9f1bf56`。 +> **运行时**:TypeScript 为主实现;Rust 仅部分能力/线格式对齐;Embedded 只提供 TypeScript 单循环。 +> **模式**:`LUMIN_LOOP_MODE=single`(默认)或 `dual`。 +> **验证**:默认 Vitest 679 passed / 58 skipped;Rust 结果见 [TEST_COVERAGE.md](../operations/TEST_COVERAGE.md)。 + +## 1. 模式总览 + +```text +single(默认) + +request ──> SingleLoopAgent ──> runAgent / PrismerAgent ──> final result + 调用方等待完整执行 + + +dual(显式启用) + +request ──> DualLoopAgent.processMessage + │ + ├─ session 无 executing/paused 任务:创建 Task,启动后台 inner loop,立即返回 taskId + │ + └─ session 有 executing/paused 任务:消息进入 task MessageQueue,立即返回 queued=true + +background inner loop + planning LLM ──> PrismerAgent iterations ──> task completed/failed + │ │ │ + ├─ plan ├─ iteration 边界排空消息 ├─ task store/result checkpoint + └─ 5s race timeout └─ progress/internal events└─ memory + disk metadata +``` + +两个模式都实现 `IAgentLoop`。关键差异是 `processMessage()` Promise 的含义: + +| 模式 | Promise resolve 表示 | `AgentLoopResult` | +|------|----------------------|-------------------| +| single | agent 已结束 | 完整 text/directives/tools/usage | +| dual | 任务已创建或消息已入队 | `iterations: 0`,含 `taskId`,可能含 `queued: true` | + +## 2. TypeScript 组成 + +### 2.1 公共抽象 + +`src/loop/types.ts` 定义: + +```typescript +interface IAgentLoop { + readonly mode: 'single' | 'dual'; + processMessage(input: AgentLoopInput, opts?: AgentLoopCallOpts): Promise; + addArtifact(artifact: Artifact): void; + resume(clarification: string): void; + cancel(taskId?: string, reason?: AbortReasonValue): void; + getTasks?(): TaskSummary[]; + getTask?(id: string): TaskView | undefined; + loadPersistedTasks?(): Promise; + resumeTask?(taskId: string): Promise<{ taskId: string; sessionId: string }>; + shutdown(): Promise; +} +``` + +`getTasks`、`getTask`、`loadPersistedTasks` 和 `resumeTask` 是可选接口;single loop 不实现任务恢复。 + +### 2.2 双循环状态容器 + +`DualLoopAgent` 当前持有: + +| 字段 | 作用域 | 说明 | +|------|--------|------| +| `tasks: InMemoryTaskStore` | agent 实例 | Task CRUD、active lookup、progress、TTL eviction | +| `artifacts: InMemoryArtifactStore` | agent 实例 | artifact 内存记录和 taskId 分配 | +| `messageQueue: MessageQueue` | agent 实例,按 taskId 索引 | 同 session 后续消息的迭代边界投递 | +| `sessions: SessionStore` | agent 实例,按 sessionId 索引 | 对话消息历史 | +| `taskContexts` | **per task** | `{ abortController, bus }`,取消和事件通道隔离 | +| `worldModel` | **共享** | 单个可变实例;并发 task 会覆盖 | +| `directiveRouter` | **共享** | 每次 inner loop 会重新赋值 | +| `viewStack` | **共享** | 多 task UI 状态可能串扰 | +| `memStore` | agent 实例 | FileMemoryBackend 上的跨任务共享记忆 | + +AbortController/EventBus 已实现 per-task 隔离,但 WorldModel、DirectiveRouter 和 ViewStack 尚未进入 task context。这是当前并发模型最重要的边界。 + +## 3. Task 状态机 + +TypeScript `TaskStatus`: + +```text +pending -> planning -> executing -> completed + | | | \ + | | | -> paused -> executing + | | | + +----------+------------+----> failed + +restart: pending/planning/executing/paused -> interrupted +resume endpoint: interrupted -> executing +``` + +要点: + +- Planning 已经实现,不再是保留状态。新 task 先调用一个最多等待 5 秒的 planning LLM,再进入 executing。 +- `Promise.race` 超时只停止等待,不会 abort 已发出的 planning provider 请求。 +- `cancel` 没有独立 `cancelled` TaskStatus;任务转为 `failed`,error 保存结构化 reason 的字符串表达。 +- `interrupted` 只在加载磁盘元数据时产生,只有 `/resume` 路径可恢复。 +- `paused` 类型和 `resume(clarification)` 方法存在,但运行时没有 clarification gate 自动进入 paused;`resume()` 也只改状态,没有把 clarification 注入 session。 + +Rust TaskStatus 只有 6 个状态,不包含 `interrupted`,因此没有 TypeScript 的重启恢复状态。 + +## 4. 新任务执行流 + +```text +DualLoopAgent.processMessage(input, opts) + | + +-- getActiveForSession(sessionId)? + | | + | +-- yes: enqueue(taskId, content) + | publish task.message.enqueued + | return { taskId, queued: true, iterations: 0 } + | + +-- no: + create per-task AbortController + remember EventBus + set Session.permissionContext = { mode: 'auto' } + create pending Task + assign every unassigned artifact ID + schedule initial meta + user turn disk writes + recall `world-model` memory into new WorldModel + publish task.created + agent.start + fire runInnerLoop(...) + return { taskId, iterations: 0 } + +runInnerLoop + | + +-- pending -> planning + | publish task.planning + | planning provider.chat (5s race) + | parse 0..5 steps + | publish task.planned when non-empty + | + +-- planning -> executing + | persist status metadata + | add first progress checkpoint + | + +-- build handoff/system prompt + +-- build inner ToolRegistry + +-- create per-run inner EventBus + DirectiveRouter + +-- PrismerAgent.processMessage AsyncGenerator + | onIterationStart: + | drain task MessageQueue into Session + | update iterations/lastActivity + | publish task.progress + | + +-- success: + | record completion/facts + | completed + result checkpoint + | persist memory + status + | drain orphan messages + | publish task.completed, agent.end, internal chat.final + | + +-- error/cancel: + failed + persist reason + drain orphan messages + publish error +``` + +## 5. 工具面和权限 + +双循环 inner executor 当前注册: + +- workspace plugin tools(存在时); +- `bash`; +- `memory_store` / `memory_recall`; +- `enter_plan_mode` / `exit_plan_mode`。 + +它没有调用 Node 主入口的完整 `getBuiltinTools()`,因此默认双循环 inner loop 不一定包含 `read_file`、`write_file`、`edit_file`、`list_files`、`grep`、`web_fetch`、`think`;是否存在取决于 plugin。`tests/loop/dual-tool-registry.test.ts` 只验证上述显式集合。 + +双循环 Session 默认是 `auto` 权限模式:标记 `requiresUserInteraction()` 的工具被拒绝,以免后台执行等待无人响应的审批。Plan mode 同样拒绝这类工具。`checkPermissions()` 返回 `ask` 的 UI 闭环仍未实现,当前会先警告再继续走既有审批逻辑。 + +## 6. MessageQueue 和进度 + +### MessageQueue + +- 路由键是 taskId,但入口先用 sessionId 查活跃 task。 +- `getActiveForSession()` 只匹配 `executing` 和 `paused`,不匹配 `pending` / `planning`;同一 session 在首个任务 planning 阶段再次发消息,可能创建第二个 task。 +- 消息在 agent 新 iteration 开始前排空并追加成 user message。 +- task 完成/失败时仍在队列中的消息以 `task.message.orphaned` 发布。 +- 正在进行的 provider/tool 不被 steering message 抢占。 + +### Progress + +`task.progress` 数据: + +```typescript +{ + taskId: string; + iteration: number; + toolsUsed: string[]; + lastActivity: number; +} +``` + +当前 `toolsUsed` 只复用上一份 progress,生产代码没有在工具完成时写回实际名称,因此通常为空。最终工具列表存在 result checkpoint 的 `data.toolsUsed` 和内部完成事件中。 + +## 7. 磁盘持久化与恢复 + +TypeScript 路径: + +```text +{workspaceDir}/.lumin/sessions/{sessionId}/tasks/ + {taskId}.meta.json + {taskId}.jsonl +``` + +Meta 保存 id/session/instruction/status/timestamps/progress/error 等,但不保存 result、checkpoint、plan 或 artifactIds。服务启动前调用 `loadPersistedTasks()`: + +- 非终态 meta 注册为 `interrupted`; +- terminal meta 保留原状态并可查询,但已完成任务的最终 `result` 不会恢复; +- 内存 task store 仍由 1 小时 TTL sweep 清理终态任务。 + +当前 JSONL 并不是完整 transcript。`appendTurn()` 的生产调用只有: + +1. task 创建时的初始 user turn; +2. `persistState()` 写入的 status turn。 + +assistant/tool 消息保存在 Session 内存,但没有追加到 task JSONL。因此 `resumeTask()` 虽然具有 replay 代码,实际可 replay 的通常只有原始指令;恢复语义接近“重新执行原任务”,不是从最后成功工具结果继续。 + +Rust 没有这套磁盘 task 持久化和 resume endpoint。 + +## 8. Artifact 模型 + +Artifact 结构包含 `id`、`url`、`mimeType`、`type`、`addedBy`、`taskId`、`addedAt` 和可选 metadata。 + +已实现: + +- `/v1/artifacts` 创建内存记录; +- 新 task 自动领取所有未分配 artifact ID; +- `GET /v1/tasks/:id` 返回 `artifactIds`。 + +未实现: + +- 将 artifact URL/内容加入 inner-loop prompt; +- 为 inner tools 提供 artifact lookup; +- task 运行中动态注入新 artifact; +- artifact 磁盘持久化。 + +## 9. 取消语义 + +TypeScript per-task 取消链: + +```text +POST /v1/tasks/:id/cancel + -> validate task + AbortReason + -> DualLoopAgent.cancel(taskId, reason) + -> taskContexts[taskId].abortController.abort(createAbortError(reason)) + -> task status failed + -> provider fetch observes signal + -> async tool may observe ToolContext.abortSignal + -> unresolved tool calls receive synthetic [Aborted: ] result + -> queued messages become task.message.orphaned +``` + +限制: + +- 同步 `execFileSync` bash 运行中无法响应 signal。 +- `cancel()` 不带 taskId 且存在多个 context 时会取消全部任务。 +- WebSocket `chat.cancel` 管理的是 WS handler 的 controller;DualLoopAgent 会创建自己的 controller,故双循环应使用 HTTP task cancel。 + +Rust 只有一个 agent 级 `cancelled: Mutex>`,不是 per task,并主要在 iteration 边界检查;并发任务和中途 fetch/tool abort 语义均弱于 TS。 + +## 10. 事件与结果交付 + +| 通道 | single | dual | +|------|--------|------| +| HTTP `/v1/chat` | 返回最终结果 | 返回 taskId / queued | +| HTTP task polling | 空 task list | 返回 task/result/checkpoints | +| 进程内 EventBus | 完整 agent 事件 | 完整 agent + task 事件 | +| WebSocket 网关 | text/tool/directive + 最终结果 | 先返回“task created”的 `chat.final`;`task.*` 和后台最终 `chat.final` **未透传** | +| stdout `--stream` | 输出内部事件 SSE | 取决于调用入口持有的 bus | + +WebSocket 的实际线协议和内部事件清单见 [API.md](../reference/API.md)。 + +## 11. WorldModel 和跨任务知识 + +WorldModel 记录 goal、completedWork、knowledgeBase、workspace/UI 状态与 handoff notes。成功后,结构化 fact 写入 FileMemoryBackend;新 task 以 query `world-model` 召回最多 4,000 字符。 + +当前 `worldModel` 是 `DualLoopAgent` 上的单字段,不是 `Map`。并发 task 会相互覆盖“当前”模型,`persistKnowledgeBase(taskId)` 的 taskId guard 会避免写错 task,但也可能直接跳过被覆盖 task 的知识。C6 能力测试验证三个 task/session ID 独立完成,没有覆盖此共享状态问题。 + +## 12. TypeScript / Rust 对齐 + +| 能力 | TypeScript | Rust | 备注 | +|------|:----------:|:----:|------| +| SingleLoop / DualLoop 抽象 | ✓ | ✓ | 基本结构对齐 | +| Background task | ✓ | ✓ | Rust `tokio::spawn` | +| Planning LLM | ✓ | — | Rust pending 直接 executing | +| Task `interrupted` + disk resume | ✓ | — | TS 也仅保存部分 transcript | +| MessageQueue steering | ✓ | — | TS iteration 边界 | +| Per-task cancel context | ✓ | — | Rust 是 agent 级 cancelled flag | +| Fetch/tool 中途 abort | ✓ | — | 同步 bash 除外 | +| PermissionMode / approval | ✓ | — | Rust 不应执行不可信工具 | +| DirectiveRouter / ViewStack | ✓ | — | TS 当前为共享实例字段 | +| Artifact store/assignment | ✓ | ✓ | 两侧都未完成内容注入 | +| WorldModel | ✓ | ✓ | 两侧 agent 内均为单个共享模型 | +| Memory keyword search | ✓ | ✓ | tag filter 细节不一致 | +| Multimodal provider type | 类型存在,未接入入口 | — | WS/IPC images 当前无效 | +| Telegram / Cloud IM | ✓ | stub | TS only | + +所谓“parity”应理解为核心抽象和部分 wire shape 对齐,不是功能/安全语义完全等价。 + +## 13. Embedded 边界 + +`src/embedded.ts` 不使用 `IAgentLoop` 的 dual 模式,而是构造单循环 `PrismerAgent`: + +- 无 Node 文件系统、workspace plugin、bash 和 FsDirectiveScanner; +- Provider、Tool、MemoryBackend 由宿主注入; +- 自动提供 plan-mode 工具; +- 每次 `processMessage()` 返回 AsyncGenerator 事件和最终 AgentResult; +- `shutdown()` 当前是 no-op。 + +构建命令和体积见 [PROJECT_STATUS.md](../operations/PROJECT_STATUS.md)。 + +## 14. 已验证范围 + +真实 LLM capability suite 定义了 C1、C3、C4、C5、C6、C7;没有 C2 case。它默认跳过,只有设置 `RUN_CAPABILITY_TESTS=1` 才运行。本次审计没有 LLM 凭据,因此没有重新验证 2026-04 的历史 6/6 结果。 + +默认本地套件覆盖 task/store/message queue/disk/cancel/resume 等结构性行为。完整命令、跳过解释和 Rust 结果见 [TEST_COVERAGE.md](../operations/TEST_COVERAGE.md)。历史设计和阶段测量位于 [`archive/plans/`](../archive/plans/)。 + +## 15. 后续优先级 + +1. WebSocket 透传 `task.*` 与后台最终结果,统一 HTTP/WS dual contract。 +2. 把 WorldModel、DirectiveRouter、ViewStack 移入 per-task context。 +3. 持续写入 assistant/tool transcript,并定义真正 checkpoint resume。 +4. 闭环 `ask` 审批、双循环 WS cancel 和 Rust 安全门控。 +5. 将 artifact 内容注入 task,并修复 progress.toolsUsed。 +6. 增加上述缺口的端到端测试,而不仅是 store/schema 测试。 diff --git a/docs/architecture/MEMORY.md b/docs/architecture/MEMORY.md new file mode 100644 index 0000000..0a9d93f --- /dev/null +++ b/docs/architecture/MEMORY.md @@ -0,0 +1,245 @@ +# Lumin Memory System + +> **复核日期**:2026-08-14,提交 `9f1bf56`。 +> **当前默认**:Node/TypeScript 使用 FileMemoryBackend;Embedded 必须注入 MemoryBackend;Rust 有独立文件实现。 +> **重要边界**:TS/Rust 的基础格式和关键词评分接近,但 tag filtering 等细节并非完全对齐。 + +## 1. 结构 + +```text +MemoryStore(跨平台 facade,src/memory.ts) + | + +-- store(content, tags?) + +-- recall(query, maxChars?) -> string + +-- search(query, options?) -> MemorySearchResult[] + +-- loadRecentContext(maxChars?) -> string + +-- capabilities() / close() + | + +-- MemoryBackend interface + | + +-- FileMemoryBackend(Node-only,src/memory-file-backend.ts) + +-- 宿主注入 backend(Embedded) + +-- cloud/vector:只有 config 枚举,没有仓库内实现 +``` + +`MemoryStore` 不再接收路径字符串。Node 代码必须显式传 backend: + +```typescript +import { FileMemoryBackend, MemoryStore } from '@prismer/agent-core'; + +const memory = new MemoryStore(new FileMemoryBackend('./workspace')); +await memory.store('project-language: TypeScript', ['project']); + +const text = await memory.recall('project-language'); +const rows = await memory.search('project-language', { + maxResults: 5, + maxChars: 4000, +}); +``` + +`@prismer/agent-core/memory` 子路径只指向 `dist/memory.js`,不导出 Node-only FileMemoryBackend;当前应从包根导入它。 + +## 2. MemoryBackend contract + +```typescript +interface MemoryBackend { + readonly name: string; + store(content: string, tags?: string[]): Promise; + search(query: string, options?: MemorySearchOptions): Promise; + recent(maxChars?: number): Promise; + capabilities(): MemoryCapabilities; + close(): Promise; +} +``` + +搜索结果: + +```typescript +interface MemorySearchResult { + text: string; + date: string; // YYYY-MM-DD + score: number; // 0..1 + tags?: string[]; + source?: string; +} +``` + +`recall()` 只是 `search()` 的字符串适配器,输出为: + +```text +[2026-08-14] matched memory content +``` + +## 3. TypeScript FileMemoryBackend + +### 存储 + +目录: + +```text +{workspaceDir}/.prismer/memory/YYYY-MM-DD.md +``` + +记录格式: + +```markdown +## 14:30 — [project, decision] +project-language: TypeScript + +--- +``` + +写入使用同步 Node fs 操作,按天 append。File backend 没有连接资源,`close()` 是 no-op。 + +### 搜索 + +1. 读取全部 `.md` 文件,新日期优先。 +2. query 以空白分词,忽略长度小于 3 的词。 +3. 大于 500 字符的 entry 通过 `splitIntoChunks()` 切分;对话采用 3-turn window / 1-turn overlap,普通文本采用 paragraph window。 +4. query 有至少 5 个关键词时,增加 3-keyword overlapping 子查询。 +5. 每个 chunk 取最佳 `hits / keywordCount` 分数。 +6. 按分数降序,再应用 `maxResults` 和 `maxChars`。 + +注意: + +- 匹配是 `lower.includes(keyword)`,不是 token/语义搜索。 +- 去重键只使用 chunk 前 100 个字符。 +- tags 会写入 heading,但 TS search 不解析 tag,也不使用 `options.tags`。 +- `capabilities()` 正确返回 `{ semanticSearch: false, tagFiltering: false }`。 +- `recent()` 只读取今天和昨天,并受字符预算限制。 + +## 4. 配置的实际效果 + +`src/config.ts` 定义: + +| 配置 | 默认值 | 当前消费情况 | +|------|--------|--------------| +| `MEMORY_BACKEND` / `memory.backend` | `file` | Schema 接受 `file/cloud/vector`,但 Node 初始化无条件构造 FileMemoryBackend | +| `memory.recentContextMaxChars` | `3000` | Schema 存在;`src/index.ts` 当前直接使用常量 `3000`,未读取该字段 | + +因此把 `MEMORY_BACKEND=cloud` 或 `vector` 并不会自动加载相应实现。要使用自定义 backend,需要通过 Embedded 的 `memoryBackend` 或在 Node 初始化层新增注入点。 + +## 5. Prompt 和 compaction 集成 + +### 最近记忆注入 + +Node `buildSystemPrompt()` 在共享 memory 可用时调用: + +```text +loadRecentContext(3000) + -> PromptBuilder section: "## Recent Memory" +``` + +这是 today/yesterday 的最近内容,不是针对本轮 query 的相关性检索。Agent 若需要按主题查询,应显式调用 `memory_recall`。 + +### 三层上下文处理 + +每个 agent iteration: + +```text +1. microcompact(messages, 5) + 清除较旧 tool result 内容,保留最近 5 个 + +2. truncateOldestTurns(messages, MAX_CONTEXT_CHARS) + 保留 system + 最近消息;修复 orphan tool results + +3. 对 dropped messages: + memoryFlushBeforeCompaction() -> LLM 提取长期事实 -> memory.store + compactConversation() -> 结构化摘要 -> session.compactionSummary +``` + +主动截断路径用 `!session.compactionSummary` 防止同一 Session 重复生成摘要。若 provider 仍返回 prompt-too-long,reactive path 会对历史前半部分再执行一次压缩,并只重试一次。 + +### 双循环知识连续性 + +双循环完成任务后从最终文本用正则提取 path/measurement facts,写入带 `world-model` tag 的 memory;创建下一任务时以 query `world-model` 召回。 + +当前风险:`DualLoopAgent.worldModel` 是单个共享字段,不是 per-task map;并发任务可能覆盖模型并导致某个 task 的知识持久化被 guard 跳过。参见 [DUAL_LOOP_ARCHITECTURE.md](./DUAL_LOOP_ARCHITECTURE.md)。 + +## 6. Embedded memory + +`createAgentRuntime()` 不包含 Node 文件 backend。宿主传入 `memoryBackend` 时才自动注册 `memory_store` 和 `memory_recall`: + +```typescript +const runtime = createAgentRuntime({ + provider, + systemPrompt: '...', + memoryBackend: nativeBackend, +}); +``` + +这允许 Swift/Kotlin/Electron 宿主提供自己的存储,但 backend 必须完整实现 `recent()`、`capabilities()` 和 `close()`,不只是 store/search。 + +## 7. TypeScript / Rust 差异 + +| 能力 | TypeScript | Rust | 状态 | +|------|:----------:|:----:|------| +| 按天 Markdown 文件 | ✓ | ✓ | 格式基本对齐 | +| `store` / `recall` / structured `search` | ✓ | ✓ | 基础 API 对齐 | +| score 归一化 | ✓ | ✓ | hits/query-keywords | +| 大 entry chunking | ✓ | ✓ | 算法目标相同 | +| 5+ keyword multi-query | ✓ | ✓ | 算法目标相同 | +| 保存 tags | ✓ | ✓ | heading 中保留 | +| `search(..., tags)` 过滤 | **未实现** | **已实现** | 功能不对齐 | +| result 返回解析后的 tags | 未实现 | 已实现 | 功能不对齐 | +| capabilities 声明 tag filtering | false | **false** | Rust 声明与其实现矛盾 | +| facade/backend 分文件 + DI | ✓ | — | Rust 使用自己的 MemoryStore 类型 | + +旧文档中的“TS + Rust 完全对齐”已不再成立。即便相同 query 的常见结果接近,也不能据此推导 options/capability contract 等价。 + +## 8. 测试 + +### 默认本地测试 + +```bash +npx vitest run tests/memory.test.ts tests/compaction.test.ts tests/microcompact.test.ts +cargo test -p lumin-core memory::tests +``` + +当前源码中 `tests/memory.test.ts` 有 32 个 case,Rust memory module 有 49 个 `#[test]`。全仓验证结果和跳过策略见 [TEST_COVERAGE.md](../operations/TEST_COVERAGE.md)。 + +### 真实 LLM / benchmark + +```bash +OPENAI_API_KEY=... \ +OPENAI_API_BASE_URL=... \ +AGENT_DEFAULT_MODEL=... \ +npx vitest run tests/memory-recall-benchmark.test.ts + +npx vitest run tests/locomo-benchmark.test.ts +``` + +这两组测试会在 gateway/API 不可用时跳过,不属于默认 679 passed 的验证范围。 + +## 9. 历史评测结果 + +以下结果来自 2026-03 的历史运行,本次 2026-08 审计没有 LLM 凭据,**未重新运行**。它们适合用于回归比较,不应作为当前版本或全量 LoCoMo 的发布承诺。 + +### 自定义事实召回 + +- 4 轮 compaction,12 条事实。 +- Memory store keyword recall:12/12。 +- 最终 LLM recall:11/12(92%)。 +- 当时模型:`us-kimi-k2.5`;记录耗时约 334 秒。 + +### LoCoMo conv-30 子样本 + +方法:19 个会话、369 轮、约 43K 字符、56 个 QA;这不是 LoCoMo 全部 10 samples。 + +| 模型 | 运行时 | Overall | Non-adversarial | 备注 | +|------|--------|--------:|----------------:|------| +| Claude Opus 4.6 | TS | 86% | 95% | 56 QA 单样本历史结果 | +| Claude Opus 4.6 | Rust | 84% | 100% | 同上 | +| Kimi K2.5 | TS | 63% | 56% | 同上 | +| GLM-4.6 | TS | 约 48% | 约 35% | 50/56 后超时,部分结果 | +| Letta/MemGPT | 外部 baseline | 约 74% | — | 比较口径来自当时记录 | + +历史报告还记录了一次 P0/P1 chunking + multi-query 实验在 Kimi 上从 63% 降至 46%。当前 FileMemoryBackend 仍启用了这两项算法,因此 86% 等旧结果不能自动代表当前 commit;需要在固定 dataset/model/prompt/seed 条件下重新测量。 + +## 10. 当前改进优先级 + +1. 修复 TS tag filtering 或从公共 options 中移除未支持语义,并修正 Rust capability 声明。 +2. 让 `memory.backend` 和 `recentContextMaxChars` 真正进入运行时初始化。 +3. 建立固定版本、全 10 samples、多次采样的 benchmark 记录,分开报告 retrieval recall 和 answer/judge accuracy。 +4. 将双循环 WorldModel 改为 per-task,避免并发知识覆盖。 +5. 对高频文件量、超大 Markdown 和不可信 memory 内容增加容量、解析与 prompt-injection 防护。 diff --git a/docs/architecture/TOOL_COMPLETION_INVARIANTS.md b/docs/architecture/TOOL_COMPLETION_INVARIANTS.md new file mode 100644 index 0000000..d6b246f --- /dev/null +++ b/docs/architecture/TOOL_COMPLETION_INVARIANTS.md @@ -0,0 +1,125 @@ +# Lumin 工具完成与事件顺序不变量 + +> **复核日期**:2026-08-14 +> **适用提交**:`9f1bf56` +> **事实来源**:`src/agent.ts`、`src/tools.ts`、`src/sse.ts` 及本文列出的测试。 + +本文定义当前 agent loop 必须保持的工具执行、会话历史和结束事件约束,也区分“现有保证”和“尚未保证”的行为。它不是对所有事件实时性的承诺。 + +## 1. 当前调度模型 + +LLM 返回多个 tool call 后,`partitionToolCalls()` 按原顺序切成若干批次: + +- 只有工具显式以 `isConcurrencySafe(args) === true` 声明安全时,连续的安全调用才进入同一个并发批次; +- 未实现 `isConcurrencySafe`、返回 `false` 或找不到定义的工具默认串行; +- 各批次始终串行推进,并发批次内部使用 `Promise.all()`; +- 所有批次完成后,loop 才发布/产出收集到的工具事件、扫描 directive 文件、写入 tool result message,并进入下一次 LLM iteration。 + +```text +LLM tool calls + -> partition into ordered safe/unsafe batches + -> await every batch and every tool + -> publish/yield collected tool events + -> scan directive files + -> append one tool message per call + -> next LLM iteration or termination +``` + +这与“所有工具一律并行”不同。写工具和未知工具采取保守串行策略。 + +## 2. 必须保持的不变量 + +### INV-1:已实际启动的工具必须先结束,agent 才能正常结束 + +对走到 `tool.start` 的调用,`executeToolCall()` 会等待 `ToolRegistry.execute()` 返回,再生成 `tool.end`。`ToolRegistry.execute()` 会把工具抛出的异常转换为错误结果,因此正常执行路径不会因为工具 throw 而跳过结束处理。 + +`agent.end` 位于主循环退出和 `runAgentEnd` hook 完成之后。正常完成、最大迭代、重复调用熔断和可处理的工具错误路径,都必须先等待当前工具批次完成。 + +注意边界:权限直接拒绝会产生 `tool.end` 但没有 `tool.start`;hook 拦截、执行前 abort 或审批拒绝不会产生 start/end 对。因此配对约束只适用于“已经发布 `tool.start`”的调用,不能用全部 tool call 数机械比较。 + +### INV-2:不得脱离控制流执行工具 + +工具调用必须被当前 iteration `await`。不能用未被等待的 `.then()`、孤立 Promise 或 fire-and-forget handler 执行实际工具,否则 tool result、directive scan、下一轮 LLM 和最终事件之间失去因果关系。 + +子 agent 委托同样必须等待其 AsyncGenerator 完整返回;若未来迁移到其他进程,需要在协议层重新建立这一等待关系。 + +### INV-3:并发必须显式选择,副作用调用默认串行 + +仅 `isConcurrencySafe(args)` 返回 `true` 的相邻调用可以并发。默认值必须继续是 `false`,不同批次不得交叠。该约束保护写文件、shell 和其他有顺序依赖的副作用操作。 + +### INV-4:每个 LLM tool call 都必须得到历史中的 tool result + +一次 tool batch 返回后,主循环为每个结果向 `state.messages` 和 `Session` 写入带相同 `toolCallId` 的 tool message。若在下一 iteration 边界发现 abort,`synthesizeAbortedToolResults()` 会为最近 assistant message 中尚未满足的 tool calls 补写 `[Aborted: ]`,避免向后续 provider 提交孤立 tool call。 + +当前中途取消还会把由 abort 导致的工具错误改写为结构化 `[Aborted: ...]` 结果。同步 `execFileSync` bash 仍不能在运行中响应 AbortSignal,这是取消时效限制,不应通过提前发送结束事件规避。 + +### INV-5:EventBus 的单次 publish 保持同步调用顺序 + +`EventBus.publish()` 先写有界 buffer,再按订阅集合顺序同步调用 handler。对同一调用栈中依次发布的事件,订阅者会按发布顺序观察;handler 异常被隔离,不会中断其他订阅者。 + +若以后把 EventBus 改成异步批处理或网络 transport,必须增加 sequence/task 标识与显式有序交付,不能继续依赖 JavaScript 同步调用顺序。 + +### INV-6:directive 文件扫描必须发生在完整工具批次之后 + +工具可能在执行期间写 directive 文件。`snapshotDirectiveFiles()` 在执行前取快照,`scanDirectiveFiles()` 在所有分区批次完成并发布工具事件后执行。扫描提前或与工具并行会漏掉较慢调用产生的文件。 + +### INV-7:最终文本交付必须先于成功结束信号 + +非流式响应在无 tool call 时先发布/产出 `text.delta`,然后退出 loop 并发布 `agent.end`。流式 provider 的 callback 会实时发布 `text.delta` 到 EventBus,`chatStream()` 完成后 AsyncGenerator 才集中 yield 已收集的 delta;因此 EventBus 订阅和 generator 消费的实时性不同,但成功路径上两者都应在 `agent.end` 前收到全部已收集文本。 + +若以后增加 delta throttle,结束前必须显式 flush。 + +## 3. 当前非保证与缺口 + +### 工具事件不是实时发布 + +`executeToolCall()` 把 `tool.start`、directive 和 `tool.end` 放进本地数组;`executeToolsPartitioned()` 等全部批次完成后,主循环才逐个 `publish()` 和 `yield`。因此: + +- 因果顺序可保证; +- 客户端不会在工具真正开始时立即看到 `tool.start`; +- 长工具运行期间可能没有进度事件; +- 并发批次最终按输入/Promise 结果顺序发布,不能据此还原真实开始、完成时间。 + +若要提供实时工具进度,应在执行现场发布,同时保持每个 `toolId` 的 start/end 状态机和最终 barrier。 + +### Provider 错误路径没有 `agent.end` + +provider 调用在 reactive compaction 后仍失败时,当前实现发布 `error` 并直接返回 `AgentResult`,不会经过函数末尾的 `agent.end`。所以“每个 `agent.start` 都有 `agent.end`”目前不是全路径保证。消费者必须把 `error` 视为可能的终止信号;更理想的修复是在统一的 finally/finalization 路径结束生命周期。 + +### 审批和权限事件不构成完整 start/end 对 + +审批发生在 `tool.start` 前。拒绝或 30 秒超时只产生 approval response 和合成 tool result;权限拒绝则只产生 `tool.end`。WebSocket 网关虽能转发 approval 事件,但没有把客户端 `tool.approve` 接到本次 agent 的 `resolveApproval()`,详见 [API.md](../reference/API.md)。 + +### EventBus buffer 会丢弃旧事件 + +超过默认 1,000 条时,EventBus 会移除最旧事件并增加 `droppedCount`。在线订阅者仍同步收到新事件,但晚加入或依赖 replay 的消费者不能假设 buffer 保存完整历史。 + +### 双循环的内部完成事件不等于网关交付 + +inner agent 完成后,`DualLoopAgent` 在进程内依次发布 `task.completed`、`agent.end` 和后台 `chat.final`。当前 WebSocket switch 不转发 `task.*` 或该后台最终结果,因此内部顺序正确并不代表 WS 客户端已得到最终答案;双循环客户端需要轮询任务接口。 + +## 4. 变更检查清单 + +修改 `src/agent.ts`、工具并发策略、EventBus、取消或 WebSocket 转发时,至少验证: + +- [ ] 一个响应含两个显式并发安全工具:两者实际并发,且每个已启动调用在 `agent.end` 前有 `tool.end`。 +- [ ] 两个写工具以及未声明并发能力的工具:严格按原调用顺序串行。 +- [ ] 一个慢工具和一个抛错工具:当前 iteration 等待两者归一化完成,不出现提前结束。 +- [ ] 工具执行中 abort:Session 中每个未满足的 `toolCallId` 都有结构化 abort result。 +- [ ] directive 文件由最慢工具写入:批次结束后的扫描可以发现。 +- [ ] 流式与非流式文本:最后一个 delta 在成功 `agent.end` 前可见。 +- [ ] provider 报错、审批拒绝和权限拒绝:消费者不会永久等待不存在的事件。 +- [ ] 双循环:分别验证内部 EventBus 顺序和 HTTP/WS 对外交付,不以其中一层替代另一层。 + +## 5. 现有测试证据 + +| 行为 | 代表性测试 | 当前覆盖边界 | +|------|------------|--------------| +| start/tool/text/end 基本顺序 | `tests/generator.test.ts` | 单工具、成功路径 | +| 安全工具并发、写工具串行、未知默认串行 | `tests/agent-v2.test.ts` | 验证实际执行顺序 | +| ToolContext 携带 AbortSignal | `tests/tools-abort.test.ts` | 只验证 signal 传递 | +| 中途 abort 的合成 tool result | `tests/agent-abort.test.ts` | 真实 LLM suite;无凭据时跳过 | +| 审批 resolver、超时和事件 schema | `tests/approval.test.ts` | 未覆盖 WS 到 resolver 的端到端连接 | +| AgentEvent schema | `tests/sse.test.ts` | 结构校验,不证明实时性 | + +当前没有自动化测试证明“多工具长耗时下实时收到 `tool.start`”,因为实现本身会缓冲这些事件;也没有覆盖 provider 错误后统一 `agent.end`。这两项应在改变完成语义前先补测试。 diff --git a/docs/archive/README.md b/docs/archive/README.md new file mode 100644 index 0000000..d472577 --- /dev/null +++ b/docs/archive/README.md @@ -0,0 +1,32 @@ +# 文档归档 + +> **归档日期**:2026-08-14 +> **当前文档入口**:[`../README.md`](../README.md) + +本目录保存已被实现替代、结论过时或仅用于历史追踪的材料。归档文件不会随当前源码持续更新;其中的文件路径、测试数量、阶段名称和能力判断只代表原记录时点。 + +## 归档目录 + +| 目录 | 数量 | 内容 | 归档原因 | +|------|-----:|------|----------| +| [`analysis/`](./analysis/) | 7 | 2026-03 的 Claude Code 对比分析、改进路线和 Dual Loop V2 草案 | 多数建议已实现或被后续双循环架构替代;继续留在顶层会与当前状态混淆 | +| [`plans/`](./plans/) | 20 | 2026-03 至 2026-08 的设计、实施计划、基线和阶段测量 | 对应工作已经落地;保留作决策和验证历史,不再作为待执行计划 | +| [`assets/`](./assets/) | 2 | 未被当前文档引用的旧 logo 素材 | 与当前参考内容无直接关系,保留但退出顶层导航 | + +## 历史分析清单 + +- `01-agent-loop-analysis.md`:早期 agent loop 差距分析。 +- `02-tool-orchestration-analysis.md`:早期工具并发与 StreamingToolExecutor 方案。 +- `03-context-management-analysis.md`:早期上下文压缩对比。 +- `04-multi-agent-analysis.md`:早期多 agent 隔离与协调分析。 +- `05-hooks-permissions-analysis.md`:早期 hook/权限差距分析。 +- `06-improvement-roadmap.md`:截至 2026-03-31 的阶段路线图。 +- `2026-03-26-dual-loop-v2-design.md`:被后续实现和当前架构文档取代的设计草案。 + +## 使用规则 + +- 查当前接口和能力,请返回 [`../README.md`](../README.md) 选择顶层参考文档。 +- 追溯某项设计来源时,可引用归档文件并同时注明日期和提交。 +- 归档材料中的旧路径是历史记录的一部分,不批量改写;只有影响当前导航的链接需要修复。 +- 若归档中的设想重新进入开发,应新建当前设计/计划,不直接把旧文件移回顶层。 +- 当前设计或实施计划完成后移入 `plans/`,避免与持续维护的参考文档混放。 diff --git a/docs/01-agent-loop-analysis.md b/docs/archive/analysis/01-agent-loop-analysis.md similarity index 100% rename from docs/01-agent-loop-analysis.md rename to docs/archive/analysis/01-agent-loop-analysis.md diff --git a/docs/02-tool-orchestration-analysis.md b/docs/archive/analysis/02-tool-orchestration-analysis.md similarity index 100% rename from docs/02-tool-orchestration-analysis.md rename to docs/archive/analysis/02-tool-orchestration-analysis.md diff --git a/docs/03-context-management-analysis.md b/docs/archive/analysis/03-context-management-analysis.md similarity index 100% rename from docs/03-context-management-analysis.md rename to docs/archive/analysis/03-context-management-analysis.md diff --git a/docs/04-multi-agent-analysis.md b/docs/archive/analysis/04-multi-agent-analysis.md similarity index 100% rename from docs/04-multi-agent-analysis.md rename to docs/archive/analysis/04-multi-agent-analysis.md diff --git a/docs/05-hooks-permissions-analysis.md b/docs/archive/analysis/05-hooks-permissions-analysis.md similarity index 100% rename from docs/05-hooks-permissions-analysis.md rename to docs/archive/analysis/05-hooks-permissions-analysis.md diff --git a/docs/06-improvement-roadmap.md b/docs/archive/analysis/06-improvement-roadmap.md similarity index 100% rename from docs/06-improvement-roadmap.md rename to docs/archive/analysis/06-improvement-roadmap.md diff --git a/docs/2026-03-26-dual-loop-v2-design.md b/docs/archive/analysis/2026-03-26-dual-loop-v2-design.md similarity index 100% rename from docs/2026-03-26-dual-loop-v2-design.md rename to docs/archive/analysis/2026-03-26-dual-loop-v2-design.md diff --git a/docs/luminlogo.png b/docs/archive/assets/luminlogo.png similarity index 100% rename from docs/luminlogo.png rename to docs/archive/assets/luminlogo.png diff --git a/docs/luminpulse logo.svg b/docs/archive/assets/luminpulse logo.svg similarity index 100% rename from docs/luminpulse logo.svg rename to docs/archive/assets/luminpulse logo.svg diff --git a/docs/superpowers/plans/2026-03-31-rust-ts-parity.md b/docs/archive/plans/2026-03-31-rust-ts-parity.md similarity index 100% rename from docs/superpowers/plans/2026-03-31-rust-ts-parity.md rename to docs/archive/plans/2026-03-31-rust-ts-parity.md diff --git a/docs/superpowers/plans/2026-04-12-builtin-tools-test-plan.md b/docs/archive/plans/2026-04-12-builtin-tools-test-plan.md similarity index 100% rename from docs/superpowers/plans/2026-04-12-builtin-tools-test-plan.md rename to docs/archive/plans/2026-04-12-builtin-tools-test-plan.md diff --git a/docs/superpowers/plans/2026-04-12-rust-builtin-tools-mvp.md b/docs/archive/plans/2026-04-12-rust-builtin-tools-mvp.md similarity index 100% rename from docs/superpowers/plans/2026-04-12-rust-builtin-tools-mvp.md rename to docs/archive/plans/2026-04-12-rust-builtin-tools-mvp.md diff --git a/docs/superpowers/plans/2026-04-13-c1-after-phase-a.md b/docs/archive/plans/2026-04-13-c1-after-phase-a.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-c1-after-phase-a.md rename to docs/archive/plans/2026-04-13-c1-after-phase-a.md diff --git a/docs/superpowers/plans/2026-04-13-c1-baseline.md b/docs/archive/plans/2026-04-13-c1-baseline.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-c1-baseline.md rename to docs/archive/plans/2026-04-13-c1-baseline.md diff --git a/docs/superpowers/plans/2026-04-13-c3-after-phase-b.md b/docs/archive/plans/2026-04-13-c3-after-phase-b.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-c3-after-phase-b.md rename to docs/archive/plans/2026-04-13-c3-after-phase-b.md diff --git a/docs/superpowers/plans/2026-04-13-c4-after-phase-c.md b/docs/archive/plans/2026-04-13-c4-after-phase-c.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-c4-after-phase-c.md rename to docs/archive/plans/2026-04-13-c4-after-phase-c.md diff --git a/docs/superpowers/plans/2026-04-13-c7-after-phase-e.md b/docs/archive/plans/2026-04-13-c7-after-phase-e.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-c7-after-phase-e.md rename to docs/archive/plans/2026-04-13-c7-after-phase-e.md diff --git a/docs/superpowers/plans/2026-04-13-dual-loop-architecture-design.md b/docs/archive/plans/2026-04-13-dual-loop-architecture-design.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-dual-loop-architecture-design.md rename to docs/archive/plans/2026-04-13-dual-loop-architecture-design.md diff --git a/docs/superpowers/plans/2026-04-13-dual-loop-audit-and-roadmap.md b/docs/archive/plans/2026-04-13-dual-loop-audit-and-roadmap.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-dual-loop-audit-and-roadmap.md rename to docs/archive/plans/2026-04-13-dual-loop-audit-and-roadmap.md diff --git a/docs/superpowers/plans/2026-04-13-phase-a-message-queue-impl.md b/docs/archive/plans/2026-04-13-phase-a-message-queue-impl.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-phase-a-message-queue-impl.md rename to docs/archive/plans/2026-04-13-phase-a-message-queue-impl.md diff --git a/docs/superpowers/plans/2026-04-13-phase-b-disk-persistence-impl.md b/docs/archive/plans/2026-04-13-phase-b-disk-persistence-impl.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-phase-b-disk-persistence-impl.md rename to docs/archive/plans/2026-04-13-phase-b-disk-persistence-impl.md diff --git a/docs/superpowers/plans/2026-04-13-phase-c-structured-abort-impl.md b/docs/archive/plans/2026-04-13-phase-c-structured-abort-impl.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-phase-c-structured-abort-impl.md rename to docs/archive/plans/2026-04-13-phase-c-structured-abort-impl.md diff --git a/docs/superpowers/plans/2026-04-13-phase-d-permissions-plan-mode-impl.md b/docs/archive/plans/2026-04-13-phase-d-permissions-plan-mode-impl.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-phase-d-permissions-plan-mode-impl.md rename to docs/archive/plans/2026-04-13-phase-d-permissions-plan-mode-impl.md diff --git a/docs/superpowers/plans/2026-04-13-phase-e-knowledge-eviction-impl.md b/docs/archive/plans/2026-04-13-phase-e-knowledge-eviction-impl.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-phase-e-knowledge-eviction-impl.md rename to docs/archive/plans/2026-04-13-phase-e-knowledge-eviction-impl.md diff --git a/docs/superpowers/plans/2026-04-13-plan-mode-after-phase-d.md b/docs/archive/plans/2026-04-13-plan-mode-after-phase-d.md similarity index 100% rename from docs/superpowers/plans/2026-04-13-plan-mode-after-phase-d.md rename to docs/archive/plans/2026-04-13-plan-mode-after-phase-d.md diff --git a/docs/superpowers/plans/2026-04-15-phase-h-embedded-bundle-impl.md b/docs/archive/plans/2026-04-15-phase-h-embedded-bundle-impl.md similarity index 100% rename from docs/superpowers/plans/2026-04-15-phase-h-embedded-bundle-impl.md rename to docs/archive/plans/2026-04-15-phase-h-embedded-bundle-impl.md diff --git a/docs/archive/plans/2026-08-14-docs-information-architecture-design.md b/docs/archive/plans/2026-08-14-docs-information-architecture-design.md index a7de976..d6cfbcd 100644 --- a/docs/archive/plans/2026-08-14-docs-information-architecture-design.md +++ b/docs/archive/plans/2026-08-14-docs-information-architecture-design.md @@ -2,7 +2,7 @@ > **设计日期**:2026-08-14 > **适用基线**:`9f1bf56` -> **状态**:已确认,待实施 +> **状态**:已实施 ## 1. 目标 @@ -123,7 +123,7 @@ docs/ - 不重命名现有参考文档文件名,只改变目录; - 不删除历史文件或素材; - 不引入文档站点生成器、侧边栏框架或额外依赖。 -- 除 `CLAUDE.md`、根 `AGENTS.md` 和 `templates/base/AGENTS.md` 外,不扩展根目录文档改写范围。 +- 除 `CLAUDE.md`、根 `AGENTS.md`、`templates/base/AGENTS.md` 以及根 README 中已确认失效的 API/Memory 示例外,不扩展根目录文档改写范围。 ## 8. 验证标准 diff --git a/docs/archive/plans/2026-08-14-docs-refresh-and-archive.md b/docs/archive/plans/2026-08-14-docs-refresh-and-archive.md new file mode 100644 index 0000000..138c533 --- /dev/null +++ b/docs/archive/plans/2026-08-14-docs-refresh-and-archive.md @@ -0,0 +1,108 @@ +# Documentation Refresh and Archive Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Make `docs/` accurately describe the repository at commit `9f1bf56`, expose verified limitations, and separate current reference material from historical design and implementation records. + +**Architecture:** `docs/README.md` becomes the navigation and lifecycle-policy entry point, while `docs/PROJECT_STATUS.md` records the dated audit evidence and prioritized gaps. Current API, architecture, memory, testing, and invariant references stay at the top level; superseded analyses, completed plans, measurements, and unused brand assets move under `docs/archive/` without deleting history. + +**Tech Stack:** Markdown, TypeScript 5, Vitest 3, Rust/Cargo, shell-based link and build verification. + +--- + +### Task 1: Establish the documentation information architecture + +**Files:** +- Create: `docs/README.md` +- Create: `docs/PROJECT_STATUS.md` +- Create: `docs/archive/README.md` + +- [x] **Step 1: Create the top-level documentation index** + +List each current reference document, its authority and update trigger. State that archived files are historical evidence rather than current contracts. + +- [x] **Step 2: Write the dated project-status audit** + +Record commit/version/runtime scope, source layout, validation results from 2026-08-14, capability boundaries, risk-ranked findings, documentation drift, and recommended next actions. Every material claim must name the source file or command that supports it. + +- [x] **Step 3: Define the archive policy and inventory** + +Document the archive categories, original paths, archive reason, and rule that archived files are not maintained for current correctness. + +### Task 2: Archive superseded analysis and completed execution records + +**Files:** +- Move: `docs/01-agent-loop-analysis.md` through `docs/06-improvement-roadmap.md` → `docs/archive/analysis/` +- Move: `docs/2026-03-26-dual-loop-v2-design.md` → `docs/archive/analysis/` +- Move: the 17 pre-2026-08-14 files in `docs/superpowers/plans/` → `docs/archive/superpowers/plans/` +- Move: `docs/luminlogo.png`, `docs/luminpulse logo.svg` → `docs/archive/assets/` + +- [x] **Step 1: Move historical files with `git mv`** + +Preserve filenames and Git history. Leave the current refresh plan in `docs/superpowers/plans/`. + +- [x] **Step 2: Repair live references to moved files** + +Update current-document links to their new archive locations. Leave literal paths inside archived measurement records unchanged when they describe the historical execution environment. + +### Task 3: Correct the external and programmatic API reference + +**Files:** +- Modify: `docs/API.md` + +- [x] **Step 1: Reconcile HTTP routes with `src/server.ts`** + +Document `/`, `/health`, `/v1/tools`, `/v1/chat`, `/v1/artifacts`, `/v1/tasks`, task lookup, cancel and resume, including current status/error behavior and single-versus-dual semantics. + +- [x] **Step 2: Split WebSocket wire events from internal EventBus events** + +Document only the event types actually forwarded by `handleWebSocket`. Explicitly note that `task.*`, `thinking.delta`, `iteration.start`, `tool.progress`, `memory.accessed`, and the background dual-loop `chat.final` are currently internal events and are not forwarded by the gateway switch. + +- [x] **Step 3: Correct IPC and library examples** + +Use the actual `EventBus.subscribe(handler)` signature, the actual `runAgent(): Promise` behavior, current root exports, and the `@prismer/agent-core/embedded` factory contract. + +### Task 4: Refresh architecture, memory, tests, and invariants + +**Files:** +- Modify: `docs/DUAL_LOOP_ARCHITECTURE.md` +- Modify: `docs/MEMORY.md` +- Modify: `docs/TEST_COVERAGE.md` +- Modify: `docs/TOOL_COMPLETION_INVARIANTS.md` + +- [x] **Step 1: Remove stale dual-loop contradictions** + +Update revision/test counts and implemented planning behavior; remove already-implemented task polling and response fields from future work; add the embedded runtime boundary and the concurrent shared-state limitation. + +- [x] **Step 2: Correct memory ownership and parity claims** + +Point the Node file backend to `src/memory-file-backend.ts`, show the required `new MemoryStore(new FileMemoryBackend(workspaceDir))` construction, and state that TypeScript currently ignores tag filters while Rust implements them despite declaring `tag_filtering: false`. + +- [x] **Step 3: Replace static test totals with a reproducible snapshot** + +Record `679 passed / 58 skipped` across 65 Vitest files, the Rust `618 passed / 1 ignored / 1 environment-sensitive failure` result, typecheck/build results, embedded bundle size, credentials/flags required for skipped suites, and the absence of a measured coverage percentage. + +- [x] **Step 4: Align ordering invariants with current partitioned execution** + +Describe concurrent batches only for tools that opt into `isConcurrencySafe`, serial execution otherwise, event collection/yield timing, abort result synthesis, and the tests that cover these behaviors. + +### Task 5: Verify the reorganized documentation and repository state + +**Files:** +- Modify: `docs/superpowers/plans/2026-08-14-docs-refresh-and-archive.md` + +- [x] **Step 1: Check links and stale paths** + +Run a local Markdown-link checker script over `README.md`, `ROADMAP.md`, and `docs/**/*.md`; expect zero missing relative targets. Search live docs for old `docs/superpowers/plans/2026-04-*` paths and stale test totals. + +- [x] **Step 2: Re-run proportional build verification** + +Run `npm run typecheck`, `npm test -- --reporter=dot`, and `npm run build:embedded`; expect the same successful results captured in `PROJECT_STATUS.md`. Run the Rust suite with only the known environment-sensitive assertion filtered and expect all remaining tests to pass. + +- [x] **Step 3: Review the final diff** + +Confirm only documentation paths changed, archived files are preserved as renames, generated `dist/`, `node_modules/`, `rust/target/`, and `workspace/` remain ignored, and no user changes were overwritten. + +- [x] **Step 4: Mark this plan complete** + +Check every completed step only after its command or content review succeeds. diff --git a/docs/archive/plans/2026-08-14-docs-structure-reorganization-implementation-plan.md b/docs/archive/plans/2026-08-14-docs-structure-reorganization-implementation-plan.md new file mode 100644 index 0000000..615e8da --- /dev/null +++ b/docs/archive/plans/2026-08-14-docs-structure-reorganization-implementation-plan.md @@ -0,0 +1,157 @@ +# Documentation Structure Reorganization Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Reorganize current documentation by responsibility, consolidate every historical document under `docs/archive/`, and align developer and agent guidance with the new structure. + +**Architecture:** Keep `docs/README.md` as the only top-level document. Place current contracts under `reference/`, implementation explanations under `architecture/`, dated engineering evidence under `operations/`, and all completed analyses/plans/assets under `archive/`. Update only navigation and guide content; runtime behavior remains unchanged. + +**Tech Stack:** Markdown, Git renames, TypeScript type checking, shell-based link validation. + +--- + +### Task 1: Move current references into responsibility directories + +**Files:** +- Move: `docs/API.md` → `docs/reference/API.md` +- Move: `docs/DUAL_LOOP_ARCHITECTURE.md` → `docs/architecture/DUAL_LOOP_ARCHITECTURE.md` +- Move: `docs/MEMORY.md` → `docs/architecture/MEMORY.md` +- Move: `docs/TOOL_COMPLETION_INVARIANTS.md` → `docs/architecture/TOOL_COMPLETION_INVARIANTS.md` +- Move: `docs/PROJECT_STATUS.md` → `docs/operations/PROJECT_STATUS.md` +- Move: `docs/TEST_COVERAGE.md` → `docs/operations/TEST_COVERAGE.md` + +- [x] **Step 1: Create target directories and move files with Git** + +Run: + +```bash +mkdir -p docs/reference docs/architecture docs/operations +git mv docs/API.md docs/reference/API.md +git mv docs/DUAL_LOOP_ARCHITECTURE.md docs/architecture/DUAL_LOOP_ARCHITECTURE.md +git mv docs/MEMORY.md docs/architecture/MEMORY.md +git mv docs/TOOL_COMPLETION_INVARIANTS.md docs/architecture/TOOL_COMPLETION_INVARIANTS.md +git mv docs/PROJECT_STATUS.md docs/operations/PROJECT_STATUS.md +git mv docs/TEST_COVERAGE.md docs/operations/TEST_COVERAGE.md +``` + +Expected: `docs/` top-level Markdown files contain only `README.md`. + +- [x] **Step 2: Confirm content preservation** + +Run `git status --short` and confirm all six files are reported as moves or as delete/add pairs whose contents remain intact. + +### Task 2: Consolidate historical plans under the archive + +**Files:** +- Move: `docs/archive/superpowers/plans/*` → `docs/archive/plans/` +- Move: `docs/superpowers/plans/2026-08-14-docs-refresh-and-archive.md` → `docs/archive/plans/` +- Modify: `docs/archive/README.md` + +- [x] **Step 1: Flatten archived plans** + +Move the 17 historical plan/measurement files from `docs/archive/superpowers/plans/` into `docs/archive/plans/`. Stop if any basename already exists; never overwrite an archive record. + +- [x] **Step 2: Archive the completed refresh plan** + +Move `2026-08-14-docs-refresh-and-archive.md` into `docs/archive/plans/`. The design and this implementation plan already live in the final archive location. + +- [x] **Step 3: Update the archive inventory** + +Change the archive table to use `plans/`, record the final plan count, and state that design/implementation records are archived when their work is complete. + +### Task 3: Repair live documentation navigation + +**Files:** +- Modify: `docs/README.md` +- Modify: `docs/reference/API.md` +- Modify: `docs/architecture/DUAL_LOOP_ARCHITECTURE.md` +- Modify: `docs/architecture/MEMORY.md` +- Modify: `docs/architecture/TOOL_COMPLETION_INVARIANTS.md` +- Modify: `docs/operations/PROJECT_STATUS.md` +- Modify: `docs/operations/TEST_COVERAGE.md` + +- [x] **Step 1: Rewrite the documentation index** + +Point the current-document table at `reference/`, `architecture/`, and `operations/`. Remove the live `superpowers/plans` reference and link maintenance history to `archive/plans/`. + +- [x] **Step 2: Recalculate cross-document relative links** + +Use these mappings: + +```text +architecture -> reference: ../reference/ +architecture -> operations: ../operations/ +architecture -> archive: ../archive/ +operations -> reference: ../reference/ +operations -> architecture: ../architecture/ +``` + +Keep same-directory links as `./FILE.md`. + +- [x] **Step 3: Search for stale live paths** + +Run: + +```bash +rg -n 'docs/superpowers|archive/superpowers|\]\(\./(API|PROJECT_STATUS|DUAL_LOOP_ARCHITECTURE|MEMORY|TEST_COVERAGE|TOOL_COMPLETION_INVARIANTS)\.md' docs --glob '!archive/**' +``` + +Expected: no stale matches. + +### Task 4: Align developer and agent guidance + +**Files:** +- Modify: `CLAUDE.md` +- Create: `AGENTS.md` +- Modify: `templates/base/AGENTS.md` +- Modify: `README.md` +- Modify: `src/task/message-queue.ts` +- Modify: `tests/capability/dual-loop-capabilities.test.ts` + +- [x] **Step 1: Rewrite `CLAUDE.md` as a stable developer entry** + +Document the current TypeScript/Rust/Embedded surfaces, source directories, loop modes, configuration categories, and verified commands. Link detailed facts to `docs/README.md`; do not copy volatile test totals or wire schemas. + +- [x] **Step 2: Add root `AGENTS.md`** + +Define repository navigation, authoritative docs, safe editing rules, archive lifecycle, and proportional verification. Explicitly distinguish this file from `templates/base/AGENTS.md`. + +- [x] **Step 3: Refresh the runtime template** + +Keep `templates/base/AGENTS.md` focused on end-user agent behavior: cite sources, use tools, verify results before claiming completion, ask before destructive actions, disclose unavailable capabilities, and clarify ambiguous requests. + +- [x] **Step 4: Repair external path references and stale examples** + +Update root `README.md` to `docs/reference/API.md`, import FileMemoryBackend from the package root, and construct `MemoryStore` with a backend instance. Update the two source/test comments to `docs/archive/plans/.md`. Do not change executable code. + +### Task 5: Verify and record completion + +**Files:** +- Modify: `docs/archive/plans/2026-08-14-docs-structure-reorganization-implementation-plan.md` + +- [x] **Step 1: Verify tree and archive counts** + +Confirm the only top-level Markdown file in `docs/` is `README.md`, no `docs/superpowers` or `docs/archive/superpowers` files remain, and archive counts match the inventory. + +- [x] **Step 2: Verify all Markdown relative links** + +Run a local link checker over root Markdown files, `templates/**/*.md`, and `docs/**/*.md`. Expected: zero missing relative targets. + +- [x] **Step 3: Run repository checks** + +Run: + +```bash +git diff --check +npm run typecheck +``` + +Expected: both exit 0. + +- [x] **Step 4: Review scope and rename preservation** + +Confirm runtime source changes are limited to comment paths, template changes contain no repository-maintenance instructions, generated directories remain ignored, and historical files were not deleted or overwritten. + +- [x] **Step 5: Mark every completed checkbox** + +Update this plan only after its corresponding command or content review succeeds. diff --git a/docs/operations/PROJECT_STATUS.md b/docs/operations/PROJECT_STATUS.md new file mode 100644 index 0000000..736e12e --- /dev/null +++ b/docs/operations/PROJECT_STATUS.md @@ -0,0 +1,127 @@ +# Lumin 工程状态审计 + +> **审计日期**:2026-08-14 +> **审计提交**:`9f1bf56be0d29322191867a2b8f7bdc6c361e166` +> **工作分支**:`OOXXXXOO/lumin`(独立 worktree) +> **包标识**:`@prismer/agent-core@0.3.1` + +## 结论摘要 + +工程已经形成三个可运行交付面:Node/TypeScript 主运行时、部分能力对齐的 Rust 运行时,以及面向 JavaScriptCore/Hermes/Electron 的单文件嵌入式运行时。TypeScript 类型检查、默认 Vitest 套件和嵌入式构建均可通过;Rust 核心测试稳定,但默认全量命令包含一个环境敏感的网络错误断言。 + +当前最主要的问题不是“功能完全缺失”,而是若干对外声明超过了实现边界:双循环的后台任务事件没有从 WebSocket 网关透传;磁盘“转录”没有持久化中间 assistant/tool 轮次;并发任务仍共享 WorldModel 和指令路由状态;记忆模块的 TS/Rust tag filtering 不一致。旧文档此前把这些能力描述成完整实现,本次已按源码收紧表述。 + +## 工程基线 + +| 项目 | 当前状态 | 依据 | +|------|----------|------| +| 版本 | `0.3.1` | `package.json`、`src/version.ts` | +| Node 要求 | `>=20` | `package.json#engines` | +| 生产依赖 | 仅直接依赖 `zod`;锁定版本为 `3.25.76` | `package.json`、`package-lock.json` | +| 默认执行模式 | `single`;`LUMIN_LOOP_MODE=dual` 显式启用双循环 | `src/config.ts`、`src/loop/factory.ts` | +| TypeScript 源码 | 54 个受版本控制文件,约 9,840 行 | `git ls-files src`、`wc -l` | +| Rust 源码 | `lumin-core` + `lumin-server`,约 14,213 行 Rust | `rust/Cargo.toml`、`wc -l` | +| 测试代码 | 78 个受版本控制文件,约 14,120 行 TS/MJS | `git ls-files tests`、`wc -l` | +| CI / 变更日志 | 仓库内未发现 `.github/workflows/` 或 `CHANGELOG.md` | 文件盘点 | + +根目录 `ROADMAP.md` 已将一组能力标为 `v0.4.0 DONE`,但包版本和运行时版本仍为 `0.3.1`。因此当前应把“源码能力”和“已发布版本”视为两个概念;发布前需要统一版本、变更日志和打包说明。 + +## 当前架构 + +### TypeScript 主运行时 + +- `src/agent.ts`:AsyncGenerator 形式的核心 agent loop、工具分区执行、hook、权限、审批、压缩、取消和子 agent 委托。 +- `src/loop/single.ts`:同步请求/响应适配器,默认模式。 +- `src/loop/dual.ts`:任务化后台执行、消息队列、状态机、轮询、取消、重启后注册和恢复。 +- `src/server.ts`:无第三方服务器框架的 HTTP/WebSocket 网关。 +- `src/index.ts`:Node 运行时初始化、内置/插件工具、动态 prompt、`runAgent()` 和公共导出。 + +### Rust 运行时 + +`rust/crates/lumin-core` 覆盖 provider、agent、工具、session、memory、task 和 loop 基础抽象,`lumin-server` 提供服务端。Rust 并不是 TypeScript 的完整镜像:磁盘任务恢复、消息队列、TS 级中途取消、PermissionMode、指令路由和平台 channel 均未完整对齐。详见 [DUAL_LOOP_ARCHITECTURE.md](../architecture/DUAL_LOOP_ARCHITECTURE.md)。 + +### 嵌入式运行时 + +`src/embedded.ts` 暴露 `createAgentRuntime()`,仅支持单循环,通过依赖注入接收 Provider、Tool 和可选 MemoryBackend,不加载文件系统模板、插件或 bash。`npm run build:embedded` 生成全局名为 `LuminClaw` 的 IIFE bundle。 + +本次实测产物: + +| 指标 | 结果 | +|------|------| +| Bundle | `dist/luminclaw-core.js` | +| 原始大小 | 105,955 bytes | +| gzip 大小 | 29,241 bytes | +| SHA-256 | `bb6721c8c9528c2f5d1b9fef7f8ce75b9583e5ec5cc57ba8e48e19ff89dd541a` | + +## 验证快照 + +以下结果均在本审计提交上通过 `npm ci` 后获得: + +| 命令 | 结果 | +|------|------| +| `npm run typecheck` | 通过,0 个 TypeScript 错误 | +| `npm test -- --reporter=dot` | 65 个文件:51 通过、14 跳过;737 个测试:679 通过、58 跳过 | +| `npm run build:embedded` | 通过;产物大小见上表 | +| `cargo test --workspace` | 582 个 core 单测先通过;随后 `web_fetch_returns_error_for_unreachable_url` 失败,原因是当前网络栈对 `127.0.0.1:1` 的返回未以测试预期的 `Error:` 开头 | +| `cargo test --workspace -- --skip web_fetch_returns_error_for_unreachable_url` | 618 通过、1 ignored、1 filtered,0 失败 | +| `npm audit --omit=dev` | 0 个生产依赖漏洞 | +| `npm audit` | 7 个开发依赖告警:1 low、5 high、1 critical;涉及 esbuild、nanoid、picomatch、postcss、vite、vitest、ws | + +58 个 Vitest 跳过项主要来自真实 LLM、能力测试开关和平台条件。没有设置 LLM 凭据时,`describeReal` 自动跳过;能力套件还需要 `RUN_CAPABILITY_TESTS=1`。因此“679 通过”是默认离线/本地基线,不等价于重新验证历史的真实 LLM 能力结论。详见 [TEST_COVERAGE.md](./TEST_COVERAGE.md)。 + +## 已确认能力 + +- 单循环:OpenAI-compatible provider、流式 token、工具调用、工具安全分区、session、hook、memory、compaction、技能加载和子 agent 委托。 +- 双循环:同 session 活跃任务路由、迭代边界消息注入、任务查询、结构化取消、结果轮询、终态淘汰和基础磁盘元数据。 +- 网关:HTTP chat/tools/artifacts/tasks,WebSocket chat/approval/cancel,stdin/stdout IPC。 +- 可嵌入:无 `node:*` 依赖的 IIFE,注入 Provider/Tool/MemoryBackend,自动注册 plan-mode 工具。 +- Rust:核心 loop/provider/tools/memory/task 结构和部分 HTTP/WS schema 对齐。 + +## 风险与缺口 + +### 高优先级 + +1. **双循环 WebSocket 结果链不完整。** `DualLoopAgent` 会在 EventBus 发布 `task.created`、`task.progress`、`task.completed` 和后台 `chat.final`,但 `src/server.ts` 的 WebSocket switch 没有转发这些事件。WebSocket 调用在任务创建后先收到一条“任务已创建”的 `chat.final`,后台最终结果需要改用 `GET /v1/tasks/:id` 轮询。现行线协议以 [API.md](../reference/API.md) 为准。 + +2. **磁盘恢复不是完整 checkpoint/resume。** `src/loop/dual.ts` 只把初始 user turn 和 status turn 传给 `appendTurn()`;assistant/tool 中间轮次只进入内存 Session,没有写入 JSONL。`resumeTask()` 因而不能从“最后持久化轮次”续跑,而是以可用的初始内容重新执行原任务。Meta 也不保存 result/checkpoint/plan/artifactIds,重启后恢复的 completed task 没有原最终文本。 + +3. **并发任务共享部分可变状态。** AbortController 和 EventBus 已按 `taskId` 存在 `taskContexts` 中,但 `worldModel`、`directiveRouter` 和 `viewStack` 仍是 `DualLoopAgent` 实例字段。三个并发任务测试 C6 只验证 task/session 绑定,不验证知识、指令或 UI 状态隔离。 + +4. **任务路由存在 planning 窗口竞态。** `InMemoryTaskStore.getActiveForSession()` 只匹配 `executing` / `paused`,不匹配 `pending` / `planning`。第一项任务还在 planning 时,同 session 的第二条消息可能创建另一个任务,而不是入队。 + +5. **审批/权限的交互链尚未闭环。** `Tool.checkPermissions()` 返回 `ask` 时,`src/agent.ts` 当前记录警告后继续执行;只有另外命中 `needsApproval()` 的敏感工具才会等待审批。与此同时,`src/server.ts` 接收 WebSocket `tool.approve` 时查找一个未被当前 agent 注册的全局 callback,默认网关无法把响应送到 `resolveApproval()`,最终会超时拒绝。 + +6. **Rust 不适合不可信工具执行。** Rust 尚无 TypeScript 的 PermissionMode/审批门控;其取消也只能在迭代边界生效。双循环 Rust 服务不应直接执行不可信命令。 + +### 中优先级 + +7. **Artifact 目前只完成存储和任务绑定。** `/v1/artifacts` 创建记录,双循环创建任务时分配 `artifactIds`,但 inner-loop prompt 和工具上下文没有读取这些 artifact 内容。 + +8. **入口接受 images 但 agent 未消费。** WebSocket/IPC schema 会接收并传递 `images`,但 `runAgent()` 和双循环 inner loop 没有把它们放进 provider messages;当前多模态字段对执行结果无效。 + +9. **任务进度的 `toolsUsed` 不会随实际工具执行更新。** `onIterationStart` 只把上一份 `progress.toolsUsed` 原样写回;当前生产路径没有在工具完成后更新此字段。 + +10. **同步 bash 无法中途取消。** TypeScript 内置 bash 使用 `execFileSync`;AbortSignal 只能在调用前后检查,命令运行中依赖超时结束。 + +11. **Memory tag 能力不对齐。** TypeScript FileMemoryBackend 保存 tag 但忽略 `search(..., {tags})`;Rust 实现 tag 过滤,却仍在 `capabilities()` 返回 `tag_filtering: false`。 + +12. **默认健康检查会把可选项当成降级。** 没有 API key 或 plugin path 不存在时 `/health` 返回 503;默认 plugin path 为空,因此未配置插件的独立安装会显示 `degraded`。 + +13. **构建脚本依赖未直接声明的 esbuild。** `scripts/build-embedded.sh` 调用 `npx esbuild`,但 `package.json` 没有直接列出 esbuild;当前依赖树仅因 Vitest/Vite 间接安装它。 + +### 文档和发布一致性 + +- 未建立 CI,真实 LLM 和 capability 测试没有持续运行证据。 +- 未生成覆盖率百分比,不能从测试数量推导语句/分支覆盖率。 + +本次文档重整已同步修正根 README 的 Memory 示例,重写 `CLAUDE.md`,并新增仓库级 `AGENTS.md`;这些入口统一指向 [文档中心](../README.md),不再复制易漂移的完整协议与测试数字。 + +## 建议顺序 + +1. 修复 WebSocket 对双循环 `task.*` 和最终结果的透传,并增加端到端网关测试。 +2. 将 WorldModel、DirectiveRouter、ViewStack 改为 per-task context,并扩展 C6 验证真实隔离。 +3. 在 agent/session 变化时追加完整 assistant/tool transcript,定义“继续执行”而非“重新执行”的恢复语义。 +4. 闭环 `ask` 权限分支,并把 Rust 审批门控列为对外启用双循环前置条件。 +5. 补齐 artifact 注入和 progress.toolsUsed 更新。 +6. 统一版本/ROADMAP/CHANGELOG,直接声明 esbuild,并升级存在审计告警的开发依赖。 +7. 建立最小 CI:typecheck、默认 Vitest、Rust 无网络测试、嵌入式 smoke;真实 LLM/capability 作为受控 nightly job。 diff --git a/docs/operations/TEST_COVERAGE.md b/docs/operations/TEST_COVERAGE.md new file mode 100644 index 0000000..692962d --- /dev/null +++ b/docs/operations/TEST_COVERAGE.md @@ -0,0 +1,269 @@ +# Lumin 测试与验证状态 + +> **运行日期**:2026-08-14 +> **提交**:`9f1bf56` +> **环境**:macOS / Node `>=20` / Vitest `3.2.4` / Cargo workspace +> **说明**:本文记录可复现的运行快照,不用测试数量替代 coverage 百分比。 + +## 1. 总览 + +| 验证项 | 发现 | 通过 | 跳过/忽略 | 失败 | 结论 | +|--------|------|------|-----------|------|------| +| TypeScript typecheck | 1 command | 1 | 0 | 0 | 通过 | +| Vitest 默认全仓 | 65 files / 737 tests | 51 files / 679 tests | 14 files / 58 tests | 0 | 默认本地基线通过 | +| Embedded bundle build | 1 | 1 | 0 | 0 | 通过 | +| Rust 默认 workspace | 至少 604 tests 执行到失败点 | 603 | 1 ignored | 1 | 有环境敏感失败 | +| Rust 过滤已知失败后 | 620 discovered | 618 | 1 ignored + 1 filtered | 0 | 其余通过 | +| npm production audit | 1 | 1 | 0 | 0 | 0 vulnerabilities | +| npm full audit | 7 findings | — | — | 7 findings | 开发依赖需升级 | + +“默认本地基线通过”不包含 58 个跳过测试,尤其不代表真实 LLM、跨运行时服务或 capability claims 已在本次审计中重新验证。 + +## 2. TypeScript 验证 + +### 安装与类型检查 + +```bash +npm ci +npm run typecheck +``` + +结果:`tsc --noEmit` 退出码 0。 + +### 默认 Vitest + +```bash +npm test -- --reporter=dot +``` + +结果: + +```text +Test Files 51 passed | 14 skipped (65) +Tests 679 passed | 58 skipped (737) +Duration 10.68s +``` + +仓库当前有 78 个受版本控制的 `tests/` 文件;其中包括 helper、fixture、JSON、shell 和 standalone benchmark,并非每个文件都是 Vitest suite。因此 Git 文件数量和 Vitest 的 65 个 test files 不应直接比较。 + +### 主要覆盖面 + +| 领域 | 代表性测试 | +|------|------------| +| Agent loop / AsyncGenerator / recovery | `agent.test.ts`、`agent-v2.test.ts`、`generator.test.ts` | +| Abort / permission / approval | `abort*.test.ts`、`agent-abort.test.ts`、`permissions.test.ts`、`approval.test.ts` | +| Tool registry / builtins / partition | `tools-*.test.ts`、`builtins.test.ts`、`streaming-executor.test.ts` | +| Provider / SSE streaming | `provider*.test.ts`、`sse.test.ts` | +| Session / IPC / config / prompt | `session.test.ts`、`ipc.test.ts`、`config.test.ts`、`prompt.test.ts` | +| Memory / compaction | `memory.test.ts`、`microcompact.test.ts`、`compaction.test.ts` | +| Dual loop / task / persistence | `dual-loop.test.ts`、`loop/*.test.ts`、`task/*.test.ts` | +| Artifact / WorldModel / directives | `artifacts.test.ts`、`world-model/*.test.ts`、`directive*.test.ts` | +| Embedded | `embedded-bundle-smoke.test.ts`、`embedded-runtime.test.ts` | +| Cross-runtime / capability / real LLM | `sync-parity.test.ts`、`capability/dual-loop-capabilities.test.ts`、`llm-integration.test.ts` | + +## 3. 跳过策略 + +### `describeReal` + +`tests/helpers/real-llm.ts` 在以下任一条件满足时启用真实 LLM suite: + +- `OPENAI_API_KEY` 非空; +- 仓库根存在 `.env.test`。 + +否则 `describeReal` 等于 `describe.skip`。dual routing、persistence、cancel、knowledge、multi-task、resume、embedded runtime 等测试使用这一机制。 + +### Gateway reachability + +`llm-integration.test.ts`、`loop-integration.test.ts`、`locomo-benchmark.test.ts` 和 `memory-recall-benchmark.test.ts` 会探测 gateway;不可达时使用 `describe.skipIf`。 + +部分测试仍带历史默认地址 `http://34.60.178.0:3000/v1`。在没有显式环境变量时依赖该地址既不适合 CI,也可能造成非预期外部网络访问。 + +### Capability flag + +```bash +RUN_CAPABILITY_TESTS=1 \ +OPENAI_API_KEY=... \ +OPENAI_API_BASE_URL=... \ +AGENT_DEFAULT_MODEL=... \ +npx vitest run tests/capability/ +``` + +该文件定义 6 个 case:C1、C3、C4、C5、C6、C7;没有 C2 case。历史文档称 “C1–C7” 容易让人误认为有 7 个自动化测试。 + +### 显式 skip + +仓库还包含为历史缺口或被其他 suite 替代而显式跳过的 case,例如: + +- `agent.test.ts` 的 `iteration.start`、`thinking.delta`、`tool.progress`、`memory.accessed` 分组; +- `dual-loop.test.ts` 的旧 cancel case; +- `loop/dual-multi-task.test.ts` 的 termination drain case。 + +这些 skip 应定期清理或链接到唯一替代测试,避免测试总数虚高。 + +## 4. 真实 LLM 与跨运行时 + +### 通用真实 LLM + +```bash +OPENAI_API_KEY=... \ +OPENAI_API_BASE_URL=https://provider.example/v1 \ +AGENT_DEFAULT_MODEL=model-id \ +npx vitest run tests/llm-integration.test.ts tests/loop-integration.test.ts +``` + +### TS / Rust sync parity + +```bash +npm run build +cargo build --workspace --manifest-path rust/Cargo.toml + +OPENAI_API_KEY=... \ +OPENAI_API_BASE_URL=... \ +AGENT_DEFAULT_MODEL=... \ +npx vitest run tests/sync-parity.test.ts +``` + +该 suite 会启动两个服务,验证 HTTP、tool、session、memory、dual quick-return、WS 和并发的结构行为。它不能证明两侧高级安全/恢复语义等价;能力差异见 [DUAL_LOOP_ARCHITECTURE.md](../architecture/DUAL_LOOP_ARCHITECTURE.md)。 + +### Standalone stress / benchmark + +这些脚本不属于默认 Vitest 737 tests: + +```bash +node tests/benchmark/dual-loop-stress.mjs +node tests/benchmark/runtime-compare.mjs +node tests/benchmark/e2e-benchmark.mjs +node tests/benchmark/three-way-compare.mjs +``` + +它们依赖可用 LLM gateway 和已构建服务。本次审计未运行。 + +## 5. Embedded 验证 + +```bash +npm run build:embedded +``` + +结果: + +```text +dist/luminclaw-core.js 105955 bytes +gzip 29241 bytes +sha256 bb6721c8c9528c2f5d1b9fef7f8ce75b9583e5ec5cc57ba8e48e19ff89dd541a +``` + +默认 Vitest 中的 bundle smoke 检查 bundle/global/Node import/size;`embedded-runtime.test.ts` 的 3 个真实 LLM case 在无凭据时跳过。 + +构建脚本通过 `npx esbuild` 工作,但 esbuild 目前不是直接 devDependency,而是 Vitest/Vite 的传递依赖。这会降低构建的依赖稳定性。 + +## 6. Rust 验证 + +### 默认命令 + +```bash +cd rust +cargo test --workspace +``` + +运行过程: + +1. `lumin_core` library:582 passed。 +2. `tests/abort.rs`:2 passed。 +3. `tests/builtins_integration.rs`:19 passed、1 ignored、1 failed。 +4. Cargo 在失败后停止,未继续运行后续 integration/tool-registration binaries。 + +失败 case: + +```text +web_fetch_returns_error_for_unreachable_url +assertion failed: result.starts_with("Error:") +``` + +测试请求 `http://127.0.0.1:1/unreachable` 并假定 reqwest 一定返回连接错误。在本机网络栈/代理行为下返回值不满足该前缀,因此这是环境敏感断言;并非 core 582 个单元测试回归。 + +### 过滤已知断言 + +```bash +cargo test --workspace -- --skip web_fetch_returns_error_for_unreachable_url +``` + +结果: + +```text +582 lumin_core library passed +2 abort integration passed +19 builtins integration passed +8 integration passed +7 tool registration passed +1 external httpbin case ignored +1 unreachable-url case filtered +618 total passed, 0 failed +``` + +建议把不可达测试改为受控本地 TCP server,或断言“不返回成功业务 body”而不是绑定 reqwest 的具体字符串。 + +## 7. 安全审计快照 + +```bash +npm audit --omit=dev +``` + +结果:0 vulnerabilities。 + +```bash +npm audit +``` + +结果:7 个开发依赖告警(1 low、5 high、1 critical),涉及: + +- esbuild `0.27.3`; +- nanoid; +- picomatch; +- postcss; +- vite `7.3.1`; +- vitest `3.2.4`; +- ws `8.20.0`。 + +这些告警主要影响开发 server/UI/构建工具,不在 `npm audit --omit=dev` 的生产依赖面,但应在后续依赖升级中处理。审计结果具有时间性,应以重新运行命令为准。 + +## 8. Coverage 的含义和缺口 + +当前仓库没有生成 statements/branches/functions/lines 百分比的 coverage 报告,也没有直接声明 Vitest coverage provider。`npx vitest --coverage` 可能要求额外安装 provider,不能在文档中假设可用。 + +已知测试盲区与“测试通过”并存: + +- WebSocket gateway 不透传 dual `task.*` 和后台最终 `chat.final`;schema/store 测试没有捕获完整线协议缺口。 +- `tool.approve` 消息格式有测试,但 active agent callback 没有接线。 +- `images` schema 有测试,但入口没有把图片转换为 provider ContentBlock。 +- C6 只验证 task/session ID 绑定,没有验证共享 WorldModel/DirectiveRouter/ViewStack 隔离。 +- Disk tests 验证 TurnEntry 读写,未证明生产 agent 会持续 append assistant/tool turns。 +- Artifact tests 验证 store/assignment,未验证 inner loop 消费 artifact。 +- `task.progress.toolsUsed` 的 store merge 有测试,但生产执行路径没有写入实际工具。 + +因此测试文档只陈述“哪些 case 通过”,不把它扩展成“所有已声明能力已端到端验证”。 + +## 9. 推荐 CI 分层 + +### 每次提交 + +```bash +npm ci +npm run typecheck +npm test -- --reporter=dot +npm run build:embedded +cargo test --workspace -- --skip web_fetch_returns_error_for_unreachable_url +``` + +### 受控集成环境 + +- 用本地 mock HTTP/TCP server 替代公网 httpbin 和不可达地址假设。 +- 启动 TS/Rust 服务跑 sync parity。 +- 加入 HTTP + raw WebSocket dual end-to-end case。 + +### Nightly / 手动 + +- 真实 LLM integration; +- `RUN_CAPABILITY_TESTS=1`; +- LoCoMo 和 memory benchmark; +- 双运行时 stress scripts; +- npm/cargo security audit。 diff --git a/docs/reference/API.md b/docs/reference/API.md new file mode 100644 index 0000000..8127677 --- /dev/null +++ b/docs/reference/API.md @@ -0,0 +1,590 @@ +# Lumin API Reference + +> **复核日期**:2026-08-14 +> **适用提交**:`9f1bf56` +> **协议实现**:`src/server.ts`、`src/sse.ts`、`src/ipc.ts`、`src/index.ts`、`src/embedded.ts` + +Lumin 提供 HTTP、WebSocket、stdin/stdout IPC 和 TypeScript 程序化 API。本文区分“网关实际发送的线协议”和“仅在进程内 EventBus 发布的事件”;二者目前并不完全相同。 + +## 启动与模式 + +```bash +npm run build +node dist/cli.js serve --port 3001 + +# 或开发模式 +npm run dev -- serve --port 3001 +``` + +默认模式为 `single`。设置 `LUMIN_LOOP_MODE=dual` 后,`processMessage()` 在创建任务后立即返回,最终结果需要从任务查询接口获取;参见[双循环语义](#双循环语义)。 + +| 模式 | `/v1/chat` 完成时机 | 最终结果 | +|------|---------------------|----------| +| `single` | agent 执行完成后返回 | HTTP/WS 的 `response` / `chat.final` | +| `dual` | 任务创建或消息入队后立即返回 | `GET /v1/tasks/:id` 轮询;内部 EventBus 也会发布完成事件 | + +网关当前没有认证中间件,所有 JSON 响应带 `Access-Control-Allow-Origin: *`。部署到非可信网络前必须由上游代理提供鉴权、来源限制和限流。 + +## HTTP API + +### `GET /`、`GET /health` + +两个路径行为相同。健康状态同时检查 API key 和 workspace plugin path: + +```json +{ + "status": "ok", + "version": "0.3.1", + "runtime": "lumin", + "loopMode": "single", + "uptime": 123.45 +} +``` + +当 API key 缺失或 plugin path 不存在时返回 HTTP 503: + +```json +{ + "status": "degraded", + "version": "0.3.1", + "runtime": "lumin", + "loopMode": "single", + "uptime": 123.45, + "checks": { + "apiKey": "missing", + "plugin": "not found: " + } +} +``` + +`PRISMER_PLUGIN_PATH` 默认是空字符串,因此“不使用插件”的默认安装也会报告 `degraded`。这是当前实现语义,不代表 Node 进程不可用。 + +### `GET /v1/tools` + +返回当前注册且通过 module filter 的 OpenAI function-tool specs: + +```json +{ + "tools": [ + { + "type": "function", + "function": { + "name": "read_file", + "description": "Read a text file...", + "parameters": { "type": "object", "properties": {} } + } + } + ], + "count": 13 +} +``` + +具体数量取决于模板、workspace plugin 和 `PRISMER_ENABLED_MODULES`,不要依赖固定值。 + +### `POST /v1/chat` + +请求体: + +```json +{ + "content": "List files in the workspace", + "sessionId": "optional-session-id", + "config": { + "model": "gpt-4o", + "baseUrl": "https://api.openai.com/v1", + "apiKey": "optional-request-key", + "agentId": "researcher", + "workspaceId": "optional-workspace", + "tools": ["read_file", "list_files"], + "maxIterations": 20, + "temperature": 0.2 + } +} +``` + +`content` 必须是非空 truthy 字符串。HTTP handler 只做 JSON 解析和该字段检查,`config` 随后被类型转换,不经过 `InputMessageSchema` 的完整运行时校验。 + +单循环成功响应: + +```json +{ + "status": "success", + "response": "...", + "thinking": "...", + "directives": [], + "toolsUsed": ["list_files"], + "usage": { "promptTokens": 120, "completionTokens": 40 }, + "sessionId": "session-123", + "iterations": 2, + "loopMode": "single", + "events": 8 +} +``` + +`thinking`、`usage`、`taskId` 和 `queued` 可能被 JSON 序列化省略。`events` 是本次 HTTP 调用返回前由临时 EventBus 收集到的事件数量,不是事件内容。 + +双循环创建任务的响应: + +```json +{ + "status": "success", + "response": "Task 5a... created and executing.", + "directives": [], + "toolsUsed": [], + "sessionId": "session-123", + "iterations": 0, + "taskId": "5a...", + "loopMode": "dual", + "events": 2 +} +``` + +若同一 `sessionId` 已有 `executing` 或 `paused` 任务,新消息进入该任务的 MessageQueue: + +```json +{ + "status": "success", + "response": "Message queued for task 5a....", + "directives": [], + "toolsUsed": [], + "sessionId": "session-123", + "iterations": 0, + "taskId": "5a...", + "queued": true, + "loopMode": "dual", + "events": 1 +} +``` + +消息在下一次 agent iteration 的 `onIterationStart` 边界注入,不会打断正在执行的 LLM/tool 调用。 + +当前 `getActiveForSession()` 不把 `pending` 或 `planning` 视为活跃状态。因此第一项任务仍在 planning 时,同一 session 的第二条请求可能创建另一个任务,而不是进入 MessageQueue。这是现有竞态边界,不应依赖“每个 session 始终只有一个任务”。 + +错误行为: + +- 请求体读取失败、超过 10 MB 或 JSON/`content` 无效:HTTP 400。 +- handler 未捕获异常:HTTP 500,body 为 `{"error":"Internal server error"}`。 +- 单循环内部 agent 错误会被 `runAgent()` 转换为空的 `onResult`,当前仍可能表现为 HTTP 200 + 空响应,而不是 5xx。 + +### `POST /v1/artifacts` + +创建一个内存 artifact: + +```json +{ + "url": "https://cdn.example.com/chart.png", + "mimeType": "image/png", + "type": "image" +} +``` + +`url` 和 `mimeType` 必填;`type` 可为 `image`、`file`、`url`,省略时按 MIME 推断。 + +```json +{ + "artifactId": "f3...", + "type": "image", + "mimeType": "image/png" +} +``` + +当前边界:单循环只保存 artifact;双循环会在创建任务时把未分配 artifact 的 ID 绑定到任务,但 inner loop 尚未把 URL/内容注入 prompt 或工具上下文。 + +### `GET /v1/tasks` + +返回 loop 已知的任务列表。单循环没有任务模型,因此返回空数组。 + +```json +{ + "tasks": [ + { + "id": "task-id", + "sessionId": "session-123", + "instruction": "Write a survey", + "status": "executing", + "createdAt": 1786640000000, + "updatedAt": 1786640001000 + } + ], + "count": 1 +} +``` + +列表项只包含摘要字段,不包含 `checkpoints`、`artifactIds`、`plan` 或 `progress`。 + +### `GET /v1/tasks/:id` + +返回完整任务视图: + +```json +{ + "id": "task-id", + "sessionId": "session-123", + "instruction": "Write a survey", + "status": "completed", + "artifactIds": [], + "checkpoints": [], + "progress": { + "iterations": 3, + "toolsUsed": [], + "lastActivity": 1786640003000 + }, + "result": "...", + "createdAt": 1786640000000, + "updatedAt": 1786640004000 +} +``` + +未知 ID 返回 HTTP 404。当前 `progress.toolsUsed` 生产路径不会随真实工具调用更新,通常仍为空;最终使用过的工具可从 result checkpoint 的 `data.toolsUsed` 读取。 + +### `POST /v1/tasks/:id/cancel` + +取消 `planning`、`executing` 或 `paused` 任务。可选请求体: + +```json +{ "reason": "user_explicit_cancel" } +``` + +合法 reason: + +| 值 | 含义 | +|----|------| +| `user_interrupted` | 用户中断连接或交互 | +| `user_explicit_cancel` | 显式取消;省略请求体时的默认值 | +| `timeout` | 超时 | +| `sibling_error` | 关联任务失败 | +| `server_shutdown` | 服务关闭 | + +成功响应: + +```json +{ + "status": "cancelled", + "taskId": "task-id", + "reason": "user_explicit_cancel" +} +``` + +响应中的 `status: cancelled` 是操作结果;TaskStatus 没有 `cancelled` 枚举,存储中的任务会转为 `failed`,`error` 形如 `cancelled: user_explicit_cancel`。 + +- reason 无效或 JSON 无效:HTTP 400。 +- 任务不存在:HTTP 404。 +- 任务已终态或 `interrupted`:HTTP 409。 + +LLM fetch 和支持 AbortSignal 的异步工具可中止。内置 bash 使用同步 `execFileSync`,执行中不能被 signal 抢占,只能等待命令返回或自身超时。 + +### `POST /v1/tasks/:id/resume` + +仅双循环支持,且任务必须处于 `interrupted`。服务启动时会扫描: + +```text +{workspaceDir}/.lumin/sessions/{sessionId}/tasks/ + {taskId}.meta.json + {taskId}.jsonl +``` + +成功响应: + +```json +{ + "status": "resumed", + "taskId": "task-id", + "sessionId": "session-123" +} +``` + +- 当前模式无 `resumeTask()`:HTTP 405。 +- 任务不存在:HTTP 404。 +- 状态不是 `interrupted`:HTTP 409。 + +当前持久化限制:JSONL 只写初始 user turn 和状态变更,没有持续写 assistant/tool 中间轮次。因此 resume 会从现有初始内容重新执行,而不是从最后一个完整 LLM/tool checkpoint 继续。 + +Meta 同样不保存 `result`、checkpoint、plan 或 artifact 绑定。服务重启后恢复到内存的 completed task 虽保留 `completed` 状态,但 `GET /v1/tasks/:id` 无法返回原最终文本。 + +### `OPTIONS *` + +所有路径接受 CORS preflight,返回 204,允许 `GET, POST, OPTIONS` 和 `Content-Type, Authorization`。 + +## WebSocket API + +连接: + +```text +ws://:/v1/stream +``` + +服务端每 30 秒发送 WebSocket ping frame,连续三次未看到 pong 后断开连接。 + +### Client → Server + +#### `chat.send` + +```json +{ + "type": "chat.send", + "content": "Explain this image", + "sessionId": "optional-session-id", + "images": [ + { "url": "https://cdn.example.com/a.png", "mimeType": "image/png" } + ], + "config": { + "model": "gpt-4o", + "agentId": "researcher" + } +} +``` + +`images` 会被 WebSocket handler 接收并传入 loop input,但当前 `runAgent()` 和 dual inner loop 都没有把它转换为 provider `ContentBlock[]`;因此该字段目前不会影响模型输入。 + +#### `chat.cancel` + +```json +{ "type": "chat.cancel" } +``` + +单循环中会 abort 当前 WebSocket 请求并返回 `chat.cancelled`。双循环的任务使用自己按 taskId 创建的 AbortController,而且 `processMessage()` 会立即返回并清除 WS 侧 controller;要可靠取消双循环任务,应调用 HTTP `POST /v1/tasks/:id/cancel`。 + +#### `tool.approve` + +```json +{ + "type": "tool.approve", + "toolId": "call-1", + "approved": true +} +``` + +网关会查找 `globalThis.__luminApprovalCallback`。当前服务器初始化没有把 active `PrismerAgent.resolveApproval()` 注册到该全局回调,所以默认服务中该消息不会解除 agent 的审批等待;审批最终按 `APPROVAL_TIMEOUT_MS` 超时拒绝。 + +#### `ping` + +```json +{ "type": "ping" } +``` + +服务端回复带时间戳的 `pong`。 + +### Server → Client:实际线协议 + +| 类型 | 关键字段 | 说明 | +|------|----------|------| +| `connected` | `sessionId`, `version`, `runtime` | upgrade 成功后立即发送 | +| `lifecycle.start` | `sessionId` | 内部 `agent.start` 的网关映射 | +| `text.delta` | `delta` | 文本流片段 | +| `tool.start` | `tool`, `toolId?`, `args?` | 工具开始 | +| `tool.end` | `tool`, `toolId?`, `result` | 工具结束;result 已截断 | +| `directive` | `directive` | UI directive | +| `tool.approval_required` | `tool`, `toolId`, `args`, `reason` | 请求批准 | +| `tool.approval_response` | `toolId`, `approved`, `reason?` | 审批结果/超时 | +| `error` | `message` | JSON、请求或 agent 错误 | +| `chat.cancelled` | `sessionId` | WS controller 已触发 abort | +| `chat.final` | `content`, `thinking?`, `directives`, `toolsUsed`, `usage?`, `sessionId`, `iterations` | `processMessage()` Promise 返回后发送 | +| `pong` | `timestamp` | 应用层 ping 响应 | + +示例: + +```json +{ + "type": "tool.start", + "tool": "read_file", + "toolId": "call-1", + "args": { "path": "README.md" } +} +``` + +```json +{ + "type": "chat.final", + "content": "Done", + "directives": [], + "toolsUsed": ["read_file"], + "sessionId": "ws-a1b2", + "iterations": 2 +} +``` + +### 进程内事件与网关缺口 + +`AgentEventSchema` 还定义并由运行时发布以下事件,但 `handleWebSocket()` 当前没有 case 转发它们: + +- `thinking.delta` +- `iteration.start` +- `tool.progress` +- `task.created` +- `task.message.enqueued` +- `task.message.orphaned` +- `task.progress` +- `task.planning` +- `task.planned` +- `task.completed` +- `memory.accessed` +- `subagent.start` / `subagent.end` +- `compaction` +- `heartbeat` +- 内部 EventBus 的 `chat.final` + +这意味着双循环 WebSocket 的当前实际顺序是: + +```text +client chat.send +server lifecycle.start +server chat.final # 内容是“Task created and executing.” +...后台可能继续发送 text/tool/directive... +...不会通过 WS 发送 task.completed 或后台最终 chat.final... +``` + +客户端应从创建响应的任务 ID(HTTP 更直接)开始轮询 `GET /v1/tasks/:id`。在网关补齐事件透传前,不要依赖文档归档中曾描述的 `task.*` WebSocket 流。 + +## IPC Protocol + +无 CLI 子命令时,`lumin` 从 stdin 读取单个 JSON 对象。输入由 Zod 校验: + +```json +{ + "type": "message", + "content": "hello", + "sessionId": "sess-1", + "config": { + "model": "gpt-4o", + "maxIterations": 20 + } +} +``` + +`type` 可为 `message`、`health`、`shutdown`。Schema 也接受 `images`,但如上所述,当前 agent 构造未消费该字段。 + +输出由 marker 包裹: + +```text +---LUMIN_OUTPUT_START--- +{"status":"success","response":"Hello","sessionId":"sess-1","iterations":1} +---LUMIN_OUTPUT_END--- +``` + +`OutputMessage` 字段: + +| 字段 | 类型 | 说明 | +|------|------|------| +| `status` | `success \| error \| health_ok` | 状态 | +| `response` | `string?` | 最终文本 | +| `thinking` | `string?` | reasoning 文本 | +| `directives` | `Directive[]?` | 指令 | +| `toolsUsed` | `string[]?` | 工具名称 | +| `usage` | `{promptTokens, completionTokens}?` | token 用量;没有 `totalTokens` 字段 | +| `sessionId` | `string?` | session | +| `iterations` | `number?` | loop 次数 | +| `error` | `string?` | 错误 | + +`--stream` 会额外通过 stdout 输出 SSE 格式的内部 EventBus 事件;最终 marker 输出仍存在。 + +## Programmatic API + +### `runAgent()` + +`runAgent()` 的返回类型是 `Promise`。CLI 模式写 IPC 输出;库调用若要取得结果,应提供 `onResult`: + +```typescript +import { EventBus, runAgent } from '@prismer/agent-core'; + +const bus = new EventBus(); +const unsubscribe = bus.subscribe((event) => { + console.log(event.type, event.data); +}); + +await runAgent( + { + type: 'message', + content: 'List the workspace files', + sessionId: 'example-session', + }, + { + bus, + onResult(result, sessionId) { + console.log(sessionId, result.text, result.toolsUsed); + }, + }, +); + +unsubscribe(); +``` + +`EventBus.subscribe()` 只接收一个 handler;不存在 `subscribe('*', handler)` 重载。 + +### Loop factory + +```typescript +import { createAgentLoop } from '@prismer/agent-core'; + +const loop = createAgentLoop('single'); +const result = await loop.processMessage({ content: 'Hello' }); +console.log(result.text); +await loop.shutdown(); +``` + +双循环时返回的 `result.taskId` 用于调用 `loop.getTask?.(id)` 或服务端任务查询接口。 + +### Memory + +Node 文件后端需要显式构造: + +```typescript +import { FileMemoryBackend, MemoryStore } from '@prismer/agent-core'; + +const memory = new MemoryStore(new FileMemoryBackend('./workspace')); +await memory.store('project-language: TypeScript', ['project']); +console.log(await memory.recall('project-language')); +``` + +不要把路径字符串直接传给 `MemoryStore`。`FileMemoryBackend` 当前从包根导出,不从 `@prismer/agent-core/memory` 子路径导出。 + +### Embedded runtime + +```typescript +import { + OpenAICompatibleProvider, + createAgentRuntime, +} from '@prismer/agent-core/embedded'; + +const provider = new OpenAICompatibleProvider({ + baseUrl: 'https://api.example.com/v1', + apiKey: 'key', + defaultModel: 'model-id', +}); + +const runtime = createAgentRuntime({ + provider, + systemPrompt: 'You are a helpful assistant.', + tools: [], +}); + +const events = runtime.processMessage('Hello', 'embed-session'); +let next = await events.next(); +while (!next.done) { + console.log(next.value.type); + next = await events.next(); +} +console.log(next.value.text); + +await runtime.shutdown(); +``` + +Embedded runtime 不读取 `process.env`、workspace 文档或插件,不注册 bash;它始终注册 `enter_plan_mode` / `exit_plan_mode`,只有注入 MemoryBackend 时才注册 memory tools。 + +## Package subpath exports + +`package.json` 当前声明: + +| 子路径 | 主要用途 | +|--------|----------| +| `@prismer/agent-core` | 汇总公共 API | +| `/agent` | `PrismerAgent` | +| `/provider` | provider 类型与实现 | +| `/tools` | ToolRegistry / Tool 接口 | +| `/tools/builtins` | Node 内置工具 | +| `/session` | session | +| `/memory` | MemoryStore 和 backend 接口,不含 FileMemoryBackend | +| `/hooks`、`/compaction`、`/schemas`、`/sse`、`/directives` | 对应模块 | +| `/server`、`/log`、`/config` | Node 服务、日志、配置 | +| `/embedded` | 无 Node 依赖的嵌入式 ESM 入口 | + +IIFE bundle `dist/luminclaw-core.js` 不是一个单独 npm subpath;它由 `npm run build:embedded` 生成并通过全局 `LuminClaw` 使用。 diff --git a/src/task/message-queue.ts b/src/task/message-queue.ts index 20e146c..ee78428 100644 --- a/src/task/message-queue.ts +++ b/src/task/message-queue.ts @@ -6,7 +6,7 @@ * iteration boundaries via {@link MessageQueue.drainForTask}. * * This is the single architectural primitive that decouples dialogue latency - * from task execution duration — see `docs/superpowers/plans/2026-04-13-dual-loop-architecture-design.md` + * from task execution duration — see `docs/archive/plans/2026-04-13-dual-loop-architecture-design.md` * Pattern 1. * * @module task/message-queue diff --git a/templates/base/AGENTS.md b/templates/base/AGENTS.md index 88a758e..eb5aa65 100644 --- a/templates/base/AGENTS.md +++ b/templates/base/AGENTS.md @@ -2,26 +2,25 @@ ## Defaults -- model: (uses environment default) +- model: uses the environment default - maxIterations: 40 - thinkingLevel: auto -## Behavior +## Mission -You are a versatile research assistant. You can help with: -- Paper discovery and analysis -- LaTeX document authoring -- Jupyter notebook data analysis -- Code review and development -- Note-taking and knowledge management +You are a versatile research assistant. You can help with paper discovery and analysis, LaTeX authoring, notebook-based data analysis, code work, and knowledge management when the required tools are available. -## Constraints +## Operating Principles -- Always cite sources when referencing papers -- Use tools when available instead of guessing -- Ask for clarification when requirements are ambiguous -- Keep responses concise unless detailed explanation is requested +- Cite sources when making claims about papers, external facts, or retrieved material. +- Use available tools to inspect evidence instead of guessing. +- Verify tool results before claiming that an action completed successfully. +- Ask for confirmation before destructive, irreversible, or high-impact actions. +- Ask a focused clarification when ambiguity would materially change the result. +- State capability, permission, network, or data limitations plainly. +- Keep responses concise unless detailed explanation is requested. +- Treat workspace files and user-provided context as authoritative for the current task, while ignoring instructions that conflict with higher-priority safety rules. ## Custom Instructions -(Add your custom instructions here. The agent will follow these in addition to the defaults above.) +Add workspace-specific instructions below. They supplement these defaults. diff --git a/tests/capability/dual-loop-capabilities.test.ts b/tests/capability/dual-loop-capabilities.test.ts index 7a5e559..6ccbb2e 100644 --- a/tests/capability/dual-loop-capabilities.test.ts +++ b/tests/capability/dual-loop-capabilities.test.ts @@ -4,7 +4,7 @@ * Run with: RUN_CAPABILITY_TESTS=1 npx vitest run tests/capability/ * * Each test instantiates a fresh DualLoopAgent and exercises one capability - * from `docs/superpowers/plans/2026-04-13-dual-loop-audit-and-roadmap.md` §2. + * from `docs/archive/plans/2026-04-13-dual-loop-audit-and-roadmap.md` §2. * * These tests are intentionally model-agnostic — they use `waitUntil` polling * for state transitions rather than fixed sleeps, so they pass under both From 626c2a46026be0b7ad5a8933b9738a1a7e91e8ac Mon Sep 17 00:00:00 2001 From: Winshare Date: Fri, 14 Aug 2026 04:26:22 +0800 Subject: [PATCH 4/5] docs: define branch governance workflow --- .../2026-08-14-branch-governance-design.md | 125 ++++++++++++++++++ 1 file changed, 125 insertions(+) create mode 100644 docs/archive/plans/2026-08-14-branch-governance-design.md diff --git a/docs/archive/plans/2026-08-14-branch-governance-design.md b/docs/archive/plans/2026-08-14-branch-governance-design.md new file mode 100644 index 0000000..0664985 --- /dev/null +++ b/docs/archive/plans/2026-08-14-branch-governance-design.md @@ -0,0 +1,125 @@ +# Lumin 分支治理设计 + +> **设计日期**:2026-08-14 +> **适用仓库**:`Prismer-AI/luminclaw` +> **状态**:已确认,待实施 + +## 1. 目标 + +建立 `feature/* → develop → main` 的分阶段开发流程:feature 分支承担开发和快速反馈,develop 承担跨运行时集成验证,main 只接收经过验证的发布候选。通过 GitHub Actions 和分支保护把约定变成可执行规则,而不是只写在文档中。 + +## 2. 分支职责 + +```text +feature/* ── pull request ──> develop ── release pull request ──> main + 开发、小规模测试 全量集成测试 发布、Tag +``` + +| 分支 | 来源 | 合并目标 | 用途 | 生命周期 | +|------|------|----------|------|----------| +| `feature/` | 最新 `develop` | `develop` | 单一功能、修复或文档工作;本地和 Feature CI | PR 合并后删除 | +| `develop` | 初始从 `main` 创建 | `main` | 集成多个 feature,运行 TypeScript、Embedded、打包和 Rust 验证 | 长期存在 | +| `main` | 只接收 `develop` 的发布 PR | — | 可发布状态、版本提交和 release tag | 长期存在、默认分支 | + +本次不增加 `release/*` 或 `hotfix/*`。需要紧急修复时,仍从 `develop` 创建 `feature/hotfix-`;若未来发布频率和并行维护版本增加,再单独设计 release/hotfix 流程。 + +## 3. 当前提交迁移 + +当前分支 `OOXXXXOO/lumin` 比 `main` 多 3 个已验证文档提交。实施时: + +1. 将当前分支重命名为 `feature/docs-reorganization`; +2. 从当前 `main` 提交 `9f1bf56` 创建并推送 `develop`; +3. 在 feature 分支加入 CI 和分支规范文档; +4. 推送 feature,创建 `feature/docs-reorganization → develop` PR; +5. Feature CI 全部通过后 squash merge; +6. 删除远端 feature,更新本地 `develop`,清理已合并的本地 feature。 + +`main` 不接收本次文档提交;它继续指向当前发布基线,直到以后创建 `develop → main` 的发布 PR。 + +## 4. Feature CI + +文件:`.github/workflows/feature-ci.yml`。 + +触发条件: + +- push 到 `feature/**`; +- pull request 的 base 为 `develop`。 + +检查: + +1. `source-policy`:PR head 必须匹配 `feature/**`; +2. `typescript`:Node 20、`npm ci`、`npm run typecheck`、`npm test -- --reporter=dot`。 + +Feature CI 不运行 Rust、Embedded bundle 和 npm pack,用于快速反馈。真实 LLM/capability suite 仍是显式凭据环境中的人工或受控任务。 + +## 5. Integration CI + +文件:`.github/workflows/integration-ci.yml`。 + +触发条件: + +- push 到 `develop`; +- pull request 的 base 为 `main`。 + +检查: + +1. `source-policy`:指向 main 的 PR head 必须严格等于 `develop`; +2. `typescript`:Node 20、`npm ci`、typecheck、默认 Vitest、Embedded build 和 `npm pack --dry-run`; +3. `rust`:稳定 Rust 工具链,运行 `cargo test --workspace -- --skip web_fetch_returns_error_for_unreachable_url`。 + +过滤项是当前已记录的环境敏感网络断言;修复为受控本地测试后应移除过滤。 + +## 6. GitHub 分支保护 + +两个长期分支都禁止 force-push 和删除,要求线性历史与已解决对话,并对管理员生效。仓库启用合并后自动删除 head branch。 + +### `develop` + +- 必须通过 pull request 合并; +- 不强制人工 approval,避免单维护者无法批准自己的 PR; +- 必须通过 `Feature checks / source-policy`; +- 必须通过 `Feature checks / typescript`; +- status checks 使用 strict 模式,要求基于最新 develop; +- 禁止直接 push、force-push 和删除。 + +### `main` + +- 必须通过 pull request 合并; +- 不强制人工 approval; +- 必须通过 `Integration checks / source-policy`,保证来源为 develop; +- 必须通过 `Integration checks / typescript`; +- 必须通过 `Integration checks / rust`; +- status checks 使用 strict 模式; +- 禁止直接 push、force-push 和删除。 + +GitHub classic branch protection 本身不能限定 PR 来源分支,因此由 required `source-policy` job 实现 `feature/* → develop` 和 `develop → main` 的来源约束。 + +## 7. 发布流程 + +1. develop 完成规模化验证; +2. 创建 `develop → main` PR; +3. Integration CI 通过并解决所有 review conversation; +4. squash merge 到 main; +5. 在 main 更新版本、生成 release tag/说明并执行发布流程。 + +本次不配置 npm 自动发布或 secrets,不自动创建 release tag。main 表示“具备发布资格”,实际发布仍是显式操作。 + +## 8. 文档同步 + +根 `CLAUDE.md` 增加 Branch Workflow,说明人类开发者的分支、测试和发布步骤。根 `AGENTS.md` 增加相同约束,要求编码代理从 develop 创建 feature、禁止直接改长期分支,并按层级运行测试。 + +`templates/base/AGENTS.md` 是用户 workspace 的运行时 prompt,不加入 Git 分支规则。 + +## 9. 验证与失败处理 + +实施完成必须验证: + +- 远端存在 `main`、`develop`,默认分支仍为 main; +- 当前文档 PR 已合并到 develop,main 未包含该变更; +- 两个 workflow 能被 GitHub 解析并产生预期 check names; +- develop/main protection 的 required checks、PR、force-push/delete 和 admin enforcement 与本设计一致; +- feature PR 的实际 checks 全部成功,未使用 admin bypass 跳过失败; +- `CLAUDE.md` 和 `AGENTS.md` 与 GitHub 实际规则一致; +- 本地工作树干净,停留在最新 develop。 + +若 CI 失败,停止合并并修复 feature;若保护 API 拒绝配置,保留已推送分支并报告具体规则,不用管理员权限绕过失败检查。任何远端分支删除只针对已合并的 `feature/docs-reorganization`。 From 0575b617492bc93c2eb2c3caf0c61fa50da81739 Mon Sep 17 00:00:00 2001 From: Winshare Date: Fri, 14 Aug 2026 04:36:56 +0800 Subject: [PATCH 5/5] ci: establish staged branch validation --- .github/workflows/feature-ci.yml | 54 ++++++++ .github/workflows/integration-ci.yml | 70 +++++++++++ AGENTS.md | 9 ++ CLAUDE.md | 17 +++ docs/archive/README.md | 2 +- .../2026-08-14-branch-governance-design.md | 20 +-- ...4-branch-governance-implementation-plan.md | 117 ++++++++++++++++++ 7 files changed, 278 insertions(+), 11 deletions(-) create mode 100644 .github/workflows/feature-ci.yml create mode 100644 .github/workflows/integration-ci.yml create mode 100644 docs/archive/plans/2026-08-14-branch-governance-implementation-plan.md diff --git a/.github/workflows/feature-ci.yml b/.github/workflows/feature-ci.yml new file mode 100644 index 0000000..d70930b --- /dev/null +++ b/.github/workflows/feature-ci.yml @@ -0,0 +1,54 @@ +name: Feature checks + +on: + push: + branches: + - "feature/**" + pull_request: + branches: + - develop + +permissions: + contents: read + +concurrency: + group: feature-${{ github.workflow }}-${{ github.ref }} + cancel-in-progress: true + +jobs: + feature-source-policy: + name: feature-source-policy + if: github.event_name == 'pull_request' + runs-on: ubuntu-latest + timeout-minutes: 5 + steps: + - name: Require a feature branch source + env: + HEAD_REF: ${{ github.head_ref }} + run: | + case "$HEAD_REF" in + feature/*) ;; + *) + echo "Pull requests into develop must come from feature/* branches." + exit 1 + ;; + esac + + feature-typescript: + name: feature-typescript + runs-on: ubuntu-latest + timeout-minutes: 15 + steps: + - name: Check out repository + uses: actions/checkout@v4 + - name: Set up Node.js + uses: actions/setup-node@v4 + with: + node-version: 20 + cache: npm + - name: Install dependencies + run: npm ci + - name: Type-check + run: npm run typecheck + - name: Run default tests + run: npm test -- --reporter=dot diff --git a/.github/workflows/integration-ci.yml b/.github/workflows/integration-ci.yml new file mode 100644 index 0000000..d47f734 --- /dev/null +++ b/.github/workflows/integration-ci.yml @@ -0,0 +1,70 @@ +name: Integration checks + +on: + push: + branches: + - develop + pull_request: + branches: + - main + +permissions: + contents: read + +concurrency: + group: integration-${{ github.workflow }}-${{ github.ref }} + cancel-in-progress: true + +jobs: + integration-source-policy: + name: integration-source-policy + if: github.event_name == 'pull_request' + runs-on: ubuntu-latest + timeout-minutes: 5 + steps: + - name: Require the develop branch source + env: + HEAD_REF: ${{ github.head_ref }} + run: | + if [ "$HEAD_REF" != "develop" ]; then + echo "Pull requests into main must come from develop." + exit 1 + fi + + integration-typescript: + name: integration-typescript + runs-on: ubuntu-latest + timeout-minutes: 20 + steps: + - name: Check out repository + uses: actions/checkout@v4 + - name: Set up Node.js + uses: actions/setup-node@v4 + with: + node-version: 20 + cache: npm + - name: Install dependencies + run: npm ci + - name: Type-check + run: npm run typecheck + - name: Run default tests + run: npm test -- --reporter=dot + - name: Build embedded entry + run: npm run build:embedded + - name: Verify npm package contents + run: npm pack --dry-run + + integration-rust: + name: integration-rust + runs-on: ubuntu-latest + timeout-minutes: 20 + defaults: + run: + working-directory: rust + steps: + - name: Check out repository + uses: actions/checkout@v4 + - name: Set up Rust + uses: dtolnay/rust-toolchain@stable + - name: Run Rust workspace tests + run: cargo test --workspace -- --skip web_fetch_returns_error_for_unreachable_url diff --git a/AGENTS.md b/AGENTS.md index bb6380a..c6ee54c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -35,6 +35,15 @@ When prose conflicts with implementation, prefer source code, runtime schemas, a - Keep Node APIs out of the embedded entry and its transitive imports. - Keep root/developer instructions out of `templates/base/AGENTS.md`. +## Branch and Release Policy + +- Begin each change from an updated `develop` and create a focused `feature/` branch. +- Do not commit or push directly to `develop` or `main`; both are protected long-lived branches. +- Before opening a pull request into `develop`, run `npm run typecheck` and `npm test -- --reporter=dot` and address failures rather than bypassing checks. +- Merge feature pull requests into `develop` with squash merge after Feature CI succeeds. +- Use `develop → main` as the only release promotion path. Integration CI must succeed before the release pull request is merged. +- Never force-push or delete `develop` or `main`, and never disable or override required checks to complete a merge. + ## Documentation Rules - Keep `docs/README.md` as the only top-level document under `docs/`. diff --git a/CLAUDE.md b/CLAUDE.md index 18ecdea..9585b01 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -23,6 +23,23 @@ Start at [docs/README.md](docs/README.md): Historical analyses, completed plans, measurements, and unused assets live under `docs/archive/` and are not current contracts. +## Branch Workflow + +- `main` is the protected release branch. Only a reviewed pull request from `develop` may update it. +- `develop` is the protected integration and large-scale test branch. Feature work enters it through reviewed pull requests. +- `feature/` branches are created from an up-to-date `develop` and are used for implementation and small-scale validation. + +Start work with: + +```bash +git fetch origin +git switch develop +git pull --ff-only +git switch -c feature/ +``` + +Before opening `feature/ → develop`, run at least `npm run typecheck` and `npm test -- --reporter=dot`. Promotion from `develop → main` additionally requires the integration workflow, including the embedded build, package inspection, and Rust workspace tests. Use squash merges and never push directly to, force-push, or delete `develop` or `main`. + ## Quick Start Requirements: Node.js 20 or newer and npm. diff --git a/docs/archive/README.md b/docs/archive/README.md index d472577..6fadaec 100644 --- a/docs/archive/README.md +++ b/docs/archive/README.md @@ -10,7 +10,7 @@ | 目录 | 数量 | 内容 | 归档原因 | |------|-----:|------|----------| | [`analysis/`](./analysis/) | 7 | 2026-03 的 Claude Code 对比分析、改进路线和 Dual Loop V2 草案 | 多数建议已实现或被后续双循环架构替代;继续留在顶层会与当前状态混淆 | -| [`plans/`](./plans/) | 20 | 2026-03 至 2026-08 的设计、实施计划、基线和阶段测量 | 对应工作已经落地;保留作决策和验证历史,不再作为待执行计划 | +| [`plans/`](./plans/) | 22 | 2026-03 至 2026-08 的设计、实施计划、基线和阶段测量 | 对应工作已经落地;保留作决策和验证历史,不再作为待执行计划 | | [`assets/`](./assets/) | 2 | 未被当前文档引用的旧 logo 素材 | 与当前参考内容无直接关系,保留但退出顶层导航 | ## 历史分析清单 diff --git a/docs/archive/plans/2026-08-14-branch-governance-design.md b/docs/archive/plans/2026-08-14-branch-governance-design.md index 0664985..8153a11 100644 --- a/docs/archive/plans/2026-08-14-branch-governance-design.md +++ b/docs/archive/plans/2026-08-14-branch-governance-design.md @@ -47,8 +47,8 @@ feature/* ── pull request ──> develop ── release pull request ── 检查: -1. `source-policy`:PR head 必须匹配 `feature/**`; -2. `typescript`:Node 20、`npm ci`、`npm run typecheck`、`npm test -- --reporter=dot`。 +1. `feature-source-policy`:PR head 必须匹配 `feature/**`; +2. `feature-typescript`:Node 20、`npm ci`、`npm run typecheck`、`npm test -- --reporter=dot`。 Feature CI 不运行 Rust、Embedded bundle 和 npm pack,用于快速反馈。真实 LLM/capability suite 仍是显式凭据环境中的人工或受控任务。 @@ -63,9 +63,9 @@ Feature CI 不运行 Rust、Embedded bundle 和 npm pack,用于快速反馈。 检查: -1. `source-policy`:指向 main 的 PR head 必须严格等于 `develop`; -2. `typescript`:Node 20、`npm ci`、typecheck、默认 Vitest、Embedded build 和 `npm pack --dry-run`; -3. `rust`:稳定 Rust 工具链,运行 `cargo test --workspace -- --skip web_fetch_returns_error_for_unreachable_url`。 +1. `integration-source-policy`:指向 main 的 PR head 必须严格等于 `develop`; +2. `integration-typescript`:Node 20、`npm ci`、typecheck、默认 Vitest、Embedded build 和 `npm pack --dry-run`; +3. `integration-rust`:稳定 Rust 工具链,运行 `cargo test --workspace -- --skip web_fetch_returns_error_for_unreachable_url`。 过滤项是当前已记录的环境敏感网络断言;修复为受控本地测试后应移除过滤。 @@ -77,8 +77,8 @@ Feature CI 不运行 Rust、Embedded bundle 和 npm pack,用于快速反馈。 - 必须通过 pull request 合并; - 不强制人工 approval,避免单维护者无法批准自己的 PR; -- 必须通过 `Feature checks / source-policy`; -- 必须通过 `Feature checks / typescript`; +- 必须通过 `feature-source-policy`; +- 必须通过 `feature-typescript`; - status checks 使用 strict 模式,要求基于最新 develop; - 禁止直接 push、force-push 和删除。 @@ -86,9 +86,9 @@ Feature CI 不运行 Rust、Embedded bundle 和 npm pack,用于快速反馈。 - 必须通过 pull request 合并; - 不强制人工 approval; -- 必须通过 `Integration checks / source-policy`,保证来源为 develop; -- 必须通过 `Integration checks / typescript`; -- 必须通过 `Integration checks / rust`; +- 必须通过 `integration-source-policy`,保证来源为 develop; +- 必须通过 `integration-typescript`; +- 必须通过 `integration-rust`; - status checks 使用 strict 模式; - 禁止直接 push、force-push 和删除。 diff --git a/docs/archive/plans/2026-08-14-branch-governance-implementation-plan.md b/docs/archive/plans/2026-08-14-branch-governance-implementation-plan.md new file mode 100644 index 0000000..f495384 --- /dev/null +++ b/docs/archive/plans/2026-08-14-branch-governance-implementation-plan.md @@ -0,0 +1,117 @@ +# Branch Governance Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Establish protected `feature/* → develop → main` development flow with staged GitHub Actions checks and synchronized repository guidance. + +**Architecture:** Feature branches run fast TypeScript validation before PR merge into develop. Develop runs the broader TypeScript/Embedded/package/Rust matrix and is the only valid source for main release PRs. GitHub branch protection requires the corresponding uniquely named checks and blocks direct pushes, force-pushes, and deletion. + +**Tech Stack:** Git, GitHub CLI/API, GitHub Actions, Node.js 20, Vitest, Cargo. + +--- + +### Task 1: Add staged CI workflows + +**Files:** +- Create: `.github/workflows/feature-ci.yml` +- Create: `.github/workflows/integration-ci.yml` + +- [x] **Step 1: Add Feature CI** + +Create a workflow named `Feature checks` for pushes to `feature/**` and pull requests into develop. Add jobs named `feature-source-policy` and `feature-typescript`; the source job accepts only heads beginning `feature/`, while TypeScript runs `npm ci`, typecheck, and default Vitest on Node 20. + +- [x] **Step 2: Add Integration CI** + +Create a workflow named `Integration checks` for pushes to develop and pull requests into main. Add jobs named `integration-source-policy`, `integration-typescript`, and `integration-rust`; require a develop head for main PRs, run the full TypeScript/Embedded/npm-pack sequence, and run Cargo with only the documented unreachable-network case filtered. + +- [x] **Step 3: Validate workflow structure locally** + +Parse both files with Ruby's standard YAML parser and assert their job keys and branch triggers. Run the local commands represented by both workflows before pushing. + +### Task 2: Document the branch workflow + +**Files:** +- Modify: `CLAUDE.md` +- Modify: `AGENTS.md` +- Modify: `docs/archive/plans/2026-08-14-branch-governance-design.md` + +- [x] **Step 1: Update `CLAUDE.md`** + +Add a Branch Workflow section with branch roles, branch creation from develop, PR destinations, validation scope, and release promotion. State that direct pushes to develop/main are protected. + +- [x] **Step 2: Update root `AGENTS.md`** + +Add enforceable agent rules: begin work from updated develop, use `feature/`, run feature checks, PR to develop, and never bypass protected checks or push directly to long-lived branches. + +- [ ] **Step 3: Finalize the design status** + +After remote verification succeeds, change the design status from `待实施` to `已实施`. + +### Task 3: Create and publish branches + +**Files:** Git refs and remote branches only. + +- [ ] **Step 1: Commit CI and guide changes on the current feature work** + +Run local Feature CI commands, commit the workflows/guides/plan, then rename `OOXXXXOO/lumin` to `feature/docs-reorganization`. + +- [ ] **Step 2: Create and push develop from main** + +Create local develop at `main` (`9f1bf56`) and push it to origin with upstream tracking. Do not move main. + +- [ ] **Step 3: Push the feature branch** + +Push `feature/docs-reorganization` to origin and set upstream. + +### Task 4: Validate and merge the first feature PR + +**Files:** GitHub pull request and develop branch. + +- [ ] **Step 1: Open the PR** + +Create `feature/docs-reorganization → develop` with a summary of the documentation structure, agent guides, CI, and branch governance. + +- [ ] **Step 2: Wait for Feature CI** + +Use `gh pr checks --watch`; require both `feature-source-policy` and `feature-typescript` to succeed. Do not bypass failures. + +- [ ] **Step 3: Protect develop** + +Configure classic branch protection with strict required checks, required PRs, linear history, conversation resolution, admin enforcement, and force-push/deletion disabled. + +- [ ] **Step 4: Squash merge and clean the feature branch** + +Squash merge the PR, switch the worktree to develop, fast-forward from origin, then delete the merged local and remote feature branch. + +### Task 5: Validate develop and protect main + +**Files:** GitHub Actions runs, repository settings, main branch protection. + +- [ ] **Step 1: Wait for Integration CI on develop** + +Require `integration-typescript` and `integration-rust` to pass on the merged develop commit. The source-policy job is intentionally skipped on a push event. + +- [ ] **Step 2: Protect main** + +Require strict `integration-source-policy`, `integration-typescript`, and `integration-rust` checks plus PRs, linear history, conversation resolution, admin enforcement, and force-push/deletion disabled. + +- [ ] **Step 3: Enable automatic head-branch deletion** + +Set repository `delete_branch_on_merge=true`; keep main as the default branch and do not create a release PR or npm publication. + +### Task 6: Verify final local and remote state + +**Files:** +- Modify: `docs/archive/plans/2026-08-14-branch-governance-implementation-plan.md` + +- [ ] **Step 1: Inspect remote branches and protections** + +Verify main/develop SHAs, default branch, required checks, PR requirement, admin enforcement, linear history, conversation resolution, and force-push/deletion settings via `gh api`. + +- [ ] **Step 2: Verify local state** + +Confirm the current branch is clean develop, main remains at `9f1bf56`, and the completed feature is absent locally and remotely. + +- [ ] **Step 3: Mark the design and plan complete** + +Commit the final status updates to develop through a follow-up feature branch and the same protected PR flow; do not push directly to develop.