diff --git a/README.md b/README.md index 0aa9017..b6bc643 100644 --- a/README.md +++ b/README.md @@ -89,6 +89,7 @@ Rule of thumb: if you build and operate your own agent in production, use a trac - **Global search** — One search box across all seven platforms at once, multi-keyword AND matching, colored platform badges per hit — including prompts recovered from sessions that Claude Code's cleanup already deleted - **Session insights** — Aggregate analytics dashboard with tool stats, error clustering and daily trends - **Evidence-backed failure events (React UI)** — Groups unresolved failures by the same tool, complete arguments and call's user turn, with repeated operations first, first/last evidence jumps and every original result retained. Successful results split groups; missing arguments stay separate. Execution success requires an explicit zero exit code or OMP-native completion evidence. These are review groups, not root-cause diagnoses or proof of task failure. Local rules, no LLM. [Try the synthetic walkthrough and read the boundaries](docs/diagnostics.md). +- **Follow-up evidence candidates** — See later calls differing only in `i`, or same-turn modifications of the same explicitly identified file. Each has a result status, matching rationale and evidence jump; candidates never automatically resolve the failure. [Matching boundaries](docs/diagnostics.md#follow-up-evidence-candidates). - **Local review queue** — Record follow-up, expected-failure or alternative-verification notes in your browser. Evidence changes invalidate the old review; manual labels never rewrite automatic outcomes. No account or review backend. [Review workflow and storage limits](docs/diagnostics.md#local-review-workflow). - **Review portability** — Preview and download current-session review notes, then import only exact evidence matches into empty local slots. Existing notes are never overwritten; stale/unmatched records are skipped. JSON files are unencrypted and contain your written notes, not automatically copied logs. [Transfer limits](docs/diagnostics.md#transfer-reviews-between-browsers). - **Narrow-screen session workflow** — Below 768px, switch between the session list and full-width content without losing the current review draft; platform tabs scroll horizontally, and evidence jumps keep navigation visible. Desktop retains the two-column layout. [Scope and tested viewports](docs/diagnostics.md#narrow-screen-session-workflow). diff --git a/README.zh-CN.md b/README.zh-CN.md index 2597c67..cb11d14 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -55,6 +55,7 @@ LangSmith、Langfuse 这类观测平台面向的是*你自己写的* agent:接 ## 功能特性 - **有证据的失败事件(React UI)** — 将同一调用所在用户轮次、同工具、完整同参数的待复查失败分组,重复最多的操作优先展示;可跳转首末及每条原始证据。同参成功切断分组,缺少参数不合并。执行成功采用明确零退出码或 OMP 原生完成证据;事件不等于根因或任务失败。本地规则,无 LLM。[合成演示与判定边界](docs/diagnostics.md#中文使用指南)。 +- **后续相关操作** — 展示仅 `i` 参数不同的调用,以及同轮次、明确同文件的后续修改;标明关系依据、五类结果状态并可跳转证据。候选不自动关闭事件、不代表原问题已修复。[匹配边界](docs/diagnostics.md#后续相关操作)。 - **本机复核队列** — 用必填依据标记“需跟进”“预期失败”“其他验证已通过”,仅存当前浏览器;新证据使旧标记失效,人工判断不改写自动结果。无需账号或复核后端。[使用方式与存储边界](docs/diagnostics.md#本机复核闭环)。 - **复核迁移** — 预览并导出当前会话的有效复核,导入只接受完整证据匹配且本地为空的记录;不覆盖已有笔记,跳过过期或不匹配记录。明文 JSON 含手写依据,不自动复制日志。[迁移边界](docs/diagnostics.md#迁移复核记录)。 - **窄屏会话复核** — 小于 768px 时切换“会话列表 / 返回内容”,正文获得完整宽度,切换列表不清空当前复核草稿;平台栏横向滚动,证据跳转保留顶部导航,桌面继续双栏。[验收范围](docs/diagnostics.md#窄屏操作)。 diff --git a/claims.json b/claims.json index 17ffc6c..bf59505 100644 --- a/claims.json +++ b/claims.json @@ -102,16 +102,16 @@ }, { "id": "test-count", - "claim": "233 tests pass on Node's built-in test runner, the count docs/ROADMAP.md records for `npm test`.", - "value": "233", - "metric": "passing node:test cases (# tests 233 / # pass 233 / # fail 0)", - "method": "npm test → node --test test/*.test.js, run in the claims job after npm ci, and the TAP summary is asserted. The roadmap sentence ('233 tests on Node's built-in runner (`npm test`, 2026-09-23)') is verified by the run, not read back from the prose.", + "claim": "255 tests pass on Node's built-in test runner, the count docs/ROADMAP.md records for `npm test`.", + "value": "255", + "metric": "passing node:test cases (# tests 255 / # pass 255 / # fail 0)", + "method": "npm test → node --test test/*.test.js, run in the claims job after npm ci, and the TAP summary is asserted. The roadmap sentence ('255 tests on Node's built-in runner (`npm test`, 2026-09-23)') is verified by the run, not read back from the prose.", "repro": "npm test 2>&1 | grep -E '^# (tests|pass|fail)'", "evidence": "docs/ROADMAP.md", "as_of": "2026-09-13", "check": { "cmd": "npm test 2>&1 | grep -E '^# (tests|pass|fail)'", - "expect": { "contains": ["# tests 233", "# pass 233", "# fail 0"] }, + "expect": { "contains": ["# tests 255", "# pass 255", "# fail 0"] }, "timeout": 120 } }, @@ -238,7 +238,7 @@ "check": { "cmd": "node scripts/claims-receipts.mjs tests-node-only", "expect": { - "equals": "17 files in test/ · 14 distinct requires: 11 node builtins, 3 relative, 0 third-party" + "equals": "18 files in test/ · 14 distinct requires: 11 node builtins, 3 relative, 0 third-party" }, "timeout": 60 } diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md index 4c9894e..dc445a8 100644 --- a/docs/ROADMAP.md +++ b/docs/ROADMAP.md @@ -8,7 +8,7 @@ - **Session browser** with tool-call inspection, trace/waterfall view, spawn tracking and message timeline - **Prompt tooling** — extraction (noise filtered), template clustering with outcome attribution, Claude-powered rewrites, and a prompt library that installs entries as native slash commands - **Global search** across all platforms, insights dashboard, incremental session backup -- **React + Vite frontend** served by an Express backend; 233 tests on Node's built-in runner (`npm test`, 2026-09-23), CI on Node 22 +- **React + Vite frontend** served by an Express backend; 255 tests on Node's built-in runner (`npm test`, 2026-09-23), CI on Node 22 - **Evidence-backed failure events and local review** with full-result invalidation, evidence navigation and narrow-screen session layout ## Current priorities @@ -19,7 +19,7 @@ The product direction is **review the coding-agent sessions you already have, wi | --- | --- | --- | | P0 | Make the new workflow immediately testable | A demo-only entry opens a clearly synthetic case: 7 pending records in 2 events, all evidence accessible, local review does not rewrite automatic results. Preserve the existing default demo and samples. | | P1 | Current-session review portability implemented | Preview-only import and explicit plaintext download; exact identity/evidence matching, no overwrites, bounded schema and partial-failure reporting. Validate with synthetic migration and publish after CI; whole-history backup and path remapping remain out of scope. | -| P2 | Validate daily usefulness with the maintainer's own sessions | Record reviewed/follow-up/expected/alternative-verification counts and timed review tasks using a fixed rubric. Keep measurements local, separate unknowns and stale labels, and publish only consented aggregate evidence. Do not infer precision or time saved from event compression. | +| P2 | Explain what happened after a failure, then measure usefulness | Show only-i and same-turn/same-file follow-up candidates with explicit status and evidence, never automatic recovery. Validate unchanged diagnostic outcomes on frozen logs. Then measure owner review tasks using a fixed rubric; separate unknowns and stale labels, and do not infer precision or time saved from candidate counts. | | P3 | Make releases reproducible for contributors | Keep clean-install tests, generated fixtures, documentation claims and release/package verification aligned. Add browser regression automation when it can run deterministically without personal logs. | No launch dates or star-count targets are promised. Progress is gated on these observable outcomes. Physical-device/keyboard coverage and complex Trace/analytics layouts remain separate work, not implied by the session-screen checks. diff --git a/docs/diagnostics-verification.md b/docs/diagnostics-verification.md index 94ecc64..742133c 100644 --- a/docs/diagnostics-verification.md +++ b/docs/diagnostics-verification.md @@ -83,3 +83,15 @@ Verified using synthetic notes and isolated browser contexts only: - Browser migration sent zero backend write requests; automatic failure/recovery counts remained unchanged. At 360px, the transfer panel had no horizontal overflow. This does not verify real-device file pickers, simultaneous cross-tab transaction safety or actual productivity gains. The file is unencrypted human text, not signed evidence; only current exact matches can be restored. Detailed local run artifacts remain ignored and are not published. + +## Follow-up evidence acceptance + +The follow-up increment adds evidence-only relationships, not automatic recovery. There are 22 new tests in `test/follow-up-evidence.test.js`, 125 combined diagnostic/review tests, and 255 tests in the full suite. The ordinary and hosted-demo UI builds pass; earlier counts above are historical release receipts. + +Synthetic tests cover all five candidate states, only-top-level-`i` differences, complete argument comparison, exact tools, same-turn edit/write paths, explicit relative-path directories, missing/ambiguous fields, pre-existing parallel calls, missing results, call-turn vs result-turn, deduplication, full candidate output in review fingerprints, and unchanged exact-match recovery. The scripted demo `node scripts/demo-follow-up.cjs` exercises seven only-`i` candidates, one same-file candidate and excluded file/turn counterexamples. The browser checked all five labels, five-item initial display plus loading the rest, original result jumps, no page-navigation movement, stale review on newly appended candidate and no narrow-screen overflow. + +A fixed, previously inspected set of 15 real sessions (five each from OMP, Codex and Claude Code) with 2,024 tool results was re-read by frozen-length/hash verification. Automatic results stayed at **142 failures, 141 pending records, 76 events and 1 recovered record**, including unchanged event members. An independent relation enumeration agreed with every candidate. It found **3 only-`i` candidates in OMP: one failure, one running result and one success**; Codex/Claude had none under these strict rules. No same-file relation qualified in this set: relevant observed paths lacked explicit absolute context or the calls belonged to later user turns. The code did not weaken criteria to inflate matches. + +Storage keys stayed stable. Only the three events with added candidates acquired new review fingerprints; no-candidate events retained their previous fingerprint contract. No real review labels were persisted. Raw logs, parameters and detailed sample identifiers remain private; the regression is not an independent benchmark, global accuracy estimate or time-saving measurement. + +The hosted synthetic sample now includes one later successful edit of the same absolute file path with different arguments. It remains **2 events / 7 pending records / 1 recovered record**. Browser checks verified the relation caveat, source result, desktop/mobile rendering and zero static-demo API requests. diff --git a/docs/diagnostics.md b/docs/diagnostics.md index 52262fe..2c6bdb9 100644 --- a/docs/diagnostics.md +++ b/docs/diagnostics.md @@ -32,6 +32,37 @@ Run `npm start` from the built checkout. Select a session and open its **消息* Start with repeated operations, open the latest failure, and inspect the surrounding calls. A card is a reason to review the transcript—not an instruction to rerun a possibly destructive command. Equivalent commands, alternate verification and external fixes still require your judgment. +## Follow-up evidence candidates + +When a failure event has matching later operations, expand **后续相关操作**. These are leads for your review, **not automatic recovery or proof that the original task passed**. Each candidate shows its relationship, call/result message positions, call's user turn, tool outcome and source evidence, arguments, and a jump to the original result. Five are shown initially; the rest can be loaded without filtering out unsuccessful results. + +Two relationships are supported: + +| Relationship | Required evidence | Not inferred | +| --- | --- | --- | +| Only top-level `i` differs | Same exact tool, equal non-empty remaining arguments, but different full arguments. Object key order is ignored; values and array order are retained. May cross user turns within this transcript. | That `i` is semantically irrelevant, or that the later call fixes the earlier problem. | +| Same-file modification | Explicit `edit`, `Edit`, `write`, `Write` or `MultiEdit` calls, same call-origin user turn, changed parameters, and the same literal `path`/`file_path` plus recorded working-directory fields. | Paths hidden in shell commands/patches, different turns, symlinks, path normalization, or implicit directory changes. | + +Both require a recorded result and a call started **after the event's last pending failure**. Results from pre-existing parallel calls, orphan results, missing arguments and calls without results are excluded. Same-file matching requires absolute paths, or identical explicit absolute `cwd`/`workdir`/`working_directory` for relative paths; ambiguous fields are rejected. Different directory field names are not treated as aliases. + +Status labels distinguish **success (tool result), failure, running, cancelled/stopped and unknown**. Source fields are shown, not inferred from optimistic prose. A running background job is not success. Platform adapters and automatic failure/recovery matching remain unchanged, including preserving `i` in automatic retry comparison. A same-file edit that returned success may have changed something else entirely. + +New or changed candidate evidence requires human review again: the event's storage key remains stable, but the review fingerprint includes complete candidate results and arguments. Existing notes become stale only for affected events; events without candidates retain their old fingerprint contract. Review-transfer files still contain only hashes and handwritten notes, not raw candidate logs. + +### Try all five states locally + +From a built source checkout: + +```sh +node scripts/demo-follow-up.cjs +``` + +Select its OMP `[Synthetic]` session. The first bash failure has seven only-`i` candidates covering the five outcome states; an edit failure has one same-file candidate. An unrelated file and a different user turn are excluded. **Three automatic failure events remain**, even when a candidate is successful. Save a review and type `n` in the terminal to append another candidate: that review becomes stale, without changing automatic failure/recovery counts. Type `q` or Ctrl-C to stop the isolated demo. + +The hosted **Try diagnostics** sample also includes a successful same-file edit with changed arguments. Its existing seven pending failure records and two events do not disappear. Both demos are synthetic; no logged command is executed. + +![Synthetic follow-up evidence](../screenshots/follow-up-evidence.png) + ## Local review workflow The automatic result and your judgment are separate. The header always reports automatic events and pending failure records. The review filters show which of those events you have inspected: @@ -122,6 +153,7 @@ These checks cover navigation, session messages and the review workflow—not co node --test test/diagnostic-events.test.js test/diagnostics.test.js test/omp-outcome.test.js node --test test/diagnostic-reviews.test.js node --test test/review-transfer.test.js +node --test test/follow-up-evidence.test.js npm test npm run build:ui npm run lint @@ -181,6 +213,18 @@ The local frozen regression set contained 30 sessions and 9,076 tool results. Gr 本轮的 373 个真实冻结事件只使用**内存中的合成测试标记**检验隔离与失效,没有替你判断真实事件,也没有把这些测试标记写成真实复核。完整证据见 [本机复核验收](diagnostics-verification.md)。 +## 后续相关操作 + +失败事件下若有候选,可展开“后续相关操作”,查看匹配依据、五类状态、调用/结果位置、参数摘要及原始结果。**它不自动关闭事件,不代表任务已通过。** + +- “仅 i 不同”:同工具,只有顶层 `i` 不同,其余完整参数相同,可跨当前会话内的用户轮次;不等于认定 `i` 无语义影响。 +- “同文件修改”:仅限明确 edit/write/MultiEdit 工具、同调用轮次、相同记录路径及显式工作目录;允许不同修改参数,不推断它们是同一修复。 +- 都必须在事件最后一次失败之后发起且已有结果。此前已启动的并行调用、无结果、缺参数、孤立结果不参与。 +- 相对路径必须有相同的显式绝对工作目录;没有就不猜。不同字段名、隐式 cwd、软链、跨轮次修改、命令里的文件名都不自动关联。 +- 成功、失败、执行中、取消、未知均有标签和字段依据。候选新增或变化会使受影响旧复核过期;不会改变自动失败/恢复计数,也不会把原始日志加入迁移文件。 + +源码运行 `node scripts/demo-follow-up.cjs` 可体验五类状态、无关文件反例和实时追加候选。页面里保存复核后,终端输入 `n` 验证重审;`q` 退出。在线 Demo 的合成诊断案例也有一条同文件成功修改,但仍保留原来的失败事件。 + ## 窄屏操作 - 小于 768px 时,点击“会话列表”进入导航,选择当前或其他会话后返回正文;“返回内容”或 Escape 也可关闭列表。桌面仍是双栏。 diff --git a/docs/releases/v1.20.0.md b/docs/releases/v1.20.0.md new file mode 100644 index 0000000..34e2946 --- /dev/null +++ b/docs/releases/v1.20.0.md @@ -0,0 +1,34 @@ +# v1.20.0 — Follow the evidence after a failure + +Failure events now show **later related operations** so you can review what happened next without treating every successful-looking result as a fix. + +## New + +- **Only-`i` candidates:** the same exact tool with only top-level `i` differing and every other argument preserved. +- **Same-file candidates:** explicit edit/write tools targeting the same recorded file and directory context in the same call-origin user turn, with changed parameters. +- **Five result states:** success (tool result), failure, running, cancelled/stopped and unknown, with field-source evidence and a jump to the original result. No success-only filtering; load all candidates. +- **Review freshness:** adding or changing candidate evidence makes affected human reviews stale. Unaffected events retain their fingerprint contract; storage keys remain stable. +- **Try it:** the hosted diagnostics sample includes a same-file follow-up edit. A source checkout also provides `node scripts/demo-follow-up.cjs` for all five states and live candidate append. + +## Important boundaries + +Candidates **do not close failure events or prove task correctness**. They must have a recorded result and a call started after the event's last pending failure. Existing automatic failure/recovery matching—including `i`—is unchanged. + +Same-file matching does not guess implicit working directories, normalize paths, inspect patch/shell text or link across user turns. Relative paths need identical explicit absolute directory information. Only-`i` candidates may cross turns within the viewed transcript, but the tool/remaining arguments must match exactly. + +Old human notes on events that gain candidates require re-review. Notes on events without candidates are unaffected. Exported reviews still contain hashes and human notes, not raw candidate logs. No new model calls, dependencies, backend endpoints or uploads. + +## Validation + +- 255 Node tests passed, including 22 new candidate tests and 125 focused diagnostic/review tests. +- Normal/demo builds and lint passed with existing findings unchanged; generated fixture parity and claims checks pass. +- Frozen real-log regression preserved 142 failures, 141 pending records, 76 events and 1 recovery across 15 sessions / 2,024 tool results. Three strict candidates were found: one failed, one running, one successful. This is not an accuracy or productivity benchmark. +- Synthetic browser tests cover all states, full candidate access, original evidence jumps, stale reviews after append and narrow-screen layout. + +[Usage and limits](https://github.com/alloevil/AgentXRay/blob/master/docs/diagnostics.md#follow-up-evidence-candidates) · [Verification receipt](https://github.com/alloevil/AgentXRay/blob/master/docs/diagnostics-verification.md) + +```sh +npx @alloevil/agent-xray@1.20.0 +``` + +Requires Node 22.13+ (22.15+ for compressed DeepSeek Harness logs). diff --git a/frontend/demo/sample-logs/omp/-demo-diagnostics/2026-09-23T08-00-00-000Z_0199demo-diagnostics.jsonl b/frontend/demo/sample-logs/omp/-demo-diagnostics/2026-09-23T08-00-00-000Z_0199demo-diagnostics.jsonl index 94a1105..decadd5 100644 --- a/frontend/demo/sample-logs/omp/-demo-diagnostics/2026-09-23T08-00-00-000Z_0199demo-diagnostics.jsonl +++ b/frontend/demo/sample-logs/omp/-demo-diagnostics/2026-09-23T08-00-00-000Z_0199demo-diagnostics.jsonl @@ -22,3 +22,5 @@ {"type":"message","id":"demo-background-call","timestamp":"2026-09-23T08:01:10.000Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"demo-background","name":"bash","arguments":{"command":"npm run verify-config","cwd":"/demo/diagnostics","async":true}}]}} {"type":"message","id":"demo-background-result","timestamp":"2026-09-23T08:01:11.000Z","message":{"role":"toolResult","toolCallId":"demo-background","toolName":"bash","isError":false,"details":{"async":{"state":"running","jobId":"synthetic-background"}},"content":[{"type":"text","text":"Synthetic verification was backgrounded. There is no completion result in this sample."}]}} {"type":"message","id":"demo-review-summary","timestamp":"2026-09-23T08:01:12.000Z","message":{"role":"assistant","content":[{"type":"text","text":"Synthetic review exercise: inspect the six edit results and the search error, then record a follow-up note with evidence. The earlier test success does not validate the edits; a running background job is not a success. This static sample does not append live results."}]}} +{"type":"message","id":"demo-related-edit-call","timestamp":"2026-09-23T08:01:15.000Z","message":{"role":"assistant","content":[{"type":"toolCall","id":"demo-related-edit","name":"edit","arguments":{"path":"/demo/diagnostics/config.ts","oldText":"timeout: 15","newText":"timeout: 30"}}]}} +{"type":"message","id":"demo-related-edit-result","timestamp":"2026-09-23T08:01:16.000Z","message":{"role":"toolResult","toolCallId":"demo-related-edit","toolName":"edit","isError":false,"content":[{"type":"text","text":"Synthetic edit completed with a different oldText argument. The file path matches the failed edits, but this is only a related operation, not proof that the original task passed."}]}} diff --git a/frontend/src/demo/fixtures.json b/frontend/src/demo/fixtures.json index a13e65e..ccba545 100644 --- a/frontend/src/demo/fixtures.json +++ b/frontend/src/demo/fixtures.json @@ -95,16 +95,16 @@ { "id": "0199demo-diagnostics", "timestamp": "2026-09-23T08:00:00.000Z", - "lastActivity": "2026-09-23T08:01:12.000Z", - "messageCount": 22, + "lastActivity": "2026-09-23T08:01:16.000Z", + "messageCount": 24, "userCount": 1, - "assistantCount": 11, - "toolCallCount": 10, - "toolResultCount": 10, + "assistantCount": 12, + "toolCallCount": 11, + "toolResultCount": 11, "topTools": [ { "name": "edit", - "count": 6 + "count": 7 }, { "name": "bash", @@ -1659,6 +1659,62 @@ "toolName": null, "details": null, "isError": false + }, + { + "id": "demo-related-edit-call", + "timestamp": "2026-09-23T08:01:15.000Z", + "role": "assistant", + "content": [], + "usage": null, + "model": null, + "provider": null, + "toolCallId": null, + "toolName": null, + "details": null, + "isError": false + }, + { + "id": "demo-related-edit", + "timestamp": "2026-09-23T08:01:15.000Z", + "role": "toolCall", + "content": [], + "usage": null, + "model": null, + "provider": null, + "toolCallId": "demo-related-edit", + "toolName": "edit", + "details": { + "path": "/demo/diagnostics/config.ts", + "oldText": "timeout: 15", + "newText": "timeout: 30" + }, + "isError": false + }, + { + "id": "demo-related-edit-result", + "timestamp": "2026-09-23T08:01:16.000Z", + "role": "toolResult", + "content": [ + { + "type": "text", + "text": "Synthetic edit completed with a different oldText argument. The file path matches the failed edits, but this is only a related operation, not proof that the original task passed." + } + ], + "usage": null, + "model": null, + "provider": null, + "toolCallId": "demo-related-edit", + "toolName": "edit", + "details": null, + "isError": false, + "ompOutcome": { + "state": "success", + "evidence": [ + "isError=false", + "content (tool result, not task acceptance)" + ], + "warnings": [] + } } ] }, @@ -2299,9 +2355,9 @@ }, "omp": { "totalSessions": 2, - "totalMessages": 32, - "totalToolCalls": 15, - "errorRate": 0.4667, + "totalMessages": 34, + "totalToolCalls": 16, + "errorRate": 0.4375, "totalCost": 0, "tokenUsage": { "input": 13000, @@ -2311,9 +2367,9 @@ "toolStats": [ { "name": "edit", - "calls": 8, + "calls": 9, "errors": 6, - "errorRate": 0.75, + "errorRate": 0.6666666666666666, "avgDurationMs": 0 }, { @@ -2450,7 +2506,7 @@ "date": "2026-09-23", "sessions": 1, "errors": 7, - "toolCalls": 10, + "toolCalls": 11, "cost": 0 } ] @@ -2571,7 +2627,7 @@ "id": "0199demo-diagnostics", "file": "2026-09-23T08-00-00-000Z_0199demo-diagnostics.jsonl", "timestamp": "2026-09-23T08:00:00.000Z", - "lastActivity": "2026-09-23T08:01:12.000Z", + "lastActivity": "2026-09-23T08:01:16.000Z", "slug": null, "title": "[Synthetic] Failure review: 7 records → 2 events", "promptCount": 1, diff --git a/frontend/src/views/sessions/RelatedOperations.tsx b/frontend/src/views/sessions/RelatedOperations.tsx new file mode 100644 index 0000000..bce5fb5 --- /dev/null +++ b/frontend/src/views/sessions/RelatedOperations.tsx @@ -0,0 +1,60 @@ +import { useState } from 'react'; +import type { RelatedOperation } from './diagnostics'; +import { formatDate, messageAnchorId } from './lib'; + +const STATES = { + success: '成功(工具结果)', failure: '失败', running: '执行中', cancelled: '取消 / 停止', unknown: '未知', +}; +const COLORS = { + success: 'text-emerald-400', failure: 'text-destructive', running: 'text-amber-500', + cancelled: 'text-muted-foreground', unknown: 'text-muted-foreground', +}; + +export function RelatedOperations({ operations, onJump }: { + operations: RelatedOperation[]; + onJump: (id: string) => void; +}) { + const [visible, setVisible] = useState(5); + return ( +
+ 后续相关操作 · {operations.length} 条候选(非恢复结论) +

+ 这些调用在本事件最后一次失败之后发起;仅按记录参数关联,未执行新的验证。成功只描述该工具结果,不证明原问题已修复。 +

+
    + {operations.slice(0, visible).map((operation) => { + const anchor = messageAnchorId(operation.message); + return ( +
  1. +
    + {operation.toolName} + {STATES[operation.state]} +
    +

    + {operation.relation === 'only-i' + ? '关联依据:同工具,只有顶层 i 参数不同,其余参数完全一致。' + : '关联依据:同一调用轮次、相同记录文件路径及显式目录信息;修改参数不同,不代表同一修复。'} +

    +

    + {operation.userTurn ? `第 ${operation.userTurn} 个用户轮次` : '首个用户消息之前'} · {formatDate(operation.message.timestamp)} + {' · '}调用消息 #{operation.callIndex + 1} → 结果消息 #{operation.index + 1} +

    +

    + {operation.argumentsText.slice(0, 240)}{operation.argumentsText.length > 240 ? '…' : ''} +

    +
    {operation.evidence}
    + {anchor ? ( + + ) :

    日志缺少定位标识

    } +
  2. + ); + })} +
+ {visible < operations.length ? ( + + ) : null} +
+ ); +} diff --git a/frontend/src/views/sessions/SessionDiagnostics.tsx b/frontend/src/views/sessions/SessionDiagnostics.tsx index 2abb461..42f043a 100644 --- a/frontend/src/views/sessions/SessionDiagnostics.tsx +++ b/frontend/src/views/sessions/SessionDiagnostics.tsx @@ -4,6 +4,7 @@ import { diagnoseSession, type FailureDiagnostic, type FailureEvent } from './di import { formatDate, messageAnchorId } from './lib'; import { DiagnosticReview, useEventReviews } from './DiagnosticReview'; import { ReviewTransferPanel } from './ReviewTransferPanel'; +import { RelatedOperations } from './RelatedOperations'; import { REVIEW_LABELS, reviewState, type EventReview, type ReviewState, type ReviewStatus } from './diagnostic-reviews'; const REASONS = { @@ -96,6 +97,7 @@ function EventCard({ event, onJump, review, ready, onReview }: { ))} ) : null} + {event.relatedOperations.length ? : null} @@ -132,6 +134,8 @@ export function SessionDiagnostics({ messages, onScrollToMessage, reviewScope }: 错误内容可能不同,也可能包含并行调用;分组不代表相同根因或串行重试。时间跨度不是执行耗时或浪费时间。 恢复仍只匹配失败之后发起的同参成功调用,不限定用户轮次。命令执行须有明确零退出码,OMP 也采用原生完成字段; 执行中、取消和未知状态不算成功,其他工具须有未报错结果。不推断等价命令、隐式工作目录或后台任务关联。 + 后续候选只按“仅 i 不同”或“同轮次同文件的编辑/写入”关联;相对路径缺少明确绝对工作目录不匹配, + 无结果或在最后一次失败之前发起的调用不列入。候选不关闭事件,新增或变化的候选证据会使旧人工复核过期。

diff --git a/frontend/src/views/sessions/diagnostic-reviews.ts b/frontend/src/views/sessions/diagnostic-reviews.ts index 96c3fd8..153cc81 100644 --- a/frontend/src/views/sessions/diagnostic-reviews.ts +++ b/frontend/src/views/sessions/diagnostic-reviews.ts @@ -48,11 +48,19 @@ export async function createReviewIdentity(scope: string, event: FailureEvent): } const first = event.failures[0]; const identity = [scope, event.toolName, args, event.userTurn, first.index, first.message.id, first.message.toolCallId]; + const evidence: unknown[] = [1, identity, event.failures.map((failure) => ({ + index: failure.index, reason: failure.reason, message: failure.message, + }))]; + if (event.relatedOperations?.length) { + evidence.push(event.relatedOperations.map((operation) => ({ + index: operation.index, callIndex: operation.callIndex, userTurn: operation.userTurn, + relation: operation.relation, state: operation.state, toolName: operation.toolName, + argumentsText: operation.argumentsText, message: operation.message, + }))); + } const [key, fingerprint] = await Promise.all([ digest(identity), - digest([1, identity, event.failures.map((failure) => ({ - index: failure.index, reason: failure.reason, message: failure.message, - }))]), + digest(evidence), ]); return { storageKey: `${REVIEW_PREFIX}${key}`, fingerprint }; } diff --git a/frontend/src/views/sessions/diagnostics.ts b/frontend/src/views/sessions/diagnostics.ts index 64969bc..ed7b626 100644 --- a/frontend/src/views/sessions/diagnostics.ts +++ b/frontend/src/views/sessions/diagnostics.ts @@ -8,6 +8,8 @@ interface RecordedCall { key: string | null; index: number; userTurn: number; + onlyIKey: string | null; + fileKey: string | null; } export interface FailureDiagnostic { @@ -26,6 +28,25 @@ export interface FailureEvent { userTurn: number; failures: FailureDiagnostic[]; spanMs: number | null; + relatedOperations: RelatedOperation[]; +} + +export interface RelatedOperation { + message: SessionMessage; + index: number; + callIndex: number; + userTurn: number; + toolName: string; + argumentsText: string; + relation: 'only-i' | 'same-file'; + state: 'success' | 'failure' | 'running' | 'cancelled' | 'unknown'; + evidence: string; +} + +interface RecordedResult { + message: SessionMessage; + index: number; + call: RecordedCall; } function eventSpan(failures: FailureDiagnostic[]): number | null { @@ -93,6 +114,55 @@ function outcome(result: SessionMessage, call: RecordedCall | undefined): Outcom return result.isError === false && result.content !== null ? 'success' : 'unknown'; } +function argumentObject(value: unknown): Record | null { + return value !== null && typeof value === 'object' && !Array.isArray(value) ? value as Record : null; +} + +function onlyIKey(name: string, args: unknown): string | null { + const object = argumentObject(args); + if (!object) return null; + const entries = Object.entries(object).filter(([key]) => key !== 'i'); + return entries.length ? JSON.stringify([name, ordered(Object.fromEntries(entries))]) : null; +} + +function fileKey(name: string, args: unknown): string | null { + if (!['edit', 'Edit', 'write', 'Write', 'MultiEdit'].includes(name)) return null; + const object = argumentObject(args); + if (!object) return null; + const paths = ['path', 'file_path'].filter((key) => key in object).map((key) => object[key]); + if (!paths.length || paths.some((value) => typeof value !== 'string' || !value || value !== paths[0])) return null; + const filename = paths[0] as string; + const directories = ['cwd', 'workdir', 'working_directory'].filter((key) => key in object).map((key) => [key, object[key]]); + if (directories.some(([, value]) => typeof value !== 'string' || !value)) return null; + if (new Set(directories.map(([, value]) => value)).size > 1) return null; + const absolute = (value: string) => /^(?:\/|[A-Za-z]:[\\/]|\\\\)/.test(value); + if (!absolute(filename) && (!directories.length || !absolute(directories[0][1] as string))) return null; + return JSON.stringify([filename, directories]); +} + +function relatedState(message: SessionMessage, call: RecordedCall): RelatedOperation['state'] { + if (message.ompOutcome) return message.ompOutcome.state; + const code = exitCode(message, call); + if (message.isError === true || (code !== null && code !== 0)) return 'failure'; + if (['running', 'pending', 'in_progress'].includes(String(message.details?.status))) return 'running'; + if (['cancelled', 'canceled'].includes(String(message.details?.status))) return 'cancelled'; + if (message.details?.status != null && !['ok', 'success', 'complete', 'completed'].includes(String(message.details.status))) return 'unknown'; + return outcome(message, call); +} + +function relatedEvidence(message: SessionMessage, call: RecordedCall, state: RelatedOperation['state']): string { + const code = exitCode(message, call); + const sources = message.ompOutcome?.evidence || [ + ...(message.isError ? ['日志标记 isError=true'] : []), + ...(code !== null ? [`退出码 ${code}`] : []), + ...(state === 'running' ? [`details.status=${message.details?.status}`] : []), + ...(state === 'cancelled' ? [`details.status=${message.details?.status}`] : []), + ...(code === null && state === 'success' ? ['isError=false;工具结果未报错(不等于任务通过)'] : []), + ...(state === 'unknown' ? ['缺少明确完成状态'] : []), + ]; + return [...sources, resultText(message)].join('\n').slice(0, 500); +} + export function diagnoseSession(messages: SessionMessage[]) { const calls = new Map(); const unresolved = new Map(); @@ -100,13 +170,17 @@ export function diagnoseSession(messages: SessionMessage[]) { const recovered = new Set(); const eventKeys = new Map(); const generations = new Map(); + const failureCalls = new Map(); + const resultsByI = new Map(); + const resultsByFile = new Map(); let userTurn = 0; function recordCall(id: string | null | undefined, name: string | null | undefined, value: unknown, index: number) { if (!id) return; const args = parseArguments(value); const key = name && args !== null && args !== undefined ? JSON.stringify([name, ordered(args)]) : null; - calls.set(id, { name: name || '未知工具', args, key, index, userTurn }); + calls.set(id, { name: name || '未知工具', args, key, index, userTurn, + onlyIKey: name ? onlyIKey(name, args) : null, fileKey: name ? fileKey(name, args) : null }); if (key) { for (const failure of unresolved.get(key) || []) { if (index > failure.index) failure.reason = 'unconfirmed-retry'; @@ -126,6 +200,15 @@ export function diagnoseSession(messages: SessionMessage[]) { } if (message.role !== 'toolResult') return; const call = message.toolCallId ? calls.get(message.toolCallId) : undefined; + if (call?.key) { + const recorded = { message, index, call }; + for (const [key, bucket] of [[call.onlyIKey, resultsByI], [call.fileKey, resultsByFile]] as const) { + if (!key) continue; + const list = bucket.get(key) || []; + list.push(recorded); + bucket.set(key, list); + } + } const result = outcome(message, call); if (result === 'failure') { const code = exitCode(message, call); @@ -143,6 +226,7 @@ export function diagnoseSession(messages: SessionMessage[]) { reason: call?.key ? 'no-success' : 'missing-call', }; failures.push(failure); + if (call?.key) failureCalls.set(failure, call); eventKeys.set(failure, { key: call?.key ? JSON.stringify([call.userTurn, call.key, generations.get(call.key) || 0]) : `orphan-${index}`, userTurn: call?.userTurn ?? userTurn, @@ -177,11 +261,33 @@ export function diagnoseSession(messages: SessionMessage[]) { userTurn: group.userTurn, failures: [failure], spanMs: null, + relatedOperations: [], }); } } const events = [...grouped.values()]; - for (const event of events) event.spanMs = eventSpan(event.failures); + for (const event of events) { + event.spanMs = eventSpan(event.failures); + const latest = event.failures[event.failures.length - 1]; + const original = failureCalls.get(latest); + if (!original) continue; + const candidates = new Map(); + const buckets = [ + ['only-i', original.onlyIKey ? resultsByI.get(original.onlyIKey) : undefined], + ['same-file', original.fileKey ? resultsByFile.get(original.fileKey) : undefined], + ] as const; + for (const [relation, results] of buckets) { + for (const { message, index, call: subsequent } of results || []) { + if (subsequent.index <= latest.index || subsequent.key === original.key || candidates.has(index)) continue; + if (relation === 'same-file' && subsequent.userTurn !== original.userTurn) continue; + const state = relatedState(message, subsequent); + candidates.set(index, { message, index, callIndex: subsequent.index, userTurn: subsequent.userTurn, + toolName: subsequent.name, argumentsText: JSON.stringify(subsequent.args), relation, state, + evidence: relatedEvidence(message, subsequent, state) }); + } + } + event.relatedOperations = [...candidates.values()].sort((left, right) => left.index - right.index); + } events.sort((left, right) => right.failures.length - left.failures.length || left.failures[0].index - right.failures[0].index); return { diff --git a/intent.md b/intent.md index c91854b..affec28 100644 --- a/intent.md +++ b/intent.md @@ -58,6 +58,20 @@ Hosted walkthrough acceptance: 216 tests pass, including raw-to-bundled fixture Portability acceptance: 17 new transfer tests and 103 focused tests pass. Browser tests exercised an actual 2-record download/import into an isolated browser, a different-origin restore, stale and cross-tab preview invalidation, no-overwrite conflicts, schema/size rejection, literal HTML notes and partial write failure/retry. Prepare v1.19.0 and verify the actual registry package and Pages deployment after protected checks, without mutating the prior v1.18.0 release. +## Follow-up evidence candidates (approved 2026-09-24) + +- Add two explainable candidate relations, not semantic recovery: same tool with only top-level `i` differing; or a later explicit edit/write of the same recorded file in the same call-origin user turn. +- Candidates must have a recorded result and a call started after the event's last pending failure. Exclude pre-existing parallel calls, same-signature repeats, orphan/missing arguments and unrecognized modification tools. Record order, not wall-clock guessing, determines "later". +- Same-file tools are `edit`, `Edit`, `write`, `Write`, `MultiEdit`; support `path`/`file_path` only when unambiguous. Compare paths literally; require absolute paths or identical explicit absolute working directories for relative paths. Preserve cwd/workdir/working_directory distinctions. No filesystem resolution, symlink guessing, command parsing, cross-turn same-file links or cross-session/child joins. +- Only-i candidates may cross user turns within the currently viewed transcript; label their call turn. Compare every other argument value, not just a text preview. Do not strip `i` from the original automatic recovery rules. +- Each candidate shows the exact relation, call/result positions, five-state result status with its source fields, an output excerpt and a jump to the original result. Running/unknown/cancelled never display as successful verification. Show five initially, with all remaining candidates reachable and no silent success-only filtering. +- Failure/event/recovery counts and event membership must remain byte-for-byte equivalent on the frozen 15-session/2,024-result audit. Candidate success does not close an event, prove causality or certify the task. +- New/changed candidate evidence invalidates human review fingerprints for affected events; storage keys remain stable. Events with no candidates retain the previous fingerprint input to avoid unrelated invalidation. Transfer still contains only hashes and human notes, not raw candidate evidence. +- Validate synthetic counterexamples and the frozen real-log prefixes; preserve raw logs privately, output only sample numbers, relation/status counts and line numbers. Real data previously inspected is regression evidence, not a blind accuracy or time-saving benchmark. +- Update the isolated walkthrough and documentation. Do not introduce dependencies, model calls, new APIs, automatic commands, arbitrary error parsing, or change platform adapters. + +Acceptance results: 255 tests and 125 focused checks pass. The frozen 15-session/2,024-result regression retains every original failure/event member and all counts. Strict matching finds three only-i candidates (failure/running/success); no same-file candidate qualifies in that real set because paths lack explicit directory context or turns differ. Synthetic same-file and five-state browser checks, load-more, result jumps and candidate-triggered review staleness pass. Prepare v1.20.0 via the authorized protected PR/release workflow, verify registry installation and public Pages before claiming release completion; do not publish private logs or raw evaluation artifacts. + ## Non-goals - No new platform, dependency, model call, account, telemetry, cloud log storage or automatic command execution. diff --git a/package-lock.json b/package-lock.json index fb2b321..14cd0b8 100644 --- a/package-lock.json +++ b/package-lock.json @@ -1,12 +1,12 @@ { "name": "@alloevil/agent-xray", - "version": "1.19.0", + "version": "1.20.0", "lockfileVersion": 3, "requires": true, "packages": { "": { "name": "@alloevil/agent-xray", - "version": "1.19.0", + "version": "1.20.0", "license": "MIT", "dependencies": { "express": "^4.21.2" diff --git a/package.json b/package.json index beeb021..5c3c83e 100644 --- a/package.json +++ b/package.json @@ -1,6 +1,6 @@ { "name": "@alloevil/agent-xray", - "version": "1.19.0", + "version": "1.20.0", "description": "Web dashboard for viewing AI agent session logs — supports OpenClaw, Codex, Claude Code, Hermes, OMP, DeepSeek Harness, and Gemini CLI", "main": "server.js", "bin": { diff --git a/screenshots/follow-up-evidence.png b/screenshots/follow-up-evidence.png new file mode 100644 index 0000000..943d861 Binary files /dev/null and b/screenshots/follow-up-evidence.png differ diff --git a/scripts/demo-follow-up.cjs b/scripts/demo-follow-up.cjs new file mode 100644 index 0000000..74641a6 --- /dev/null +++ b/scripts/demo-follow-up.cjs @@ -0,0 +1,97 @@ +const fs = require('node:fs/promises'); +const path = require('node:path'); +const readline = require('node:readline'); +const { startServer } = require('../test/helpers'); + +async function main() { + const server = await startServer(); + const directory = path.join(server.home, '.omp/agent/sessions/synthetic-follow-up'); + const id = '01990000-0000-7000-8000-000000000177'; + const file = path.join(directory, `2026-09-24T08-00-00_${id}.jsonl`); + let tick = 0; + const stamp = () => new Date(Date.UTC(2026, 8, 24, 8) + tick++ * 1000).toISOString(); + const message = (recordId, payload) => ({ type: 'message', id: recordId, timestamp: stamp(), message: payload }); + const operation = (callId, name, args, state) => [ + message(`${callId}-call`, { + role: 'assistant', + content: [{ type: 'toolCall', id: callId, name, arguments: args }], + }), + message(`${callId}-result`, { + role: 'toolResult', + toolCallId: callId, + toolName: name, + isError: state === 'failure', + details: + state === 'running' || state === 'cancelled' + ? { async: { state } } + : state === 'unknown' + ? {} + : { exitCode: state === 'success' ? 0 : 1, wallTimeMs: 10 }, + content: [{ type: 'text', text: `Synthetic ${state} result for ${callId}. Nothing was executed.` }], + }), + ]; + const bashArgs = { command: 'synthetic test', i: 'initial' }; + const editArgs = { path: '/synthetic/config.ts', oldText: 'before', newText: 'after' }; + const records = [ + { type: 'session', id, timestamp: stamp(), cwd: '/synthetic/project' }, + message('synthetic-user', { + role: 'user', + content: [{ type: 'text', text: '[Synthetic] 后续相关操作:候选不等于恢复' }], + }), + ...operation('failed-bash', 'bash', bashArgs, 'failure'), + ...operation('failed-edit', 'edit', editArgs, 'failure'), + ...operation('same-file', 'edit', { ...editArgs, oldText: 'different anchor' }, 'success'), + ...operation('unrelated-file', 'edit', { ...editArgs, path: '/synthetic/other.ts' }, 'success'), + ]; + const states = ['running', 'cancelled', 'unknown', 'success', 'success', 'success', 'failure']; + states.forEach((state, index) => + records.push(...operation(`candidate-${index + 1}`, 'bash', { ...bashArgs, i: `next-${index}` }, state)) + ); + records.push( + message('next-user', { + role: 'user', + content: [{ type: 'text', text: 'Synthetic next turn: same-file changes should not be linked across turns.' }], + }) + ); + records.push(...operation('other-turn', 'edit', { ...editArgs, newText: 'next turn' }, 'success')); + try { + await fs.mkdir(directory, { recursive: true }); + await fs.writeFile(file, `${records.map((record) => JSON.stringify(record)).join('\n')}\n`); + } catch (error) { + await server.stop(); + throw error; + } + console.log(`Synthetic-only follow-up demo: ${server.base}`); + console.log(`Session API: ${server.base}/api/omp/sessions/${id}`); + console.log('OMP 合成会话:3 个自动事件;初始 bash 事件有 7 条不同状态候选,edit 事件有 1 条同文件候选。'); + console.log('输入 n:追加仅 i 不同的成功候选,使旧复核过期;q 或 Ctrl-C 清理退出。'); + const input = readline.createInterface({ input: process.stdin }); + let appended = 0; + const stop = async () => { + input.close(); + await server.stop(); + process.exit(0); + }; + input.on('line', async (line) => { + if (line.trim() === 'q') return stop(); + if (line.trim() !== 'n') return; + appended++; + try { + const next = operation(`appended-${appended}`, 'bash', { ...bashArgs, i: `new-evidence-${appended}` }, 'success'); + await fs.appendFile(file, `${next.map((record) => JSON.stringify(record)).join('\n')}\n`); + console.log( + 'PASS: appended candidate evidence; automatic failures unchanged, affected review must become stale.' + ); + } catch (error) { + console.error(error.message); + await stop(); + } + }); + process.on('SIGINT', stop); + process.on('SIGTERM', stop); +} + +main().catch((error) => { + console.error(error.message); + process.exitCode = 1; +}); diff --git a/test/diagnostic-events.test.js b/test/diagnostic-events.test.js index aaab659..da4a9e1 100644 --- a/test/diagnostic-events.test.js +++ b/test/diagnostic-events.test.js @@ -242,6 +242,9 @@ test('hosted synthetic walkthrough yields two events and preserves seven failure ); assert.equal(report.events[0].spanMs, 50000); assert.equal(report.events[1].toolName, 'web_search'); + assert.equal(report.events[0].relatedOperations.length, 1); + assert.equal(report.events[0].relatedOperations[0].relation, 'same-file'); + assert.equal(report.events[0].relatedOperations[0].state, 'success'); assert.equal(report.events[1].failures[0].message.isError, false); assert.equal(detail.messages.find((message) => message.id === 'demo-background-result').ompOutcome.state, 'running'); coverage(report); diff --git a/test/follow-up-evidence.test.js b/test/follow-up-evidence.test.js new file mode 100644 index 0000000..e26778d --- /dev/null +++ b/test/follow-up-evidence.test.js @@ -0,0 +1,317 @@ +const { test, before } = require('node:test'); +const assert = require('node:assert/strict'); +const { readFileSync } = require('node:fs'); +const path = require('node:path'); +const { stripTypeScriptTypes } = require('node:module'); + +let diagnoseSession; +let reviews; +before(async () => { + const load = async (file) => { + const source = readFileSync(path.join(__dirname, '../frontend/src/views/sessions', file), 'utf8'); + return import(`data:text/javascript;base64,${Buffer.from(stripTypeScriptTypes(source)).toString('base64')}`); + }; + ({ diagnoseSession } = await load('diagnostics.ts')); + reviews = await load('diagnostic-reviews.ts'); +}); + +const call = (id, args, toolName = 'bash') => ({ role: 'toolCall', id, toolCallId: id, toolName, details: args }); +const result = (id, state = 'success', options = {}) => ({ + role: 'toolResult', + id: `${id}-result`, + toolCallId: id, + timestamp: null, + isError: state === 'failure', + details: null, + content: [{ type: 'text', text: `Synthetic ${id} result` }], + ompOutcome: { state, evidence: [`synthetic.state=${state}`], warnings: [] }, + ...options, +}); +const user = { role: 'user', content: [{ type: 'text', text: 'Synthetic next task' }] }; +const failed = (args = { command: 'synthetic test', i: 'first' }, tool = 'bash') => [ + call('first', args, tool), + result('first', 'failure'), +]; +const candidates = (messages) => diagnoseSession(messages).events[0].relatedOperations; + +test('only-i follow-up success is evidence, not automatic recovery', () => { + const messages = [...failed(), call('retry', { i: 'second', command: 'synthetic test' }), result('retry')]; + const report = diagnoseSession(messages); + assert.equal(report.failureCount, 1); + assert.equal(report.recoveredCount, 0); + assert.equal(report.failures.length, 1); + assert.equal(report.events[0].relatedOperations.length, 1); + const entry = report.events[0].relatedOperations[0]; + assert.equal(entry.relation, 'only-i'); + assert.equal(entry.state, 'success'); + assert.equal(entry.message, messages[3]); + assert.equal(entry.callIndex, 2); + assert.equal(entry.index, 3); + assert.match(entry.evidence, /synthetic.state=success/); +}); + +for (const state of ['failure', 'running', 'cancelled', 'unknown']) { + test(`only-i candidate preserves ${state} instead of claiming verification`, () => { + const rows = candidates([ + ...failed(), + call('later', { command: 'synthetic test', i: 'second' }), + result('later', state), + ]); + assert.equal(rows[0].state, state); + assert.equal(rows[0].relation, 'only-i'); + }); +} + +test('only-i matching compares full arguments and exact tools, including nested values', () => { + const original = { command: 'synthetic test', config: { FIRST: 1, SECOND: 2 }, i: 'first' }; + const rows = candidates([ + ...failed(original), + call('changed', { ...original, config: { FIRST: 1, SECOND: 3 }, i: 'next' }), + result('changed'), + call('other-tool', { ...original, i: 'next' }, 'shell'), + result('other-tool'), + call('match', JSON.stringify({ i: 'next', config: { SECOND: 2, FIRST: 1 }, command: 'synthetic test' })), + result('match'), + ]); + assert.deepEqual( + rows.map((row) => row.message.toolCallId), + ['match'] + ); +}); + +test('only-i without meaningful remaining parameters or equal arguments is not a relation', () => { + assert.equal(candidates([...failed({ i: 'first' }), call('next', { i: 'next' }), result('next')]).length, 0); + assert.equal( + candidates([...failed(), call('same', { command: 'synthetic test', i: 'first' }), result('same', 'running')]) + .length, + 0 + ); + assert.equal( + candidates([ + ...failed({ command: ['a', 'b'], i: 'first' }), + call('different', { command: ['b', 'a'], i: 'next' }), + result('different'), + ]).length, + 0 + ); +}); + +test('adding or removing top-level i is supported and can cross call user turns', () => { + const rows = candidates([ + ...failed({ command: 'synthetic test' }), + user, + call('next', { command: 'synthetic test', i: 'next' }), + result('next'), + ]); + assert.equal(rows.length, 1); + assert.equal(rows[0].userTurn, 1); + assert.equal(rows[0].relation, 'only-i'); +}); + +test('same-file modifications need the same turn and explicitly resolved path context', () => { + const args = { path: '/synthetic/file.ts', oldText: 'first', newText: 'second' }; + const rows = candidates([ + ...failed(args, 'edit'), + call('write', { file_path: '/synthetic/file.ts', content: 'different content' }, 'Write'), + result('write'), + call('unrelated', { path: '/synthetic/other.ts', content: 'different content' }, 'write'), + result('unrelated'), + call('read', { path: '/synthetic/file.ts' }, 'read'), + result('read'), + user, + call('next-turn', { ...args, newText: 'third' }, 'edit'), + result('next-turn'), + ]); + assert.deepEqual( + rows.map((row) => row.message.toolCallId), + ['write'] + ); + assert.equal(rows[0].relation, 'same-file'); +}); + +test('relative paths require equal explicit absolute cwd and aliases are not guessed', () => { + for (const [first, second, expected] of [ + [{ path: 'file.ts' }, { path: 'file.ts', content: 'next' }, 0], + [{ path: 'file.ts', cwd: '/synthetic' }, { path: 'file.ts', cwd: '/synthetic', content: 'next' }, 1], + [{ path: 'file.ts', cwd: '/one' }, { path: 'file.ts', cwd: '/two', content: 'next' }, 0], + [{ path: 'file.ts', cwd: '/one' }, { path: 'file.ts', workdir: '/one', content: 'next' }, 0], + [{ path: '/file.ts', cwd: '/one' }, { path: '/file.ts', cwd: '/two', content: 'next' }, 0], + [{ path: '/a/../file.ts' }, { path: '/file.ts', content: 'next' }, 0], + [{ path: 'file.ts', cwd: 'relative' }, { path: 'file.ts', cwd: 'relative', content: 'next' }, 0], + [{ path: '/one', file_path: '/two' }, { path: '/one', content: 'next' }, 0], + ]) { + assert.equal(candidates([...failed(first, 'edit'), call('next', second, 'edit'), result('next')]).length, expected); + } +}); + +test('same-file relation does not infer paths from shell commands or patch text', () => { + assert.equal( + candidates([ + ...failed({ command: 'edit /synthetic/file.ts' }), + call('next', { path: '/synthetic/file.ts', content: 'next' }, 'write'), + result('next'), + ]).length, + 0 + ); +}); + +test('earlier-started parallel calls and missing results do not become follow-up evidence', () => { + const args = { command: 'synthetic test', i: 'next' }; + const rows = candidates([ + call('first', { command: 'synthetic test', i: 'first' }), + call('parallel', args), + result('first', 'failure'), + result('parallel'), + call('pending', args), + result('orphan'), + ]); + assert.equal(rows.length, 0); +}); + +test('candidates start after the last failure, not merely the first event member', () => { + const args = { path: '/synthetic/file.ts', content: 'first' }; + const report = diagnoseSession([ + ...failed(args, 'write'), + call('middle', { ...args, content: 'middle' }, 'write'), + result('middle'), + call('repeat', args, 'write'), + result('repeat', 'failure'), + call('last', { ...args, content: 'last' }, 'write'), + result('last'), + ]); + assert.equal(report.events[0].failures.length, 2); + assert.deepEqual( + report.events[0].relatedOperations.map((row) => row.message.toolCallId), + ['last'] + ); +}); + +test('call user turn, not delayed result turn, controls same-file relation', () => { + const args = { path: '/synthetic/file.ts', content: 'first' }; + const rows = candidates([ + ...failed(args, 'write'), + call('later', { ...args, content: 'next' }, 'write'), + user, + result('later'), + ]); + assert.equal(rows[0].relation, 'same-file'); + assert.equal(rows[0].userTurn, 0); +}); + +test('a result matching both relations is listed once with the more specific only-i reason', () => { + const args = { path: '/synthetic/file.ts', content: 'first', i: 'first' }; + const rows = candidates([...failed(args, 'write'), call('later', { ...args, i: 'next' }, 'write'), result('later')]); + assert.equal(rows.length, 1); + assert.equal(rows[0].relation, 'only-i'); +}); + +test('five-state evidence follows explicit source fields for non-OMP logs too', () => { + const examples = [ + [{ details: { exitCode: 0 } }, 'success'], + [{ details: { exitCode: 2 } }, 'failure'], + [{ details: { status: 'running', exitCode: 0 } }, 'running'], + [{ details: null }, 'unknown'], + [ + { details: null, content: [{ type: 'text', text: 'Wall time: 1 seconds\nExit code: 0\nOutput:\nSynthetic' }] }, + 'success', + ], + ]; + for (const [overrides, expected] of examples) { + const rows = candidates([ + ...failed(), + call('later', { command: 'synthetic test', i: 'next' }), + result('later', 'success', { ompOutcome: undefined, isError: false, ...overrides }), + ]); + assert.equal(rows[0].state, expected); + assert.ok(rows[0].evidence.length); + } +}); + +test('all candidates stay ordered and raw input is not mutated', () => { + const messages = failed(); + for (let index = 0; index < 12; index++) + messages.push(call(`next-${index}`, { command: 'synthetic test', i: index }), result(`next-${index}`)); + const before = JSON.stringify(messages); + const report = diagnoseSession(messages); + assert.equal(report.events[0].relatedOperations.length, 12); + assert.equal(report.recoveredCount, 0); + assert.deepEqual( + report.events[0].relatedOperations.map((row) => row.index), + Array.from({ length: 12 }, (_, index) => index * 2 + 3) + ); + assert.equal(JSON.stringify(messages), before); +}); + +test('new candidates invalidate old human notes but preserve storage keys', async () => { + const messages = failed(); + const before = diagnoseSession(messages).events[0]; + const after = diagnoseSession([...messages, call('next', { command: 'synthetic test', i: 'next' }), result('next')]) + .events[0]; + const original = await reviews.createReviewIdentity('scope', before); + const changed = await reviews.createReviewIdentity('scope', after); + assert.equal(original.storageKey, changed.storageKey); + assert.notEqual(original.fingerprint, changed.fingerprint); + const stored = new Map(); + const storage = { getItem: (key) => stored.get(key) ?? null, setItem: (key, value) => stored.set(key, value) }; + reviews.saveReview(storage, original, 'expected', 'Synthetic old review'); + assert.equal(reviews.reviewState(reviews.readReview(storage, changed)), 'unreviewed'); +}); + +test('candidate output beyond displayed preview participates in review fingerprint', async () => { + const messages = [ + ...failed(), + call('next', { command: 'synthetic test', i: 'next' }), + result('next', 'success', { content: [{ type: 'text', text: 'x'.repeat(700) }] }), + ]; + const before = await reviews.createReviewIdentity('scope', diagnoseSession(messages).events[0]); + messages[3].content[0].text += 'changed'; + const after = await reviews.createReviewIdentity('scope', diagnoseSession(messages).events[0]); + assert.notEqual(before.fingerprint, after.fingerprint); +}); + +test('no candidates preserves prior fingerprint contract and exact recovery remains unchanged', async () => { + const event = diagnoseSession(failed()).events[0]; + const legacy = { ...event }; + delete legacy.relatedOperations; + assert.deepEqual( + await reviews.createReviewIdentity('scope', event), + await reviews.createReviewIdentity('scope', legacy) + ); + const report = diagnoseSession([ + ...failed(), + call('exact', { command: 'synthetic test', i: 'first' }), + result('exact'), + ]); + assert.equal(report.recoveredCount, 1); + assert.equal(report.events.length, 0); +}); + +test('explicit cancellation and unrecognized status do not turn into a successful candidate', () => { + for (const [status, expected] of [ + ['cancelled', 'cancelled'], + ['canceled', 'cancelled'], + ['future-state', 'unknown'], + ]) { + const rows = candidates([ + ...failed({ path: '/synthetic/read.txt', i: 'first' }, 'Read'), + call('next', { path: '/synthetic/read.txt', i: 'next' }, 'Read'), + result('next', 'success', { ompOutcome: undefined, details: { status }, isError: false }), + ]); + assert.equal(rows[0].state, expected); + } +}); + +test('embedded calls link to their own result while missing arguments are not inferred', () => { + const embedded = (id, input) => ({ role: 'assistant', content: [{ type: 'toolCall', id, name: 'Edit', input }] }); + const args = { file_path: '/synthetic/file.ts', old_string: 'a', new_string: 'b' }; + const rows = candidates([ + embedded('first', args), + result('first', 'failure'), + embedded('later', { ...args, new_string: 'c' }), + result('later'), + ]); + assert.equal(rows.length, 1); + assert.equal(rows[0].message.toolCallId, 'later'); + assert.equal(rows[0].relation, 'same-file'); + assert.equal(candidates([...failed(null, 'Edit'), embedded('later', args), result('later')]).length, 0); +});