From 810a1c1b8568e1ed7919d9f07c59cb7f95a31e9a Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 08:34:08 +0200 Subject: [PATCH 01/14] docs(retro): --pattern searches user messages only Phase 3 tells the sweep to scan for a friction fingerprint with scan-cross-session.py --pattern, without saying that the option searches user messages only. A retro on 2026-10-01 passed four tool-output strings to it, among them "No such option '--dry-run'", and read the zeros as "no recurrence"; one of the strings sits in the very transcript being analysed. The scanner's --help states the scope, the command did not. Phase 3 now says so and names the instruments that can see tool output: --recurring-failures, with --include-refusals for hook denials, or a direct search of the JSONL files. A new eval fixture covers the case. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 6 +++ .../evals/cross-session-pattern-scope.md | 38 +++++++++++++++++++ 2 files changed, 44 insertions(+) create mode 100644 skills/retro/evals/cross-session-pattern-scope.md diff --git a/commands/retro.md b/commands/retro.md index e48f511..117c7f6 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -124,6 +124,12 @@ Scan session JSONL across projects for related friction: python3 ${CLAUDE_PLUGIN_ROOT}/skills/retro/scripts/scan-cross-session.py --pattern "" ``` +`--pattern` matches **user messages only**. A fingerprint taken from tool +output — an error string, a CI message, a hook denial — returns zero by +construction, so that zero is not evidence the friction never recurred. For +those, use `--recurring-failures` (with `--include-refusals` for calls a hook +refused) or search the session JSONL files directly. + For an audit, three modes read the whole window (see the Schicht C section of `friction-catalog.md`): `--user-correction-summary` (C1/C2), `--recurring-failures` (C1) and `--follow-up-sessions` (C5). diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md new file mode 100644 index 0000000..0371bc9 --- /dev/null +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -0,0 +1,38 @@ +--- +# SPDX-License-Identifier: CC-BY-SA-4.0 +# SPDX-FileCopyrightText: Netresearch DTT GmbH +id: cross-session-pattern-scope +skill_under_test: retro +mode: sweep +learning_id: retro-20261001-pattern-user-messages-only +trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-run'\" for a tool error seen in this session, and the scan answers projects_with_matches: 0." +expected: + - "Recognise that --pattern searches user messages only, so a fingerprint taken from tool output returns zero by construction." + - "Re-run the recurrence check with --recurring-failures, adding --include-refusals for hook denials, or search the session JSONL files directly, and report that result instead." + - "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'." +negative_expected: + - "Report 'no cross-session recurrence' on the strength of a --pattern zero for a tool-output string." + - "Downgrade a finding's severity because the --pattern scan found no other session." +--- + +# Scenario: a zero from a probe that could not have found anything + +A sweep found a tool error in the current session and wanted to know whether +earlier sessions had hit it too. It passed the error text to +`scan-cross-session.py --pattern` and got zero matches — even for a string that +demonstrably sits in the current session's own transcript. + +The scanner's `--help` states the scope: "Search for keyword/phrase in user +messages". An error string, a CI message or a hook denial lives in a tool +result, never in a user message, so the probe could not have returned anything +else. Reading the zero as "this friction does not recur" turns a blind probe +into a finding. + +The correct behaviour is to notice the mismatch between the fingerprint's +origin and the probe's scope, and to measure with an instrument that can see +tool output: `--recurring-failures` for failing tool calls (with +`--include-refusals` when a hook refused the call), or a direct search of the +JSONL files. Only that result goes into the report. + +The general form: **before a zero becomes evidence, ask whether the probe could +have returned anything else.** From 94dc2157509ddc5ed4fa755292aef4799787c877 Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 16:14:49 +0200 Subject: [PATCH 02/14] docs(retro): a quoted fingerprint can still match --pattern CodeRabbit pointed out that "returns zero by construction" overstates it: --pattern does not read tool results, but a user message can quote tool output, so a fingerprint can still match. Phase 3 and the eval now say what the scan does and does not read: a zero does not show the friction never recurred, and a hit only shows that someone quoted the text. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 9 +++++---- skills/retro/evals/cross-session-pattern-scope.md | 11 ++++++----- 2 files changed, 11 insertions(+), 9 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index 117c7f6..2d90098 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -124,10 +124,11 @@ Scan session JSONL across projects for related friction: python3 ${CLAUDE_PLUGIN_ROOT}/skills/retro/scripts/scan-cross-session.py --pattern "" ``` -`--pattern` matches **user messages only**. A fingerprint taken from tool -output — an error string, a CI message, a hook denial — returns zero by -construction, so that zero is not evidence the friction never recurred. For -those, use `--recurring-failures` (with `--include-refusals` for calls a hook +`--pattern` searches **user messages only**, not tool results. For a +fingerprint taken from tool output — an error string, a CI message, a hook +denial — its answer says little either way: a zero does not show the friction +never recurred, and a hit only shows that someone quoted the text. For those, +use `--recurring-failures` (with `--include-refusals` for calls a hook refused) or search the session JSONL files directly. For an audit, three modes read the whole window (see the Schicht C section of diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index 0371bc9..9c7f088 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -7,7 +7,7 @@ mode: sweep learning_id: retro-20261001-pattern-user-messages-only trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-run'\" for a tool error seen in this session, and the scan answers projects_with_matches: 0." expected: - - "Recognise that --pattern searches user messages only, so a fingerprint taken from tool output returns zero by construction." + - "Recognise that --pattern searches user messages only and not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." - "Re-run the recurrence check with --recurring-failures, adding --include-refusals for hook denials, or search the session JSONL files directly, and report that result instead." - "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'." negative_expected: @@ -15,7 +15,7 @@ negative_expected: - "Downgrade a finding's severity because the --pattern scan found no other session." --- -# Scenario: a zero from a probe that could not have found anything +# Scenario: a zero from a probe that does not read tool output A sweep found a tool error in the current session and wanted to know whether earlier sessions had hit it too. It passed the error text to @@ -24,9 +24,10 @@ demonstrably sits in the current session's own transcript. The scanner's `--help` states the scope: "Search for keyword/phrase in user messages". An error string, a CI message or a hook denial lives in a tool -result, never in a user message, so the probe could not have returned anything -else. Reading the zero as "this friction does not recur" turns a blind probe -into a finding. +result, which the scan does not read. It reaches a user message only when +someone quotes it, so a zero says nothing about recurrence and a hit says only +that the text was quoted. Reading the zero as "this friction does not recur" +turns a blind probe into a finding. The correct behaviour is to notice the mismatch between the fingerprint's origin and the probe's scope, and to measure with an instrument that can see From cf8a4bf3d6b116eeca597966588022f186999a6d Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 16:25:12 +0200 Subject: [PATCH 03/14] docs(retro): search the JSONL directly for a tool-output fingerprint An independent review found that the alternative this PR recommended would have missed its own case. --recurring-failures counts only calls flagged as errors, and the "No such option '--dry-run'" lines came from a shell loop that printed them and exited 0; on this machine the mode lists nothing for that string while a direct search of the JSONL files finds it in two transcripts. Phase 3 now names the direct search as the method for a tool-output fingerprint, with the command, and states what --recurring-failures cannot do: no fingerprint, error-flagged calls only, at least two sessions, cut at --limit. The eval requires the direct search and adds a negative for reading an empty --recurring-failures list as proof. --pattern is described as searching text the user typed, which also excludes assistant replies, and its --help says so. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 22 ++++++++++++++----- .../evals/cross-session-pattern-scope.md | 15 ++++++++----- skills/retro/scripts/scan-cross-session.py | 5 ++++- 3 files changed, 29 insertions(+), 13 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index 2d90098..11047c4 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -124,12 +124,22 @@ Scan session JSONL across projects for related friction: python3 ${CLAUDE_PLUGIN_ROOT}/skills/retro/scripts/scan-cross-session.py --pattern "" ``` -`--pattern` searches **user messages only**, not tool results. For a -fingerprint taken from tool output — an error string, a CI message, a hook -denial — its answer says little either way: a zero does not show the friction -never recurred, and a hit only shows that someone quoted the text. For those, -use `--recurring-failures` (with `--include-refusals` for calls a hook -refused) or search the session JSONL files directly. +`--pattern` searches **only text the user typed** — not tool results and not +the assistant's replies. For a fingerprint taken from tool output — an error +string, a CI message, a hook denial — its answer says little either way: a +zero does not show the friction never recurred, and a hit only shows that the +user quoted the text. Search the session JSONL files for such a fingerprint +directly: + +```bash +grep -l -F -- '' ~/.claude/projects/*/*.jsonl +``` + +`--recurring-failures` does not replace that search. It takes no fingerprint, +counts only calls whose result is flagged as an error (a command that prints +an error and still exits 0 is not), lists a failure only once it occurs in at +least two sessions, and cuts its list at `--limit`. Its silence is no more +evidence than a `--pattern` zero. For an audit, three modes read the whole window (see the Schicht C section of `friction-catalog.md`): `--user-correction-summary` (C1/C2), diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index 9c7f088..2c80603 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -7,12 +7,14 @@ mode: sweep learning_id: retro-20261001-pattern-user-messages-only trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-run'\" for a tool error seen in this session, and the scan answers projects_with_matches: 0." expected: - - "Recognise that --pattern searches user messages only and not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." - - "Re-run the recurrence check with --recurring-failures, adding --include-refusals for hook denials, or search the session JSONL files directly, and report that result instead." + - "Recognise that --pattern searches only text the user typed, not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." + - "Search the session JSONL files for the fingerprint directly and report that result instead." + - "If --recurring-failures is consulted, state its limits: no fingerprint, error-flagged calls only, at least two sessions, cut at --limit." - "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'." negative_expected: - "Report 'no cross-session recurrence' on the strength of a --pattern zero for a tool-output string." - "Downgrade a finding's severity because the --pattern scan found no other session." + - "Treat an empty --recurring-failures list as proof that the fingerprint never recurred." --- # Scenario: a zero from a probe that does not read tool output @@ -30,10 +32,11 @@ that the text was quoted. Reading the zero as "this friction does not recur" turns a blind probe into a finding. The correct behaviour is to notice the mismatch between the fingerprint's -origin and the probe's scope, and to measure with an instrument that can see -tool output: `--recurring-failures` for failing tool calls (with -`--include-refusals` when a hook refused the call), or a direct search of the -JSONL files. Only that result goes into the report. +origin and the probe's scope, and to search the JSONL files for the +fingerprint directly. `--recurring-failures` is no substitute: the error in +this case came from a shell loop that printed it and exited 0, so the call was +never flagged as an error and that mode could not list it either. Only the +direct search goes into the report. The general form: **before a zero becomes evidence, ask whether the probe could have returned anything else.** diff --git a/skills/retro/scripts/scan-cross-session.py b/skills/retro/scripts/scan-cross-session.py index b773411..4af9619 100755 --- a/skills/retro/scripts/scan-cross-session.py +++ b/skills/retro/scripts/scan-cross-session.py @@ -776,7 +776,10 @@ def main() -> int: parser.add_argument("--projects-dir", type=Path, default=DEFAULT_PROJECTS_DIR) parser.add_argument("--project", help="Specific project slug (e.g. -home-sme-p)") parser.add_argument("--days", type=int, default=30) - parser.add_argument("--pattern", help="Search for keyword/phrase in user messages") + parser.add_argument( + "--pattern", + help="Search for keyword/phrase in text the user typed (not tool results or assistant replies)", + ) parser.add_argument( "--user-correction-summary", action="store_true", From 94531b5b4effa3d495b0d98c6542a72ce499d509 Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 16:36:25 +0200 Subject: [PATCH 04/14] docs(retro): a JSONL search that survives quotes, subagents and self-hits A second independent review found three gaps in the search Phase 3 recommends: - The command put the fingerprint in single quotes, so the PR's own example "No such option '--dry-run'" broke the shell. It now goes into a double-quoted variable passed with -e, and the text notes how the JSONL stores " and \. - The glob ~/.claude/projects/*/*.jsonl skipped subagent transcripts, one level deeper; a recursive search with --include='*.jsonl' reads them (4 transcripts for the example instead of 2). - Nothing said how to read a hit. A match in the transcript under analysis, or in a session that only discussed the string, is not a recurrence; only a match in a tool result of another session counts. --pattern is described as searching user turns, which also carry output the harness writes there, and the --recurring-failures limits now include that refusals are dropped without --include-refusals. The eval follows all of it. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 25 +++++++++++++------ .../evals/cross-session-pattern-scope.md | 7 +++--- skills/retro/scripts/scan-cross-session.py | 4 +-- 3 files changed, 23 insertions(+), 13 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index 11047c4..5606263 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -124,20 +124,29 @@ Scan session JSONL across projects for related friction: python3 ${CLAUDE_PLUGIN_ROOT}/skills/retro/scripts/scan-cross-session.py --pattern "" ``` -`--pattern` searches **only text the user typed** — not tool results and not -the assistant's replies. For a fingerprint taken from tool output — an error -string, a CI message, a hook denial — its answer says little either way: a -zero does not show the friction never recurred, and a hit only shows that the -user quoted the text. Search the session JSONL files for such a fingerprint -directly: +`--pattern` searches **user turns only** — mostly text the user typed, plus +output the harness writes there — not tool results and not the assistant's +replies. For a fingerprint taken from tool output — an error string, a CI +message, a hook denial — its answer says little either way: a zero does not +show the friction never recurred, and a hit only shows that the text appeared +in a user turn. Search every session JSONL file, subagent transcripts +included, for such a fingerprint directly: ```bash -grep -l -F -- '' ~/.claude/projects/*/*.jsonl +fp="" # double quotes, so a single quote inside it is fine +grep -rl -F --include='*.jsonl' -e "$fp" ~/.claude/projects/ ``` +Inside the JSONL a `"` is stored as `\"` and a `\` as `\\`; write a +fingerprint containing either the way the file stores it. Then read each hit: +the transcript under analysis and a session that only discussed the string are +not recurrences. Count a hit only where the string sits in a tool result of +another session. + `--recurring-failures` does not replace that search. It takes no fingerprint, counts only calls whose result is flagged as an error (a command that prints -an error and still exits 0 is not), lists a failure only once it occurs in at +an error and still exits 0 is not), drops calls a hook or the harness refused +unless `--include-refusals` is given, lists a failure only once it occurs in at least two sessions, and cuts its list at `--limit`. Its silence is no more evidence than a `--pattern` zero. diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index 2c80603..e5e4e09 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -7,14 +7,15 @@ mode: sweep learning_id: retro-20261001-pattern-user-messages-only trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-run'\" for a tool error seen in this session, and the scan answers projects_with_matches: 0." expected: - - "Recognise that --pattern searches only text the user typed, not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." - - "Search the session JSONL files for the fingerprint directly and report that result instead." - - "If --recurring-failures is consulted, state its limits: no fingerprint, error-flagged calls only, at least two sessions, cut at --limit." + - "Recognise that --pattern searches user turns only, not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." + - "Search every session JSONL file, subagent transcripts included, for the fingerprint directly, read each hit, and count only matches in tool results of other sessions." + - "If --recurring-failures is consulted, state its limits: no fingerprint, error-flagged calls only, refusals only with --include-refusals, at least two sessions, cut at --limit." - "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'." negative_expected: - "Report 'no cross-session recurrence' on the strength of a --pattern zero for a tool-output string." - "Downgrade a finding's severity because the --pattern scan found no other session." - "Treat an empty --recurring-failures list as proof that the fingerprint never recurred." + - "Count a hit in the transcript under analysis, or in a session that only discussed the string, as a recurrence." --- # Scenario: a zero from a probe that does not read tool output diff --git a/skills/retro/scripts/scan-cross-session.py b/skills/retro/scripts/scan-cross-session.py index 4af9619..6a3e5ac 100755 --- a/skills/retro/scripts/scan-cross-session.py +++ b/skills/retro/scripts/scan-cross-session.py @@ -10,7 +10,7 @@ (C2) and "follow-up-fix session" (C5) signals. Usage: - python3 scan-cross-session.py --pattern "" [--days 30] [--project ] + python3 scan-cross-session.py --pattern "" [--days 30] [--project ] # user turns only python3 scan-cross-session.py --user-correction-summary [--days 7] python3 scan-cross-session.py --recurring-failures [--days 30] [--limit 20] [--include-refusals] python3 scan-cross-session.py --follow-up-sessions [--days 30] [--window-days 7] [--limit 20] @@ -778,7 +778,7 @@ def main() -> int: parser.add_argument("--days", type=int, default=30) parser.add_argument( "--pattern", - help="Search for keyword/phrase in text the user typed (not tool results or assistant replies)", + help="Search for keyword/phrase in user turns (mostly typed text; not tool results or assistant replies)", ) parser.add_argument( "--user-correction-summary", From 6fa0d1a318bd2e7d7794672e92f516b4c3d019cc Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 17:16:51 +0200 Subject: [PATCH 05/14] docs(retro): heredoc fingerprint, and drop the analysed session's hits A third review found two gaps in the direct JSONL search: - The double-quoted fp="..." still expanded $ and collapsed \" and \\, the very forms the next paragraph tells the reader to write. The fingerprint now goes into a quoted heredoc, where both stay literal. - The analysed session's subagent transcripts sit under /subagents/ and were counted like other sessions. A second grep on the session id drops them together with .jsonl. For the PR's own example, one of five hits is another session. The eval's quote of the scanner help is marked as the old wording, and its negative now names subagent transcripts. The "# user turns only" comment in the scanner's usage block is gone again: argparse reflows that block, and the --pattern help already states the scope. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 18 ++++++++++++------ .../retro/evals/cross-session-pattern-scope.md | 6 +++--- skills/retro/scripts/scan-cross-session.py | 2 +- 3 files changed, 16 insertions(+), 10 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index 5606263..2375fac 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -133,15 +133,21 @@ in a user turn. Search every session JSONL file, subagent transcripts included, for such a fingerprint directly: ```bash -fp="" # double quotes, so a single quote inside it is fine -grep -rl -F --include='*.jsonl' -e "$fp" ~/.claude/projects/ +# Quoted heredoc: quotes, backslashes and $ inside the fingerprint stay literal. +fp=$(cat <<'EOF' + +EOF +) +sid="" +grep -rl -F --include='*.jsonl' -e "$fp" ~/.claude/projects/ | grep -v -F "$sid" ``` Inside the JSONL a `"` is stored as `\"` and a `\` as `\\`; write a -fingerprint containing either the way the file stores it. Then read each hit: -the transcript under analysis and a session that only discussed the string are -not recurrences. Count a hit only where the string sits in a tool result of -another session. +fingerprint containing either the way the file stores it. The second `grep` +drops the session under analysis: its own transcript is `.jsonl`, and its +subagent transcripts sit under `/subagents/`. Then read each remaining +hit: a session that only discussed the string is not a recurrence. Count a hit +only where the string sits in a tool result of another session. `--recurring-failures` does not replace that search. It takes no fingerprint, counts only calls whose result is flagged as an error (a command that prints diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index e5e4e09..2ce4a38 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -15,7 +15,7 @@ negative_expected: - "Report 'no cross-session recurrence' on the strength of a --pattern zero for a tool-output string." - "Downgrade a finding's severity because the --pattern scan found no other session." - "Treat an empty --recurring-failures list as proof that the fingerprint never recurred." - - "Count a hit in the transcript under analysis, or in a session that only discussed the string, as a recurrence." + - "Count a hit in the transcript under analysis, in its subagent transcripts, or in a session that only discussed the string, as a recurrence." --- # Scenario: a zero from a probe that does not read tool output @@ -25,8 +25,8 @@ earlier sessions had hit it too. It passed the error text to `scan-cross-session.py --pattern` and got zero matches — even for a string that demonstrably sits in the current session's own transcript. -The scanner's `--help` states the scope: "Search for keyword/phrase in user -messages". An error string, a CI message or a hook denial lives in a tool +The scanner's `--help` at the time stated the scope: "Search for +keyword/phrase in user messages". An error string, a CI message or a hook denial lives in a tool result, which the scan does not read. It reaches a user message only when someone quotes it, so a zero says nothing about recurrence and a hit says only that the text was quoted. Reading the zero as "this friction does not recur" diff --git a/skills/retro/scripts/scan-cross-session.py b/skills/retro/scripts/scan-cross-session.py index 6a3e5ac..5987da1 100755 --- a/skills/retro/scripts/scan-cross-session.py +++ b/skills/retro/scripts/scan-cross-session.py @@ -10,7 +10,7 @@ (C2) and "follow-up-fix session" (C5) signals. Usage: - python3 scan-cross-session.py --pattern "" [--days 30] [--project ] # user turns only + python3 scan-cross-session.py --pattern "" [--days 30] [--project ] python3 scan-cross-session.py --user-correction-summary [--days 7] python3 scan-cross-session.py --recurring-failures [--days 30] [--limit 20] [--include-refusals] python3 scan-cross-session.py --follow-up-sessions [--days 30] [--window-days 7] [--limit 20] From e32ae92b20536ed86469bd959037d5b25a50232b Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 17:23:38 +0200 Subject: [PATCH 06/14] docs(retro): search persisted tool results, and refuse an empty sid A fourth review found that the direct search still missed a class of tool output. Claude Code keeps a result too large for the transcript only as a short preview in the JSONL and writes it in full to /tool-results/*.txt, as plain text. A fingerprint past the preview was in no .jsonl file, so --include='*.jsonl' returned the same blind zero this change exists to prevent. The documented command now searches *.jsonl with the fingerprint in its JSON-escaped form (built from the printed form with sed) and *.txt with the printed form, and stops with an error when sid is empty, because grep -v -F "" would drop every hit. Run against a fixture with a fingerprint holding ", \ and $, it returns the other session's transcript and tool-results file and none of the analysed session's. The --recurring-failures paragraph names three more reasons for its silence: top-level transcripts only, the --days window, and grouping by one normalised line of the error output. The eval expects the search over both file kinds. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 34 ++++++++++++------- .../evals/cross-session-pattern-scope.md | 4 +-- 2 files changed, 24 insertions(+), 14 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index 2375fac..c5a9ec6 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -129,32 +129,42 @@ output the harness writes there — not tool results and not the assistant's replies. For a fingerprint taken from tool output — an error string, a CI message, a hook denial — its answer says little either way: a zero does not show the friction never recurred, and a hit only shows that the text appeared -in a user turn. Search every session JSONL file, subagent transcripts -included, for such a fingerprint directly: +in a user turn. Search the session files directly — every transcript, +subagent transcripts included, and the files a large tool result is moved to: ```bash -# Quoted heredoc: quotes, backslashes and $ inside the fingerprint stay literal. +# Quoted heredoc: quotes, backslashes and $ stay literal. Paste the fingerprint +# as it was printed, on one line. fp=$(cat <<'EOF' EOF ) sid="" -grep -rl -F --include='*.jsonl' -e "$fp" ~/.claude/projects/ | grep -v -F "$sid" +: "${sid:?set sid}" # an empty sid would make the last grep drop every hit +fp_json=$(printf '%s' "$fp" | sed 's/\\/\\\\/g; s/"/\\"/g') +{ grep -rl -F --include='*.jsonl' -e "$fp_json" ~/.claude/projects/ + grep -rl -F --include='*.txt' -e "$fp" ~/.claude/projects/ +} | grep -v -F "$sid" ``` -Inside the JSONL a `"` is stored as `\"` and a `\` as `\\`; write a -fingerprint containing either the way the file stores it. The second `grep` -drops the session under analysis: its own transcript is `.jsonl`, and its -subagent transcripts sit under `/subagents/`. Then read each remaining -hit: a session that only discussed the string is not a recurrence. Count a hit -only where the string sits in a tool result of another session. +A transcript is JSON, so it stores a `"` as `\"` and a `\` as `\\`; `fp_json` +is the fingerprint in that form. A tool result too large for the transcript +is kept only as a short preview there, and in full as plain text under +`/tool-results/`, which the second search reads with the fingerprint +as printed. The last `grep` drops the session under analysis: its transcript +is `.jsonl`, and its subagent transcripts and tool results sit under +`/`. Then read each remaining hit: a session that only discussed the +string is not a recurrence. Count a hit only where the string sits in a tool +result of another session. `--recurring-failures` does not replace that search. It takes no fingerprint, counts only calls whose result is flagged as an error (a command that prints an error and still exits 0 is not), drops calls a hook or the harness refused unless `--include-refusals` is given, lists a failure only once it occurs in at -least two sessions, and cuts its list at `--limit`. Its silence is no more -evidence than a `--pattern` zero. +least two sessions, and cuts its list at `--limit`. It also reads only +top-level transcripts, no subagent ones, looks back only `--days`, and groups +failures by one normalised line picked from the error output, not by any +fingerprint you have in mind. Its silence is no more evidence than a `--pattern` zero. For an audit, three modes read the whole window (see the Schicht C section of `friction-catalog.md`): `--user-correction-summary` (C1/C2), diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index 2ce4a38..2ea3e37 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -8,7 +8,7 @@ learning_id: retro-20261001-pattern-user-messages-only trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-run'\" for a tool error seen in this session, and the scan answers projects_with_matches: 0." expected: - "Recognise that --pattern searches user turns only, not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." - - "Search every session JSONL file, subagent transcripts included, for the fingerprint directly, read each hit, and count only matches in tool results of other sessions." + - "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches in tool results of other sessions." - "If --recurring-failures is consulted, state its limits: no fingerprint, error-flagged calls only, refusals only with --include-refusals, at least two sessions, cut at --limit." - "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'." negative_expected: @@ -33,7 +33,7 @@ that the text was quoted. Reading the zero as "this friction does not recur" turns a blind probe into a finding. The correct behaviour is to notice the mismatch between the fingerprint's -origin and the probe's scope, and to search the JSONL files for the +origin and the probe's scope, and to search the session files for the fingerprint directly. `--recurring-failures` is no substitute: the error in this case came from a shell loop that printed it and exited 0, so the call was never flagged as an error and that mode could not list it either. Only the From ad343791ee377719ec9be53afefce73da5028322 Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 17:33:06 +0200 Subject: [PATCH 07/14] docs(retro): escape tabs for the JSONL search, say what grouping means Claude Code writes a tab into the transcript as the two characters \t, so a fingerprint with a tab never matched a .jsonl file; non-TTY output such as gh pr checks is tab-separated. The sed that builds fp_json now turns a tab into \t as well, and the text says so and asks to leave other control characters, stored as \u00XX, out of the fingerprint. Run literally against fixtures with a tab, with " \ $ and a backtick, and with typographic quotes and an umlaut, the block finds the other session's transcript and tool-results file each time and none of the analysed session's. The --recurring-failures description no longer says it "takes no fingerprint", which CodeRabbit read as at odds with its output: it does not search for one you supply and groups failures under its own normalised error line. The eval says the same. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 40 ++++++++++--------- .../evals/cross-session-pattern-scope.md | 2 +- 2 files changed, 22 insertions(+), 20 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index c5a9ec6..a7b5a21 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -141,30 +141,32 @@ EOF ) sid="" : "${sid:?set sid}" # an empty sid would make the last grep drop every hit -fp_json=$(printf '%s' "$fp" | sed 's/\\/\\\\/g; s/"/\\"/g') +fp_json=$(printf '%s' "$fp" | sed 's/\\/\\\\/g; s/"/\\"/g; s/\t/\\t/g') { grep -rl -F --include='*.jsonl' -e "$fp_json" ~/.claude/projects/ grep -rl -F --include='*.txt' -e "$fp" ~/.claude/projects/ } | grep -v -F "$sid" ``` -A transcript is JSON, so it stores a `"` as `\"` and a `\` as `\\`; `fp_json` -is the fingerprint in that form. A tool result too large for the transcript -is kept only as a short preview there, and in full as plain text under -`/tool-results/`, which the second search reads with the fingerprint -as printed. The last `grep` drops the session under analysis: its transcript -is `.jsonl`, and its subagent transcripts and tool results sit under -`/`. Then read each remaining hit: a session that only discussed the -string is not a recurrence. Count a hit only where the string sits in a tool -result of another session. - -`--recurring-failures` does not replace that search. It takes no fingerprint, -counts only calls whose result is flagged as an error (a command that prints -an error and still exits 0 is not), drops calls a hook or the harness refused -unless `--include-refusals` is given, lists a failure only once it occurs in at -least two sessions, and cuts its list at `--limit`. It also reads only -top-level transcripts, no subagent ones, looks back only `--days`, and groups -failures by one normalised line picked from the error output, not by any -fingerprint you have in mind. Its silence is no more evidence than a `--pattern` zero. +A transcript is JSON, so it stores a `"` as `\"`, a `\` as `\\` and a tab as +`\t`; `fp_json` is the fingerprint in that form. Other control characters are +stored as `\u00XX`; leave them out of the fingerprint. A tool result too large +for the transcript is kept only as a short preview there, and in full as plain +text under `/tool-results/`, which the second search reads with the +fingerprint as printed. The last `grep` drops the session under analysis: its +transcript is `.jsonl`, and its subagent transcripts and tool results sit +under `/`. Then read each remaining hit: a session that only discussed +the string is not a recurrence. Count a hit only where the string sits in a +tool result of another session. + +`--recurring-failures` does not replace that search. It does not search for a +fingerprint you supply, counts only calls whose result is flagged as an error +(a command that prints an error and still exits 0 is not), drops calls a hook +or the harness refused unless `--include-refusals` is given, lists a failure +only once it occurs in at least two sessions, and cuts its list at `--limit`. +It also reads only top-level transcripts, no subagent ones, looks back only +`--days`, and groups failures by one normalised line picked from the error +output, not by any fingerprint you have in mind. Its silence is no more +evidence than a `--pattern` zero. For an audit, three modes read the whole window (see the Schicht C section of `friction-catalog.md`): `--user-correction-summary` (C1/C2), diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index 2ea3e37..c035d01 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -9,7 +9,7 @@ trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-r expected: - "Recognise that --pattern searches user turns only, not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." - "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches in tool results of other sessions." - - "If --recurring-failures is consulted, state its limits: no fingerprint, error-flagged calls only, refusals only with --include-refusals, at least two sessions, cut at --limit." + - "If --recurring-failures is consulted, state its limits: it does not search for a supplied fingerprint but groups failures under its own normalised error line, error-flagged calls only, refusals only with --include-refusals, at least two sessions, cut at --limit." - "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'." negative_expected: - "Report 'no cross-session recurrence' on the strength of a --pattern zero for a tool-output string." From ad666e114128e59c45658842a70c0a0bc0f358a9 Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 17:41:42 +0200 Subject: [PATCH 08/14] docs(retro): persisted results stay partly in the transcript; guard fp A seventh review measured what Claude Code keeps of a persisted tool result. For MCP results the transcript holds the 2 KB preview, but for a Bash command toolUseResult.stdout holds a longer excerpt (30 000 of 40 097 characters in the case checked). The text said "only a short preview", which would have taught a reader to distrust a genuine transcript hit; it now says the transcript keeps part of the result. Also: control characters other than the tab have escapes of their own (\n, \r, \u00XX), not only \u00XX; an empty fingerprint now stops the block like an empty session id, instead of matching every file; and the eval says tool output reaches a user turn when quoted or written there by the harness, as Phase 3 does, rewrapped to the file's width. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 11 ++++++----- skills/retro/evals/cross-session-pattern-scope.md | 12 ++++++------ 2 files changed, 12 insertions(+), 11 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index a7b5a21..0589122 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -140,7 +140,7 @@ fp=$(cat <<'EOF' EOF ) sid="" -: "${sid:?set sid}" # an empty sid would make the last grep drop every hit +: "${fp:?set fp}" "${sid:?set sid}" # an empty sid would drop every hit fp_json=$(printf '%s' "$fp" | sed 's/\\/\\\\/g; s/"/\\"/g; s/\t/\\t/g') { grep -rl -F --include='*.jsonl' -e "$fp_json" ~/.claude/projects/ grep -rl -F --include='*.txt' -e "$fp" ~/.claude/projects/ @@ -148,10 +148,11 @@ fp_json=$(printf '%s' "$fp" | sed 's/\\/\\\\/g; s/"/\\"/g; s/\t/\\t/g') ``` A transcript is JSON, so it stores a `"` as `\"`, a `\` as `\\` and a tab as -`\t`; `fp_json` is the fingerprint in that form. Other control characters are -stored as `\u00XX`; leave them out of the fingerprint. A tool result too large -for the transcript is kept only as a short preview there, and in full as plain -text under `/tool-results/`, which the second search reads with the +`\t`; `fp_json` is the fingerprint in that form. Other control characters get +escapes of their own (`\n`, `\r`, `\u00XX`); leave them out of the +fingerprint. A tool result too large for the transcript is kept there only in +part (a short preview, for a Bash command a longer excerpt) and in full as +plain text under `/tool-results/`, which the second search reads with the fingerprint as printed. The last `grep` drops the session under analysis: its transcript is `.jsonl`, and its subagent transcripts and tool results sit under `/`. Then read each remaining hit: a session that only discussed diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index c035d01..cc2d1bd 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -25,12 +25,12 @@ earlier sessions had hit it too. It passed the error text to `scan-cross-session.py --pattern` and got zero matches — even for a string that demonstrably sits in the current session's own transcript. -The scanner's `--help` at the time stated the scope: "Search for -keyword/phrase in user messages". An error string, a CI message or a hook denial lives in a tool -result, which the scan does not read. It reaches a user message only when -someone quotes it, so a zero says nothing about recurrence and a hit says only -that the text was quoted. Reading the zero as "this friction does not recur" -turns a blind probe into a finding. +The scanner's `--help` at the time stated the scope: "Search for keyword/phrase +in user messages". An error string, a CI message or a hook denial lives in a +tool result, which the scan does not read. It reaches a user turn only when +someone quotes it or the harness writes it there, so a zero says nothing about +recurrence and a hit says only that the text appeared in a user turn. Reading +the zero as "this friction does not recur" turns a blind probe into a finding. The correct behaviour is to notice the mismatch between the fingerprint's origin and the probe's scope, and to search the session files for the From f63df5570459bcd970e536c36837fa69a2b7bdd7 Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 17:50:05 +0200 Subject: [PATCH 09/14] docs(retro): a probe's own echo is not a recurrence An eighth review found that the counting rule let the search count itself. --pattern prints the pattern back in its JSON output, and that output is a tool result of the session running the retro; under /retro outcome the running session is not the analysed one, so the block lists it and the rule counted it as another session whose tool result holds the string. The same happened to every later retro probing a string an earlier retro had probed. The rule now names the running session and any tool result that only repeats a probe (--pattern output, an earlier run of this search) as not a recurrence, and counts only a tool result that produced the string. The eval's expectation and negative say the same. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 8 +++++--- skills/retro/evals/cross-session-pattern-scope.md | 4 ++-- 2 files changed, 7 insertions(+), 5 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index 0589122..88c75df 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -155,9 +155,11 @@ part (a short preview, for a Bash command a longer excerpt) and in full as plain text under `/tool-results/`, which the second search reads with the fingerprint as printed. The last `grep` drops the session under analysis: its transcript is `.jsonl`, and its subagent transcripts and tool results sit -under `/`. Then read each remaining hit: a session that only discussed -the string is not a recurrence. Count a hit only where the string sits in a -tool result of another session. +under `/`. Then read each remaining hit. Not a recurrence: a session that +only discussed the string, the session running this retro, and a tool result +that only repeats a probe for the string, such as `--pattern` output, which +echoes its pattern, or an earlier run of this search. Count a hit only where +the string sits in a tool result of another session that produced it. `--recurring-failures` does not replace that search. It does not search for a fingerprint you supply, counts only calls whose result is flagged as an error diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index cc2d1bd..b45da16 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -8,14 +8,14 @@ learning_id: retro-20261001-pattern-user-messages-only trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-run'\" for a tool error seen in this session, and the scan answers projects_with_matches: 0." expected: - "Recognise that --pattern searches user turns only, not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." - - "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches in tool results of other sessions." + - "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches in tool results of other sessions that produced the string." - "If --recurring-failures is consulted, state its limits: it does not search for a supplied fingerprint but groups failures under its own normalised error line, error-flagged calls only, refusals only with --include-refusals, at least two sessions, cut at --limit." - "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'." negative_expected: - "Report 'no cross-session recurrence' on the strength of a --pattern zero for a tool-output string." - "Downgrade a finding's severity because the --pattern scan found no other session." - "Treat an empty --recurring-failures list as proof that the fingerprint never recurred." - - "Count a hit in the transcript under analysis, in its subagent transcripts, or in a session that only discussed the string, as a recurrence." + - "Count a hit in the transcript under analysis, in its subagent transcripts, in the session running the retro, in a session that only discussed the string, or in a tool result that only repeats a probe for it (such as --pattern output), as a recurrence." --- # Scenario: a zero from a probe that does not read tool output From 259088a3a61fb3ce1a4e52bd975ce4e7d6329fb7 Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 17:57:43 +0200 Subject: [PATCH 10/14] docs(retro): an earlier run of the search matches through its command The hit-reading sentence listed an earlier run of this search as a tool result that repeats a probe. The block prints file paths only, and a Bash tool result does not store the command, so such a run matches through the command in the assistant's tool call, not through a tool result. The sentence now says that. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 9 +++++---- 1 file changed, 5 insertions(+), 4 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index 88c75df..4831678 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -156,10 +156,11 @@ plain text under `/tool-results/`, which the second search reads with the fingerprint as printed. The last `grep` drops the session under analysis: its transcript is `.jsonl`, and its subagent transcripts and tool results sit under `/`. Then read each remaining hit. Not a recurrence: a session that -only discussed the string, the session running this retro, and a tool result -that only repeats a probe for the string, such as `--pattern` output, which -echoes its pattern, or an earlier run of this search. Count a hit only where -the string sits in a tool result of another session that produced it. +only discussed the string, the session running this retro, a tool result that +only repeats a probe for the string, such as `--pattern` output, which echoes +its pattern, and an earlier run of this search, which matches through its own +command rather than a tool result. Count a hit only where the string sits in a +tool result of another session that produced it. `--recurring-failures` does not replace that search. It does not search for a fingerprint you supply, counts only calls whose result is flagged as an error From 2ea0036e600f9a2636fd18360398e63a9434578e Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 18:00:35 +0200 Subject: [PATCH 11/14] docs(retro): a quoting file is no recurrence; eval names the earlier run The previous commit made "an earlier run of this search" its own exclusion, but the eval's negative still covered it only as a probe's tool result; it now names the earlier run. Both lists also exclude a tool result that shows a file quoting the string, such as a diff or a PR body. That is the commonest harmless hit: two review subagents of this PR hit the example fingerprint only by reading the PR body and the diff, and every session that reads this page after the merge will. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 9 +++++---- skills/retro/evals/cross-session-pattern-scope.md | 2 +- 2 files changed, 6 insertions(+), 5 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index 4831678..b05b910 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -157,10 +157,11 @@ fingerprint as printed. The last `grep` drops the session under analysis: its transcript is `.jsonl`, and its subagent transcripts and tool results sit under `/`. Then read each remaining hit. Not a recurrence: a session that only discussed the string, the session running this retro, a tool result that -only repeats a probe for the string, such as `--pattern` output, which echoes -its pattern, and an earlier run of this search, which matches through its own -command rather than a tool result. Count a hit only where the string sits in a -tool result of another session that produced it. +shows a file or text quoting the string (a diff, a PR body, this page), a tool +result that only repeats a probe for the string, such as `--pattern` output, +which echoes its pattern, and an earlier run of this search, which matches +through its own command rather than a tool result. Count a hit only where the +string sits in a tool result of another session that produced it. `--recurring-failures` does not replace that search. It does not search for a fingerprint you supply, counts only calls whose result is flagged as an error diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index b45da16..d0ef8c4 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -15,7 +15,7 @@ negative_expected: - "Report 'no cross-session recurrence' on the strength of a --pattern zero for a tool-output string." - "Downgrade a finding's severity because the --pattern scan found no other session." - "Treat an empty --recurring-failures list as proof that the fingerprint never recurred." - - "Count a hit in the transcript under analysis, in its subagent transcripts, in the session running the retro, in a session that only discussed the string, or in a tool result that only repeats a probe for it (such as --pattern output), as a recurrence." + - "Count a hit in the transcript under analysis, in its subagent transcripts, in the session running the retro, in a session that only discussed the string, in a tool result that shows a file quoting it, in a tool result that only repeats a probe for it (such as --pattern output), or in an earlier run of this search, as a recurrence." --- # Scenario: a zero from a probe that does not read tool output From ee5c285e6381dca792cf3f22b02a882e7b8f8282 Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 18:02:59 +0200 Subject: [PATCH 12/14] docs(retro): say what counts as a hit before what does not The previous commit excluded "a tool result that shows a file or text quoting the string". Read literally, that also drops a CI log shown by gh run view --log or cat build.log, which is how other sessions see the CI messages Phase 3 names as typical fingerprints, and "a tool result that produced it" pointed the same way. A recurring CI failure could have been reported as no recurrence. The rule now states first what counts: the string as output of the failure itself, including the log of the failing run shown by a tool. The exclusions follow as a list, the quoting case narrowed to a document that quotes the string (a diff, a PR body, a skill or eval file). The eval and the PR body use the same wording. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 22 +++++++++++++------ .../evals/cross-session-pattern-scope.md | 4 ++-- 2 files changed, 17 insertions(+), 9 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index b05b910..38b29d5 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -155,13 +155,21 @@ part (a short preview, for a Bash command a longer excerpt) and in full as plain text under `/tool-results/`, which the second search reads with the fingerprint as printed. The last `grep` drops the session under analysis: its transcript is `.jsonl`, and its subagent transcripts and tool results sit -under `/`. Then read each remaining hit. Not a recurrence: a session that -only discussed the string, the session running this retro, a tool result that -shows a file or text quoting the string (a diff, a PR body, this page), a tool -result that only repeats a probe for the string, such as `--pattern` output, -which echoes its pattern, and an earlier run of this search, which matches -through its own command rather than a tool result. Count a hit only where the -string sits in a tool result of another session that produced it. +under `/`. Then read each remaining hit. + +A hit counts when the string sits in a tool result of another session as +output of the failure itself: a command's own output, or the log of the +failing run shown by a tool (`gh run view --log`, `glab ci trace`, +`cat build.log`). It does not count when it sits in: + +- a session that only discussed the string; +- the session running this retro; +- a tool result showing a document that quotes the string (a diff, a PR body, + a skill or eval file); +- a tool result that only echoes a probe for the string, such as `--pattern` + output, which repeats its pattern; +- an earlier run of this search, which matches through its own command, not + through a tool result. `--recurring-failures` does not replace that search. It does not search for a fingerprint you supply, counts only calls whose result is flagged as an error diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index d0ef8c4..89ea56d 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -8,14 +8,14 @@ learning_id: retro-20261001-pattern-user-messages-only trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-run'\" for a tool error seen in this session, and the scan answers projects_with_matches: 0." expected: - "Recognise that --pattern searches user turns only, not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." - - "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches in tool results of other sessions that produced the string." + - "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches where a tool result of another session shows the string as output of the failure itself, including a CI log of the failing run." - "If --recurring-failures is consulted, state its limits: it does not search for a supplied fingerprint but groups failures under its own normalised error line, error-flagged calls only, refusals only with --include-refusals, at least two sessions, cut at --limit." - "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'." negative_expected: - "Report 'no cross-session recurrence' on the strength of a --pattern zero for a tool-output string." - "Downgrade a finding's severity because the --pattern scan found no other session." - "Treat an empty --recurring-failures list as proof that the fingerprint never recurred." - - "Count a hit in the transcript under analysis, in its subagent transcripts, in the session running the retro, in a session that only discussed the string, in a tool result that shows a file quoting it, in a tool result that only repeats a probe for it (such as --pattern output), or in an earlier run of this search, as a recurrence." + - "Count a hit in the transcript under analysis, in its subagent transcripts, in the session running the retro, in a session that only discussed the string, in a tool result showing a document that quotes it (a diff, a PR body, a skill or eval file), in a tool result that only repeats a probe for it (such as --pattern output), or in an earlier run of this search, as a recurrence." --- # Scenario: a zero from a probe that does not read tool output From f98f3828634ecc299073ec5a17f6b8f100910680 Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 18:05:01 +0200 Subject: [PATCH 13/14] docs(retro): a hook's refusal counts as output of the failure Phase 3 names a hook denial as a typical fingerprint, but the counting rule listed only a command's own output and the log of a failing run. A denied call never ran and has no log; its tool result holds the refusal ("PreToolUse:Bash hook error: ..."), so a literal reader would have rejected every hit of a recurring denial. The rule and the eval now list a hook's or the harness's refusal of the call as the third form, and the PR body states the same rule. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 7 ++++--- skills/retro/evals/cross-session-pattern-scope.md | 2 +- 2 files changed, 5 insertions(+), 4 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index 38b29d5..5777f7c 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -158,9 +158,10 @@ transcript is `.jsonl`, and its subagent transcripts and tool results sit under `/`. Then read each remaining hit. A hit counts when the string sits in a tool result of another session as -output of the failure itself: a command's own output, or the log of the -failing run shown by a tool (`gh run view --log`, `glab ci trace`, -`cat build.log`). It does not count when it sits in: +output of the failure itself: a command's own output, a hook's or the +harness's refusal of the call, or the log of the failing run shown by a tool +(`gh run view --log`, `glab ci trace`, `cat build.log`). It does not count +when it sits in: - a session that only discussed the string; - the session running this retro; diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index 89ea56d..a82c2bc 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -8,7 +8,7 @@ learning_id: retro-20261001-pattern-user-messages-only trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-run'\" for a tool error seen in this session, and the scan answers projects_with_matches: 0." expected: - "Recognise that --pattern searches user turns only, not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." - - "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches where a tool result of another session shows the string as output of the failure itself, including a CI log of the failing run." + - "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches where a tool result of another session shows the string as output of the failure itself: a command's output, a hook's or the harness's refusal, or a CI log of the failing run." - "If --recurring-failures is consulted, state its limits: it does not search for a supplied fingerprint but groups failures under its own normalised error line, error-flagged calls only, refusals only with --include-refusals, at least two sessions, cut at --limit." - "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'." negative_expected: From 9cbad8c15e2c82cf7926a8d6507b2ef774c57c46 Mon Sep 17 00:00:00 2001 From: Sebastian Mendel Date: Thu, 1 Oct 2026 18:07:38 +0200 Subject: [PATCH 14/14] docs(retro): a refusal counts by its reason, not the command it quotes Many hook refusals quote the refused command back; a review counted 34 of 136 Bash refusals carrying 40 or more characters of it verbatim. A commit message or PR comment that merely contained the fingerprint, refused by a gate, would have counted as a recurrence of the error. The counted form is now the refusal's reason, not the command it quotes back, in Phase 3, the eval and the PR body. Learning-Id: retro-20261001-pattern-user-messages-only Assisted-by: claude-code:claude-opus-5-5 Agent-Session: https://claude.ai/code/session_01VpqNFMhbu5Ei961Ca4ZXuL Agent-Host: 32116e Signed-off-by: Sebastian Mendel --- commands/retro.md | 6 +++--- skills/retro/evals/cross-session-pattern-scope.md | 2 +- 2 files changed, 4 insertions(+), 4 deletions(-) diff --git a/commands/retro.md b/commands/retro.md index 5777f7c..2d6f4ec 100644 --- a/commands/retro.md +++ b/commands/retro.md @@ -159,9 +159,9 @@ under `/`. Then read each remaining hit. A hit counts when the string sits in a tool result of another session as output of the failure itself: a command's own output, a hook's or the -harness's refusal of the call, or the log of the failing run shown by a tool -(`gh run view --log`, `glab ci trace`, `cat build.log`). It does not count -when it sits in: +harness's refusal of the call (its reason, not the command it quotes back), +or the log of the failing run shown by a tool (`gh run view --log`, +`glab ci trace`, `cat build.log`). It does not count when it sits in: - a session that only discussed the string; - the session running this retro; diff --git a/skills/retro/evals/cross-session-pattern-scope.md b/skills/retro/evals/cross-session-pattern-scope.md index a82c2bc..d59512c 100644 --- a/skills/retro/evals/cross-session-pattern-scope.md +++ b/skills/retro/evals/cross-session-pattern-scope.md @@ -8,7 +8,7 @@ learning_id: retro-20261001-pattern-user-messages-only trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-run'\" for a tool error seen in this session, and the scan answers projects_with_matches: 0." expected: - "Recognise that --pattern searches user turns only, not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred." - - "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches where a tool result of another session shows the string as output of the failure itself: a command's output, a hook's or the harness's refusal, or a CI log of the failing run." + - "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches where a tool result of another session shows the string as output of the failure itself: a command's output, a hook's or the harness's refusal (its reason, not the command it quotes back), or a CI log of the failing run." - "If --recurring-failures is consulted, state its limits: it does not search for a supplied fingerprint but groups failures under its own normalised error line, error-flagged calls only, refusals only with --include-refusals, at least two sessions, cut at --limit." - "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'." negative_expected: