Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
58 changes: 58 additions & 0 deletions commands/retro.md
Original file line number Diff line number Diff line change
Expand Up @@ -124,6 +124,64 @@ Scan session JSONL across projects for related friction:
python3 ${CLAUDE_PLUGIN_ROOT}/skills/retro/scripts/scan-cross-session.py --pattern "<fingerprint>"
```

`--pattern` searches **user turns only** — mostly text the user typed, plus
output the harness writes there — not tool results and not the assistant's
replies. For a fingerprint taken from tool output — an error string, a CI
message, a hook denial — its answer says little either way: a zero does not
show the friction never recurred, and a hit only shows that the text appeared
in a user turn. Search the session files directly — every transcript,
subagent transcripts included, and the files a large tool result is moved to:

```bash
# Quoted heredoc: quotes, backslashes and $ stay literal. Paste the fingerprint
# as it was printed, on one line.
fp=$(cat <<'EOF'
<fingerprint>
EOF
)
sid="<id of the session under analysis>"
: "${fp:?set fp}" "${sid:?set sid}" # an empty sid would drop every hit
fp_json=$(printf '%s' "$fp" | sed 's/\\/\\\\/g; s/"/\\"/g; s/\t/\\t/g')
{ grep -rl -F --include='*.jsonl' -e "$fp_json" ~/.claude/projects/
grep -rl -F --include='*.txt' -e "$fp" ~/.claude/projects/
} | grep -v -F "$sid"
```

A transcript is JSON, so it stores a `"` as `\"`, a `\` as `\\` and a tab as
`\t`; `fp_json` is the fingerprint in that form. Other control characters get
escapes of their own (`\n`, `\r`, `\u00XX`); leave them out of the
fingerprint. A tool result too large for the transcript is kept there only in
part (a short preview, for a Bash command a longer excerpt) and in full as
plain text under `<sid>/tool-results/`, which the second search reads with the
fingerprint as printed. The last `grep` drops the session under analysis: its
transcript is `<sid>.jsonl`, and its subagent transcripts and tool results sit
under `<sid>/`. Then read each remaining hit.

A hit counts when the string sits in a tool result of another session as
output of the failure itself: a command's own output, a hook's or the
harness's refusal of the call (its reason, not the command it quotes back),
or the log of the failing run shown by a tool (`gh run view --log`,
`glab ci trace`, `cat build.log`). It does not count when it sits in:

- a session that only discussed the string;
- the session running this retro;
- a tool result showing a document that quotes the string (a diff, a PR body,
a skill or eval file);
- a tool result that only echoes a probe for the string, such as `--pattern`
output, which repeats its pattern;
- an earlier run of this search, which matches through its own command, not
through a tool result.

`--recurring-failures` does not replace that search. It does not search for a
fingerprint you supply, counts only calls whose result is flagged as an error
(a command that prints an error and still exits 0 is not), drops calls a hook
or the harness refused unless `--include-refusals` is given, lists a failure
only once it occurs in at least two sessions, and cuts its list at `--limit`.
It also reads only top-level transcripts, no subagent ones, looks back only
`--days`, and groups failures by one normalised line picked from the error
output, not by any fingerprint you have in mind. Its silence is no more
evidence than a `--pattern` zero.

For an audit, three modes read the whole window (see the Schicht C section of
`friction-catalog.md`): `--user-correction-summary` (C1/C2),
`--recurring-failures` (C1) and `--follow-up-sessions` (C5).
Expand Down
43 changes: 43 additions & 0 deletions skills/retro/evals/cross-session-pattern-scope.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,43 @@
---
# SPDX-License-Identifier: CC-BY-SA-4.0
# SPDX-FileCopyrightText: Netresearch DTT GmbH
id: cross-session-pattern-scope
skill_under_test: retro
mode: sweep
learning_id: retro-20261001-pattern-user-messages-only
trigger: "Phase 3 runs scan-cross-session.py --pattern \"No such option '--dry-run'\" for a tool error seen in this session, and the scan answers projects_with_matches: 0."
expected:
- "Recognise that --pattern searches user turns only, not tool results, so its answer for a fingerprint taken from tool output does not establish whether the friction recurred."
- "Search the session files directly — every JSONL transcript, subagent transcripts included, and the plain-text tool-results files a large output is moved to — leave out the analysed session, read each hit, and count only matches where a tool result of another session shows the string as output of the failure itself: a command's output, a hook's or the harness's refusal (its reason, not the command it quotes back), or a CI log of the failing run."
- "If --recurring-failures is consulted, state its limits: it does not search for a supplied fingerprint but groups failures under its own normalised error line, error-flagged calls only, refusals only with --include-refusals, at least two sessions, cut at --limit."
- "Say in the report that the first zero was uninformative rather than counting it as 'no recurrence'."
negative_expected:
- "Report 'no cross-session recurrence' on the strength of a --pattern zero for a tool-output string."
- "Downgrade a finding's severity because the --pattern scan found no other session."
- "Treat an empty --recurring-failures list as proof that the fingerprint never recurred."
- "Count a hit in the transcript under analysis, in its subagent transcripts, in the session running the retro, in a session that only discussed the string, in a tool result showing a document that quotes it (a diff, a PR body, a skill or eval file), in a tool result that only repeats a probe for it (such as --pattern output), or in an earlier run of this search, as a recurrence."
---

# Scenario: a zero from a probe that does not read tool output

A sweep found a tool error in the current session and wanted to know whether
earlier sessions had hit it too. It passed the error text to
`scan-cross-session.py --pattern` and got zero matches — even for a string that
demonstrably sits in the current session's own transcript.

The scanner's `--help` at the time stated the scope: "Search for keyword/phrase
in user messages". An error string, a CI message or a hook denial lives in a
tool result, which the scan does not read. It reaches a user turn only when
someone quotes it or the harness writes it there, so a zero says nothing about
recurrence and a hit says only that the text appeared in a user turn. Reading
the zero as "this friction does not recur" turns a blind probe into a finding.

The correct behaviour is to notice the mismatch between the fingerprint's
origin and the probe's scope, and to search the session files for the
fingerprint directly. `--recurring-failures` is no substitute: the error in
this case came from a shell loop that printed it and exited 0, so the call was
never flagged as an error and that mode could not list it either. Only the
direct search goes into the report.

The general form: **before a zero becomes evidence, ask whether the probe could
have returned anything else.**
5 changes: 4 additions & 1 deletion skills/retro/scripts/scan-cross-session.py
Original file line number Diff line number Diff line change
Expand Up @@ -776,7 +776,10 @@ def main() -> int:
parser.add_argument("--projects-dir", type=Path, default=DEFAULT_PROJECTS_DIR)
parser.add_argument("--project", help="Specific project slug (e.g. -home-sme-p)")
parser.add_argument("--days", type=int, default=30)
parser.add_argument("--pattern", help="Search for keyword/phrase in user messages")
parser.add_argument(
"--pattern",
help="Search for keyword/phrase in user turns (mostly typed text; not tool results or assistant replies)",
)
parser.add_argument(
"--user-correction-summary",
action="store_true",
Expand Down
Loading