Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/hook/INDEX.md
Original file line number Diff line number Diff line change
Expand Up @@ -144,7 +144,7 @@ Fail-open on infrastructure errors by design.
| [external-write-falsify-check](../../hooks/advisory-nudge/external-write-falsify-check/spec.md) | PreToolUse (opt-in) | Warn before posting hypothesis-stage text or unverified applied-on-branch claims (#656) to PR / issue / Slack / Notion |
| [external-api-literal-trigger](../../hooks/advisory-nudge/external-api-literal-trigger/spec.md) | PreToolUse | Advisory nudge when ALL_CAPS enum candidates or 3-part SQL identifiers are written without prior retrieval verification |
| [output-block-falsify-advisory](../../hooks/advisory-nudge/output-block-falsify-advisory/spec.md) | PreToolUse | Output-block falsification gate before surfacing `(Recommended)` options or bulk-action commands — T1 (exact `(Recommended)` marker without `Falsified:` line) and T2 (confidence-anchoring tokens) both emit `ask` (T1 was a hard deny until #899); on `Bash`, a blast-radius two-factor predicate (irreversible verb + shared-surface token in one command segment) also emits `ask` (#1010) |
| [source-citation-probe-gate](../../hooks/advisory-nudge/source-citation-probe-gate/spec.md) | PreToolUse | Advisory when an external-write body (gh / Slack / Notion) cites source facts (file:line, inline-code call syntax, test-semantics claims) with no read-probe in the recent transcript and no in-body `Probe:` / `[verified]` basis — issue #830 |
| [source-citation-probe-gate](../../hooks/advisory-nudge/source-citation-probe-gate/spec.md) | PreToolUse | Advisory when an external-write body (gh / Slack / Notion) cites source facts (file:line, inline-code call syntax, test-semantics claims) with no read-probe anywhere in the session's transcript and no in-body `Probe:` / `[verified]` basis — issue #830 |
| [composed-command-gate](../../hooks/advisory-nudge/composed-command-gate/spec.md) | PreToolUse | Advisory when an external-write body's fenced blocks carry `$` command lines with no counterpart among this session's Bash calls — the shape where the pasted output is genuine and only the command line above it was composed; clears on head-binary + 60% operand overlap, on env-var/placeholder substitution, or on `[transcribed]` — issue #1117 |
| [caller-probe-gate](../../hooks/advisory-nudge/caller-probe-gate/spec.md) | PreToolUse | Advisory when an external-write body asserts a code defect (fix( title / bug label / defect token) while citing code, with no call-site search in the recent transcript and no in-body `Caller-probe:` basis — a Read of the cited file deliberately does not clear — issue #906 |
| [unenforced-step-advisory](../../hooks/advisory-nudge/unenforced-step-advisory/spec.md) | PreToolUse | Advisory naming the MANDATORY workflow step that has no gate of its own, at the action that precedes it — code-reviewer dispatch before a content commit, pre-PR / pre-merge rebase (only when `HEAD..<base>` is non-empty), open-PR enumeration before `git worktree add` / cmux dispatch; keyed on transcript absence, never blocks in default mode — issue #1064 |
Expand Down
169 changes: 111 additions & 58 deletions hooks/advisory-nudge/source-citation-probe-gate/impl.py
Original file line number Diff line number Diff line change
Expand Up @@ -28,11 +28,12 @@
Anti-bypass ported from output-block-falsify-advisory (PR #796):
unfilled scaffold placeholders (`<command>` / `<observed>` / `<...>`
/ `<output>`) or empty evidence after the arrow do NOT count.
Arm B (transcript, last 400 JSONL lines): a Read tool_use whose
file_path basename matches the cited basename, or a read-tool Bash
command (grep/rg/sed/cat/head/tail/awk/nl) containing it. T2 clears
on the called function name appearing in a transcript Bash command
or Read path; T3 clears on a pytest run / test-file basename.
Arm B (transcript, whole session): a Read tool_use whose file_path
basename matches the cited basename, a read-tool Bash command
(grep/rg/sed/cat/head/tail/awk/nl) containing it, or that command's
output naming `<basename>:<line>`. T2 clears on the called function
name appearing in a Bash command, a Read path or a read-tool
command's output; T3 clears on a pytest run / test-file basename.

Exits 0 by default — advisory, not block. Set
`PRAXIS_SOURCE_CITATION_STRICT=1` (literal "1" only) to convert into a
Expand All @@ -44,7 +45,6 @@
"""
from __future__ import annotations

import json
import os
import re
import sys
Expand All @@ -57,7 +57,7 @@
safe_tokenize,
)
from _payload import read_payload # type: ignore[import-not-found] # noqa: E402
from _transcript import TRANSCRIPT_SCAN_LINES, tail_lines # type: ignore[import-not-found] # noqa: E402
from _transcript import iter_transcript # type: ignore[import-not-found] # noqa: E402


# Shared surface detection + body extraction now lives in
Expand Down Expand Up @@ -173,60 +173,113 @@ def _detect_citations(body: str) -> list[tuple[str, str, str]]:
_TEST_PROBE_RE = re.compile(r"\bpytest\b|\btest_\w+\.\w+")


def _recent_probes(transcript_path: str) -> tuple[list[str], list[str]]:
"""Return (bash_commands, read_file_paths) from the last N JSONL lines.
# Only lines carrying a tool call or a tool result can clear a citation, so
# every other line is rejected before `json.loads` (see `iter_transcript`).
_PROBE_NEEDLES = ('"tool_use"', '"tool_result"')

Extends external-write-falsify-check's `_recent_bash_commands` to also
collect Read tool_use file_paths — a Read IS a read-probe for Arm B.
# The line number a T1 sample cites: `impl.py:42` / `impl.py#L42`.
_CITED_LINE_RE = re.compile(r"(?::|#L)(\d{1,6})$")


def _output_line_re(sample: str, basename: str) -> re.Pattern[str] | None:
"""How a read command's output names the cited line: `grep -n` prints
`<path>/<basename>:<line>:`, and a context line `<basename>-<line>-`."""
m = _CITED_LINE_RE.search(sample)
if not m:
return None
return re.compile(rf"(?<![\w.-]){re.escape(basename)}[:-]{m.group(1)}\b")


def _result_text(content: object) -> str:
if isinstance(content, str):
return content
if isinstance(content, list):
return "\n".join(
b["text"] for b in content if isinstance(b, dict) and isinstance(b.get("text"), str)
)
return ""


def _clears_by_command(tier: str, clear_key: str, cmd: str) -> bool:
if tier == "file:line":
return bool(_READ_TOOL_RE.search(cmd)) and clear_key in cmd
if tier == "call-syntax":
return clear_key in cmd
return bool(_TEST_PROBE_RE.search(cmd))


def _clears_by_read_path(tier: str, clear_key: str, path: str) -> bool:
basename = path.rsplit("/", 1)[-1]
if tier == "file:line":
return basename == clear_key
if tier == "call-syntax":
return clear_key in path
return basename.startswith("test_")


def _clears_by_output(tier: str, clear_key: str, line_re: re.Pattern[str] | None, text: str) -> bool:
if tier == "file:line":
return line_re is not None and bool(line_re.search(text))
if tier == "call-syntax":
return clear_key in text
return False


def _probed_citations(transcript_path: str, citations: list[tuple[str, str, str]]) -> set[int]:
"""Indices of `citations` that a read-probe anywhere in the session covers.

The whole session is scanned, not a tail window: a body is routinely
written long after the investigation that read what it cites, and a
400-line tail missed that read in 13 of 16 sampled fires (issue #1541).

A read command's output counts too: `grep -rn <symbol> <dir>` reads the
cited line although the file name appears only in what it printed. Only
the result of a read-tool call is consulted, paired to it by id, so the
output of an unrelated command clears nothing.
"""
if not transcript_path or not os.path.isfile(transcript_path):
return [], []
cmds: list[str] = []
read_paths: list[str] = []
for line in tail_lines(transcript_path, TRANSCRIPT_SCAN_LINES):
line = line.strip()
if not line:
continue
try:
entry = json.loads(line)
except (json.JSONDecodeError, ValueError):
continue
if not isinstance(entry, dict):
continue
return set()
pending = {
i: (tier, key, _output_line_re(sample, key) if tier == "file:line" else None)
for i, (tier, sample, key) in enumerate(citations)
}
read_call_ids: set[str] = set()

for entry in iter_transcript(transcript_path, _PROBE_NEEDLES):
msg = entry.get("message") or {}
if not isinstance(msg, dict) or msg.get("role") != "assistant":
if not isinstance(msg, dict):
continue
for block in (msg.get("content") or []):
if not isinstance(block, dict) or block.get("type") != "tool_use":
continue
inp = block.get("input") or {}
if not isinstance(inp, dict):
if not isinstance(block, dict):
continue
if block.get("name") == "Bash":
cmd = inp.get("command", "")
if isinstance(cmd, str) and cmd.strip():
cmds.append(cmd)
elif block.get("name") == "Read":
fp = inp.get("file_path", "")
if isinstance(fp, str) and fp.strip():
read_paths.append(fp)
return cmds, read_paths


def _is_probed(tier: str, clear_key: str, cmds: list[str], read_paths: list[str]) -> bool:
"""True if the transcript contains a read-probe covering this citation."""
if tier == "file:line":
basename = clear_key
if any(fp.rsplit("/", 1)[-1] == basename for fp in read_paths):
return True
return any(_READ_TOOL_RE.search(cmd) and basename in cmd for cmd in cmds)
if tier == "call-syntax":
name = clear_key
return any(name in cmd for cmd in cmds) or any(name in fp for fp in read_paths)
# test-semantics: clears on any pytest run / test-file probe.
if any(_TEST_PROBE_RE.search(cmd) for cmd in cmds):
return True
return any(fp.rsplit("/", 1)[-1].startswith("test_") for fp in read_paths)
if block.get("type") == "tool_use" and msg.get("role") == "assistant":
inp = block.get("input") or {}
if not isinstance(inp, dict):
continue
if block.get("name") == "Bash":
cmd = inp.get("command", "")
if not isinstance(cmd, str) or not cmd.strip():
continue
if _READ_TOOL_RE.search(cmd) and isinstance(block.get("id"), str):
read_call_ids.add(block["id"])
for i, (tier, key, _) in list(pending.items()):
if _clears_by_command(tier, key, cmd):
del pending[i]
elif block.get("name") == "Read":
fp = inp.get("file_path", "")
if not isinstance(fp, str) or not fp.strip():
continue
for i, (tier, key, _) in list(pending.items()):
if _clears_by_read_path(tier, key, fp):
del pending[i]
elif block.get("type") == "tool_result" and block.get("tool_use_id") in read_call_ids:
text = _result_text(block.get("content"))
for i, (tier, key, line_re) in list(pending.items()):
if text and _clears_by_output(tier, key, line_re, text):
del pending[i]
if not pending:
break
return set(range(len(citations))) - set(pending)


# ---------------------------------------------------------------------------
Expand All @@ -235,7 +288,7 @@ def _is_probed(tier: str, clear_key: str, cmds: list[str], read_paths: list[str]

ADVISORY_MESSAGE = (
"REMINDER (External-Surface Write / Source-Citation Probe): body cites "
"source facts ({samples}) with no read-probe found in the recent "
"source facts ({samples}) with no read-probe found in this session's "
"transcript.\n"
"file:line, exact call syntax, and test-semantics claims are "
"recall-prone — re-read the cited site (Read / grep -n) before "
Expand Down Expand Up @@ -295,11 +348,11 @@ def main() -> int:
return 0

# Arm B: per-citation transcript probe scan.
cmds, read_paths = _recent_probes(transcript_path)
probed = _probed_citations(transcript_path, citations)
unprobed = [
(tier, sample)
for tier, sample, clear_key in citations
if not _is_probed(tier, clear_key, cmds, read_paths)
for i, (tier, sample, _) in enumerate(citations)
if i not in probed
]
if not unprobed:
return 0
Expand Down
39 changes: 33 additions & 6 deletions hooks/advisory-nudge/source-citation-probe-gate/spec.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,7 +7,8 @@ PreToolUse(Bash) advisory
that fires when an external-write body (PR/issue bodies and comments) cites
**source facts** — `file:line` references,
exact call syntax, or test-semantics claims — with **no read-probe** found in
the recent transcript or in the body itself.
the session's transcript (the whole session, not a tail window) or in the
body itself.

It enforces the *Information Accuracy* rule's "checkmark = citation" clause
([`ETHOS.md` → Rules praxis carries](../../../ETHOS.md#rules-praxis-carries)) at the external-write surface: a source-fact
Expand Down Expand Up @@ -64,21 +65,39 @@ convention, issue #907):
`<observed>`, `<...>`, `<output>`) or with empty evidence after the
arrow (`→` / `->`) does NOT clear.

**Arm B (transcript, last 400 JSONL lines):**
**Arm B (transcript, whole session):**

- T1 clears when a `Read` tool_use file_path basename equals the cited
basename, or a read-tool Bash command (`grep` / `rg` / `sed` / `cat` /
`head` / `tail` / `awk` / `nl`) contains it.
basename, a read-tool Bash command (`grep` / `rg` / `sed` / `cat` /
`head` / `tail` / `awk` / `nl`) contains it, or that command's output
names `<basename>:<line>` or `<basename>-<line>` for the cited line — the
output arm covers a directory grep (`grep -rn foo hooks/`) whose command
never names the file.
- T2 clears when the called function name appears in any transcript Bash
command or Read file_path.
command, a Read file_path, or a read-tool command's output.
- T3 clears on a `pytest` run or a `test_*` file basename in a transcript
Bash command / Read.

The scan reads the whole session through `_lib/_transcript.iter_transcript`,
pre-filtered to lines carrying `"tool_use"` or `"tool_result"`, and stops as
soon as every citation has cleared. Until #1541 it read the last 400 JSONL
lines only; a PR body is routinely written long after the reads it rests on,
and in the 16 fires #1538 sampled that tail missed the read in 13. In the
other 3 the cited file appeared in no read command line or Read path at all,
only in a read command's output. The replayed
fires are pinned, scrubbed to aliases, under
`tests/fixtures/source-citation-probe-gate/replay-1541/`.

Cost: the scan runs only for a `gh` external write whose body carries a
citation. On the largest local transcript (133 MB) one full pass took about
0.3 s against the hook's 5 s timeout; the slowest of the 16 replayed calls
took about 0.4 s end to end (0.38 s and 0.40 s on two runs).

## Response

```text
REMINDER (External-Surface Write / Source-Citation Probe): body cites
source facts ({up to 3 samples}) with no read-probe found in the recent
source facts ({up to 3 samples}) with no read-probe found in this session's
transcript.
file:line, exact call syntax, and test-semantics claims are recall-prone —
re-read the cited site (Read / grep -n) before publishing, then cite inline
Expand Down Expand Up @@ -118,6 +137,14 @@ block). Set `PRAXIS_SOURCE_CITATION_STRICT=1` to convert into a hard block
`internal.corp:8080` has extension `corp` (not in the denylist) and a
digit run after the colon — it matches T1. The denylist covers the common
public TLDs only.
- **A read of the file clears every line of it.** T1 clears on the file
having been read, not on the cited line having been seen:
`sed -n '50,60p' dag.py` clears `dag.py:14`. Only the output arm checks
the line number.
- **A compound command's output is read as a whole.** When a Bash call
pairs a read tool with `echo` (`grep -n x a.py; echo a.py:9`), the whole
output counts as read-tool output, so the echoed `a.py:9` clears. A
command that is only `echo` does not clear.
- **T2 is the weakest detector.** Call syntax without a `.` / `[` in the
argument list (`foo(x)`) is deliberately not matched, and a matching
function name anywhere in any transcript Bash command clears it — recall
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
- the test asserts the result
Loading
Loading