Skip to content

Both arms omitted provenance: report never said findings were read, not observed #80

Description

@twistedmelonman

Both arms of a run produced a technical report that never distinguished findings read off tool output from findings produced by running something. The reviewer caught it as shared residue. In the surface it was written for, a message to the engineer who owns the broken component, that distinction was the difference between a defensible report and a confidently-wrong one.

What happened

Run 2026-09-10T13-12-40. Surface: direct message to a colleague reporting a Terraform plan failure in code he owns. Both arms kept every fact, number, and file reference. Neither said the author had not run the tool.

The reviewer's shared-residue line, verbatim:

both omit the source fact that Andrew has not run terraform himself (only "I don't have state access"), so the reader cannot tell the findings are plan-output-derived rather than hands-on

Arm A wrote "I don't have state access, so that one is for whoever does." Arm B dropped the point entirely. Neither is false, and neither answers the question the recipient actually has: did you observe this, or did you read it?

Why it matters beyond one run

The eventual sent version needed an explicit line ("I haven't run any terraform against this myself; it's read off the Atlantis output plus the module source") and it had to be added by hand, after the skill had returned text described as send-ready.

This is a substance failure, not a phrasing one, so compression does not reach it. Groups V, W, and Z all pass a sentence that omits provenance: it is short, it names an actor, and it states an outcome. Nothing in the current rule set asks whether a claim's basis is visible to the reader.

The failure mode is asymmetric. A report that omits "I ran this and saw it" reads as appropriately cautious. A report that omits "I did not run this, I read the output" reads as first-hand observation, which is a stronger claim than the author can support. Only the second costs credibility, and only when the recipient acts on it and finds out.

Related: smartwatermelon/personify#44 covers voice fidelity from a corpus; this is orthogonal, about epistemic status rather than voice.

Suggested fix

Unverified as a rule change — this is one observed run, not a pattern across the evidence set. Worth checking against ~/.claude/personify-evidence/ for other technical-report runs before acting.

Options, roughly in order of how much they change:

  1. Add provenance to the Work register section as a check rather than a new group: when a claim is about system behavior, is it visible whether the author observed it or read it? Cheapest, and it fits where the substance rules already live.
  2. Fold it into group W. W already asks "who did this, and are they in the sentence?" The extension is "and did they do it, or read that it happened?" Risk: W is the highest-priority group and broadening it may dilute it.
  3. Leave it to the reviewer. It caught the problem here. That is reactive per run, and it only helps when the user reads the residue line, which the quiet default hides.

Option 1 looks right, but the evidence-set check should come first.

Reproduction

Not reproduced. Single observed instance, recorded in full at ~/.claude/personify-evidence/2026-09-10T13-12-40.md — input, both arms, and the reviewer's four lines are in that record.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or requestunverifiedFinding asserts a fact that was never checked

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions