Admission refuses a sentinel pointing the wrong way - #109
Merged
Conversation
`START` is the graph's entry and `END` its exit, but `_check_endpoints` accepted both in either role — the sentinel test did not look at which side of the edge it was on. So a proposal carrying `END -> x` or `x -> START` was admitted, and then died in `Materializer` with `StateGraph`'s own "END cannot be a start node" / "START cannot be an end node". The run does not proceed either way. What was wrong is where the failure was charged and what the planner was told. `GovernedLoop` counts a `MaterializationError` as an execution failure, against `max_consecutive_execution_failures` (2), rather than as a rejection against `max_consecutive_rejections` (3) — so a planner got fewer retries for a mistake admission is supposed to catch than for one it does catch, and two in a row ended the run as `EXECUTION_FAILED`, a stop reason claiming the graph ran when nothing had. And a rejection is meant to be data. `feedback()` hands the planner codes and remedies; what it got here was prose assembled from an exception, with no code, no remedy, and nothing on the `admission` event's failed-check list — because admission had not failed. The prompt in proposal.py already tells models the rule; the gate is what did not hold when a model ignored it. Both endpoints are still reported rather than the first, matching every other check. The rejection rides `Check.REGISTRY` with code `sentinel_wrong_direction` and a remedy naming the side the sentinel belongs on. Two of the three new tests go red without the fix; the third is the guard that the normal shape still admits, which must stay green either way. Closes #108 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Shashankss1205
added a commit
that referenced
this pull request
Aug 13, 2026
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Shashankss1205
added a commit
that referenced
this pull request
Aug 13, 2026
…ning (#112) * A sandboxed child that dies without sending names the tool it was running `poll()` returns true when the pipe is readable, and a closed pipe is readable. So a child that died without sending — `os._exit`, a segfault, an OOM-kill, anything that kills the process rather than raising inside `spec.fn` — fell through the timeout guard into `recv()` and raised a bare `EOFError('')`. `run` has two failure shapes and that was neither: `SandboxViolation` for a confinement breach, `RuntimeError(f"tool {name!r} failed: ...")` for the tool's own exception. An EOFError with an empty message is attributable to nothing. Downstream is where it bit. `AgentNode`'s loop catches it under its blanket `except Exception` and renders `f"TOOL_ERROR: {exc}"` — and `str(EOFError(''))` is `''`, so the model was handed `TOOL_ERROR:` with nothing after the colon. It was told its call failed and given no way to tell why, which tool, or whether a retry could help; the same text went into the trace, so the audit trail could not explain the failure either. The `_CALL_SHAPE_ERROR` hint below that clause cannot fire on it, since it matches on message text and there is none. This is the shape of a defect already closed here once: a curtailed phase writing `[budget_exhausted] ` with nothing after it. Same empty message, one layer down. Now a `RuntimeError` naming the tool and the exit code, with a negative one rendered as the signal that killed it — which is what tells an OOM-kill apart from a deliberate `_exit`. Deliberately not a `SandboxViolation`: a child dying is not evidence it tried to escape confinement, and a violation is a specific accusation that lands in the trace as one. The three new tests go red without the fix, with the EOFError raising out of `multiprocessing/connection.py` exactly as reported. The normal return, the tool-raises and the timeout paths are untouched. Closes #111 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * Rebase on main: #109's three admission tests move the count to 2,151 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Merged
Shashankss1205
added a commit
that referenced
this pull request
Aug 13, 2026
Two defects closed since 0.1.6, both found by re-verifying a stale bug backlog against main rather than by a report: - admission accepted `END` as an edge source and `START` as a target, so a graph that cannot be built was admitted and failed in materialisation — charged to the execution-failure allowance rather than the rejection one, and reaching the planner as prose instead of a code and a remedy (#108/#109). - a sandboxed child that died without sending escaped as a bare `EOFError('')`, which the agent loop rendered to the model as `TOOL_ERROR:` and nothing else (#111/#112). The version moves in the two places CI compares and the six where prose states it. `tests/test_deep_dive.py` and `tests/test_cookbook_*.py` assert all but the README line, which is how they stay right. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #108.
The gap
_check_endpointstestedendpoint in _SENTINELSwithout looking at which side of the edge it was on.STARTis the graph's entry andENDits exit, so:Why it mattered
The run doesn't proceed either way. What was wrong is where the failure was charged and what the planner was told.
GovernedLooptreats aMaterializationErroras an execution failure —max_consecutive_execution_failures(2) — rather than a rejection —max_consecutive_rejections(3). A planner got fewer retries for a mistake admission is supposed to catch than for one it does catch, and two in a row ended the run asLoopStop.EXECUTION_FAILED: a stop reason claiming the graph ran and failed, when nothing ran.And a rejection is meant to be data.
feedback()hands the planner codes and remedies. What it got instead:No code, no remedy, and nothing on the
admissiontrace event's failed-check list — because admission hadn't failed. The prompt inproposal.pyalready tells models the rule; the gate is what didn't hold when a model ignored it.After
Both endpoints are still reported rather than only the first, matching every other check's posture.
Verification
tests/test_admission.py. Two go red without the fix; the third is the guard that the normalSTART -> … -> ENDshape still admits, which must stay green either way.failed_checks() == (Check.REGISTRY,)and relies on the fact that every registered kind's factory in that file is_explode— anything that built the graph anyway would raise rather than fail quietly.uv run pytestgreen ·uv run ruff check .clean.The deep-dive selected-count moved 2,145 → 2,148 for the three new tests, which
tests/test_deep_dive.pycaught on the first full run.Out of scope
Whether a
MaterializationErrorshould count against the rejection allowance in general — that is a separate question about the loop. Closing this removes the only known way to reach it from a well-formed proposal.🤖 Generated with Claude Code