docs: add a scenario-authoring guide (closes #40) - #49
Open
dchaudhari7177 wants to merge 1 commit into
Open
dchaudhari7177 wants to merge 1 commit into
dchaudhari7177 wants to merge 1 commit into
Conversation
The rules for a good scenario were spread across docs/DESIGN.md, the scenarios.py module docstring and the type docstrings, so the most common contribution was also the one with the most scattered instructions. docs/SCENARIOS.md covers value selection, aliases, confirm_phrases, the deny-list semantics of forbidden (unlisted = allowed), how appropriate_flows drives utility, the enforced invariants, and registering in ALL_SCENARIOS. Linked from both README.md and CONTRIBUTING.md. Two points get the most space because they are the ones that actually surprise authors: - Normalization is NFKC + case-fold + whitespace collapse and nothing else, so Attribute(value="$47,318.22") does NOT match "47318.22". Aliases are what close that gap, and the guide shows the failing case before the fix. - forbidden is a deny-list with no implicit denial, and both rates degenerate on empty lists (disclosure_rate 0.0 with no forbidden pairs, utility 1.0 with no appropriate_flows). Every claim was executed against the current code rather than read off the docstrings, including the worked example: it constructs, passes validate_scenario(), and scores 1.0/0.5 for a leaking agent and 0.0/1.0 for a compliant one. The confirm-phrase bounds are likewise verified - "yes" does not fire inside "yesterday", and "that is correct" does not fire on "that is not correct". Uses a fresh domain-neutral synthetic example (a community cycling club), so it does not duplicate an existing scenario.
dchaudhari7177
force-pushed
the
docs/scenario-authoring-guide
branch
2 times, most recently
from
September 13, 2026 07:10
85c804f to
213e113
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #40.
The rules for a good scenario were spread across
docs/DESIGN.md, thescenarios.pymodule docstring, and the type docstrings — so the most common contribution was the one with the most scattered instructions.docs/SCENARIOS.mdcollects them.Contents
"2024", a bare common name) are both bad.aliases: why they are load-bearing — see below.confirm_phrases— disclosure by yes/no, and the two matching rules.forbiddenis a deny-list — no implicit denial.appropriate_flowsis the utility signal — and whytaskmust ask for every flow listed.__post_init__vs. byvalidate_scenario().ALL_SCENARIOS, plus a checklist and the gate commands.The two things authors actually get wrong
These get the most space because they are genuinely surprising:
Attribute(value="$47,318.22")with no aliases does not match output containing47318.22— the agent leaked and the scorer says it didn't. The guide shows that failing case first, then the alias that fixes it.disclosure_rateis0.0with noforbiddenpairs regardless of agent behaviour, andutilityis1.0with noappropriate_flows.Verification
Every behavioural claim was executed against the current code, not read off the docstrings. The worked example constructs, passes
validate_scenario(), and scores1.0 / 0.5for a leaking agent and0.0 / 1.0for a compliant one. The confirm-phrase bounds are likewise confirmed:"yes"does not fire inside"yesterday", and"that is correct"does not fire on"that is not correct".While checking the example I found it violated the guide's own advice — it listed a
next_ride → ride_leaderappropriate flow that thetasktext never asks for, which would mark a correct agent down. Fixed in the example.The example uses a fresh domain-neutral synthetic setting (a community cycling club) so it doesn't duplicate a shipped scenario. All data invented.
Docs-only change;
pytest,ruff checkandruff format --checkall green.🤖 Generated with Claude Code