Skip to content

docs: add a scenario-authoring guide (closes #40) - #49

Open
dchaudhari7177 wants to merge 1 commit into
bamdadd:mainfrom
dchaudhari7177:docs/scenario-authoring-guide
Open

dchaudhari7177 wants to merge 1 commit into
bamdadd:mainfrom
dchaudhari7177:docs/scenario-authoring-guide

Conversation

@dchaudhari7177

Copy link
Copy Markdown
Contributor

Closes #40.

The rules for a good scenario were spread across docs/DESIGN.md, the scenarios.py module docstring, and the type docstrings — so the most common contribution was the one with the most scattered instructions. docs/SCENARIOS.md collects them.

Contents

  • What makes a scenario worth adding — the one-value/two-recipients/opposite-verdicts asymmetry, framed as two sentences an author should be able to finish before writing code.
  • Anatomy — a complete worked example.
  • Choosing a value — meaningful-but-unique; why opaque hex and coincidence-prone values ("2024", a bare common name) are both bad.
  • aliases: why they are load-bearing — see below.
  • confirm_phrases — disclosure by yes/no, and the two matching rules.
  • forbidden is a deny-list — no implicit denial.
  • appropriate_flows is the utility signal — and why task must ask for every flow listed.
  • Invariants — which are enforced in __post_init__ vs. by validate_scenario().
  • Registering it in ALL_SCENARIOS, plus a checklist and the gate commands.

The two things authors actually get wrong

These get the most space because they are genuinely surprising:

  1. Normalization is NFKC + case-fold + whitespace-collapse and nothing else. So Attribute(value="$47,318.22") with no aliases does not match output containing 47318.22 — the agent leaked and the scorer says it didn't. The guide shows that failing case first, then the alias that fixes it.
  2. Both rates degenerate on empty lists — disclosure_rate is 0.0 with no forbidden pairs regardless of agent behaviour, and utility is 1.0 with no appropriate_flows.

Verification

Every behavioural claim was executed against the current code, not read off the docstrings. The worked example constructs, passes validate_scenario(), and scores 1.0 / 0.5 for a leaking agent and 0.0 / 1.0 for a compliant one. The confirm-phrase bounds are likewise confirmed: "yes" does not fire inside "yesterday", and "that is correct" does not fire on "that is not correct".

While checking the example I found it violated the guide's own advice — it listed a next_ride → ride_leader appropriate flow that the task text never asks for, which would mark a correct agent down. Fixed in the example.

The example uses a fresh domain-neutral synthetic setting (a community cycling club) so it doesn't duplicate a shipped scenario. All data invented.

Docs-only change; pytest, ruff check and ruff format --check all green.

🤖 Generated with Claude Code

The rules for a good scenario were spread across docs/DESIGN.md, the
scenarios.py module docstring and the type docstrings, so the most common
contribution was also the one with the most scattered instructions.

docs/SCENARIOS.md covers value selection, aliases, confirm_phrases, the
deny-list semantics of forbidden (unlisted = allowed), how appropriate_flows
drives utility, the enforced invariants, and registering in ALL_SCENARIOS.
Linked from both README.md and CONTRIBUTING.md.

Two points get the most space because they are the ones that actually surprise
authors:

- Normalization is NFKC + case-fold + whitespace collapse and nothing else, so
  Attribute(value="$47,318.22") does NOT match "47318.22". Aliases are what
  close that gap, and the guide shows the failing case before the fix.
- forbidden is a deny-list with no implicit denial, and both rates degenerate
  on empty lists (disclosure_rate 0.0 with no forbidden pairs, utility 1.0 with
  no appropriate_flows).

Every claim was executed against the current code rather than read off the
docstrings, including the worked example: it constructs, passes
validate_scenario(), and scores 1.0/0.5 for a leaking agent and 0.0/1.0 for a
compliant one. The confirm-phrase bounds are likewise verified - "yes" does not
fire inside "yesterday", and "that is correct" does not fire on "that is not
correct".

Uses a fresh domain-neutral synthetic example (a community cycling club), so it
does not duplicate an existing scenario.
@dchaudhari7177
dchaudhari7177 force-pushed the docs/scenario-authoring-guide branch 2 times, most recently from 85c804f to 213e113 Compare September 13, 2026 07:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Write a scenario-authoring guide (choosing values, aliases, forbidden vs appropriate)

1 participant