Skip to content

Multi-wave terminology rename across the golden/canary/results JSON - #7

Merged
YauheniPo merged 1 commit into
mainfrom
rename
Jul 28, 2026
Merged

YauheniPo merged 1 commit into
mainfrom
rename

Conversation

@YauheniPo

Copy link
Copy Markdown
Owner

contracts, the runtime judge agents, and all consuming code/tests/docs,
driven by the words being unclear to a non-ML-background reader:

  • k1/k2 -> min_score_supported_output_facts/min_score_expected_output_facts (CLI flags, manifest defaults, /aissert:eval args)

  • extracted_facts -> output_facts; golden_facts/golden_fact_id(s) -> reference_facts/reference_fact_id(s) (golden-set + canary contracts)

  • agents/judge-precision.md -> agents/judge-supported-output-facts.md,

  • k1/k2 -> min_score_supported_output_facts/min_score_expected_output_facts (CLI flags, manifest defaults, /aissert:eval args)

  • extracted_facts -> output_facts; golden_facts/golden_fact_id(s) -> reference_facts/reference_fact_id(s) (golden-set + canary contracts)

  • agents/judge-precision.md -> agents/judge-supported-output-facts.md, agents/judge-recall.md -> agents/judge-expected-output-facts.md (git mv, rubric prose updated to match)

  • total_extracted/total_golden -> total_output_facts/total_reference_facts (RunMetrics, results.json, report.md, error strings) — closes the gap left when the input-contract fields were renamed but this output-side pair wasn't

  • fixed two pre-existing schema-doc bugs surfaced along the way: golden-set-schema.md's field table said "must be 2", canary-schema.md's example showed schema_version 1, neither tracked the shared constant

Bumps shared SCHEMA_VERSION 1->4 across golden manifest, canary manifest,
and results.json (breaking field renames, not additive). Real data files
rebumped: golden/example/manifest.json (set_version 1.0.4),
canary/manifest.json.

Also:

  • target_skill is now optional in /aissert:eval, defaulting from the golden set's manifest.json; explicit target_skill still cross-checked against it as a safety net
  • hook_bump_golden_version.py: new PostToolUse hook, auto-bumps a golden set's set_version when items or manifest.json are edited directly
  • hook_stop_verify.py: prefer .venv/bin/pytest over bare pytest (wasn't on PATH, broke the Stop hook on every turn)
  • README: documented the GitHub-marketplace install path for other users (/plugin marketplace add /aissert)

132/132 tests pass; wiki re-anchored to this commit.

Known gap, not closed here: judge rubric wording changed (not just the
file name/label), which per knowledge/domains/change-playbooks.md
requires a live canary re-run before trusting the next real eval's
numbers — not executable from this environment.

Summary

Checklist

  • pytest tests/ -q passes.
  • python3 scripts/build_plugin_zip.py passes.
  • No real/proprietary golden data is present in the repo tree.
  • Contract changes update skills/aissert/references/.
  • Design deviations update DESIGN.md.
  • User-facing changes update README.md or related docs.

Release Impact

  • Patch (default for fix:, docs:, test:, ci:, chore:, refactor:).
  • Minor (feat:).
  • Major (feat!: or other breaking change).

  contracts, the runtime judge agents, and all consuming code/tests/docs,
  driven by the words being unclear to a non-ML-background reader:

  - k1/k2 -> min_score_supported_output_facts/min_score_expected_output_facts
    (CLI flags, manifest defaults, /aissert:eval args)
  - extracted_facts -> output_facts; golden_facts/golden_fact_id(s) ->
    reference_facts/reference_fact_id(s) (golden-set + canary contracts)
  - agents/judge-precision.md -> agents/judge-supported-output-facts.md,

  - k1/k2 -> min_score_supported_output_facts/min_score_expected_output_facts
    (CLI flags, manifest defaults, /aissert:eval args)
  - extracted_facts -> output_facts; golden_facts/golden_fact_id(s) ->
    reference_facts/reference_fact_id(s) (golden-set + canary contracts)
  - agents/judge-precision.md -> agents/judge-supported-output-facts.md,
    agents/judge-recall.md -> agents/judge-expected-output-facts.md (git mv,
    rubric prose updated to match)
  - total_extracted/total_golden -> total_output_facts/total_reference_facts
    (RunMetrics, results.json, report.md, error strings) — closes the gap
    left when the input-contract fields were renamed but this output-side
    pair wasn't
  - fixed two pre-existing schema-doc bugs surfaced along the way:
    golden-set-schema.md's field table said "must be 2", canary-schema.md's
    example showed schema_version 1, neither tracked the shared constant

  Bumps shared SCHEMA_VERSION 1->4 across golden manifest, canary manifest,
  and results.json (breaking field renames, not additive). Real data files
  rebumped: golden/example/manifest.json (set_version 1.0.4),
  canary/manifest.json.

  Also:
  - target_skill is now optional in /aissert:eval, defaulting from the
    golden set's manifest.json; explicit target_skill still cross-checked
    against it as a safety net
  - hook_bump_golden_version.py: new PostToolUse hook, auto-bumps a golden
    set's set_version when items or manifest.json are edited directly
  - hook_stop_verify.py: prefer .venv/bin/pytest over bare `pytest` (wasn't
    on PATH, broke the Stop hook on every turn)
  - README: documented the GitHub-marketplace install path for other users
    (/plugin marketplace add <owner>/aissert)

  132/132 tests pass; wiki re-anchored to this commit.

  Known gap, not closed here: judge rubric wording changed (not just the
  file name/label), which per knowledge/domains/change-playbooks.md
  requires a live canary re-run before trusting the next real eval's
  numbers — not executable from this environment.
@github-actions

Copy link
Copy Markdown

Snapshot Release

Snapshot build: aissert--v0.10.1-SNAPSHOT-pr7

Download the plugin zip from the release page.

@YauheniPo
YauheniPo merged commit 0b45e11 into main Jul 28, 2026
4 checks passed
@YauheniPo
YauheniPo deleted the rename branch July 28, 2026 07:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant