Repository navigation
Conversation
contracts, the runtime judge agents, and all consuming code/tests/docs,
driven by the words being unclear to a non-ML-background reader:
- k1/k2 -> min_score_supported_output_facts/min_score_expected_output_facts
(CLI flags, manifest defaults, /aissert:eval args)
- extracted_facts -> output_facts; golden_facts/golden_fact_id(s) ->
reference_facts/reference_fact_id(s) (golden-set + canary contracts)
- agents/judge-precision.md -> agents/judge-supported-output-facts.md,
- k1/k2 -> min_score_supported_output_facts/min_score_expected_output_facts
(CLI flags, manifest defaults, /aissert:eval args)
- extracted_facts -> output_facts; golden_facts/golden_fact_id(s) ->
reference_facts/reference_fact_id(s) (golden-set + canary contracts)
- agents/judge-precision.md -> agents/judge-supported-output-facts.md,
agents/judge-recall.md -> agents/judge-expected-output-facts.md (git mv,
rubric prose updated to match)
- total_extracted/total_golden -> total_output_facts/total_reference_facts
(RunMetrics, results.json, report.md, error strings) — closes the gap
left when the input-contract fields were renamed but this output-side
pair wasn't
- fixed two pre-existing schema-doc bugs surfaced along the way:
golden-set-schema.md's field table said "must be 2", canary-schema.md's
example showed schema_version 1, neither tracked the shared constant
Bumps shared SCHEMA_VERSION 1->4 across golden manifest, canary manifest,
and results.json (breaking field renames, not additive). Real data files
rebumped: golden/example/manifest.json (set_version 1.0.4),
canary/manifest.json.
Also:
- target_skill is now optional in /aissert:eval, defaulting from the
golden set's manifest.json; explicit target_skill still cross-checked
against it as a safety net
- hook_bump_golden_version.py: new PostToolUse hook, auto-bumps a golden
set's set_version when items or manifest.json are edited directly
- hook_stop_verify.py: prefer .venv/bin/pytest over bare `pytest` (wasn't
on PATH, broke the Stop hook on every turn)
- README: documented the GitHub-marketplace install path for other users
(/plugin marketplace add <owner>/aissert)
132/132 tests pass; wiki re-anchored to this commit.
Known gap, not closed here: judge rubric wording changed (not just the
file name/label), which per knowledge/domains/change-playbooks.md
requires a live canary re-run before trusting the next real eval's
numbers — not executable from this environment.
Snapshot ReleaseSnapshot build: Download the plugin zip from the release page. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
contracts, the runtime judge agents, and all consuming code/tests/docs,
driven by the words being unclear to a non-ML-background reader:
k1/k2 -> min_score_supported_output_facts/min_score_expected_output_facts (CLI flags, manifest defaults, /aissert:eval args)
extracted_facts -> output_facts; golden_facts/golden_fact_id(s) -> reference_facts/reference_fact_id(s) (golden-set + canary contracts)
agents/judge-precision.md -> agents/judge-supported-output-facts.md,
k1/k2 -> min_score_supported_output_facts/min_score_expected_output_facts (CLI flags, manifest defaults, /aissert:eval args)
extracted_facts -> output_facts; golden_facts/golden_fact_id(s) -> reference_facts/reference_fact_id(s) (golden-set + canary contracts)
agents/judge-precision.md -> agents/judge-supported-output-facts.md, agents/judge-recall.md -> agents/judge-expected-output-facts.md (git mv, rubric prose updated to match)
total_extracted/total_golden -> total_output_facts/total_reference_facts (RunMetrics, results.json, report.md, error strings) — closes the gap left when the input-contract fields were renamed but this output-side pair wasn't
fixed two pre-existing schema-doc bugs surfaced along the way: golden-set-schema.md's field table said "must be 2", canary-schema.md's example showed schema_version 1, neither tracked the shared constant
Bumps shared SCHEMA_VERSION 1->4 across golden manifest, canary manifest,
and results.json (breaking field renames, not additive). Real data files
rebumped: golden/example/manifest.json (set_version 1.0.4),
canary/manifest.json.
Also:
pytest(wasn't on PATH, broke the Stop hook on every turn)132/132 tests pass; wiki re-anchored to this commit.
Known gap, not closed here: judge rubric wording changed (not just the
file name/label), which per knowledge/domains/change-playbooks.md
requires a live canary re-run before trusting the next real eval's
numbers — not executable from this environment.
Summary
Checklist
pytest tests/ -qpasses.python3 scripts/build_plugin_zip.pypasses.skills/aissert/references/.DESIGN.md.README.mdor related docs.Release Impact
fix:,docs:,test:,ci:,chore:,refactor:).feat:).feat!:or other breaking change).