feat(metrics): identical-call rate and one-to-one call pairing (1.9.0) - #2
Merged
Merged
Conversation
redundant_rate counts calls beyond the gold's per-tool expectation, so a research agent issuing several different searches scores the same as one repeating an identical call. identical_count/identical_rate isolate exact repeats (same tool and arguments, key order ignored). Additive: existing keys and composite scores are unchanged.
…and side-effect scoring Positional matching (first same-tool call at index >= i) let one missing, extra or reordered call shift every later comparison, scoring correct calls as 0. Use the best one-to-one assignment per tool instead. An empty expected and empty actual trace now scores argument_f1 1.0.
Add a What's new section to the README, a 1.9.0 CHANGELOG entry, user-guide sections on call pairing and loop detection, and render CHANGELOG.md in the Sphinx docs instead of a stale copy.
|
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
…ts it anthropic 1.8 removed temperature from messages.create, so the Claude judge raised TypeError and mypy failed in CI.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Found while evaluating a real research agent (bytedance/deer-flow) with Agent Eval Flow + Toolscore:
redundant_ratecounts calls beyond the expected per-tool count, so an agent stuck repeating one call and an agent running many different searches look the same.What changes
identical_count/identical_rateinefficiency_metrics: exact repeats (same tool, same arguments, key order ignored). Additive; composite score unchanged.argument_f11.0.argument_f1beforeredundant_rate,sequence_accuracy)anthropicSDKs (1.8 removedtemperaturefrommessages.create, so it raisedTypeError; this also made CI's mypy step fail).temperatureis passed only when the SDK accepts it.[1.9.0], user-guide sections on call pairing and loop detection;docs/changelog.rstnow rendersCHANGELOG.md(it was a stale 0.1.0 copy).Release impact
feat:commit → semantic-release cuts 1.9.0 on merge. Scores can rise for existing baselines and snapshots; the CHANGELOG and README say to re-approve after upgrading.Verification
ruff check,ruff format --check,mypy toolscoreclean with all extras installed (anthropic 1.8.0)