Skip to content

feat(metrics): identical-call rate and one-to-one call pairing (1.9.0) - #2

Merged
yotambraun merged 4 commits into
mainfrom
feat/identical-call-rate
Sep 28, 2026
Merged

yotambraun merged 4 commits into
mainfrom
feat/identical-call-rate

Conversation

@yotambraun

@yotambraun yotambraun commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Why

Found while evaluating a real research agent (bytedance/deer-flow) with Agent Eval Flow + Toolscore:

  1. redundant_rate counts calls beyond the expected per-tool count, so an agent stuck repeating one call and an agent running many different searches look the same.
  2. Argument and side-effect checks matched each expected call to the first same-tool call at the same or a later position. One missing, extra or reordered call shifted every later comparison, and correct calls scored 0.

What changes

  • Added identical_count / identical_rate in efficiency_metrics: exact repeats (same tool, same arguments, key order ignored). Additive; composite score unchanged.
  • Fixed argument F1 and side-effect validation now pair expected and actual calls one-to-one per tool by best argument match (exact assignment up to 12 calls per tool, greedy above). No tools expected and none called now gives argument_f1 1.0.
Case argument_f1 before after
One earlier call missing, the other two correct 0.00 0.67
Two same-tool calls in swapped order 0.00 1.00
Wrong attempt, then the correct call 0.00 1.00 (still penalised by redundant_rate, sequence_accuracy)
No tools expected, none called 0.00 (score 0.70) 1.00 (score 1.00)
  • Fixed the Claude LLM judge with current anthropic SDKs (1.8 removed temperature from messages.create, so it raised TypeError; this also made CI's mypy step fail). temperature is passed only when the SDK accepts it.
  • Docs: README "What's new in 1.9", CHANGELOG [1.9.0], user-guide sections on call pairing and loop detection; docs/changelog.rst now renders CHANGELOG.md (it was a stale 0.1.0 copy).

Release impact

feat: commit → semantic-release cuts 1.9.0 on merge. Scores can rise for existing baselines and snapshots; the CHANGELOG and README say to re-approve after upgrading.

Verification

  • 14 new tests, each seen failing before its change; full suite 884 passed
  • ruff check, ruff format --check, mypy toolscore clean with all extras installed (anthropic 1.8.0)
  • Sphinx build renders the new sections and the 1.9.0 changelog

redundant_rate counts calls beyond the gold's per-tool expectation, so a
research agent issuing several different searches scores the same as one
repeating an identical call. identical_count/identical_rate isolate exact
repeats (same tool and arguments, key order ignored). Additive: existing
keys and composite scores are unchanged.
…and side-effect scoring

Positional matching (first same-tool call at index >= i) let one missing,
extra or reordered call shift every later comparison, scoring correct calls
as 0. Use the best one-to-one assignment per tool instead. An empty expected
and empty actual trace now scores argument_f1 1.0.
Add a What's new section to the README, a 1.9.0 CHANGELOG entry, user-guide
sections on call pairing and loop detection, and render CHANGELOG.md in the
Sphinx docs instead of a stale copy.
@codecov-commenter

codecov-commenter commented Sep 28, 2026 •

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 73.33333% with 16 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
toolscore/metrics/arguments.py 74.46% 11 Missing and 1 partial ⚠️
toolscore/metrics/efficiency.py 50.00% 2 Missing ⚠️
toolscore/metrics/llm_judge.py 80.00% 1 Missing ⚠️
toolscore/metrics/side_effects.py 75.00% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

…ts it

anthropic 1.8 removed temperature from messages.create, so the Claude judge
raised TypeError and mypy failed in CI.
@yotambraun
yotambraun merged commit 0b81056 into main Sep 28, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants