Skip to content

fix(mcp): accurate scorecards on real MCP servers (union types, reproducibility, schema hints) - #3

Merged
yotambraun merged 3 commits into
mainfrom
fix/mcp-union-types
Sep 28, 2026
Merged

yotambraun merged 3 commits into
mainfrom
fix/mcp-union-types

Conversation

@yotambraun

Copy link
Copy Markdown
Owner

Why

Running toolscore mcp test against the official MCP reference servers (modelcontextprotocol/servers, pinned 2026.8.31 / 2026.8.18) showed several failing grades that were Toolscore's mistakes, not the servers':

Server Before After Cause
sequential-thinking crash: unhashable type: 'list' A (98%) JSON-Schema union type (["boolean", "string"]) unsupported
everything B (84%) A (93%) default / format: uri ignored when generating inputs
time F (40%) C (70%) sent "sample_timezone" although the description says "e.g., 'America/New_York'"
git 5 lint errors "missing a 'type'" 0 pydantic optional fields are anyOf: [{type: string}, {type: null}]

Also, scenario values came from the unseeded global random, so the "deterministic" scorecard could differ between runs.

Changes

  • Union types: generate for the first non-null type; wrong-type probes match none of the listed types.
  • Reproducible scenarios: one generator per tool, seeded by the tool name.
  • Linter: anyOf/oneOf (all branches typed), allOf (any branch), enum, const, $ref count as typed.
  • Generator: prefer examples, non-null default, an example quoted in the description, then a valid value for format (uri, date-time, date, email, uuid, ...).
  • CHANGELOG [Unreleased] entry, including the known limitation: servers restricted to allowed roots (filesystem, git) still reject generated paths.

Verification

  • 15 new tests, each seen failing first; full suite 899 passed.
  • ruff check, ruff format --check, mypy toolscore clean.
  • Re-ran the scorecard on the six reference servers with this build (grades above).

A JSON-Schema "type" may be a list (e.g. ["boolean", "string"], as published
by the official sequential-thinking server). The edge-case generator crashed
with 'unhashable type: list' and the value generator produced None for such
fields. Generate for the first non-null type, and pick wrong-type probes that
match none of the listed types.

Scenario values came from the unseeded global random module, so the
scorecard could differ between runs. Seed a generator per tool name.
…yped

Found by running the scorecard against the official MCP reference servers:
- the linter flagged pydantic optional fields (anyOf [string, null]) as
  "missing a 'type'" (5 false errors on mcp-server-git);
- happy-path values ignored examples, defaults, format and description
  examples, so e.g. mcp-server-time received "sample_timezone" and its
  correct rejection was scored as a server failure.
@codecov-commenter

Copy link
Copy Markdown

⚠️ Please install the 'codecov app svg image' to ensure uploads and comments are reliably processed by Codecov.

Codecov Report

❌ Patch coverage is 84.93151% with 11 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
toolscore/mcp/harness.py 66.66% 7 Missing and 3 partials ⚠️
toolscore/generators/synthetic.py 97.67% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@yotambraun
yotambraun merged commit 484a825 into main Sep 28, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants