Repository navigation
fix(mcp): accurate scorecards on real MCP servers (union types, reproducibility, schema hints) - #3
Merged
Merged
Conversation
A JSON-Schema "type" may be a list (e.g. ["boolean", "string"], as published by the official sequential-thinking server). The edge-case generator crashed with 'unhashable type: list' and the value generator produced None for such fields. Generate for the first non-null type, and pick wrong-type probes that match none of the listed types. Scenario values came from the unseeded global random module, so the scorecard could differ between runs. Seed a generator per tool name.
…yped Found by running the scorecard against the official MCP reference servers: - the linter flagged pydantic optional fields (anyOf [string, null]) as "missing a 'type'" (5 false errors on mcp-server-git); - happy-path values ignored examples, defaults, format and description examples, so e.g. mcp-server-time received "sample_timezone" and its correct rejection was scored as a server failure.
|
Codecov Report❌ Patch coverage is
📢 Thoughts on this report? Let us know! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Running
toolscore mcp testagainst the official MCP reference servers (modelcontextprotocol/servers, pinned 2026.8.31 / 2026.8.18) showed several failing grades that were Toolscore's mistakes, not the servers':unhashable type: 'list'type(["boolean", "string"]) unsupporteddefault/format: uriignored when generating inputs"sample_timezone"although the description says "e.g., 'America/New_York'"anyOf: [{type: string}, {type: null}]Also, scenario values came from the unseeded global
random, so the "deterministic" scorecard could differ between runs.Changes
anyOf/oneOf(all branches typed),allOf(any branch),enum,const,$refcount as typed.examples, non-nulldefault, an example quoted in the description, then a valid value forformat(uri, date-time, date, email, uuid, ...).[Unreleased]entry, including the known limitation: servers restricted to allowed roots (filesystem, git) still reject generated paths.Verification
ruff check,ruff format --check,mypy toolscoreclean.