Skip to content

Independent measurement of a Jev skill router on a real agent roster (and one pricing question) #12

Description

@OrMizL

Sharing an independent measurement in case it's useful, plus one question about pricing. Full data, method and every failure case: https://github.com/OrMizL/jev-skill-router-bench

Setup. 81 author-labelled turns against a real 84-skill Hermes agent roster, upstream thresholds unchanged, no tuning on the evaluation set. Results: top-1 53.6% (37/69), returned no suggestion on 36.2% (25/69), named a wrong skill on 10.1% (7/69), and correctly stayed silent on all 12 turns where no skill was needed. Latency 0.59–3.11 s from a server in Israel (the docs quote 70–500 ms for US West), cost ≈ $0.00015–0.00029 per turn.

The most transferable observation: most of the router's errors were abstentions rather than wrong named skills — 25 versus 7. That shape matters for anyone deciding whether to act on a suggestion or on its absence, and I'd not seen it characterised before.

The pricing question. Your models page states output tokens are free, and that shaped how I budgeted this. A downstream Hermes plugin vendors its own client that prices output at $0.13/Mtok, which on my workload is a 1.88× difference ($0.000292/turn vs $0.000155/turn). I have reported the discrepancy both ways rather than assuming. Could you confirm which is correct? If output is free as documented, I'll point the plugin author at their constant.

One reproducible miss, offered as a concrete example rather than a complaint: “pull up the espresso panel and check the last shot” returns gate 0.85 — clearly judging that a skill is wanted — and then names nothing, because the correct skill's own fits lands at 0.20 under a 0.40 threshold. Self-contained request, no missing context.

I also found the abstention-heavy profile is not just a threshold artefact: it comes from fits, not gate, and largely on short context-dependent turns. The benchmark calls the router with no recent_context, so part of that is likely my design rather than the model, and I've said so in the report.

Nothing here is a claim about production accuracy — it is one roster, one model, one pass, hand-selected turns, my own labels. The reproduction scripts and raw responses are in the repo.

Thanks — the typed-answer interface made this unusually easy to measure honestly, which is not something I can say about most scoring APIs.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions