Sharing an independent measurement in case it's useful, plus one question about pricing. Full data, method and every failure case: https://github.com/OrMizL/jev-skill-router-bench
Setup. 81 author-labelled turns against a real 84-skill Hermes agent roster, upstream thresholds unchanged, no tuning on the evaluation set. Results: top-1 53.6% (37/69), returned no suggestion on 36.2% (25/69), named a wrong skill on 10.1% (7/69), and correctly stayed silent on all 12 turns where no skill was needed. Latency 0.59–3.11 s from a server in Israel (the docs quote 70–500 ms for US West), cost ≈ $0.00015–0.00029 per turn.
The most transferable observation: most of the router's errors were abstentions rather than wrong named skills — 25 versus 7. That shape matters for anyone deciding whether to act on a suggestion or on its absence, and I'd not seen it characterised before.
The pricing question. Your models page states output tokens are free, and that shaped how I budgeted this. A downstream Hermes plugin vendors its own client that prices output at $0.13/Mtok, which on my workload is a 1.88× difference ($0.000292/turn vs $0.000155/turn). I have reported the discrepancy both ways rather than assuming. Could you confirm which is correct? If output is free as documented, I'll point the plugin author at their constant.
One reproducible miss, offered as a concrete example rather than a complaint: “pull up the espresso panel and check the last shot” returns gate 0.85 — clearly judging that a skill is wanted — and then names nothing, because the correct skill's own fits lands at 0.20 under a 0.40 threshold. Self-contained request, no missing context.
I also found the abstention-heavy profile is not just a threshold artefact: it comes from fits, not gate, and largely on short context-dependent turns. The benchmark calls the router with no recent_context, so part of that is likely my design rather than the model, and I've said so in the report.
Nothing here is a claim about production accuracy — it is one roster, one model, one pass, hand-selected turns, my own labels. The reproduction scripts and raw responses are in the repo.
Thanks — the typed-answer interface made this unusually easy to measure honestly, which is not something I can say about most scoring APIs.
Sharing an independent measurement in case it's useful, plus one question about pricing. Full data, method and every failure case: https://github.com/OrMizL/jev-skill-router-bench
Setup. 81 author-labelled turns against a real 84-skill Hermes agent roster, upstream thresholds unchanged, no tuning on the evaluation set. Results: top-1 53.6% (37/69), returned no suggestion on 36.2% (25/69), named a wrong skill on 10.1% (7/69), and correctly stayed silent on all 12 turns where no skill was needed. Latency 0.59–3.11 s from a server in Israel (the docs quote 70–500 ms for US West), cost ≈ $0.00015–0.00029 per turn.
The most transferable observation: most of the router's errors were abstentions rather than wrong named skills — 25 versus 7. That shape matters for anyone deciding whether to act on a suggestion or on its absence, and I'd not seen it characterised before.
The pricing question. Your models page states output tokens are free, and that shaped how I budgeted this. A downstream Hermes plugin vendors its own client that prices output at $0.13/Mtok, which on my workload is a 1.88× difference ($0.000292/turn vs $0.000155/turn). I have reported the discrepancy both ways rather than assuming. Could you confirm which is correct? If output is free as documented, I'll point the plugin author at their constant.
One reproducible miss, offered as a concrete example rather than a complaint: “pull up the espresso panel and check the last shot” returns
gate0.85 — clearly judging that a skill is wanted — and then names nothing, because the correct skill's ownfitslands at 0.20 under a 0.40 threshold. Self-contained request, no missing context.I also found the abstention-heavy profile is not just a threshold artefact: it comes from
fits, notgate, and largely on short context-dependent turns. The benchmark calls the router with norecent_context, so part of that is likely my design rather than the model, and I've said so in the report.Nothing here is a claim about production accuracy — it is one roster, one model, one pass, hand-selected turns, my own labels. The reproduction scripts and raw responses are in the repo.
Thanks — the typed-answer interface made this unusually easy to measure honestly, which is not something I can say about most scoring APIs.