Hi — while building judgekit (MIT, a runtime judgment layer over the System One API: YAML-defined tasks → native choice/score/noul calls, OpenAI-compatible fallback, per-decision cost accounting), we ran a 4-provider comparison on 130 hand-labeled Chinese samples (support-ticket routing 40 / review sentiment 30 / comment spam 30 / support urgency 30). Protocol: temperature=0, one decision call per sample, no few-shot; CI cross-checked with statsmodels; per-sample records and all misjudgments are public in the repo.
Three findings you may find relevant:
-
Jev matches a much larger LLM judge at 4.5× lower latency. Jev (jev-1.13.0, native decisions API): 97.7% (127/130), ~890 ms, ¥0.105 / 1000 decisions at list price. GLM-5.3-flash: 97.7% (127/130) at ~4.0 s. DeepSeek-flash: 96.2% at ~1.1 s — and it missed 2 urgent support tickets, the highest-cost error type in routing scenarios.
-
Your Confidence-gated routing pattern holds empirically. All 3 Jev misjudgments scored below 0.7 confidence while low-confidence outputs were only 9.2% of runs — a 0.7 gate caught 100% of errors at a 9% escalation rate, on real borderline samples (2 of the 3 are also missed by GLM and DeepSeek).
-
Determinism: 3 independent runs produced zero judgment flips (127/130 each) — relevant for anyone building reproducible pipelines on the API.
Happy to turn any of this into a cookbook entry, demo, or benchmark contribution — tell us where it would be most useful.
Hi — while building judgekit (MIT, a runtime judgment layer over the System One API: YAML-defined tasks → native choice/score/noul calls, OpenAI-compatible fallback, per-decision cost accounting), we ran a 4-provider comparison on 130 hand-labeled Chinese samples (support-ticket routing 40 / review sentiment 30 / comment spam 30 / support urgency 30). Protocol: temperature=0, one decision call per sample, no few-shot; CI cross-checked with statsmodels; per-sample records and all misjudgments are public in the repo.
Three findings you may find relevant:
Jev matches a much larger LLM judge at 4.5× lower latency. Jev (
jev-1.13.0, native decisions API): 97.7% (127/130), ~890 ms, ¥0.105 / 1000 decisions at list price. GLM-5.3-flash: 97.7% (127/130) at ~4.0 s. DeepSeek-flash: 96.2% at ~1.1 s — and it missed 2 urgent support tickets, the highest-cost error type in routing scenarios.Your Confidence-gated routing pattern holds empirically. All 3 Jev misjudgments scored below 0.7 confidence while low-confidence outputs were only 9.2% of runs — a 0.7 gate caught 100% of errors at a 9% escalation rate, on real borderline samples (2 of the 3 are also missed by GLM and DeepSeek).
Determinism: 3 independent runs produced zero judgment flips (127/130 each) — relevant for anyone building reproducible pipelines on the API.
Happy to turn any of this into a cookbook entry, demo, or benchmark contribution — tell us where it would be most useful.