You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Code, per-item results and figures for a study dissociating prompt quality from response compliance in automated prompt-engineering assessment. Four-agent MATLAB evaluator on a locally hosted Qwen 2.5-7B judge, over 498 prompts from IFEval, LMSYS-Chat-1M and WildChat.
Measure position/verbosity/assertiveness bias in an LLM-as-judge (Claude) by judging pairs in both orders. Finding: no position bias, but 75% verbosity bias and 100% assertiveness bias — judge scores gameable by length + tone.
SSIT: a label-free, gold-free test for whether LLM-judge position x verbosity bias corrections actually compose. No human labels, no model of the judge. Code, 7-judge/6-family pilot data, and pre-registered protocol for the NeurIPS 2026 JUDGe workshop paper.