You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A counterfactual battery for instruction following in vision-language-action policies: change one word of the instruction, hold the scene, and record which object the arm reaches. Ships the deconfounded demonstration generators and their confounded control.
Offline recommender metrics vs true online lift, graded against a known answer (Open Bandit Dataset). Ends in an evaluation gate that refuses to report a number when the diagnostics say it would be meaningless.
Branching rollouts show that replay-based evaluation of per-step model routing in LLM agents scores states that never occur. Efficient Reasoning Workshop @ COLM 2026.
Position-bias-aware ranking: estimating and correcting position bias to optimize Earnings Per Visitor (EPV) · simulation study · IPW & propensity modeling · Python
Off policy evaluation audit for logged bandits: measures how much of a target policy's probability mass sits outside what the log could have produced, shows that a 95 percent interval around the standard estimator covers 0.26 of the time, and returns exit 2 rather than a number when the estimate would be about a different quantity.
Counterfactual dependability framework for testing whether language-model safety mechanisms are behaviorally load-bearing under removal, misrouting, bypass, and family-ablation controls.