A research lab for off-policy evaluation, exploration, and policy-generated bias in contextual-bandit recommendation systems.
recommender-systems causal-inference importance-sampling contextual-bandits exploration-exploitation selection-bias counterfactual-learning off-policy-evaluation doubly-robust inverse-propensity-scoring
-
Updated
Sep 9, 2026 - Python