On-policy relabeling for robust search-agent distillation under rollout distribution shift.
-
Updated
May 17, 2026 - Python
On-policy relabeling for robust search-agent distillation under rollout distribution shift.
On-policy RL (VPG, PPO) algorithms from scratch
Add a description, image, and links to the on-policy-learning topic page so that developers can more easily learn about it.
To associate your repository with the on-policy-learning topic, visit your repo's landing page and select "manage topics."