Collection of my Reinforcement Learning (RL) practices including DQN, D3QN, and Adaptive Gamma, applied to the Lunar Lander and CartPole environments. 🚀🕹️
-
Updated
Oct 21, 2024 - Jupyter Notebook
Collection of my Reinforcement Learning (RL) practices including DQN, D3QN, and Adaptive Gamma, applied to the Lunar Lander and CartPole environments. 🚀🕹️
Speculative decoding runtime with rejection sampling, adaptive gamma strategy, and provable correctness guarantees. Achieves 1.41x speedup on CPU with Qwen2-0.5B/1.5B pair. Draft model generates candidates, target model verifies in single forward pass. 31/31 tests passing.
Add a description, image, and links to the adaptive-gamma topic page so that developers can more easily learn about it.
To associate your repository with the adaptive-gamma topic, visit your repo's landing page and select "manage topics."