A Deep Reinforcement Learning implementation of a Double Dueling Deep Q-Network (DD-DQN) for learning to play Flappy Bird using PyTorch and flappy-bird-gymnasium. The project combines several established improvements over the original Deep Q-Network to improve training stability, sample efficiency, and convergence in sparse-reward environments.
This repository implements a Deep Q-Learning agent capable of learning to play Flappy Bird from environment interactions. Rather than using a standard DQN, the agent incorporates modern architectural and optimization techniques to reduce value overestimation, improve representation learning, and stabilize training.
The implementation includes:
- Dueling Deep Q-Network architecture
- Double DQN optimization
- Experience Replay
- Soft Target Network Updates (Polyak Averaging)
The network separates state-value estimation from action-advantage estimation. Instead of directly predicting action values, it learns two functions:
- State Value, (V(s))
- Action Advantage, (A(s,a))
which are combined to estimate the final Q-value. This decomposition allows the model to identify valuable states without requiring accurate estimates for every possible action.
Standard DQN suffers from overestimation caused by using the same network for both action selection and evaluation. Double DQN addresses this issue by selecting actions with the online network while evaluating them using the target network, producing more reliable value estimates.
Training samples are stored in a replay buffer and randomly sampled during optimization. This breaks the temporal correlation between consecutive transitions, improves sample efficiency, and stabilizes gradient updates.
Instead of periodically copying parameters to the target network, the implementation performs incremental updates using Polyak Averaging:
where (
├── agent.py # Training loop and optimization logic
├── dqn.py # Standard and Dueling DQN architectures
├── experience_replay.py # Replay buffer implementation
├── main.py # Training entry point
├── config.yaml # Hyperparameter configuration
├── runs/ # Training logs and saved checkpoints
└── README.md
Clone the repository:
git clone https://github.com/shaahmir/flappy-bird-dqn.git
cd flappy-bird-dqnInstall the required dependencies:
pip install uv
uv sync
To begin training, run
uv run -m python main.pyThe agent will interact with the environment, collect experience, optimize the online network, perform soft updates on the target network, and periodically save model checkpoints.
Training hyperparameters are defined in config.yaml.
Common parameters include:
alpha: 0.001
gamma: 0.99
epsilon_init: 1.0
epsilon_min: 0.01
epsilon_decay: 0.995
replay_memory_size: 100000
batch_size: 64
reward_threshold: 1000
soft_update_tau: 0.005
hidden_dim: 256
dueling: true
double_dqn: true
episodes: 5000
save_every: 100- Double Deep Q-Network (Double DQN)
- Dueling Network Architecture
- Experience Replay Buffer
- Soft Target Updates (Polyak Averaging)
- Epsilon-Greedy Exploration
- PyTorch Implementation
- Configurable Hyperparameters
-
Mnih, V., et al. Human-level control through deep reinforcement learning. Nature, 2015.
-
Van Hasselt, H., Guez, A., & Silver, D. Deep Reinforcement Learning with Double Q-learning. AAAI, 2016.
-
Wang, Z., et al. Dueling Network Architectures for Deep Reinforcement Learning. ICML, 2016.
This project is released under the MIT License. See the LICENSE file for additional information.
