This is an ongoing project meant as an educational exploration of classic and deep Reinforcement Learning (RL) algorithms, from dynamic programming to tabular methods and modern policy gradient approaches. Each algorithm is implemented from scratch and demonstrated on Gymnasium environments.
This project is designed as a hands-on introduction to Reinforcement Learning. It covers a broad range of algorithms organized by family, each paired with concrete examples on classic control and navigation tasks. The goal is to provide clean, readable implementations that closely follow the original papers and textbooks, making it straightforward to study the mechanics of each method.
| Family | Algorithms |
|---|---|
| Heuristics | A* Search |
| Dynamic Programming | Value Iteration, Q-Iteration, Policy Iteration |
| Monte Carlo | First-visit and Every-visit Monte Carlo Control |
| Temporal Difference | SARSA, Expected SARSA, Q-Learning, Double Q-Learning |
| Deep RL — Policy Gradient | REINFORCE, REINFORCE with Baseline |
| Deep RL — Actor-Critic | Actor-Critic (online), A2C with normalized advantages |
| Deep RL — Value-Based | DQN, Double DQN |
| Deep RL — Clipped Surrogate | PPO (multi-worker, with GAE and mini-batch updates) |
Agents are trained and evaluated on standard Gymnasium environments:
- Discrete grid worlds: FrozenLake, CliffWalking, Taxi
- Classic control: CartPole, MountainCar, Acrobot
- Continuous observation / discrete action: LunarLander
- Card game: Blackjack
Every algorithm family follows the same layout: a core algorithms.py with the implementation, and an examples/ folder with environment-specific agent subclasses, network definitions, and one runnable training script per environment.
src/reinforcement_learning_course/
├── core.py # Abstract Agent base class
├── heuristics/
│ ├── algorithms.py # A* Search
│ └── examples/
├── dynamic_programming/
│ ├── algorithms.py # ValueIteration, QIteration, PolicyIteration
│ └── examples/
├── monte_carlo/
│ ├── algorithms.py # MonteCarlo (first-visit & every-visit)
│ └── examples/
├── temporal_difference/
│ ├── algorithms.py # SARSA, ExpectedSARSA, QLearning, DoubleQLearning
│ └── examples/
└── deep_rl/
├── reinforce/
│ ├── algorithms.py # Reinforce
│ └── examples/
├── reinforce_baseline/
│ ├── algorithms.py # ReinforceBaseline
│ └── examples/
├── actor_critic/
│ ├── algorithms.py # ActorCritic (online, single-worker)
│ └── examples/
├── a2c_normalized/
│ ├── algorithms.py # A2CWorker — multi-worker A2C with normalized advantages
│ └── examples/
├── dqn/
│ ├── algorithms.py # DQN, Double DQN
│ ├── utils.py # Experience replay buffer, epsilon schedule
│ └── examples/
└── ppo/
├── algorithms.py # PPOWorker — multi-worker PPO with GAE
└── examples/
├── agents.py # Environment-specific agent subclasses
├── neural_networks.py # Network architectures
├── cartpole.py # Runnable training + test script
└── lunar_lander.py
Recommended — Poetry (installs all dependencies including PyTorch with CUDA 12.1 support):
git clone https://github.com/Chaoukia/Reinforcement-Learning-course.git
cd Reinforcement-Learning-course
poetry installThen activate the environment:
$(poetry env activate)Alternatively, using pip (installs the CPU-only build of PyTorch — for CUDA support, follow the official instructions after installation):
git clone https://github.com/Chaoukia/Reinforcement-Learning-course.git
cd Reinforcement-Learning-course
pip install -e .All examples follow the same pattern: train an agent, render a test run in the environment, and optionally save a GIF. Scripts live under src/reinforcement_learning_course/<algorithm_family>/examples/.
Every script exposes --n_test (number of test episodes) and --path_gif (save a GIF of the trained agent) arguments. Pass --help to any script to see all available options.
Trains a multi-worker PPO agent on LunarLander-v3. Workers share the actor and critic networks and synchronize gradient updates via a barrier. GAE is used to estimate advantages, which are normalized across workers before each mini-batch update.
python src/reinforcement_learning_course/deep_rl/ppo/examples/lunar_lander.py \
--n_workers 12 \
--t_max 1024 \
--n_train 10000 \
--epochs 4 \
--batch_size 64 \
--epsilon 0.2 \
--lambd 0.98 \
--gamma 0.999 \
--lr_actor 3e-4 \
--lr_critic 3e-4 \
--alpha_entropy 0.01 \
--thresh 250 \
--log_dir runs/ppo_lunarlander \
--path_gif gifs/ppo_lunarlander.gifKey arguments:
| Argument | Default | Description |
|---|---|---|
--n_workers |
12 | Number of parallel worker processes |
--t_max |
1024 | Environment steps collected per iteration per worker |
--epochs |
4 | PPO optimization epochs per collected batch |
--batch_size |
64 | Mini-batch size for gradient updates |
--epsilon |
0.2 | PPO clipping parameter |
--lambd |
0.98 | GAE lambda for advantage estimation |
--thresh |
250 | Mean return threshold for early stopping |
Trains a Deep Q-Network agent on CartPole-v1 with experience replay and a target network. Pass --double_learning yes to switch to Double DQN, which reduces overestimation bias by decoupling action selection from value estimation.
python src/reinforcement_learning_course/deep_rl/dqn/examples/cartpole.py \
--n_train 10000 \
--n_pretrain 64 \
--epsilon_start 1.0 \
--epsilon_stop 0.1 \
--decay_rate 5e-6 \
--n_learn 10 \
--tau 50 \
--batch_size 64 \
--lr 1e-4 \
--gamma 0.99 \
--thresh 400 \
--log_dir runs/dqn_cartpole \
--double_learning no \
--path_gif gifs/dqn_cartpole.gifKey arguments:
| Argument | Default | Description |
|---|---|---|
--n_pretrain |
64 | Episodes of random play to pre-fill the replay buffer |
--n_learn |
10 | Steps between consecutive Q-network updates |
--tau |
50 | Steps between consecutive target network syncs |
--double_learning |
no |
Use Double DQN (yes / no) |
--thresh |
400 | Mean return threshold for early stopping |
Trains a tabular Q-Learning agent on Taxi-v3. The --algorithm flag selects between sarsa, expected_sarsa, q_learning, and double_q_learning.
python src/reinforcement_learning_course/classical_rl/temporal_difference/examples/taxi.py \
--algorithm q_learning \
--alpha 0.1
--n_train 100000 \
--epsilon_start 1.0 \
--epsilon_stop 0.1 \
--decay_rate 1e-4 \
--gamma 0.99 \
--print_iter 1000 \
--n_test 5 \
--verbose yes \
--path_gif gifs/qlearning_taxi.gifKey arguments:
| Argument | Default | Description |
|---|---|---|
--algorithm |
q_learning |
TD algorithm: sarsa, expected_sarsa, q_learning, double_q_learning |
--alpha |
0.1 | Fixed learning rate (omit to use 1/n visit schedule) |
--epsilon_start |
1.0 | Initial exploration rate |
--epsilon_stop |
0.1 | Final exploration rate |
--decay_rate |
1e-4 | Exponential decay rate for epsilon |
python src/reinforcement_learning_course/classical_rl/dynamic_programming/examples/frozen_lake.py \
--map_name 4x4 \
--algorithm policy_iteration \
--gamma 0.99python src/reinforcement_learning_course/classical_rl/monte_carlo/examples/frozen_lake.py \
--map_name 4x4 \
--is_slippery yes \
--first_visit no \
--n_train 100000 \
--epsilon_start 1.0 \
--epsilon_stop 0.1 \
--decay_rate 1e-4python src/reinforcement_learning_course/classical_rl/heuristics/examples/frozen_lake.py \
--map_name 4x4python src/reinforcement_learning_course/deep_rl/actor_critic/examples/lunar_lander.py \
--n_train 10000 \
--t_max 5 \
--lr_policy 1e-4 \
--lr_value 1e-4 \
--alpha_entropy 0.0 \
--thresh 400
--log_dir runs/actor_critic_cartpoleAll deep RL scripts log losses and returns to TensorBoard. Launch the dashboard with:
tensorboard --logdir runs/- Sutton, R. S., & Barto, A. G. (2018). Reinforcement Learning: An Introduction (2nd ed.). MIT Press.
- Mnih, V., et al. (2015). Human-level control through deep reinforcement learning. Nature.
- van Hasselt, H., Guez, A., & Silver, D. (2016). Deep reinforcement learning with double Q-learning. AAAI.
- Williams, R. J. (1992). Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning.
- Mnih, V., et al. (2016). Asynchronous methods for deep reinforcement learning. ICML.
- Schulman, J., et al. (2017). Proximal policy optimization algorithms. arXiv.
- Schulman, J., et al. (2015). High-dimensional continuous control using generalized advantage estimation. ICLR.





