This repository contains the code used in the experiments presented in the Master's thesis "Evolutionary Algorithms in Reinforcement Learning" by Sabína Gulčíková
Faculty of Information Technology, Brno University of Technology, 2025
Email: xgulci00@stud.fit.vutbr.cz
dip/
├──cluster/
├── python_imenv.sif ---------- # pre-built singularity image with required packages
├── run_batch_experiments.sh ---------- # script for submitting a batch of experiments to the cluster
└── run_experiments.sh ---------- # script for submitting a single experiment (includes CPU time, nodes, etc.)
├──results/ ---------- # folder containing plotted results
├──src/
├── agents/ ---------- # DQN and PPO agent implementations
├── envs/ ---------- # Base and custom CartPole environments
├── evolution/ ------- # Evolution framework using genetic programming
├── experiments/ ----- # Scripts to run experiments and baselines
├── random_search/ --- # Random search module for reward shaping
├── utils/ ---------- # Configs, plotting, argument parsing, logging
├── config.py --------- # Central config file for all experiments
└── main.py ---------- # Entry point for running experiments
├──requirements.txt
└──README.md
The source code in src/agents/ was developed based on the structure and educational materials
provided in Callum McDougal's course on AI Safety, including section on Reinforcement Learning. Specifically,
boilerplate for methods and classes, such as their signatures, and certain helper methods were adapted from the course material.
However, all algorithmic logic, fitness functions, and transformations of environments were independently
written by the author, using the course as a conceptual guide.
Each file that includes adapted boilerplate contains an attribution comment at the top, clearly stating the origin of the code.
Set your current working directory to the project root (dip/).
This project requires Python 3.12. All necessary dependencies are listed in the requirements.txt file. You can install them by running:
# Create a virtual environment
python3.12 -m venv venv
# Activate the virtual environment
source venv/bin/activate # On macOS/Linux
# Install dependencies
pip install -r requirements.txtNext:
- Ensure your current working directory is still the project root (
dip/) - Make sure you've activated a virtual environment using Python 3.12.
- After installing the requirements, run the main script using one of the following commands, depending on your system setup:
python -m src.main
# or
python3 -m src.mainIf you prefer not to install dependencies manually, the cluster/ folder contains a pre-built Singularity image
(python_imenv.sif) with all required packages. You can run the project in an isolated environment using:
singularity exec cluster/python_imenv.sif python -m src.mainThis file serves as the central configuration module for the entire project. It defines all parameters used in training, evaluation, and evolutionary processes through structured Python dataclasses.
The main configuration entry point is the ExperimentConfig class. It aggregates and coordinates sub-configurations for different experiment types:
- ppo baseline:
PPOArgs– Configuration for training PPO agents - dqn baseline:
DQNArgs– Configuration for training DQN agents - evolution:
EvolutionArgs– Controls evolutionary search and genetic programming settings - random search:
RandomSearchArgs– Parameters for running random search over reward functions - manual shaping:
ManualShapingArgs– Contains a manually defined reward shaping expression. This mode is also useful for reproducibility, as theshaping_fncan be replaced by custom expression, and the user can examine its impact onto training.
-
ExperimentConfig: Top-level toggles for experiment selection, agent/environment setup, reproducibility, logging, and shared parameters. It also ensures nested sub-configs inherit key arguments such as verbosity and agent type. -
Baseline: It is recommended not to tweak these RL-specific parameters (PPOArgs and DQNArgs dataclasses), as they are very sensitive to hyperparameters and their changing might result in agent's inability to learn.PPOArgs: Defines all PPO training hyperparameters, such as rollout lengths, optimizer settings, and reward shaping options.DQNArgs: Analogous to PPOArgs but tailored for DQN, including replay buffer size, exploration schedule, and training frequency.
-
EvolutionArgs: Governs all aspects of evolutionary search: population control, mutation/crossover rates, selection strategy, logging, and the evolutionary algorithm variant (e.g. (μ + λ)). -
RandomSearchArgs: Minimal configuration to run a large number of randomly sampled reward trees with a depth cap. -
ManualShapingArgs: Provides a manually specified reward function string, allowing for deterministic shaping independent of evolution or random search.
Use this file to tweak any aspect of the experiment without modifying src/main.py or providing CLI arguments.
When src/main.py is executed without command-line arguments, all values from src/config.py are used automatically.
Optionally, one can provide command-line arguments. These were primarily intended for running automated batch scripts, but can be used manually as well.
If no arguments are provided, the script defaults to running evolution with the PPO agent, unless otherwise specified in the src/config.py file.
For quick execution, the following two can be used:
-
--mode: Selects the experiment mode. Multiple of these can be used, separated by space. Possible values:evolutionrandommanualbaseline
-
--agent_type: Selects the RL agent. Possible values:ppodqn
# defaults to evolutionary reward search with ppo
python -m src.main
# applies manual reward specified in config.py on dqn
python -m src.main --mode manual --agent_type dqn
# evolution and random rewad search on ppo, run sequentially
python -m src.main --mode evolution random --agent_type ppoResults and reward evolution are logged using Weights & Biases (wandb). To enable logging:
- Set your
WANDB API keyin theconfig.py, or use the one provided by the author. During training, the stdout will show a link to access the visualization. In order to view full configuration of run experiment, check the "ℹ️ Overview" tab in the wandb UI.
In case the project is not public, wandb restriction might be in place. Please email the author to be given the permission to view the project in wandb.
Experiments were executed on Metacentrum using job submission scripts and a Singularity image file (.sif), located in the dip/cluster directory.
To run the project on the Metacentrum cloud infrastructure, these files must be moved outside the dip/ folder, into the same level as the dip/ project directory. Your directory structure should look like this:
root/
├── dip/
├── src/
└── results/
├── python_imenv.sif
├── run_batch_experiments.sh
└── run_experiments.sh
- Modify configuration values in run_batch_experiments.sh (e.g. experiment selection, parallelization, agent type).
- Make sure your current working directory is root/.
- Submit the batch job:
./run_batch_experiments.sh
💡 Note: Ensure that the .sif file and both run_*.sh scripts are executable and accessible at the root level, otherwise job submission may fail. You might need to run permission-adjusting commands before execution, such as:
chmod +x ./run_experiments.sh
chmod +x ./run_batch_experiments.sh