Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Source Code: Evolutionary Algorithms in Reinforcement Learning

This repository contains the code used in the experiments presented in the Master's thesis "Evolutionary Algorithms in Reinforcement Learning" by Sabína Gulčíková
Faculty of Information Technology, Brno University of Technology, 2025
Email: xgulci00@stud.fit.vutbr.cz

Code Structure

dip/
├──cluster/
    ├── python_imenv.sif            ---------- # pre-built singularity image with required packages
    ├── run_batch_experiments.sh    ---------- # script for submitting a batch of experiments to the cluster
    └── run_experiments.sh          ---------- # script for submitting a single experiment (includes CPU time, nodes, etc.)
├──results/                        ---------- # folder containing plotted results 
├──src/ 
    ├── agents/  ---------- # DQN and PPO agent implementations 
    ├── envs/    ---------- # Base and custom CartPole environments 
    ├── evolution/  ------- # Evolution framework using genetic programming 
    ├── experiments/  ----- # Scripts to run experiments and baselines 
    ├── random_search/  --- # Random search module for reward shaping 
    ├── utils/   ---------- # Configs, plotting, argument parsing, logging 
    ├── config.py --------- # Central config file for all experiments 
    └── main.py  ---------- # Entry point for running experiments 
├──requirements.txt
└──README.md

Attribution for Agents Implementation

The source code in src/agents/ was developed based on the structure and educational materials provided in Callum McDougal's course on AI Safety, including section on Reinforcement Learning. Specifically, boilerplate for methods and classes, such as their signatures, and certain helper methods were adapted from the course material. However, all algorithmic logic, fitness functions, and transformations of environments were independently written by the author, using the course as a conceptual guide.

Each file that includes adapted boilerplate contains an attribution comment at the top, clearly stating the origin of the code.


⚙️ Setup & Running the Code

Set your current working directory to the project root (dip/).

Option A: virtual environment

This project requires Python 3.12. All necessary dependencies are listed in the requirements.txt file. You can install them by running:

# Create a virtual environment
python3.12 -m venv venv

# Activate the virtual environment
source venv/bin/activate  # On macOS/Linux

# Install dependencies
pip install -r requirements.txt

Next:

  1. Ensure your current working directory is still the project root (dip/)
  2. Make sure you've activated a virtual environment using Python 3.12.
  3. After installing the requirements, run the main script using one of the following commands, depending on your system setup:
python -m src.main
# or
python3 -m src.main

Option B: singularity image

If you prefer not to install dependencies manually, the cluster/ folder contains a pre-built Singularity image (python_imenv.sif) with all required packages. You can run the project in an isolated environment using:

singularity exec cluster/python_imenv.sif python -m src.main

⚙️ src/config.py

This file serves as the central configuration module for the entire project. It defines all parameters used in training, evaluation, and evolutionary processes through structured Python dataclasses.

Core Structure

The main configuration entry point is the ExperimentConfig class. It aggregates and coordinates sub-configurations for different experiment types:

  • ppo baseline: PPOArgs – Configuration for training PPO agents
  • dqn baseline: DQNArgs – Configuration for training DQN agents
  • evolution: EvolutionArgs – Controls evolutionary search and genetic programming settings
  • random search: RandomSearchArgs – Parameters for running random search over reward functions
  • manual shaping: ManualShapingArgs – Contains a manually defined reward shaping expression. This mode is also useful for reproducibility, as the shaping_fn can be replaced by custom expression, and the user can examine its impact onto training.

Individual Configuration Blocks

  • ExperimentConfig: Top-level toggles for experiment selection, agent/environment setup, reproducibility, logging, and shared parameters. It also ensures nested sub-configs inherit key arguments such as verbosity and agent type.

  • Baseline: It is recommended not to tweak these RL-specific parameters (PPOArgs and DQNArgs dataclasses), as they are very sensitive to hyperparameters and their changing might result in agent's inability to learn.

    • PPOArgs: Defines all PPO training hyperparameters, such as rollout lengths, optimizer settings, and reward shaping options.
    • DQNArgs: Analogous to PPOArgs but tailored for DQN, including replay buffer size, exploration schedule, and training frequency.
  • EvolutionArgs: Governs all aspects of evolutionary search: population control, mutation/crossover rates, selection strategy, logging, and the evolutionary algorithm variant (e.g. (μ + λ)).

  • RandomSearchArgs: Minimal configuration to run a large number of randomly sampled reward trees with a depth cap.

  • ManualShapingArgs: Provides a manually specified reward function string, allowing for deterministic shaping independent of evolution or random search.

Use this file to tweak any aspect of the experiment without modifying src/main.py or providing CLI arguments. When src/main.py is executed without command-line arguments, all values from src/config.py are used automatically.

⚠️ Command Line Arguments

Optionally, one can provide command-line arguments. These were primarily intended for running automated batch scripts, but can be used manually as well. If no arguments are provided, the script defaults to running evolution with the PPO agent, unless otherwise specified in the src/config.py file.

For quick execution, the following two can be used:

Common Options

  • --mode: Selects the experiment mode. Multiple of these can be used, separated by space. Possible values:

    • evolution
    • random
    • manual
    • baseline
  • --agent_type: Selects the RL agent. Possible values:

    • ppo
    • dqn

Example usage

# defaults to evolutionary reward search with ppo
python -m src.main

# applies manual reward specified in config.py on dqn
python -m src.main --mode manual --agent_type dqn

# evolution and random rewad search on ppo, run sequentially
python -m src.main --mode evolution random --agent_type ppo

⚠️ Logging & Visualization

Results and reward evolution are logged using Weights & Biases (wandb). To enable logging:

  1. Set your WANDB API key in the config.py, or use the one provided by the author. During training, the stdout will show a link to access the visualization. In order to view full configuration of run experiment, check the "ℹ️ Overview" tab in the wandb UI.

In case the project is not public, wandb restriction might be in place. Please email the author to be given the permission to view the project in wandb.


🖥️ HPC (Metacentrum) Usage Instructions

Experiments were executed on Metacentrum using job submission scripts and a Singularity image file (.sif), located in the dip/cluster directory.

To run the project on the Metacentrum cloud infrastructure, these files must be moved outside the dip/ folder, into the same level as the dip/ project directory. Your directory structure should look like this:

root/
├── dip/
    ├── src/
    └── results/
├── python_imenv.sif
├── run_batch_experiments.sh
└── run_experiments.sh

Submitting jobs

  1. Modify configuration values in run_batch_experiments.sh (e.g. experiment selection, parallelization, agent type).
  2. Make sure your current working directory is root/.
  3. Submit the batch job:
./run_batch_experiments.sh

💡 Note: Ensure that the .sif file and both run_*.sh scripts are executable and accessible at the root level, otherwise job submission may fail. You might need to run permission-adjusting commands before execution, such as:

chmod +x ./run_experiments.sh
chmod +x ./run_batch_experiments.sh

About

Master's thesis. Using genetic programming to evolve shaping rewards in CartPole environment.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages