This project sets up a local Apache Spark cluster using Docker Compose, with Jupyter Lab for interactive data analysis using PySpark. Everything is controlled by a centralized configuration file.
TLDR: make setup, make build, make up
- ποΈ Centralized Configuration: Control everything from
config/config.yaml - β‘ UV Package Manager: Fast Python package installation
- π§ Configurable Resources: Easily adjust memory, cores, and workers
- π¦ Custom Packages: Add your own Python packages via config
- ποΈ Easy Management: Simple commands via Makefile
- π Organized Structure: Clean folder organization
The setup includes:
- Spark Master: Coordinates the cluster and provides the Spark UI
- Spark Worker: Executes Spark applications (scalable)
- Jupyter Lab: Interactive notebook environment with PySpark
All components use Spark 3.5.0 to ensure version compatibility.
- Docker and Docker Compose installed
- At least 4GB of RAM available for containers
- Python 3.x (for configuration management)
Edit config/config.yaml to customize your environment:
# Spark Cluster Configuration
spark:
version: "3.5.0"
worker:
memory: "4g" # Increase worker memory
cores: "4" # Increase worker cores
instances: 2 # Run 2 workers
# Python Package Configuration
python:
packages:
- "pandas>=2.0.0"
- "your-custom-package>=1.0.0" # Add your packages here# Build and start everything
make up
# Or manually:
docker compose -f docker/docker-compose.yml --env-file config/.env up -d- Jupyter Lab: http://localhost:8888 (no password)
- Spark Master UI: http://localhost:8080
- Spark Worker UI: http://localhost:8081
- Open the ipynb file in VSCode
- On the top right corner, click "Select Kernel"
- Select "Kernel > Existing Jupyter Server"
- Enter "http://localhost:8888". Refer to file config/config.yaml under containers > jupyter > port
- Click "Select"
- You can now run the notebook
local_spark/
βββ config/ # ποΈ Configuration files
β βββ config.yaml # Centralized configuration (SINGLE SOURCE OF TRUTH)
β βββ .env # Environment variables (auto-generated)
βββ docker/ # π³ Docker files
β βββ Dockerfile # Custom Jupyter image with UV
β βββ docker-compose.yml # Container orchestration (auto-generated)
βββ scripts/ # π οΈ Automation scripts
β βββ setup.py # Configuration generator
βββ notebooks/ # π Your Jupyter notebooks
βββ data/ # π Data files
βββ Makefile # π― Management commands
βββ README.md # π This file
Use the Makefile for easy management:
make help # Show all commands
make setup # Generate files from config/config.yaml
make up # Start the cluster
make down # Stop the cluster
make restart # Restart everything
make logs # Show container logs
make status # Show container status
make clean # Clean up everything
make rebuild # Complete rebuildAll Python packages are defined in config/config.yaml - no separate requirements.txt needed!
python:
packages:
- "tensorflow>=2.13.0"
- "torch>=2.0.0"
- "transformers>=4.30.0"
- "your-package>=1.0.0"The Dockerfile reads packages directly from config/config.yaml during build.
Edit config/config.yaml to adjust resources:
spark:
worker:
memory: "4g" # Worker memory
cores: "4" # Worker cores
instances: 3 # Number of workers
executor:
memory: "2g" # Executor memory per job
cores: "2" # Executor cores per jobScale workers dynamically:
# Using make
make scale-workers
# Or directly with docker compose
docker compose -f docker/docker-compose.yml --env-file config/.env up -d --scale spark-worker=5Change ports in config/config.yaml:
containers:
jupyter:
port: 9999
spark_master:
ui_port: 9090
port: 7077Customize data and notebook paths:
volumes:
notebooks_path: "./my-notebooks"
data_path: "./my-data"Change Python version:
python:
version: "3.12" # Use Python 3.12In your notebook cells:
from pyspark.sql import SparkSession
spark = SparkSession.builder \
.appName("My Spark App") \
.master("spark://spark-master:7077") \
.config("spark.executor.memory", "2g") \
.config("spark.executor.cores", "2") \
.getOrCreate()See the notebooks/spark_example.ipynb for examples of:
- Creating Spark sessions
- Loading and manipulating data
- Performing aggregations
- Saving data as Parquet files
- Edit
config/config.yaml - Regenerate files:
make setup - Rebuild:
make rebuild
make config # Show current configuration
make logs # Check container logsReduce memory in config/config.yaml:
spark:
worker:
memory: "1g"
executor:
memory: "512m"Check UV installation:
docker compose -f docker/docker-compose.yml --env-file config/.env exec jupyter uv --versionAll containers use versions from config/config.yaml. Ensure compatibility:
spark:
version: "3.5.0"
python:
packages:
- "pyspark[sql]==3.5.0" # Match Spark version- Scale workers: Increase
worker.instancesin config - Optimize memory: Adjust
worker.memoryandexecutor.memory - Partition data: Use
.repartition()in your code - Monitor: Check Spark UI at http://localhost:8080
For production environments:
- Enable Spark security in config
- Use persistent volumes
- Set resource limits
- Configure monitoring
- Use external data sources (S3, HDFS)
- Check logs:
make logs - Verify config:
make config - Clean rebuild:
make rebuild - Check resources: Ensure Docker has enough memory allocated