Skip to content

About

No description, website, or topics provided.

Resources

Contributing

Stars

12 stars

Watchers

0 watching

Forks

 
 

Latest commit

 

History

279 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Docker Auto-Heal Service

Unit Tests Test Coverage Mutation score Docker Hub Pulls GHCR Pulls Docker Image Size GitHub Release License: MIT

Docker Auto-Heal watches your Docker containers and restarts the ones that fail or go unhealthy, so you don't have to. It runs as a single container with a web dashboard, a REST API, and Prometheus metrics.

Why use it

  • Automatic recovery — containers that exit with an error or fail their health check are restarted for you, with cooldowns and exponential backoff so a broken container doesn't restart in a tight loop.
  • Quarantine for flapping containers — a container that keeps failing past a configurable threshold is quarantined (auto-healing paused for it) and automatically un-quarantined once it recovers.
  • No YAML to hand-write — enable monitoring with a single Docker label, or manage it entirely from the web UI.
  • Custom health checks — beyond Docker's native HEALTHCHECK, you can define HTTP, TCP, or exec-based checks per container.
  • Observability built in — a Prometheus metrics endpoint, an event log of every restart/quarantine decision, and optional notifications to Discord, Slack, Telegram, ntfy, Gotify, Pushover, or a generic webhook.

Project status

Docker Auto-Heal is under active development. The core monitoring/restart engine, web UI, REST API, and notification system are implemented and covered by an automated test suite (see Unit Tests workflow). Configuration is stored on disk in /data and managed through the web UI or the REST API — there are currently no environment-variable settings for tuning monitoring behavior (see Configuration).

Quick Start

Run the service with the Docker socket mounted so it can see and manage your containers:

docker run -d \
  --name docker-autoheal \
  -v /var/run/docker.sock:/var/run/docker.sock:ro \
  -v ./data:/data \
  -p 3131:3131 \
  -p 9090:9090 \
  --restart unless-stopped \
  tommye123/docker-autoheal:latest

Images are also published to GitHub Container Registry as ghcr.io/tommye123/docker-autoheal:latest if you prefer GHCR over Docker Hub.

Or with Docker Compose:

services:
  autoheal:
    image: tommye123/docker-autoheal:latest
    container_name: docker-autoheal
    restart: unless-stopped
    volumes:
      - /var/run/docker.sock:/var/run/docker.sock:ro
      - ./data:/data
    ports:
      - "3131:3131"  # Web UI
      - "9090:9090"  # Prometheus metrics
    labels:
      - "autoheal=false"  # keeps the monitor unselected under the default label-based
                          # selection; see docs/user/labels.md if you enable include_all
docker compose up -d

Open http://localhost:3131 — that's the dashboard. Docker Auto-Heal is now running and ready to monitor containers.

For a from-scratch walkthrough (requirements, verifying the socket mount, testing with a sample container), see Installation.

Enabling Auto-Healing for a Container

Add the autoheal label to any container you want monitored:

services:
  webapp:
    image: nginx:alpine
    labels:
      autoheal: "true"
    healthcheck:
      test: ["CMD", "wget", "--quiet", "--tries=1", "--spider", "http://localhost"]
      interval: 30s
      timeout: 10s
      retries: 3

Docker Auto-Heal picks up labelled containers automatically — both containers that are already running when it starts, and new ones as they start. A container's Docker healthcheck (if it has one) determines when Auto-Heal considers it unhealthy; without a health check, Auto-Heal still restarts the container if it exits with a non-zero code. You can also select containers to monitor from the web UI without adding a label at all.

See Labels and Health Checks for the full picture, including restart policy, cooldowns, and quarantine behavior.

Web UI

The dashboard at http://<host>:3131 shows every container's status, lets you toggle auto-healing per container, view the event log, edit configuration, and manage custom health checks and notifications — see Usage.

Interactive API documentation (Swagger UI) is available at http://<host>:3131/docs.

Metrics

Prometheus metrics are exposed on a separate port (9090 by default) at /metrics:

curl http://localhost:9090/metrics

See Monitoring & Metrics for the full metric list.

Troubleshooting

Container not being monitored? Check it has the autoheal=true label (unless you enabled "monitor all containers"), isn't in the excluded list, and check the Events tab in the UI for the monitoring decision.

Service won't start? Confirm the Docker socket is mounted and readable, and that ports 3131 and 9090 are free.

For more, see the full Troubleshooting guide.

Documentation

For users For developers For maintainers
Installation Development setup Dependency management
Configuration Architecture Publishing
Usage Project structure Release process
Labels Frontend development
Health checks Testing
Monitoring & metrics
Notifications
Maintenance mode
Troubleshooting
Migration

The full documentation index, including historical/superseded documents, is in docs/index.md.

Contributing

Contributions are welcome. See CONTRIBUTING.md to get set up, and docs/developer/ for architecture and project structure.

License

MIT

This project was originally created by @satya-sovan. Thank you to the original author for making the project available as open source.

About

No description, website, or topics provided.

Resources

Contributing

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages