A multi-agent system for academic outreach. Scrapes university faculty pages, evaluates research alignment via LLMs, and queues highly customized cold emails for review before sending.
Scrape (Day 1) Evaluate & Draft (Day 2) Send (Day 2)
┌──────────┐ ┌──────────────┐ ┌─────────────┐ ┌─────────────┐
│ Async │───>│ Evaluator │──>│ Writer │───>│ SMTP │
│ Scraper │ │ Agent │ │ Agent │ │ Relay │
└──────────┘ └──────────────┘ └─────────────┘ └─────────────┘
│ │ │ │
v v v v
raw_html score + reasoning email draft email_sent
in postgres in postgres pending_review in postgres
Outreach state machine:
scraped → parsed → pending_review → approved → email_sent
# 1. Clone and set up
git clone <repo-url> && cd Distributed_Research_Email_Agent
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
# 2. Start PostgreSQL
docker compose up -d
# 3. Configure
cp .env.example .env
# Edit .env — set your Anthropic API key, email, research interests, SMTP creds
# 4. Seed test data (CMU + Duke faculty directories)
python -m scripts.seed
# 5. Run the scraper + agent pipeline
python -m src.main
# 6. Review drafts
python -m scripts.approve
# 7. Send approved emails (dry-run is ON by default)
python -m src.main --pipeline-only --sendsrc/
├── config.py # All settings via RA_* env vars (pydantic-settings)
├── logging_config.py # JSON structured logging with trace IDs
├── main.py # Entry point (--pipeline-only, --send)
├── db/
│ ├── pool.py # asyncpg connection pool
│ ├── schema.py # Idempotent DDL migrations
│ └── models.py # Pydantic models for all DB tables
├── scraper/
│ ├── rate_limiter.py # Per-domain token bucket rate limiter
│ ├── client.py # aiohttp client with retry + backoff
│ └── engine.py # Scraper orchestration loop
├── agents/
│ ├── models.py # Structured output schemas (EvaluationResult, DraftEmail)
│ ├── llm_client.py # Anthropic SDK wrapper with backoff
│ ├── evaluator.py # Research alignment scoring (1-10)
│ ├── writer.py # Personalized email drafting
│ └── pipeline.py # Agent orchestration loop
└── email/
└── sender.py # aiosmtplib relay with dry-run toggle
scripts/
├── seed.py # Insert test universities + scrape targets
└── approve.py # Interactive draft review CLI
All settings are controlled via environment variables prefixed with RA_. See .env.example for the full list.
Key settings:
| Variable | Description | Default |
|---|---|---|
RA_ANTHROPIC_API_KEY |
Anthropic API key | (required) |
RA_ANTHROPIC_MODEL |
Model for agent calls | claude-sonnet-4-6 |
RA_EVALUATOR_SCORE_THRESHOLD |
Min score to draft an email | 8 |
RA_USER_RESEARCH_INTERESTS |
Your interests for alignment matching | (required) |
RA_SMTP_DRY_RUN |
Log emails instead of sending | true |
RA_RATE_LIMIT_TOKENS_PER_SECOND |
Scraper rate limit per domain | 2.0 |
# Run everything (scraper + agents)
python -m src.main
# Agents only (skip scraping)
python -m src.main --pipeline-only
# Agents + auto-send approved emails
python -m src.main --pipeline-only --send
# Review and approve drafts interactively
python -m scripts.approve
# Approve all pending drafts at once
python -m scripts.approve --approve-all
# Re-seed the database
python -m scripts.seed- Python 3.10+ — fully async (
asyncio) - PostgreSQL 16 — normalized schema with enum-based state machines
- asyncpg — connection pooling and raw SQL
- aiohttp — concurrent HTTP scraping with token bucket rate limiting
- Anthropic SDK — structured output via forced tool-use
- aiosmtplib — async SMTP with STARTTLS
- pydantic v2 — strict typing for all data models and config
- BeautifulSoup4 — HTML text extraction
MIT