SignalForge is a local market intelligence pipeline for company research. It ingests SEC EDGAR 10-K filings and approved company source feeds, indexes the content locally with Qdrant, and answers research questions with cited evidence.
It is built for experimenting with retrieval, query planning, and answer generation over a local SQLite database and vector index.
- Downloads and parses SEC 10-K filings by ticker.
- Extracts key 10-K sections such as Business, Risk Factors, MD&A, and Market Risk.
- Discovers candidate company sources such as blogs, newsrooms, investor relations pages, RSS feeds, and Atom feeds.
- Lets you approve, reject, list, or manually add sources before ingestion.
- Ingests approved source articles into generic documents.
- Embeds filing and document chunks into a local Qdrant store.
- Plans research queries with DeepSeek when configured, or a local fallback when not.
- Generates cited answers from retrieved evidence.
- Includes a FastAPI backend and React research console.
- Python 3.11+
uv- Node.js and npm, only needed for the frontend
- SEC EDGAR user-agent details for filing downloads
- Optional: DeepSeek API key for LLM planning and answer generation
The default vector store uses embedded qdrant-client, so no external Qdrant server is required.
Install Python dependencies:
uv sync --extra devCreate a local .env file:
SEC_COMPANY_NAME="Your Name or App Name"
SEC_EMAIL="you@example.com"
# Optional
DEEPSEEK_API_KEY="your_deepseek_api_key"
DEEPSEEK_BASE_URL="https://api.deepseek.com"Without DEEPSEEK_API_KEY, SignalForge still supports local rule-based planning and extractive evidence output.
Ingest a recent 10-K:
uv run python -m signalforge.cli.ingest --ticker NVDA --limit 1Discover candidate company sources. Candidates are saved for review, but they are not ingested automatically:
uv run python -m signalforge.cli.discover_sources --ticker NVDA --website-domain nvidia.comReview candidate, approved, rejected, and manual sources:
uv run python -m signalforge.cli.list_sources --ticker NVDAApprove or reject discovered sources:
uv run python -m signalforge.cli.approve_source --source-id 1
uv run python -m signalforge.cli.reject_source --source-id 2Ingest approved sources:
uv run python -m signalforge.cli.ingest_sources --ticker NVDA --limit-per-source 5Build or refresh the mixed SEC/document vector index:
uv run python -m signalforge.cli.vectorizeRun the background worker loop for approved source ingestion and vectorization:
uv run python -m signalforge.workerAsk a mixed-corpus research question:
uv run python -m signalforge.cli.answer_query \
"What is NVIDIA saying recently about AI infrastructure and supply constraints?" \
--show-plan \
--show-chunksAnswers cite retrieved evidence with normalized labels. SEC evidence is labeled by ticker, filing year, filing item, and chunk. Web document evidence is labeled by source name, publication date when available, and article title.
Source discovery follows this flow:
ticker -> discover candidate sources -> approve or reject sources
-> ingest approved sources -> vectorize documents -> answer with citations
Discovery uses deterministic heuristics such as official-domain matching, common blog/newsroom/investor-relations paths, source subdomains, reachable pages, page titles, and RSS/Atom links.
Manual source registration is available for demos and edge cases:
uv run python -m signalforge.cli.add_source \
--ticker NVDA \
--name "NVIDIA Newsroom" \
--url "https://nvidianews.nvidia.com/" \
--source-type newsroom \
--ownership official \
--trust-level highThe lightweight local mode uses SQLite and embedded Qdrant paths by default:
SIGNALFORGE_DB_PATH=data/signalforge.sqlite3
SIGNALFORGE_QDRANT_PATH=data/qdrantRun the API:
uv run uvicorn signalforge.api:app --reload --port 8000Run the frontend:
cd frontend
npm install
npm run devOpen http://localhost:5173. The frontend calls http://localhost:8000 by default.
You can use the Make targets for the same local workflow:
make api
make frontend
make workerRun one background ingestion/vectorization cycle without starting the infinite worker loop:
make worker-onceInspect local index state from the terminal:
make index-stateStart Postgres and Qdrant through Docker Compose, then run local Python processes against those services:
docker compose up -d postgres qdrant
export SIGNALFORGE_DATABASE_URL=postgresql+psycopg://signalforge:signalforge@localhost:15432/signalforge
export SIGNALFORGE_QDRANT_URL=http://localhost:6333
uv run alembic -c alembic.ini upgrade head
make apiIn another shell, run the frontend:
make frontendRun a one-shot worker cycle against Postgres/Qdrant:
make worker-onceBuild each application component independently:
docker build -f Dockerfile.api -t signalforge-api .
docker build -f Dockerfile.worker -t signalforge-worker .
docker build -f frontend/Dockerfile -t signalforge-frontend ./frontendRun the API image:
docker run --rm -p 8000:8000 \
-e SIGNALFORGE_DATABASE_URL=postgresql+psycopg://user:password@host.docker.internal:5432/signalforge \
-e SIGNALFORGE_QDRANT_URL=http://host.docker.internal:6333 \
signalforge-apiRun the worker image with the same database and Qdrant configuration:
docker run --rm \
-e SIGNALFORGE_DATABASE_URL=postgresql+psycopg://user:password@host.docker.internal:5432/signalforge \
-e SIGNALFORGE_QDRANT_URL=http://host.docker.internal:6333 \
signalforge-workerRun the frontend image and configure the API URL at container startup:
docker run --rm -p 8080:80 \
-e SIGNALFORGE_API_BASE_URL=http://localhost:8000 \
signalforge-frontendOpen http://localhost:8080.
Run the full local stack:
docker compose up --buildThis starts:
- Postgres on
localhost:15432 - Qdrant on
localhost:6333 - API on
http://localhost:8000 - Frontend on
http://localhost:8080 - Worker background ingestion/vectorization loop
The stack uses named Docker volumes for Postgres data, Qdrant data, processed article text, and raw SEC filing downloads. Stop the stack while keeping data:
docker compose downStop the stack and delete persisted local Docker data:
docker compose down -vCopy .env.example to .env if you want to override ports, credentials, models,
worker settings, or optional DeepSeek/SEC values. Docker Compose reads .env
automatically.
Useful checks:
docker compose ps
curl http://localhost:8000/health
curl http://localhost:8000/api/index
docker compose logs -f api
docker compose logs -f workerEquivalent Make targets:
make compose-up
make compose-logs
make compose-downMove an existing local SQLite database into Postgres with a dry run first:
uv run python -m signalforge.cli.migrate_sqlite_to_postgres \
--sqlite-path data/signalforge.sqlite3 \
--postgres-url postgresql+psycopg://signalforge:signalforge@localhost:15432/signalforgeIf the dry run reports the expected row counts, run the migration:
uv run python -m signalforge.cli.migrate_sqlite_to_postgres \
--sqlite-path data/signalforge.sqlite3 \
--postgres-url postgresql+psycopg://signalforge:signalforge@localhost:15432/signalforge \
--executeThe migration preserves IDs and timestamps, copies document metadata into Postgres
JSONB, validates row counts after import, and resets Postgres sequences. The
target database must be empty unless you pass --replace, which deletes existing
target rows before importing:
uv run python -m signalforge.cli.migrate_sqlite_to_postgres \
--sqlite-path data/signalforge.sqlite3 \
--postgres-url postgresql+psycopg://signalforge:signalforge@localhost:15432/signalforge \
--execute \
--replace# Run database migrations for the configured database
uv run alembic -c alembic.ini upgrade head
# Ingest an existing local SEC download without hitting EDGAR
uv run python -m signalforge.cli.ingest --ticker NVDA --no-download
# Run one approved-source ingestion pass
uv run python -m signalforge.cli.ingest_sources
# Run one worker cycle: approved-source ingestion plus vectorization
uv run python -m signalforge.cli.run_worker_once
# Semantic search over indexed chunks
uv run python -m signalforge.cli.search "AI infrastructure risks" --ticker NVDA --section 1A
# Discover sources without saving candidates
uv run python -m signalforge.cli.discover_sources --ticker NVDA --website-domain nvidia.com --dry-run
# Inspect a query plan
uv run python -m signalforge.cli.plan_query "Compare NVDA and MSFT risk factors"
# Inspect database/index state
uv run python -m signalforge.cli.index_stateThe same operations are available as Make targets:
make migrate
make worker-once
make vectorize
make index-stateGET /api/index returns the local index state used by the frontend:
- Indexed filing coverage by ticker and section.
- Approved source count.
- Candidate source count.
- Web document count.
- Per-source document counts.
- Last ingestion status and completion time.
Default local paths:
- SQLite database:
data/signalforge.sqlite3 - Qdrant store:
data/qdrant - Raw filings:
data/raw - Processed filing and article text:
data/processed
Default models:
- Embeddings:
jinaai/jina-embeddings-v2-small-en - Planner:
deepseek-v4-flash - Answer generator:
deepseek-v4-flash
The API supports these environment overrides:
SIGNALFORGE_DATABASE_URLSIGNALFORGE_DB_PATHSIGNALFORGE_QDRANT_URLSIGNALFORGE_QDRANT_PATHSIGNALFORGE_COLLECTIONSIGNALFORGE_EMBEDDING_MODELSIGNALFORGE_PLANNER_MODELSIGNALFORGE_ANSWER_MODELSIGNALFORGE_CORS_ORIGINS
SIGNALFORGE_CORS_ORIGINS is a comma-separated browser origin allowlist. By default,
the API allows the local Vite dev server and a frontend container mapped to port 8080.
Worker-specific overrides:
SIGNALFORGE_WORKER_INTERVAL_SECONDSSIGNALFORGE_INGEST_LIMIT_PER_SOURCESIGNALFORGE_ENABLE_SCHEDULED_INGESTIONSIGNALFORGE_PROCESSED_DIRSIGNALFORGE_CHUNK_SIZESIGNALFORGE_CHUNK_OVERLAPSIGNALFORGE_VECTORIZE_BATCH_SIZESIGNALFORGE_LOG_LEVEL
Docker Compose-specific overrides are documented in .env.example and include:
POSTGRES_USERPOSTGRES_PASSWORDPOSTGRES_DBPOSTGRES_PORTQDRANT_PORTAPI_PORTFRONTEND_PORTSIGNALFORGE_API_BASE_URLSIGNALFORGE_IMAGE_TAG
SIGNALFORGE_DATABASE_URL takes precedence over SIGNALFORGE_DB_PATH.
SIGNALFORGE_QDRANT_URL takes precedence over SIGNALFORGE_QDRANT_PATH.
For Postgres, use a URL such as:
SIGNALFORGE_DATABASE_URL=postgresql+psycopg://user:password@localhost:5432/signalforgeFor Qdrant server mode, use:
SIGNALFORGE_QDRANT_URL=http://localhost:6333Set a custom frontend API URL with:
VITE_API_BASE_URL=http://localhost:8000 npm run devThe frontend Docker image reads SIGNALFORGE_API_BASE_URL at container startup.
The worker runs this cycle repeatedly:
load approved enabled sources -> ingest new source documents
-> chunk documents -> vectorize pending SEC/document chunks -> sleep
It uses SIGNALFORGE_WORKER_INTERVAL_SECONDS between cycles. Set
SIGNALFORGE_ENABLE_SCHEDULED_INGESTION=false to skip source ingestion and only
vectorize pending chunks. Use SIGNALFORGE_INGEST_LIMIT_PER_SOURCE to cap article
ingestion per source during demos or testing.
For operational checks, prefer a one-shot cycle before starting the loop:
uv run python -m signalforge.cli.run_worker_onceFor the Docker Compose Postgres database, create a compressed backup:
make backup-postgresThis writes:
backups/signalforge.dump
Restore that backup into the Compose Postgres service:
make restore-postgresQdrant data is stored in the signalforge_qdrant-data Docker volume. For a
simple local backup, stop the stack and archive the named volume with Docker or
your host backup tooling. For SQLite local development, back up
data/signalforge.sqlite3, data/qdrant, data/raw, and data/processed.
uv run pytest
cd frontend
npm testRun the optional Postgres persistence profile against a disposable database
whose name contains test:
SIGNALFORGE_POSTGRES_TEST_DATABASE_URL=postgresql+psycopg://user:password@localhost:5432/signalforge_test \
uv run pytest -m postgresSignalForge is intended for local research workflows. Generated answers are grounded in retrieved filing and document chunks, but they are not financial advice. Verify important claims against the original SEC filing or source document.