AI-powered code reviewer with build validation for ANY codebase
- Reviews code hierarchically: Directory → File → Function
- Makes fixes automatically: Edits source files based on analysis
- Validates with build: Runs your project's build command to test changes
- Iterates until success: If build fails, fixes errors and rebuilds
- Commits working code: Only commits when build succeeds
- ✅ Generic: Works with any language/build system
- ✅ Scalable: Function-by-function chunking handles files of any size
- ✅ Safe: Tests every change with your build system
- ✅ Autonomous: Runs for hours reviewing entire directories
- ✅ Smart: Works with any OpenAI-compatible LLM provider (vLLM, TokenHub, OpenAI, etc.)
- ✅ Parallel: Optional concurrent file processing for faster reviews
- ✅ Self-Healing: Auto-detects loops, learns from build failures
- ✅ Configurable: Multiple agent personalities via Oracle Agent Spec
| Language/Project | Build Command | Example |
|---|---|---|
| C/C++ (Make) | make -j$(nproc) |
Linux kernel |
| C/C++ (CMake) | cmake --build build -j$(nproc) |
LLVM, Qt |
| FreeBSD | sudo make -j$(sysctl -n hw.ncpu) buildworld |
FreeBSD source |
| Rust | cargo build --release |
Rustc, ripgrep |
| Go | go build ./... |
Kubernetes, Docker |
| Python | python -m pytest |
Django, Flask |
| Node.js | npm test |
React, Vue |
Any project with a build/test command works!
make check-depsYou need at least one running OpenAI-compatible LLM server. Any of these work:
| Provider | Example URL | Notes |
|---|---|---|
| vLLM | http://localhost:8000 |
GPU inference, no key needed |
| TokenHub | http://localhost:8090 |
Multi-provider router, optional key |
| OpenAI | https://api.openai.com |
Requires API key |
| Ollama (OpenAI mode) | http://localhost:11434 |
Local models |
| llama.cpp server | http://localhost:8080 |
CPU/GPU inference |
If you use TokenHub, there's a convenience launcher: make tokenhub-start
make config-initOr manually:
cp config.yaml.sample config.yaml
vim config.yamlMinimal config.yaml:
llm:
providers:
- url: "http://localhost:8000" # local vLLM — tried first
- url: "http://my-server:8090"
api_key: "my-api-key" # failover with auth
- url: "https://api.openai.com"
api_key: "sk-..." # cloud fallback
source:
root: ".."
build_command: "make -j$(nproc)" # YOUR build command
build_timeout: 600
review:
workflow: "review" # review or rewrite
persona: "personas/freebsd-angry-ai" # Choose your agentmake runAgents define the AI's review personality and focus. Each agent is configured in Oracle Agent Spec format.
| Agent | Focus | Best For |
|---|---|---|
| freebsd-angry-ai | Security, style(9), POSIX | Production audits (default) |
| freebsd-rust-rewriter | C/C++ to Rust, build integration | FreeBSD rewrite workflow |
| security-hawk | Vulnerabilities, exploits | Security-critical code |
| performance-cop | Speed, algorithms, cache | Performance optimization |
| friendly-mentor | Learning, best practices | Training, onboarding |
| example | Balanced, educational | General code review |
Each agent is defined in personas/<name>/agent.yaml:
component_type: Agent
agentspec_version: "26.1.0"
name: "Security Hawk"
description: "Paranoid security auditor"
inputs:
- title: "codebase_path"
type: "string"
outputs:
- title: "review_summary"
type: "string"
- title: "critical_vulnerabilities"
type: "integer"
system_prompt: |
You are a paranoid security auditor...
llm_config:
component_type: OpenAiCompatibleConfig
name: "{{llm_name}}"
url: "{{llm_url}}"
model_id: "{{model_id}}"# Copy an existing agent
cp -r personas/example personas/my-agent
# Edit the agent configuration
vim personas/my-agent/agent.yaml
# Validate
python3 persona_validator.py personas/my-agent
# Use in config.yaml
review:
persona: "personas/my-agent"The runner has two workflow modes. Personas still provide behavior and taste, but the workflow mode controls the system prompt, progress index, summary file, and success criteria.
| Mode | Metadata | Use For |
|---|---|---|
review |
REVIEW-INDEX.md, REVIEW-SUMMARY.md |
Defect-finding, fixes, audits |
rewrite |
REWRITE-INDEX.md, REWRITE-SUMMARY.md |
Translation, refactors, API migrations, decomposition, hardening rewrites |
Rewrite mode is broader than translation. Configure it with an objective and constraints:
review:
workflow: "rewrite"
persona: "personas/friendly-mentor"
rewrite:
preflight_build: false
selection_policy: "small_first" # Use "bottom_up" for normal long runs
objective: "Refactor selected modules without changing public behavior."
strategy: "Complete one small directory at a time."
output_policy: "Modify the source tree directly and preserve external interfaces."
constraints:
- "Preserve CLI behavior, APIs, file formats, and exit statuses."
success_criteria:
- "The configured build command succeeds."
- "Existing tests still pass."
contract:
required_changed_files:
any: ["**/*.c", "**/*.h", "**/*.py", "**/*.go", "**/*.rs"]
commands:
- name: "unit tests"
command: "{build_command}"
timeout: 300Rewrite indexes are work-unit graphs, not just flat directory checklists. The engine records generic directory units and the files directly related to each unit. Project-specific meaning comes from the selected persona and rewrite configuration, not hard-coded language or OS rules. Each unit can carry:
kind- generic directory/test metadatastage- generic ordering metadatadepends_on- earlier units that should be completed firstfiles- related source, test, manifest, and build filesbuild_command/test_command- optional validation commands supplied by metadata
Rewrite mode processes these units bottom-up by stage and dependency by default.
For smoke/e2e runs, set review.rewrite.selection_policy: "small_first" to
prefer smaller source-backed units first. SET_SCOPE shows the selected unit
metadata, and BUILD runs the configured project build command unless unit
metadata explicitly supplies a narrower command.
Rewrite contracts are project-generic. After the configured build command succeeds,
the runner can require changed-file globs, required artifact paths, build-output
markers, and arbitrary validation commands. Contract command templates can use
variables such as {unit}, {unit_dir}, {source_root}, and {build_command}.
Language-specific rewrite requirements belong in the persona and
review.rewrite.contract, not in runner code.
Personas can ship default rewrite contract checks in agent.yaml metadata, and
config.yaml contract entries are merged with those defaults. For example, the
FreeBSD Rust rewriter validates that the active unit still supports the standard
cleandir make lifecycle target after Makefile changes.
Source Tree (entire codebase)
└─ Rewrite Unit (directory or configured source unit) ← BUILD + COMMIT HERE
└─ Related files (source, tests, manifests, build glue)
└─ File
└─ Chunks (individual functions)
- SET_SCOPE - Select directory/work-unit key to work
- READ_FILE - Inspect files (chunked if large)
- EDIT_FILE / WRITE_FILE - Apply fixes or rewrites
- BUILD - Validate changes with the unit or configured build command
- Iterate - If build fails, fix and rebuild
- Commit - When build succeeds, commit changes
- Next - Move to next work unit
Files over 400 lines are automatically chunked by function:
- Reviews one function at a time
- No timeouts or memory issues
- Handles files of any size
┌─────────────────────────────────────────────────────┐
│ ai-code-reviewer/ │
│ ├─ reviewer.py (main loop) │
│ ├─ llm_client.py (OpenAI-compat LLM client) │
│ ├─ persona_validator.py (Agent Spec validation) │
│ ├─ personas/ (agent configurations) │
│ │ ├─ freebsd-angry-ai/agent.yaml │
│ │ ├─ security-hawk/agent.yaml │
│ │ └─ ... │
│ └─ config.yaml (your configuration) │
└─────────────────────────────────────────────────────┘
↓ HTTP /v1/chat/completions ↓ subprocess
LLM Provider(s) (your build command)
├─ vLLM, TokenHub, OpenAI
├─ Ollama, llama.cpp, etc.
└─ Any OpenAI-compatible server
ai-code-reviewer/
├── personas/ # Agent configurations (Agent Spec)
│ ├── freebsd-angry-ai/
│ │ ├── agent.yaml # Agent Spec configuration
│ │ └── README.md # Agent documentation
│ ├── security-hawk/
│ ├── performance-cop/
│ ├── friendly-mentor/
│ └── example/
├── reviewer.py # Main review loop
├── persona_validator.py # Agent Spec validator
├── llm_client.py # OpenAI-compatible LLM client (multi-provider)
├── build_executor.py # Build system integration
├── chunker.py # Large file handling
├── config.yaml.sample # Configuration template
├── AGENTS.md # AI agent instructions
└── docs/ # Additional documentation
source:
root: "/usr/src/linux"
build_command: "make -j$(nproc) bzImage modules"
build_timeout: 1800source:
root: "/home/user/my-rust-project"
build_command: "cargo build --release && cargo test"
build_timeout: 300source:
root: "/home/user/my-python-project"
build_command: "python -m pytest tests/"
build_timeout: 120| Target | Description |
|---|---|
make help |
Show all targets |
make check-deps |
Install dependencies |
make config-init |
Interactive setup |
make validate |
Validate LLM connection |
make run |
Run code review |
make run-verbose |
Run with verbose logging |
make test |
Run tests |
- Python 3.8+ with PyYAML
- At least one OpenAI-compatible LLM provider (vLLM, TokenHub, OpenAI, Ollama, etc.)
- Your project's build system (make, cmake, cargo, etc.)
- Git repository (for tracking changes)
- SETUP_GUIDE.md - Detailed installation guide
- AGENTS.md - AI agent instructions
- docs/ - Additional documentation
Part 2 of an ongoing chronicle. ← Part 1: NanoLang | Part 3: Aviation → Chronicle index · Ordered by first recorded AI-assisted commit.
The programmer had reached the stage of software development at which the code existed, the bugs existed, and the two had negotiated an arrangement that excluded him.
“I need another reviewer,” he announced.
Sir Reginald von Fluffington III, who was already reviewing the desk for objects that could be pushed off it, did not apply.
The obvious solution was to ask an AI to read the code. The less obvious problem was that code tended to arrive in quantities larger than a model's attention span. The programmer therefore arranged the review into directories, files, and functions, as though conducting an inspection of a very large hotel whose occupants had all denied responsibility for the plumbing.
Each reviewer needed a disposition. A friendly mentor could explain a mistake gently. A security hawk could regard it as evidence. A FreeBSD reviewer could bring the accumulated irritation of several decades to bear on a single questionable allocation. These became configurable personas, which the programmer considered an improvement over the traditional system of discovering a colleague's personality in the comments on a pull request.
Sir Reginald selected the security role by sitting on the only available input device.
There remained the question of whether the proposed fixes worked. An eloquent explanation was insufficient. The configured build command would have to run; a failed build would send the reviewer back to its work. Review and rewrite acquired separate workflows and progress records, because “I examined it” and “I replaced it” were statements the programmer preferred to keep distinguishable.
“It will use the project's own build,” he said. “Any language. Whatever proves the change.”
Sir Reginald knocked a pen off the desk. It reached the floor. The test was reproducible, the result unambiguous, and no second model was needed to judge it.
The programmer called the arrangement elegant. Sir Reginald withheld endorsement, citing insufficient tuna and a review process that still permitted the programmer to submit code.
MIT License - See LICENSE