Status: Current. This README describes the implementation available today; future work is listed separately under “Capability boundaries and roadmap.”
A distributed runtime for agents, models, and AI workloads.
Eruun currently provides a Kubernetes application and workflow runtime, standalone command jobs, and Harbor evaluation through standalone jobs or application workflows. You declare components, runtime configuration, and execution steps through the API. Eruun persists tasks, dispatches execution, reconciles Kubernetes resources, and exposes status, logs, and results.
It serves as an execution backend for self-hosted platforms, bringing application deployment, batch tasks, and evaluation under the same workspace authorization and task execution mechanisms. Its longer-term direction is an AI runtime; general-purpose agent sessions, MCP tool governance, and dedicated model-serving APIs remain planned work.
Use cases · Core concepts · Current capabilities · Runtime architecture · Capability boundaries and roadmap · Quick start · Documentation
| Scenario | What Eruun provides | What you need to provide |
|---|---|---|
| Self-hosted application platforms and internal developer platforms | Applications that group services, databases, configuration, and deployment workflows; authorization through personal or team workspaces | A Kubernetes cluster, application images, a platform UI, and your business processes |
| Multi-component application delivery and operations | Sequential or DAG workflows, approvals, version updates, start/stop operations, failure cleanup, status, and logs | Component declarations, execution steps, storage, and network configuration |
| One-off container tasks | Submit a command Job directly, inspect its status, or cancel it without creating an application |
An image, command, and inputs compatible with the workspace security policy |
| LLM evaluation | Upload native Harbor task packages, declare type: job with traits.eval as a standalone Job or application component, and retain rewards, trajectories, logs, and artifacts |
Prebuilt task images, verifiers, Runner configuration, and Secret-backed credentials when using a model |
| Bringing existing Kubernetes applications into Eruun | Read-only observation or explicit adoption through dry-run and signed plans | Clear resource ownership, a supported resource scope, and adoption key configuration |
If you only need to apply a few Kubernetes manifests, Eruun’s database, queue, and multi-node deployment add operational cost. These dependencies become more useful when you need durable tasks, workspace authorization, multi-component execution, and a common query interface.
Eruun uses an OAM-inspired “components + traits + workflows” model, with two execution entry points: application workflows and standalone workspace Jobs.
| Concept | The question it answers | Examples and boundaries |
|---|---|---|
| Workspace | Who can access resources, and where do tasks run? | Personal and team workspaces associate membership permissions with a namespace, quotas, and network policies |
| Application | Which components belong to the same application? | Groups an API, database, and configuration under one deployment and authorization boundary |
| Component | What should run or be generated? | Long-running services, stateful services, Jobs, CronJobs, ConfigMaps, and Secrets |
| Trait | How should a component run? | Storage mounts, environment variables, resource declarations, and probes attached to a component |
| Workflow | In which steps should component operations execute? | StepByStep/DAG execution, approvals, cancellation, timeouts, failure policies, and terminal callbacks |
| Standalone Job | How do you run a one-off task outside an application? | type: command tasks and type: job with traits.eval belong directly to a workspace, return a taskId, and require no appId |
For example, an application can contain a webservice API and a store database, use traits to declare PVCs, environment variables, and probes, and organize deployment through a Workflow. Evaluation uses type: job with traits.eval, submitted directly as a workspace Job or declared as a top-level application component and referenced by a Workflow jobType: deploy step. Both paths reuse the existing scheduler, execution records, and database leases.
| Capability | Available today | Detailed contract |
|---|---|---|
| Application lifecycle | Create, execute, update versions, start, stop, restart, and clean up resources; management mode constrains write access | Create and execute, Version updates, Management modes |
| Workflows | StepByStep/DAG execution, approval pause/resume, cancellation, timeouts, failure cleanup, callbacks, and execution recovery | Execution architecture, Approvals, Failure policies |
| Standalone tasks and evaluation | command container tasks; standalone or application-workflow Harbor evaluation using traits.eval, with task packages, execution progress, complete results, and independent storage destinations |
Workspace Job API, Evaluation example |
| Programmatic integration | HTTP /api/v1, gRPC v1, canonical JSON Schema, Try validation, editable specification retrieval, submission idempotency, and allowedActions |
Canonical JSON, gRPC |
| Runtime diagnostics | Application/component/task status, container information, log streams, log archives, file export, and shell execution | Status, Logs, Files and execution |
| Accounts and workspaces | Login sessions, personal/team workspaces, member roles, resource ownership, and Kubernetes workspace security baselines | Accounts and workspaces |
| Existing resource adoption | Read-only observe mode, explicit adopted ownership, and controlled reconciliation/cleanup |
Namespace import |
| Component type | Execution target |
|---|---|
webservice |
A Deployment for a long-running stateless service |
store |
A StatefulSet for a service requiring stable identity or persistent storage |
job / scheduledjob |
An application-owned Kubernetes Job / CronJob |
config / secret |
A ConfigMap / Secret |
cloudjob |
A registered Provider/Action extension in the engine; current workspace application policy rejects submission |
The component model includes the following trait fields. A field’s presence does not authorize every workspace to use it. The platform also checks workspace, network, storage, and container permissions.
| Trait | Purpose and Kubernetes mapping |
|---|---|
storage |
PVCs, ConfigMaps, Secrets, ephemeral volumes, and mount paths |
envs / envFrom |
Individual environment variables and bulk imports from Secrets/ConfigMaps |
resources |
CPU, memory, and nvidia.com/gpu requests/limits |
eval |
Harbor evaluation configuration for standalone or top-level application job declarations; shares validation and Runner construction |
probes |
Liveness, readiness, and startup probes |
securityPolicy |
Container securityContext, subject to the workspace Restricted policy |
targetWorkEnv |
Pod nodeSelector |
rollout |
Deployment/StatefulSet update strategies |
init / sidecar |
Init and sidecar containers, with a restricted subset of nested traits |
service / ingress |
Service exposure and Ingress routing, subject to workspace port and domain rules |
share |
Reuse and lifecycle policies for shared resources within a namespace |
rbac |
The engine can generate ServiceAccounts/Roles/Bindings; current workspace policy rejects additional RBAC and arbitrary ServiceAccounts |
Resource generation, reconciliation, and cleanup paths handle service and share; the shared evaluation builder handles eval. See the architecture and trait documentation for the other traits’ processing order, nesting rules, and permission boundaries. Standalone Jobs accept their own subset of traits; they cannot use a complete component configuration unchanged.
Every node runs the same eruun-server and competes for one Kubernetes Lease. The elected Leader serves HTTP/gRPC APIs, scheduling, and Controller maintenance. Other nodes execute as Workers and remain election candidates.
flowchart TB
Client[API clients] --> Service[Stable Service: selects only the Leader]
Service --> Leader[Leader: API / scheduling / maintenance]
Leader -. election .-> Lease[One Kubernetes Lease]
Worker[Worker nodes and candidates] -. election .-> Lease
Leader <--> DB[(MySQL: domain state and execution leases)]
Leader --> Queue[Redis Streams or Kafka]
Queue --> Worker
Worker <--> DB
Worker --> K8s[Kubernetes workloads]
Leader <--> K8s
Leader --> Redis[Redis: cache and application coordination]
Worker --> Redis
The default is four nodes: one Leader and three Workers. Production needs at least two nodes; a single node has no Worker capacity for new tasks. Promotion stops new task intake while previously claimed tasks continue with their original ownership and heartbeats. Losing leadership stops API/control duties before returning to Worker intake. Healthy Workers remain PodReady; the business Service selects only the Leader.
The execution path is: Leader API persists a task → Leader dispatches it → Worker claims a database lease and operates on Kubernetes → execution progress and results are saved. Acceptance, workload readiness, execution completion, and result storage are separate states. Leadership changes interrupt HTTP/gRPC connections; reconnect, check task state, and retry only under the endpoint's idempotency contract.
| Dependency | Responsibility | Boundary |
|---|---|---|
| MySQL | Applications, tasks, accounts, configuration, and execution lease/generation/token | Source of truth for durable task state and execution ownership |
| Redis | Cache, application mutation locks, cancellation signals, and default Redis Streams messaging | Still required when using Kafka |
| Kafka (optional) | Replaces Redis Streams as the workflow message transport | Does not own task state or replace database leases |
| Kubernetes | Container execution, resource reconciliation, scheduling, and isolation | Provides actual resource state; business queries primarily read database projections |
| MinIO (optional) | Stores complete evaluation result files | Database-only storage is supported; failed storage destinations can be retried independently |
The queue delivers messages at least once. Database leases and generation/token fencing prevent stale executors from overwriting newer state; external side effects still need idempotency or compensation. Multiple nodes do not automatically provide data availability: the chart’s bundled MySQL and Redis are single-replica development dependencies. See Helm deployment for production topology guidance.
- Product delivery: The repository ships the server runtime, Helm/manifests, and API examples. Clients integrate using
curlor gRPC. Eruun does not ship a client command-line application; the integrating platform provides its UI and business approval system. - Kubernetes: An existing cluster is required. Eruun is not currently a cluster installer, multi-cluster control platform, or general-purpose cloud resource manager. Existing CloudJob extensions do not make general-purpose AI Providers available.
- Permissions: Platform login permissions, workload Kubernetes identities, and container permissions are checked separately. Current workspace policy rejects privileged/host capabilities, cross-namespace requests, additional RBAC, CloudJob, and NodePort/LoadBalancer, among other restrictions. Network isolation requires a CNI that enforces NetworkPolicy. See the workspace security contract.
- Workflows: There is no general-purpose step output dataflow,
dependsOn, or conditional branching contract today. StandalonecommandJobs do not support cron, delayed execution, or automatic reruns on failure; applicationscheduledjobcomponents and Workflow schedules have their own contracts. - Evaluation:
traits.evaluses a fixed Harbor adapter. The model is the object of evaluation;agentselects the harness that executes the task. The platform pulls prebuilt images; it does not build the Dockerfiles in task packages online or accept arbitrary evaluation frameworks or Python adapters. The trait is restricted to standalone or top-level applicationjobdeclarations. Evaluation rewards and successful task execution are distinct outcomes. - AI capabilities: Resource declarations and container execution do not imply support for agent sessions/memory, MCP authorization and auditing, dedicated vLLM deployment, or GPU-aware scheduling.
| Stage | Scope |
|---|---|
| Current | The Application/Workflow capabilities, homogeneous runtime, workspace authorization, standalone command Jobs, and Harbor evaluation through both execution entry points described here |
| Next | Self-hosted agent execution contracts, MCP/CLI tool bindings, credential delegation, tool permissions and auditing, and more evaluation frameworks |
| Later | Model serving, GPU-aware scheduling, vectorization, managed AI Providers, and cloud or multi-cluster execution |
The AI runtime vision explains the direction of development. Draft / Proposal documents are not promises of existing APIs or deployment capabilities. Current behavior is defined by the implementation and the Current documents in the documentation index.
Prepare Bash, kubectl, OpenSSL, Helm, and access to a Kubernetes cluster. The cluster must support NetworkPolicy and Restricted v1.34 Pod Security, with storage available for the MySQL/Redis PVCs; see Helm deployment. Installing published images requires neither Go nor Rust.
Prepare /secure/eruun/accounts.json from the account configuration example. Replace credential and cluster network placeholders, and disable login providers you do not need. Keep the file outside the repository with mode 0600. See account and workspace deployment for the complete configuration contract.
Run from the repository root:
AUTH_CONFIG_FILE=/secure/eruun/accounts.json INSTALL_MODE=helm \
./deploy/all_in_one_install_quickstart.sh
kubectl -n eruun-system port-forward svc/eruun 8000:8000The installer deploys the unified nodes plus MySQL and Redis. On the first installation, it generates random database/cache passwords when none are supplied and passes them through temporary files with mode 0600 and Kubernetes Secrets. Subsequent installations reuse existing credentials. port-forward connects to the current Leader; restart it after a leadership change. Check from another terminal:
export ERUUN_URL=http://127.0.0.1:8000
curl --fail "$ERUUN_URL/api/v1/healthz"
curl --fail "$ERUUN_URL/api/v1/readyz"Follow the account API examples to register/log in and obtain an access token and personal or team workspace ID. The bootstrap administrator must change their password and log in again first. Set ACCESS_TOKEN and WORKSPACE_ID in the current terminal; task submission requires the member role or higher in that workspace. Health checks do not require login; business APIs require a Bearer token and workspace authorization. Browser login, OAuth, and cookie refresh also require same-site HTTPS endpoints configured as described in the account documentation.
Submit a one-off command that runs in the workspace namespace:
curl --fail-with-body -X POST "$ERUUN_URL/api/v1/jobs" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "X-Eruun-Workspace-ID: $WORKSPACE_ID" \
-H 'Content-Type: application/json' \
--data-binary @- <<'JSON'
{
"name": "hello-eruun",
"type": "command",
"spec": {
"image": "busybox:1.37.0",
"command": ["sh", "-c"],
"args": ["echo hello-eruun"],
"timeoutSeconds": 300
},
"traits": {
"securityPolicy": {"runAsUser": 1000, "runAsNonRoot": true},
"resources": {"cpu": "100m", "memory": "128Mi", "cpuLimit": "500m", "memoryLimit": "256Mi"}
}
}
JSONA successful submission returns HTTP 202, with data.taskId identifying this execution. Put that value in the command below and query until the task reaches a terminal state. Submitting again creates a new task.
TASK_ID='<returned data.taskId>'
curl --fail-with-body "$ERUUN_URL/api/v1/jobs/$TASK_ID" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "X-Eruun-Workspace-ID: $WORKSPACE_ID"To try application deployment next, validate a declaration using canonical Application JSON and Try, then call the create-and-execute API. To try evaluation, first configure the Runner, task images, and task package following the Harbor example. The default installation does not enable evaluation automatically.
Configure the server through flags or ERUUN_ environment variables; for example, --bind-addr maps to ERUUN_BIND_ADDR. See the default configuration and go run ./cmd/main.go --help for configuration details and the full flag list.
| Configuration | Purpose |
|---|---|
ERUUN_LEADER_LOCK_NAME |
Shared election Lease; defaults to eruun-runtime |
ERUUN_LEADER_SERVICE_NAME |
Cluster business Service; leave empty locally to access the elected node directly |
ERUUN_POD_NAME |
Pod name for cluster entry-point maintenance; election identity defaults to a separate UUID per process |
ERUUN_BIND_ADDR / ERUUN_GRPC_BIND_ADDR |
Local HTTP defaults to 127.0.0.1:8001; API gRPC defaults to 127.0.0.1:9001; cluster deployments use 8000 / 9000 respectively |
ERUUN_DATASTORE_URL |
The actual MySQL DSN; the password placeholder must be replaced |
ERUUN_CACHE_HOST / ERUUN_CACHE_PASSWORD |
Redis connection configuration |
ERUUN_MSG_TYPE / ERUUN_MSG_KAFKA_BROKERS |
Redis Streams by default; configure brokers when selecting Kafka, and retain Redis |
ERUUN_ENABLE_TRACING / ERUUN_JAEGER_ENDPOINT |
Enable tracing with one flag; the Jaeger endpoint controls span export |
ERUUN_AUTH_CONFIG_FILE |
Account, session, and workspace policy JSON |
ERUUN_JOBS_CONFIG_FILE |
Enables the Harbor Runner and optional MinIO; use the same configuration for all nodes |
All runtime nodes require Redis for cache, authentication, and coordination. --cache-type / ERUUN_CACHE_TYPE accepts only redis; memory fails startup validation. In --datastore-schema-mode=migrate-only, validation checks only datastore configuration.
Tracing defaults to ERUUN_ENABLE_TRACING=true; set it to false to disable tracing. This setting controls both the tracer provider and HTTP middleware, independently of the Redis/Kafka backend. Setting a Jaeger endpoint alone does not enable tracing. Without an endpoint, tracing can add trace IDs to API request logs but does not export spans. The static stack manifest enables tracing.
Migration: --auto-tracing and ERUUN_AUTO_TRACING have been removed. Remove the old setting and set --enable-tracing / ERUUN_ENABLE_TRACING to the logical OR of the previous two flags (previous defaults: enabled). The removed CLI flag or environment variable now causes startup to fail, including when the old environment variable is empty or false. Explicit CLI settings take precedence over the remaining environment setting.
Local source development requires Go 1.27, plus GNU Make when using Make targets. Start and configure MySQL, Redis, and optional Kafka using the local dependencies guide, prepare account configuration and Kubernetes access, then start a node:
go run ./cmd/main.goThis starts one candidate node. Once elected, a single node serves APIs but cannot claim new tasks. End-to-end execution requires at least two nodes sharing the namespace, Lease, database, and messaging configuration, with unique instance IDs. Local processes need distinct HTTP/gRPC addresses. Migrate the database first, then start other nodes in validation mode; see the Helm deployment contract for cluster setup.
Common development checks:
make build
go test ./... -race -cover
go vet ./...make build verifies the build and discards the binary. To produce an executable, use go build -o eruun-server ./cmd/main.go. See AGENTS.md and the documentation index for module routing, validation, and contribution rules.
| What you want to understand | Start here |
|---|---|
| Documentation status, current behavior, and code locations | Documentation index |
| Domain model, traits, permissions, and module boundaries | Architecture, Cross-layer field contracts |
| Leader/Worker nodes, execution sequences, leases, and failure recovery | Architecture diagrams, Workflow architecture, Distributed runtime |
| API and automation integration | Canonical JSON, gRPC, Account examples |
| Deployment and workspace isolation | Helm, Accounts and workspaces, Local dependencies |
| Commands, evaluation, task packages, and results | Workspace Job API, Evaluation example, Runner |
| Future AI runtime direction | Vision and roadmap (Draft / Proposal) |
Eruun is licensed under the MIT License. See LICENSE.