Skip to content
PixelCoresPublic

About

A distributed runtime for agents, models, and AI workloads.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

English | 简体中文

Eruun

Status: Current. This README describes the implementation available today; future work is listed separately under “Capability boundaries and roadmap.”

A distributed runtime for agents, models, and AI workloads.

Eruun currently provides a Kubernetes application and workflow runtime, standalone command jobs, and Harbor evaluation through standalone jobs or application workflows. You declare components, runtime configuration, and execution steps through the API. Eruun persists tasks, dispatches execution, reconciles Kubernetes resources, and exposes status, logs, and results.

It serves as an execution backend for self-hosted platforms, bringing application deployment, batch tasks, and evaluation under the same workspace authorization and task execution mechanisms. Its longer-term direction is an AI runtime; general-purpose agent sessions, MCP tool governance, and dedicated model-serving APIs remain planned work.

Use cases · Core concepts · Current capabilities · Runtime architecture · Capability boundaries and roadmap · Quick start · Documentation

Use cases

Scenario What Eruun provides What you need to provide
Self-hosted application platforms and internal developer platforms Applications that group services, databases, configuration, and deployment workflows; authorization through personal or team workspaces A Kubernetes cluster, application images, a platform UI, and your business processes
Multi-component application delivery and operations Sequential or DAG workflows, approvals, version updates, start/stop operations, failure cleanup, status, and logs Component declarations, execution steps, storage, and network configuration
One-off container tasks Submit a command Job directly, inspect its status, or cancel it without creating an application An image, command, and inputs compatible with the workspace security policy
LLM evaluation Upload native Harbor task packages, declare type: job with traits.eval as a standalone Job or application component, and retain rewards, trajectories, logs, and artifacts Prebuilt task images, verifiers, Runner configuration, and Secret-backed credentials when using a model
Bringing existing Kubernetes applications into Eruun Read-only observation or explicit adoption through dry-run and signed plans Clear resource ownership, a supported resource scope, and adoption key configuration

If you only need to apply a few Kubernetes manifests, Eruun’s database, queue, and multi-node deployment add operational cost. These dependencies become more useful when you need durable tasks, workspace authorization, multi-component execution, and a common query interface.

Core concepts

Eruun uses an OAM-inspired “components + traits + workflows” model, with two execution entry points: application workflows and standalone workspace Jobs.

Concept The question it answers Examples and boundaries
Workspace Who can access resources, and where do tasks run? Personal and team workspaces associate membership permissions with a namespace, quotas, and network policies
Application Which components belong to the same application? Groups an API, database, and configuration under one deployment and authorization boundary
Component What should run or be generated? Long-running services, stateful services, Jobs, CronJobs, ConfigMaps, and Secrets
Trait How should a component run? Storage mounts, environment variables, resource declarations, and probes attached to a component
Workflow In which steps should component operations execute? StepByStep/DAG execution, approvals, cancellation, timeouts, failure policies, and terminal callbacks
Standalone Job How do you run a one-off task outside an application? type: command tasks and type: job with traits.eval belong directly to a workspace, return a taskId, and require no appId

For example, an application can contain a webservice API and a store database, use traits to declare PVCs, environment variables, and probes, and organize deployment through a Workflow. Evaluation uses type: job with traits.eval, submitted directly as a workspace Job or declared as a top-level application component and referenced by a Workflow jobType: deploy step. Both paths reuse the existing scheduler, execution records, and database leases.

Current capabilities

Capability Available today Detailed contract
Application lifecycle Create, execute, update versions, start, stop, restart, and clean up resources; management mode constrains write access Create and execute, Version updates, Management modes
Workflows StepByStep/DAG execution, approval pause/resume, cancellation, timeouts, failure cleanup, callbacks, and execution recovery Execution architecture, Approvals, Failure policies
Standalone tasks and evaluation command container tasks; standalone or application-workflow Harbor evaluation using traits.eval, with task packages, execution progress, complete results, and independent storage destinations Workspace Job API, Evaluation example
Programmatic integration HTTP /api/v1, gRPC v1, canonical JSON Schema, Try validation, editable specification retrieval, submission idempotency, and allowedActions Canonical JSON, gRPC
Runtime diagnostics Application/component/task status, container information, log streams, log archives, file export, and shell execution Status, Logs, Files and execution
Accounts and workspaces Login sessions, personal/team workspaces, member roles, resource ownership, and Kubernetes workspace security baselines Accounts and workspaces
Existing resource adoption Read-only observe mode, explicit adopted ownership, and controlled reconciliation/cleanup Namespace import

Components and traits

Component type Execution target
webservice A Deployment for a long-running stateless service
store A StatefulSet for a service requiring stable identity or persistent storage
job / scheduledjob An application-owned Kubernetes Job / CronJob
config / secret A ConfigMap / Secret
cloudjob A registered Provider/Action extension in the engine; current workspace application policy rejects submission

The component model includes the following trait fields. A field’s presence does not authorize every workspace to use it. The platform also checks workspace, network, storage, and container permissions.

Trait Purpose and Kubernetes mapping
storage PVCs, ConfigMaps, Secrets, ephemeral volumes, and mount paths
envs / envFrom Individual environment variables and bulk imports from Secrets/ConfigMaps
resources CPU, memory, and nvidia.com/gpu requests/limits
eval Harbor evaluation configuration for standalone or top-level application job declarations; shares validation and Runner construction
probes Liveness, readiness, and startup probes
securityPolicy Container securityContext, subject to the workspace Restricted policy
targetWorkEnv Pod nodeSelector
rollout Deployment/StatefulSet update strategies
init / sidecar Init and sidecar containers, with a restricted subset of nested traits
service / ingress Service exposure and Ingress routing, subject to workspace port and domain rules
share Reuse and lifecycle policies for shared resources within a namespace
rbac The engine can generate ServiceAccounts/Roles/Bindings; current workspace policy rejects additional RBAC and arbitrary ServiceAccounts

Resource generation, reconciliation, and cleanup paths handle service and share; the shared evaluation builder handles eval. See the architecture and trait documentation for the other traits’ processing order, nesting rules, and permission boundaries. Standalone Jobs accept their own subset of traits; they cannot use a complete component configuration unchanged.

Runtime architecture

Every node runs the same eruun-server and competes for one Kubernetes Lease. The elected Leader serves HTTP/gRPC APIs, scheduling, and Controller maintenance. Other nodes execute as Workers and remain election candidates.

flowchart TB
    Client[API clients] --> Service[Stable Service: selects only the Leader]
    Service --> Leader[Leader: API / scheduling / maintenance]
    Leader -. election .-> Lease[One Kubernetes Lease]
    Worker[Worker nodes and candidates] -. election .-> Lease
    Leader <--> DB[(MySQL: domain state and execution leases)]
    Leader --> Queue[Redis Streams or Kafka]
    Queue --> Worker
    Worker <--> DB
    Worker --> K8s[Kubernetes workloads]
    Leader <--> K8s
    Leader --> Redis[Redis: cache and application coordination]
    Worker --> Redis
Loading

The default is four nodes: one Leader and three Workers. Production needs at least two nodes; a single node has no Worker capacity for new tasks. Promotion stops new task intake while previously claimed tasks continue with their original ownership and heartbeats. Losing leadership stops API/control duties before returning to Worker intake. Healthy Workers remain PodReady; the business Service selects only the Leader.

The execution path is: Leader API persists a task → Leader dispatches it → Worker claims a database lease and operates on Kubernetes → execution progress and results are saved. Acceptance, workload readiness, execution completion, and result storage are separate states. Leadership changes interrupt HTTP/gRPC connections; reconnect, check task state, and retry only under the endpoint's idempotency contract.

Dependency Responsibility Boundary
MySQL Applications, tasks, accounts, configuration, and execution lease/generation/token Source of truth for durable task state and execution ownership
Redis Cache, application mutation locks, cancellation signals, and default Redis Streams messaging Still required when using Kafka
Kafka (optional) Replaces Redis Streams as the workflow message transport Does not own task state or replace database leases
Kubernetes Container execution, resource reconciliation, scheduling, and isolation Provides actual resource state; business queries primarily read database projections
MinIO (optional) Stores complete evaluation result files Database-only storage is supported; failed storage destinations can be retried independently

The queue delivers messages at least once. Database leases and generation/token fencing prevent stale executors from overwriting newer state; external side effects still need idempotency or compensation. Multiple nodes do not automatically provide data availability: the chart’s bundled MySQL and Redis are single-replica development dependencies. See Helm deployment for production topology guidance.

Capability boundaries and roadmap

  • Product delivery: The repository ships the server runtime, Helm/manifests, and API examples. Clients integrate using curl or gRPC. Eruun does not ship a client command-line application; the integrating platform provides its UI and business approval system.
  • Kubernetes: An existing cluster is required. Eruun is not currently a cluster installer, multi-cluster control platform, or general-purpose cloud resource manager. Existing CloudJob extensions do not make general-purpose AI Providers available.
  • Permissions: Platform login permissions, workload Kubernetes identities, and container permissions are checked separately. Current workspace policy rejects privileged/host capabilities, cross-namespace requests, additional RBAC, CloudJob, and NodePort/LoadBalancer, among other restrictions. Network isolation requires a CNI that enforces NetworkPolicy. See the workspace security contract.
  • Workflows: There is no general-purpose step output dataflow, dependsOn, or conditional branching contract today. Standalone command Jobs do not support cron, delayed execution, or automatic reruns on failure; application scheduledjob components and Workflow schedules have their own contracts.
  • Evaluation: traits.eval uses a fixed Harbor adapter. The model is the object of evaluation; agent selects the harness that executes the task. The platform pulls prebuilt images; it does not build the Dockerfiles in task packages online or accept arbitrary evaluation frameworks or Python adapters. The trait is restricted to standalone or top-level application job declarations. Evaluation rewards and successful task execution are distinct outcomes.
  • AI capabilities: Resource declarations and container execution do not imply support for agent sessions/memory, MCP authorization and auditing, dedicated vLLM deployment, or GPU-aware scheduling.
Stage Scope
Current The Application/Workflow capabilities, homogeneous runtime, workspace authorization, standalone command Jobs, and Harbor evaluation through both execution entry points described here
Next Self-hosted agent execution contracts, MCP/CLI tool bindings, credential delegation, tool permissions and auditing, and more evaluation frameworks
Later Model serving, GPU-aware scheduling, vectorization, managed AI Providers, and cloud or multi-cluster execution

The AI runtime vision explains the direction of development. Draft / Proposal documents are not promises of existing APIs or deployment capabilities. Current behavior is defined by the implementation and the Current documents in the documentation index.

Quick start

1. Deploy the homogeneous runtime

Prepare Bash, kubectl, OpenSSL, Helm, and access to a Kubernetes cluster. The cluster must support NetworkPolicy and Restricted v1.34 Pod Security, with storage available for the MySQL/Redis PVCs; see Helm deployment. Installing published images requires neither Go nor Rust.

Prepare /secure/eruun/accounts.json from the account configuration example. Replace credential and cluster network placeholders, and disable login providers you do not need. Keep the file outside the repository with mode 0600. See account and workspace deployment for the complete configuration contract.

Run from the repository root:

AUTH_CONFIG_FILE=/secure/eruun/accounts.json INSTALL_MODE=helm \
  ./deploy/all_in_one_install_quickstart.sh

kubectl -n eruun-system port-forward svc/eruun 8000:8000

The installer deploys the unified nodes plus MySQL and Redis. On the first installation, it generates random database/cache passwords when none are supplied and passes them through temporary files with mode 0600 and Kubernetes Secrets. Subsequent installations reuse existing credentials. port-forward connects to the current Leader; restart it after a leadership change. Check from another terminal:

export ERUUN_URL=http://127.0.0.1:8000
curl --fail "$ERUUN_URL/api/v1/healthz"
curl --fail "$ERUUN_URL/api/v1/readyz"

2. Log in and submit a task

Follow the account API examples to register/log in and obtain an access token and personal or team workspace ID. The bootstrap administrator must change their password and log in again first. Set ACCESS_TOKEN and WORKSPACE_ID in the current terminal; task submission requires the member role or higher in that workspace. Health checks do not require login; business APIs require a Bearer token and workspace authorization. Browser login, OAuth, and cookie refresh also require same-site HTTPS endpoints configured as described in the account documentation.

Submit a one-off command that runs in the workspace namespace:

curl --fail-with-body -X POST "$ERUUN_URL/api/v1/jobs" \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "X-Eruun-Workspace-ID: $WORKSPACE_ID" \
  -H 'Content-Type: application/json' \
  --data-binary @- <<'JSON'
{
  "name": "hello-eruun",
  "type": "command",
  "spec": {
    "image": "busybox:1.37.0",
    "command": ["sh", "-c"],
    "args": ["echo hello-eruun"],
    "timeoutSeconds": 300
  },
  "traits": {
    "securityPolicy": {"runAsUser": 1000, "runAsNonRoot": true},
    "resources": {"cpu": "100m", "memory": "128Mi", "cpuLimit": "500m", "memoryLimit": "256Mi"}
  }
}
JSON

A successful submission returns HTTP 202, with data.taskId identifying this execution. Put that value in the command below and query until the task reaches a terminal state. Submitting again creates a new task.

TASK_ID='<returned data.taskId>'
curl --fail-with-body "$ERUUN_URL/api/v1/jobs/$TASK_ID" \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "X-Eruun-Workspace-ID: $WORKSPACE_ID"

To try application deployment next, validate a declaration using canonical Application JSON and Try, then call the create-and-execute API. To try evaluation, first configure the Runner, task images, and task package following the Harbor example. The default installation does not enable evaluation automatically.

Configuration and local development

Configure the server through flags or ERUUN_ environment variables; for example, --bind-addr maps to ERUUN_BIND_ADDR. See the default configuration and go run ./cmd/main.go --help for configuration details and the full flag list.

Configuration Purpose
ERUUN_LEADER_LOCK_NAME Shared election Lease; defaults to eruun-runtime
ERUUN_LEADER_SERVICE_NAME Cluster business Service; leave empty locally to access the elected node directly
ERUUN_POD_NAME Pod name for cluster entry-point maintenance; election identity defaults to a separate UUID per process
ERUUN_BIND_ADDR / ERUUN_GRPC_BIND_ADDR Local HTTP defaults to 127.0.0.1:8001; API gRPC defaults to 127.0.0.1:9001; cluster deployments use 8000 / 9000 respectively
ERUUN_DATASTORE_URL The actual MySQL DSN; the password placeholder must be replaced
ERUUN_CACHE_HOST / ERUUN_CACHE_PASSWORD Redis connection configuration
ERUUN_MSG_TYPE / ERUUN_MSG_KAFKA_BROKERS Redis Streams by default; configure brokers when selecting Kafka, and retain Redis
ERUUN_ENABLE_TRACING / ERUUN_JAEGER_ENDPOINT Enable tracing with one flag; the Jaeger endpoint controls span export
ERUUN_AUTH_CONFIG_FILE Account, session, and workspace policy JSON
ERUUN_JOBS_CONFIG_FILE Enables the Harbor Runner and optional MinIO; use the same configuration for all nodes

All runtime nodes require Redis for cache, authentication, and coordination. --cache-type / ERUUN_CACHE_TYPE accepts only redis; memory fails startup validation. In --datastore-schema-mode=migrate-only, validation checks only datastore configuration.

Tracing defaults to ERUUN_ENABLE_TRACING=true; set it to false to disable tracing. This setting controls both the tracer provider and HTTP middleware, independently of the Redis/Kafka backend. Setting a Jaeger endpoint alone does not enable tracing. Without an endpoint, tracing can add trace IDs to API request logs but does not export spans. The static stack manifest enables tracing.

Migration: --auto-tracing and ERUUN_AUTO_TRACING have been removed. Remove the old setting and set --enable-tracing / ERUUN_ENABLE_TRACING to the logical OR of the previous two flags (previous defaults: enabled). The removed CLI flag or environment variable now causes startup to fail, including when the old environment variable is empty or false. Explicit CLI settings take precedence over the remaining environment setting.

Local source development requires Go 1.27, plus GNU Make when using Make targets. Start and configure MySQL, Redis, and optional Kafka using the local dependencies guide, prepare account configuration and Kubernetes access, then start a node:

go run ./cmd/main.go

This starts one candidate node. Once elected, a single node serves APIs but cannot claim new tasks. End-to-end execution requires at least two nodes sharing the namespace, Lease, database, and messaging configuration, with unique instance IDs. Local processes need distinct HTTP/gRPC addresses. Migrate the database first, then start other nodes in validation mode; see the Helm deployment contract for cluster setup.

Common development checks:

make build
go test ./... -race -cover
go vet ./...

make build verifies the build and discards the binary. To produce an executable, use go build -o eruun-server ./cmd/main.go. See AGENTS.md and the documentation index for module routing, validation, and contribution rules.

Documentation

What you want to understand Start here
Documentation status, current behavior, and code locations Documentation index
Domain model, traits, permissions, and module boundaries Architecture, Cross-layer field contracts
Leader/Worker nodes, execution sequences, leases, and failure recovery Architecture diagrams, Workflow architecture, Distributed runtime
API and automation integration Canonical JSON, gRPC, Account examples
Deployment and workspace isolation Helm, Accounts and workspaces, Local dependencies
Commands, evaluation, task packages, and results Workspace Job API, Evaluation example, Runner
Future AI runtime direction Vision and roadmap (Draft / Proposal)

License

Eruun is licensed under the MIT License. See LICENSE.

About

A distributed runtime for agents, models, and AI workloads.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages