feat: add Tenki Cloud compute provider - #1
Conversation
|
cc @camcalaquian — opened in my fork per your note (holding upstream for your review gate). Adds Tenki as a 7th compute provider alongside E2B/Daytona/Modal/Fly.io/Docker/local (both E2B & Daytona already in-tree, per the 'prioritize competitor-integrated repos' steer). Stable features only (exec + file I/O), no volume/snapshot. Validated live vs real Tenki: provision + pandas pip-install + exec + file round-trip + clean teardown. |
camcalaquian
left a comment
There was a problem hiding this comment.
overall lgtm. once the version is updated let's work on getting this to upstream
| # Install via: pip install git+https://github.com/daytonaio/daytona-python | ||
| ] | ||
| tenki = [ | ||
| "tenki-sandbox>=0.3", |
There was a problem hiding this comment.
Done in 9354602 — bumped to tenki-sandbox>=0.3.6. Ready for the upstream step whenever you are.
|
@rishijoshi sorry there was another release of sdk yesterday, let's update to that and submit this PR to upstream. |
|
@camcalaquian thanks! Quick check before I bump — PyPI's latest Could you point me at the new version (or where it's published)? As soon as it's on PyPI I'll bump |
|
seems the updated sdk wasn't posted immediately, you should see the updated sdk now on pypi |
|
0.4.0 landed on PyPI (2026-07-17) — bumped to |
…-3656-20260804-1008 fix: correct stale resolve_output import in precedence ladder tests
…-3649-20260804-1008 fix: wire delegate_task tool to subagent runtime
…-3657-20260804-1008 fix: OpenAI-compatible ManagerInstructions schema for hierarchical process
…-3651-20260804-1008 feat: unified remember/recall/forget convenience API for Memory
…-3655-20260804-1008 fix: ASCII-safe capabilities in models CLI on Windows cp1252
…-3650-20260804-1008 feat: add praisonai init team multi-agent project scaffold
…-3654-20260804-1008 fix: migrate invalid AgentTeam kwargs to output=/execution=
…-3666-20260804-1027 docs: add minimal sequential team YAML example (fixes MervinPraison#3666)
…-3665-20260804-1027 docs: correct-signature human-review team example + regression test
…ixes MervinPraison#3669) Add the missing repo-committed environment definition joint: a lightweight loader that maps .praisonai/environment.yaml onto the existing ComputeConfig schema, plus the one new field (setup) with a live consumer. - Core (praisonaiagents): ComputeConfig.setup: List[str]; load_environment_definition() and find_environment_definition() (walks up like config discovery). YAML is lazy-imported; unknown top-level keys raise ValueError with the file path. network/backend preferences carried in existing typed fields (no new surface). - Live consumer: docker compute provider runs setup commands once post-provision, streamed to logs, with failures surfaced (raised) rather than swallowed. - No file -> byte-identical behaviour (loader returns None; callers keep defaults). Deliberately minimal per AGENTS.md (lightweight & powerful): no new CLI group, no dispatcher rewrite, no SandboxConfig bridge, no RuntimeConfig remap. Co-authored-by: MervinPraison <MervinPraison@users.noreply.github.com>
…, clean up leaked container Addresses the three Greptile P1 findings on PR MervinPraison#3671: - Wire load_environment_definition() into provision_compute() as an opt-in baseline (explicit kwargs/instance config still win) so a committed .praisonai/environment.yaml is actually applied. - Validate nested config shapes in the loader so malformed values raise a contextual ValueError (with file path) instead of incidental AttributeError/TypeError. - Tear down the just-started Docker container if package install or setup fails, so a failed provision no longer leaks an unreachable container + registry entry. Co-authored-by: Mervin Praison <MervinPraison@users.noreply.github.com>
…-3669-20260804-1506 feat: add .praisonai/environment.yaml loader + ComputeConfig.setup
…vinPraison#3672) Add functools.lru_cache to the pure model-capability predicates so the same (model_name -> bool) result is computed once per process instead of on every request/turn. Behaviour is unchanged: guards, litellm-None fallbacks and except paths are all preserved, and keys are hashable strings. Co-authored-by: MervinPraison <MervinPraison@users.noreply.github.com>
…son#3674) PraisonAIAgents was exported from praisonaiagents.agents but omitted from the root lazy-import map and __all__, so `from praisonaiagents import PraisonAIAgents` raised ImportError. Added it alongside its siblings and a regression test covering all back-compat aliases. Co-authored-by: MervinPraison <MervinPraison@users.noreply.github.com>
…contract (MervinPraison#3674) Two existing tests asserted PraisonAIAgents raises ImportError (a v4 "removed" decision). Issue MervinPraison#3674 (owner) explicitly restores it as a root-level silent alias for AgentTeam, which PR MervinPraison#3675 implements. Update those tests to assert the restored back-compat contract instead. Co-authored-by: Mervin Praison <MervinPraison@users.noreply.github.com>
…-3674-20260804-2348 fix: expose PraisonAIAgents alias from package root
…-3672-20260804-1715 fix: memoise pure model-capability lookups in LLM hot path
…-3636-20260804-0853 feat: declarative per-tool model-facing output (model_output/to_model_output) for context economy
…aison#3825) Move the workflow_runs page implementation into praisonai_bot where its consumer (host_app.py) expects it, following the established bot_health C9 shim pattern. The wrapper module becomes a backward-compat shim. Co-authored-by: MervinPraison <MervinPraison@users.noreply.github.com>
…dbox CLI (fixes MervinPraison#3826) Co-authored-by: MervinPraison <MervinPraison@users.noreply.github.com>
…son#3827) Wrapper BaseFrameworkAdapter._resolve_llm now delegates the spec->model-name precedence to core's super()._resolve_llm instead of re-implementing it, then derives base_url/api_key locally to build the PraisonAIModel provider object. Removes duplicate logic and drift risk. Behaviour-preserving. Co-authored-by: MervinPraison <MervinPraison@users.noreply.github.com>
…-3824-20260809-1619 fix: reuse image/URL helpers from replay.judge in recipe_optimizer
…-3825-20260809-1619 fix: restore orphaned workflow-runs L3 dashboard page
…-3827-20260809-1620 fix: delegate _resolve_llm model precedence to core
…-3826-20260809-1620 fix: remove superseded unreferenced argparse builders in eval and sandbox CLI
…ection scan (fixes MervinPraison#3836) - security/injection.py: _extract_strings walks iteratively bounded by bytes+ cardinality instead of a depth-4 cutoff, scans dict keys, and is cycle-safe so nested tool inputs can no longer bypass the scanner. - db/adapter.py: share one _serialize_tool_call helper across sync/async paths that persists the full tool result verbatim (no silent [:1000] truncation). - persistence/orchestrator.py: add aretrieve_knowledge/aadd_knowledge that offload sync stores via asyncio.to_thread (and await native async stores) so RAG no longer blocks the event loop. Co-authored-by: Mervin Praison <MervinPraison@users.noreply.github.com>
The before-tool injection hook allowed a tool call whenever no scanned string matched. Because _extract_strings bounded the walk by bytes / cardinality and returned a partial result, an attacker controlling tool_input could pad benign strings ahead of an injection payload so the payload was never reached, and the gate allowed the call (Greptile P1). _extract_strings_bounded now reports whether the walk was truncated, and create_hook blocks (fail-closed) on truncated, un-vettable input instead of allowing it. _extract_strings is kept as a back-compat shim. Co-authored-by: Mervin Praison <MervinPraison@users.noreply.github.com>
…MervinPraison#3838) Capability helpers (supports_structured_outputs / function_calling / parallel_function_calling / web_search / prompt_caching) delegated entirely to litellm and returned False for every model when litellm was absent, silently disabling structured-output / tool-calling / caching / web paths on lean provider-native installs. Each helper now falls back to a small, conservative, pattern-based static heuristic only when litellm is unavailable. litellm remains authoritative when installed (no behaviour change for existing deployments); no new deps, no network, no config knobs. Co-authored-by: MervinPraison <MervinPraison@users.noreply.github.com>
Address reviewer feedback on PR MervinPraison#3839: - When litellm is installed but a capability helper is missing or raises, preserve litellm as authoritative (return False) instead of overriding with the static heuristic. Heuristics now run only when litellm is genuinely absent (litellm is None), preventing unsupported params (response_format / web_search_options) from reaching provider requests. - supports_parallel_function_calling now uses a dedicated, narrower heuristic instead of the general function-calling one, so serial-only tool-calling families are not over-reported as supporting parallel calls. - Add tests for installed-litellm error path and serial-vs-parallel. Co-authored-by: Mervin Praison <MervinPraison@users.noreply.github.com>
…-3836-20260810-0829 fix: async RAG offload, non-lossy tool persistence, nesting-proof injection scan
…MervinPraison#3841) Add a canonical GATEWAY_CONFIG_VERSION stamp and a declarative LegacyConfigRule contract in core (praisonaiagents/gateway/config.py) with a single migrate_config_with_doctor executor. Wire `gateway doctor` to detect an out-of-date config and `--fix` to migrate it forward once and stamp the new version, so cross-release upgrades no longer silently drift. Co-authored-by: MervinPraison <MervinPraison@users.noreply.github.com>
…on (fixes MervinPraison#3840) Per-channel restart budgets are blind to a systemic fault that restarts every channel at once. Add a pure core FleetSupervisionPolicy breaker (praisonaiagents/gateway) enforced by the wrapper ChannelHealthMonitor, which holds restarts and records ONE gateway degraded-owner fact on the shared DegradedCapabilityRegistry when the fleet thrashes. Wire the gateway's shared registry into the single supervision owner and expose gateway.health.fleet_* YAML thresholds. Co-authored-by: MervinPraison <MervinPraison@users.noreply.github.com>
…ite (MervinPraison#3841) Address reviewer findings on PR MervinPraison#3842: - core: migrate_config_with_doctor now rejects a config stamped by a newer build (version > GATEWAY_CONFIG_VERSION) and a malformed/boolean stamp via new ConfigVersionError, so an older binary can never downgrade a newer config or coerce config_version: true to 1. Exported from gateway package. - wrapper: doctor guards the migration import (ImportError-safe for older core), surfaces the version error as a warning without migrating, and rewrites gateway.yaml atomically (tmpfile + fsync + os.replace) so an interrupted --fix can't truncate the config. --fix help/docs document config migration. - tests: newer-version rejection, boolean/non-integer stamp rejection, export. Co-authored-by: Mervin Praison <MervinPraison@users.noreply.github.com>
…guard (MervinPraison#3840) - Core: FleetSupervisionPolicy no longer records events while cooling down and clears its event window when the cooldown elapses, so the breaker re-arms without an external reset() and held (non-)restarts cannot re-trip it. - Monitor: evaluate the failing-channel fraction on every sweep (note_fleet_state now has a live caller) and re-evaluate/clear the breaker inside the monitor loop, so the degraded-owner fact clears without a status read. - Monitor: a disabled per-channel restart budget (max_restarts_per_hour=0) no longer counts idle channels as failing. - Tests: fix CI regression in test_gateway_health_block_matches_runtime_consumer (new fleet keys are genuinely consumed by from_dict), add monitor-loop cooldown recovery coverage, tighten failing-fraction assertion, correct misleading comment. 🤖 Generated with [Claude Code](https://claude.ai/code) Co-authored-by: Mervin Praison <MervinPraison@users.noreply.github.com>
…-3838-20260810-0852 fix: litellm-free capability fallback in model_capabilities
…-3840-20260810-0915 fix: fleet-level crash-loop breaker for gateway channel supervision
…-3841-20260810-0915 fix: gateway config_version stamp + doctor-driven migration
…ixes MervinPraison#3848) Replace the duplicated public-callable extraction walk in the legacy recipe/workflow tools.py path with the canonical owner in praisonai_code.tool_resolver. Adds extract_functions_from_loaded_module (single authority over the walk) and reuses it from both ToolResolver.load_functions_from_module and workflow_commands, preserving the PRAISONAI_ALLOW_LOCAL_TOOLS gate, safe-loader, own-module-origin filter, and "Loaded N tools" notice. Co-authored-by: MervinPraison <MervinPraison@users.noreply.github.com>
….4.0) Tenki renamed the PyPI package (tenki-sandbox -> tenki) and dropped the project concept from the API: - dep: tenki-sandbox>=0.4.0 -> tenki>=0.5.4; import `from tenki import Client` - create() no longer takes project_id -> removed all project_id / TENKI_PROJECT_ID handling; _resolve_ids simplified to _resolve_workspace - eliminates the earlier protobuf conflict: tenki 0.5.4 needs protobuf>=5.29.5 (vs tenki-sandbox's >=6.31), which now co-resolves with autogen-core (5.29.6)
…-3848-20260810-1623 fix: consolidate module-member tool extraction onto canonical helper
- networking: allow_outbound now keys off the real enum — only 'unrestricted' gets outbound; 'limited' (also used for --no-networking) disables it. The old check compared against a nonexistent 'restricted' value, so outbound was never actually disabled. - get_status/list_instances: a refresh() *exception* is now treated as unknown (assume running) rather than STOPPED, so a transient outage no longer hides a live, still-billing sandbox. A successful refresh with a non-RUNNING state still reports stopped. - failed-install teardown: terminate() before dropping the local handle, so a failed terminate keeps the sandbox tracked instead of leaking it (matches the shutdown path). - don't log raw pip/npm specs (can carry private-index URLs/tokens); log counts.
…KI_AUTH_TOKEN - managed_local.provision_compute() now forwards `networking` and `metadata` into ComputeConfig. Without this they were dropped, so Tenki's allow_outbound / tenki_image (the first provider to read them) always saw defaults regardless of the caller's request. (Greptile: 'networking policy is dropped') - TenkiCompute credential resolution now matches the SDK's precedence (auth_token -> TENKI_AUTH_TOKEN -> TENKI_API_KEY), so is_available agrees with what Client() resolves; docstring + is_available test updated. (CodeRabbit)
…provider # Conflicts: # src/praisonai/praisonai/integrations/managed_local.py
Align the integration-test skip guard with the credential resolution added in d2fa4df — otherwise the tests skip even when only TENKI_AUTH_TOKEN is set. (Greptile follow-up)
The upstream merge pulled in .github/workflows changes that aren't part of this feature; revert them to the fork's versions so the push doesn't require the workflow OAuth scope. Main's workflows win at merge time and are what run for the PR anyway.
- Guard Unix-only ``resource`` import so _admission.py loads on Windows; sampler self-disables gracefully when unavailable (Greptile P1 #1). - Wire ``max_rss_mb`` through BotOS + gateway.yaml via a single-knob build_memory_pressure_policy() so the feature has a live production consumer instead of a dormant param (Greptile P1 MervinPraison#3). - Surface max_rss_mb in admission stats; document resource-only soft-queue degradation honestly rather than silently misleading (Greptile P1 MervinPraison#2). - Correct MemoryPressurePolicy docstring to reference the real config key. - Add tests: Windows-safe sampler, single-knob policy, BotOS wiring, stats. Co-authored-by: Mervin Praison <MervinPraison@users.noreply.github.com>
What & why
PraisonAI already ships compute providers for E2B, Daytona, Modal, Fly.io, Docker and local (one file per vendor in
integrations/compute/). This adds Tenki Cloud as another option — disposable Linux microVMs — so managed-agent tools can run in Tenki sandboxes.What it does
TenkiComputeimplements the fullComputeProviderProtocol(provision/execute/shutdown/get_status/upload_file/download_file/list_instances), running tools in ephemeral Tenki microVMs. Sync SDK wrapped inrun_in_executor, exactly likeDaytonaCompute/E2BCompute."tenki"in the compute barrel (__init__.py), the_resolve_computefactory (managed_local.py), and the provider hint sets (managed_agents.py,hosted_agent.py).TENKI_API_KEY; optionaltenkiextra (tenki-sandbox). Auto-resolves workspace/project from the key.Feature scope
Uses only stable Tenki features — ephemeral
exec+ file I/O (base64 over exec). No volume/snapshot/template. The default stock image shipspython3; pip packages are installed on demand (config.packages). Setconfig.metadata["tenki_image"]to use a prebaked image instead.Testing
provider_name,is_available, protocol methods, nonexistent-instance handling, and the barrel export — mirroring the E2B/Daytona suites intest_cloud_compute.py.TENKI_API_KEYis set): provision → execute → file upload/download → shutdown, plus pip-install.