Repository navigation
Conversation
…d polling (kubernetes-sigs#1241) * feat(sdk): single-watch wait_for_claim_ready + faster dev port-forward polling Add wait_for_claim_ready() (sync and async) that waits for a SandboxClaim using a single watch on the claim itself, resolving on the claim's own status (bound sandbox name + forwarded Ready condition) instead of two sequential watches (claim then sandbox). The watch is anchored to the initial read's resourceVersion so no event is missed, and transparently restarts from "0" on 410 Gone. Legacy two-watch methods are kept fully intact and API-compatible. Also reduce the dev port-forward readiness polling interval from 500ms to 50ms and document SDK ready-wait latency behavior in the README. Related to kubernetes-sigs#574, kubernetes-sigs#286 Signed-off-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com> * review: goal-aware DELETED error, fail fast on terminal claim Ready reasons Signed-off-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com> * review: rename goal to wait_target, drop change-narration comments Rename the watch loop's 'goal' local to 'wait_target' in both helpers, replace the byte-identical-message comment with a terse note about the test assertion, and drop the retained-for-compatibility dual-path docstrings from _wait_for_sandbox_ready in both clients. Signed-off-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com> --------- Signed-off-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com> Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
… overwrites via optimistic-locked status writes (kubernetes-sigs#1256) Stateless fix: let the API server arbitrate stale cache views instead of tracking them in controller memory. - The claim status patch carries an optimistic lock. A 409 means the pass computed status from a view that is stale relative to some committed write on the claim (usually this controller's own earlier write, but equally any concurrent writer); the stale patch is dropped as benign, and the conflicting write's own claim watch event re-enqueues the key (getTimingPredicate admits every update; noted at both ends so a future predicate change cannot silently break it). - Ready-transition metrics record only when the pass's status view is authoritative (persisted, or no write needed), closing the duplicate-histogram-observation window durably, including across controller restarts. - A 409 on the adoption-annotation update is retried in-pass on a fresh authoritative read (APIReader) instead of failing the pass and burning the popped warm candidate; the shared fresh-base helper carries the guard/mutate/copy-back shape. - Composes with the merged NotFound handling for claims deleted mid-reconcile (treated as authoritative: nothing to persist, no later pass exists). The adoption-assignment half (optimistic-locked completeAdoption, authoritative resolution, AdoptionConflict reason) lives in a sibling PR based on main; terminal annotation-write conflicts here surface as plain reconciler errors until it lands. Squashed and rebased onto current main (composed with the merged status-write NotFound handling and the sandbox Owns-watch predicate, which filters sandbox events only and does not affect the claim-watch re-enqueue this PR relies on). Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
Copy the approvers list from clients/OWNERS so sandbox-router changes can be approved by the same group. Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
…imistic lock on the adoption patch (kubernetes-sigs#1277) Once the adoption annotation is committed on a claim, a 404/409 while patching the assigned sandbox only proves the pass's cached view of that sandbox is stale — it says nothing about who owns it now. Falling through to the next candidate at that point overwrites the committed assignment, orphans the recorded sandbox, and under sustained load amplifies into assignment-flip storms, burned warm candidates, duplicate binds and doomed writes against deleted objects. Make the adoption transfer safe and terminal instead: - completeAdoption applies the ownership patch with an optimistic lock, so a transfer computed from a stale base is rejected by the server instead of silently re-transferring an already-adopted sandbox. - resolveAdoptionCompletion resolves a completion failure against authoritative reads: accept an already-completed adoption with no write, re-patch once on the fresh base when the candidate is genuinely still pool-owned and adoptable, or clean the dead reference (annotation and, for legacy claims, the deprecated label) and end the pass. Never a retry against a deleted object, never in-pass rebinding to another candidate. - The recovery path in getOrCreateSandbox and the candidate loop both stop falling through to another candidate while the claim still references the committed one. - Terminal contention surfaces as a benign AdoptionConflict Ready reason (carrying the per-case detail) instead of a generic ReconcilerError, paced by the workqueue's per-item rate limiter; exhausted in-pass annotation retries are upgraded to the same reason. - Composes with the merged sibling: reuses authoritativeReader and updateClaimOnFreshBase (removeAssignedSandboxReference is the consumer of the helper's skip-write mode) instead of local reader-fallback duplicates. Regression tests pin the invariants: a completion conflict must be resolved on the SAME candidate, a stale re-patch must degrade to zero effective writes, and a deleted assigned sandbox (annotation- or label-referenced) must produce exactly one doomed write, authoritative cleanup and a clean re-adoption on the next pass. Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
* add mcp server * add missing parts to pyproject.toml * fix Dockerfile * fix content lengh in download file tool * improve errors * add missing field to in-clister settings * rename get_sandboxes unittest case * decode content in upload tool * add docstring to create_sandbox tool * add session_id verification for each tool that uses sandbox * fix upload/download tools docstrings * minor fixes * raise error in None seesion id * remove unneded arguments from get_sandbox function * make a default value for TTL * fix pyproject.toml and Dockefile * Add non-root user to dockerfile, small fixed * remove setuptools-scm * check for None session id in all tools and resources * extend download/upload binary file tests * small fixes * Add env variables prefix * add agent close on shutdown * add TODO about sandbox labels * Add anootated fields to tools. Add default and max timeouts. * ADD env vers prefix to README * add missing env prefix * make tools timeout less or equal * add more descriptive unicode decode error * Add notion about timeouts for long running tasks * add uvicorn version constraint * make create_sandbox tools's timeout option non-optional * add mcp server github actions workflow * add anyio as dependency * add default value for ttl in create_sandbox tool * extract session owner check * add missing authentication warning to README * remove unused import * return back the None for the ttl for create_sadnbox tool, reformat the arguments list * add tests to presubmit workflow
* add new package and main codebase * remove typing.Self * small fixes * add file upload download tests * address the rest of the comments * reformat, fix typos * minor fixes * fix outdated code from sandbox.id mathod * add missing sandbox settings options, fix sandbox label and scope label merging * fix sandbox execute shell command * remove unneded dependency * Copy the labels dict, add warning about possible race in scoped sandboxes * add modes to a file check * change sandbox scope lable prefix * improve multiple sadnboxes found error * add mode check in file upload/download tests * use literals for read and write access mode checks * remove ununsed import * check for lingering claim on sandboxnotfound error * refactor paths conversion * fix warmpool type in README * add tests in the presubmit workflow * add missing arguments to sandbox factory methods * check for new sandbox instances when getting sandbox * fix docstring * refactor sandbox id
… optional claim event/annotation flags (kubernetes-sigs#1250) Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
* docs: expand security threat model * docs: threat model updates for SandboxTemplate/WarmPool enforcement * docs: address review comments on threat model
Extract reusable baseline measurements (baselineColdStart, baselineWarmClaim, baselinePoolFill) and rewire burst recovery and startup comparison tests to use them. Add SANDBOX_LONGEVITY env var for time-bounded soak runs with adaptive batch sizing that self-tunes to the controller's refill rate — decreasing on pool depletion, increasing on recovery. Signed-off-by: vvoronko <vvoronko@redhat.com>
Add longevity soak test mode (SANDBOX_LONGEVITY) with heuristic initial batch sizing max(4, 0.5*poolSize/coldStart), static inter-batch delay coldStart*batchSize/poolSize (floor 50ms), and ±1 adaptive thresholds with wide steady zone (decrease at pool/2, increase at pool-batchSize). Refactor dumpControllerLogs into fetchControllerLogs with optional SinceTime filtering. Add DumpControllerLogsSince for per-pool scoped log capture: unconditional for regular burst, on failure or SANDBOX_DEBUG for longevity. Single API call per pod, tail from buffer. Replace CSV under_1s boolean column with RFC3339 timestamp matching controller log format for direct cross-referencing. Add batch_size column and crash-safe per-batch CSV flush. Introduce summary CSV emitted every 10 batches with p50/p95 latencies and throughput. Override workload duration to max(5, poolSize/2) in longevity mode to prevent pod accumulation. Report output under artifacts/ subdirectory. Update README with longevity mode, controller log capture, all new env vars (SANDBOX_LONGEVITY, SANDBOX_DEBUG), and updated CSV schema. Signed-off-by: vvoronko <vvoronko@redhat.com>
…h fix - Add MinReadyReplicasPredicate to predicates package: checks ReadyReplicas >= MinReady, returns error if minReady > spec.replicas - Add WaitForWarmPoolMinReady to framework/client.go: watch-based via WaitForObject (same pattern as WaitForWarmPoolReady), no polling loop - Longevity mode: wait for 2×batchSize ready instead of full pool fill, skip settle wait, fill continues in background - Fix workload CrashLoopBackOff: derive duration from cold start calibration (max(10, coldStart×5)) instead of pool-size-based formula - Move cold start measurement before template creation so longevity can compute workload duration from it - Replace CSV under_1s column with RFC3339 timestamp matching controller zap format for log correlation - Clean up interBatchDelay: Duration math instead of float64(time.Second) - Replace sort with slices (depguard lint) Signed-off-by: vvoronko <vvoronko@redhat.com>
…ficient Add ShutdownPolicy: Delete with TTLSecondsAfterFinished: 0 to every claim creation site (lifecycle, burst, baseline, benchmarks) via a shared package-level claimLifecycle var. Without this, the API default (Retain or nil lifecycle) leaves finished claims, sandboxes, and VMs alive indefinitely — a resource leak and security gap documented in kubernetes-sigs#1306. Add SANDBOX_TTL env var (default 0) to control the TTL for testing Retain-like behavior with delayed cleanup. Reduce the longevity batch coefficient from 0.5 to 0.3, giving 3x refill capacity over drain rate. The previous 0.5 factor depleted runc pool-40 within 2 minutes on a 4-worker cluster. Update README: add SANDBOX_TTL to env table, fix workload formula to max(10, coldStart×5), update coefficient references, document claim auto-cleanup across all tests, add Design Decisions section linking to kubernetes-sigs#1306. Signed-off-by: vvoronko <vvoronko@redhat.com>
vvoronko
added a commit
that referenced
this pull request
Aug 26, 2026
- Clarify that claimDefaults handles completed-workload cleanup, not running-pod crash recovery - Broaden cold-start triggers beyond pool exhaustion - Fix version skew: nil lifecycle is not Retain, it is no management - Correct Alternative #1 rationale: CRD defaulting only applies when lifecycle is explicitly set Signed-off-by: vvoronko <vvoronko@redhat.com>
vvoronko
added a commit
that referenced
this pull request
Oct 5, 2026
- Clarify that claimDefaults handles completed-workload cleanup, not running-pod crash recovery - Broaden cold-start triggers beyond pool exhaustion - Fix version skew: nil lifecycle is not Retain, it is no management - Correct Alternative #1 rationale: CRD defaulting only applies when lifecycle is explicitly set Signed-off-by: vvoronko <vvoronko@redhat.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds longevity/soak testing mode to
TestRuntimeClassBurstRecoverywith adaptivebatch sizing, and applies
ShutdownPolicy: Delete+TTL=0to all claim creationsites to prevent zombie claim/sandbox/pod accumulation.
Validated across three runtimes: runc (pool-40), gVisor (pool-40), kata-clh (pool-20)
on a 4-worker GCP cluster (4×n2-standard-8, 28 vCPU).
Commits
28ea016— baseline functions, adaptive batch cap, longevity modeExtract reusable
baselineColdStart,baselineWarmClaim,baselinePoolFill.Add
SANDBOX_LONGEVITYenv var for time-bounded soak runs with adaptive batchsizing that self-tunes to the controller's refill rate.
9a13c70— longevity mode, scoped controller logs, adaptive tuningHeuristic initial batch sizing, static inter-batch delay, ±1 adaptive thresholds
with wide steady zone. Scoped controller log capture (
DumpControllerLogsSince).Summary CSV emitted every 10 batches. Workload override and minimum pool size (20).
3791cd3— partial pool fill, watch-based MinReady, workload crash fixMinReadyReplicasPredicate+WaitForWarmPoolMinReady(watch-based, no polling).Longevity waits for 2×batchSize ready instead of full pool fill. Fix workload
CrashLoopBackOff: derive duration from cold start calibration (
max(10, coldStart×5)).bce0322— ShutdownPolicy Delete for all claims, configurable TTL, 0.3 coefficientShared
claimLifecyclevar withShutdownPolicy: Delete+TTL=0(configurablevia
SANDBOX_TTL). Applied to all 5 claim sites (lifecycle, burst, baseline,benchmarks). Batch coefficient 0.5→0.3 for 3× refill margin. See
Default ShutdownPolicy for WarmPool-sourced SandboxClaims should be Delete, not Retain kubernetes-sigs/agent-sandbox#1306 for the security rationale.
Longevity test results (2-minute runs)
New env vars
SANDBOX_LONGEVITY2m,2h)SANDBOX_DEBUGSANDBOX_TTL0Related
Test plan
go build/go vetpass