Skip to content

WIP: e2e: longevity mode, adaptive burst recovery, ShutdownPolicy Delete - #1

Closed
vvoronko wants to merge 13 commits into
mainfrom
runtime-test-longevity-debug
Closed

vvoronko wants to merge 13 commits into
mainfrom
runtime-test-longevity-debug

Conversation

@vvoronko

Copy link
Copy Markdown
Owner

Summary

Adds longevity/soak testing mode to TestRuntimeClassBurstRecovery with adaptive
batch sizing, and applies ShutdownPolicy: Delete + TTL=0 to all claim creation
sites to prevent zombie claim/sandbox/pod accumulation.

Validated across three runtimes: runc (pool-40), gVisor (pool-40), kata-clh (pool-20)
on a 4-worker GCP cluster (4×n2-standard-8, 28 vCPU).

Commits

  • 28ea016 — baseline functions, adaptive batch cap, longevity mode
    Extract reusable baselineColdStart, baselineWarmClaim, baselinePoolFill.
    Add SANDBOX_LONGEVITY env var for time-bounded soak runs with adaptive batch
    sizing that self-tunes to the controller's refill rate.

  • 9a13c70 — longevity mode, scoped controller logs, adaptive tuning
    Heuristic initial batch sizing, static inter-batch delay, ±1 adaptive thresholds
    with wide steady zone. Scoped controller log capture (DumpControllerLogsSince).
    Summary CSV emitted every 10 batches. Workload override and minimum pool size (20).

  • 3791cd3 — partial pool fill, watch-based MinReady, workload crash fix
    MinReadyReplicasPredicate + WaitForWarmPoolMinReady (watch-based, no polling).
    Longevity waits for 2×batchSize ready instead of full pool fill. Fix workload
    CrashLoopBackOff: derive duration from cold start calibration (max(10, coldStart×5)).

  • bce0322 — ShutdownPolicy Delete for all claims, configurable TTL, 0.3 coefficient
    Shared claimLifecycle var with ShutdownPolicy: Delete + TTL=0 (configurable
    via SANDBOX_TTL). Applied to all 5 claim sites (lifecycle, burst, baseline,
    benchmarks). Batch coefficient 0.5→0.3 for 3× refill margin. See
    Default ShutdownPolicy for WarmPool-sourced SandboxClaims should be Delete, not Retain kubernetes-sigs/agent-sandbox#1306 for the security rationale.

Longevity test results (2-minute runs)

Metric runc gVisor kata-clh
Cold start 1.42s 1.03s 2.38s
Warm claim 0.32s 0.32s 0.32s
Pool size 40 40 20
Initial batch 8 11 4
Inter-batch delay 283ms 283ms 476ms
Final batch 8 (held) 6 (adapted) 1 (adapted)
Total claims 839 818 108
Under 1s 94.5% 90.5% 98.1%
Throughput 7.0/s 6.8/s 0.8/s
Namespace cleanup 65s 68s 16s

New env vars

Variable Default Description
SANDBOX_LONGEVITY (unset) Go duration for soak mode (e.g. 2m, 2h)
SANDBOX_DEBUG (unset) Dump controller logs on success
SANDBOX_TTL 0 TTL seconds for claim auto-cleanup. Set higher to simulate Retain-like behavior

Related

Test plan

  • go build / go vet pass
  • runc pool-40 longevity 2min — PASS
  • gVisor pool-40 longevity 2min — PASS
  • kata-clh pool-20 longevity 2min — PASS
  • CI presubmit (pending)

aditya-shantanu and others added 13 commits July 28, 2026 17:48
…d polling (kubernetes-sigs#1241)

* feat(sdk): single-watch wait_for_claim_ready + faster dev port-forward polling

Add wait_for_claim_ready() (sync and async) that waits for a SandboxClaim
using a single watch on the claim itself, resolving on the claim's own
status (bound sandbox name + forwarded Ready condition) instead of two
sequential watches (claim then sandbox). The watch is anchored to the
initial read's resourceVersion so no event is missed, and transparently
restarts from "0" on 410 Gone. Legacy two-watch methods are kept fully
intact and API-compatible.

Also reduce the dev port-forward readiness polling interval from 500ms
to 50ms and document SDK ready-wait latency behavior in the README.

Related to kubernetes-sigs#574, kubernetes-sigs#286

Signed-off-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>

* review: goal-aware DELETED error, fail fast on terminal claim Ready reasons

Signed-off-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>

* review: rename goal to wait_target, drop change-narration comments

Rename the watch loop's 'goal' local to 'wait_target' in both helpers,
replace the byte-identical-message comment with a terse note about the
test assertion, and drop the retained-for-compatibility dual-path
docstrings from _wait_for_sandbox_ready in both clients.

Signed-off-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>

---------

Signed-off-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
… overwrites via optimistic-locked status writes (kubernetes-sigs#1256)

Stateless fix: let the API server arbitrate stale cache views instead
of tracking them in controller memory.

- The claim status patch carries an optimistic lock. A 409 means the
  pass computed status from a view that is stale relative to some
  committed write on the claim (usually this controller's own earlier
  write, but equally any concurrent writer); the stale patch is dropped
  as benign, and the conflicting write's own claim watch event
  re-enqueues the key (getTimingPredicate admits every update; noted at
  both ends so a future predicate change cannot silently break it).
- Ready-transition metrics record only when the pass's status view is
  authoritative (persisted, or no write needed), closing the
  duplicate-histogram-observation window durably, including across
  controller restarts.
- A 409 on the adoption-annotation update is retried in-pass on a fresh
  authoritative read (APIReader) instead of failing the pass and
  burning the popped warm candidate; the shared fresh-base helper
  carries the guard/mutate/copy-back shape.
- Composes with the merged NotFound handling for claims deleted
  mid-reconcile (treated as authoritative: nothing to persist, no later
  pass exists).

The adoption-assignment half (optimistic-locked completeAdoption,
authoritative resolution, AdoptionConflict reason) lives in a sibling
PR based on main; terminal annotation-write conflicts here surface as
plain reconciler errors until it lands.

Squashed and rebased onto current main (composed with the merged
status-write NotFound handling and the sandbox Owns-watch predicate,
which filters sandbox events only and does not affect the claim-watch
re-enqueue this PR relies on).

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
Copy the approvers list from clients/OWNERS so sandbox-router changes
can be approved by the same group.

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
…imistic lock on the adoption patch (kubernetes-sigs#1277)

Once the adoption annotation is committed on a claim, a 404/409 while
patching the assigned sandbox only proves the pass's cached view of
that sandbox is stale — it says nothing about who owns it now. Falling
through to the next candidate at that point overwrites the committed
assignment, orphans the recorded sandbox, and under sustained load
amplifies into assignment-flip storms, burned warm candidates,
duplicate binds and doomed writes against deleted objects.

Make the adoption transfer safe and terminal instead:

- completeAdoption applies the ownership patch with an optimistic lock,
  so a transfer computed from a stale base is rejected by the server
  instead of silently re-transferring an already-adopted sandbox.
- resolveAdoptionCompletion resolves a completion failure against
  authoritative reads: accept an already-completed adoption with no
  write, re-patch once on the fresh base when the candidate is
  genuinely still pool-owned and adoptable, or clean the dead
  reference (annotation and, for legacy claims, the deprecated label)
  and end the pass. Never a retry against a deleted object, never
  in-pass rebinding to another candidate.
- The recovery path in getOrCreateSandbox and the candidate loop both
  stop falling through to another candidate while the claim still
  references the committed one.
- Terminal contention surfaces as a benign AdoptionConflict Ready
  reason (carrying the per-case detail) instead of a generic
  ReconcilerError, paced by the workqueue's per-item rate limiter;
  exhausted in-pass annotation retries are upgraded to the same reason.
- Composes with the merged sibling: reuses authoritativeReader and
  updateClaimOnFreshBase (removeAssignedSandboxReference is the
  consumer of the helper's skip-write mode) instead of local
  reader-fallback duplicates.

Regression tests pin the invariants: a completion conflict must be
resolved on the SAME candidate, a stale re-patch must degrade to zero
effective writes, and a deleted assigned sandbox (annotation- or
label-referenced) must produce exactly one doomed write, authoritative
cleanup and a clean re-adoption on the next pass.

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
* add mcp server

* add missing parts to pyproject.toml

* fix Dockerfile

* fix content lengh in download file tool

* improve errors

* add missing field to in-clister settings

* rename get_sandboxes unittest case

* decode content in upload tool

* add docstring to create_sandbox tool

* add session_id verification for each tool that uses sandbox

* fix upload/download tools docstrings

* minor fixes

* raise error in None seesion id

* remove unneded arguments from get_sandbox function

* make a default value for TTL

* fix pyproject.toml and Dockefile

* Add non-root user to dockerfile, small fixed

* remove setuptools-scm

* check for None session id in all tools and resources

* extend download/upload binary file tests

* small fixes

* Add env variables prefix

* add agent close on shutdown

* add TODO about sandbox labels

* Add anootated fields to tools. Add default and max timeouts.

* ADD env vers prefix to README

* add missing env prefix

* make tools timeout less or equal

* add more descriptive unicode decode error

* Add notion about timeouts for long running tasks

* add uvicorn version constraint

* make create_sandbox tools's timeout option non-optional

* add mcp server github actions workflow

* add anyio as dependency

* add default value for ttl in create_sandbox tool

* extract session owner check

* add missing authentication warning to README

* remove unused import

* return back the None for the ttl for create_sadnbox tool, reformat the arguments list

* add tests to presubmit workflow
* add new package and main codebase

* remove typing.Self

* small fixes

* add file upload download tests

* address the rest of the comments

* reformat, fix typos

* minor fixes

* fix outdated code from sandbox.id mathod

* add missing sandbox settings options, fix sandbox label and scope label merging

* fix sandbox execute shell command

* remove unneded dependency

* Copy the labels dict, add warning about possible race in scoped sandboxes

* add modes to a file check

* change sandbox scope lable prefix

* improve multiple sadnboxes found error

* add mode check in file upload/download tests

* use literals for read and write access mode checks

* remove ununsed import

* check for lingering claim on sandboxnotfound error

* refactor paths conversion

* fix warmpool type in README

* add tests in the presubmit workflow

* add missing arguments to sandbox factory methods

* check for new sandbox instances when getting sandbox

* fix docstring

* refactor sandbox id
… optional claim event/annotation flags (kubernetes-sigs#1250)

Co-authored-by: Aditya Shantanu <aditya-shantanu@users.noreply.github.com>
* docs: expand security threat model

* docs: threat model updates for SandboxTemplate/WarmPool enforcement

* docs: address review comments on threat model
Extract reusable baseline measurements (baselineColdStart,
baselineWarmClaim, baselinePoolFill) and rewire burst recovery
and startup comparison tests to use them.

Add SANDBOX_LONGEVITY env var for time-bounded soak runs with
adaptive batch sizing that self-tunes to the controller's refill
rate — decreasing on pool depletion, increasing on recovery.

Signed-off-by: vvoronko <vvoronko@redhat.com>
Add longevity soak test mode (SANDBOX_LONGEVITY) with heuristic initial
batch sizing max(4, 0.5*poolSize/coldStart), static inter-batch delay
coldStart*batchSize/poolSize (floor 50ms), and ±1 adaptive thresholds
with wide steady zone (decrease at pool/2, increase at pool-batchSize).

Refactor dumpControllerLogs into fetchControllerLogs with optional
SinceTime filtering. Add DumpControllerLogsSince for per-pool scoped
log capture: unconditional for regular burst, on failure or SANDBOX_DEBUG
for longevity. Single API call per pod, tail from buffer.

Replace CSV under_1s boolean column with RFC3339 timestamp matching
controller log format for direct cross-referencing. Add batch_size
column and crash-safe per-batch CSV flush. Introduce summary CSV
emitted every 10 batches with p50/p95 latencies and throughput.

Override workload duration to max(5, poolSize/2) in longevity mode
to prevent pod accumulation. Report output under artifacts/ subdirectory.

Update README with longevity mode, controller log capture, all new
env vars (SANDBOX_LONGEVITY, SANDBOX_DEBUG), and updated CSV schema.

Signed-off-by: vvoronko <vvoronko@redhat.com>
…h fix

- Add MinReadyReplicasPredicate to predicates package: checks
  ReadyReplicas >= MinReady, returns error if minReady > spec.replicas
- Add WaitForWarmPoolMinReady to framework/client.go: watch-based via
  WaitForObject (same pattern as WaitForWarmPoolReady), no polling loop
- Longevity mode: wait for 2×batchSize ready instead of full pool fill,
  skip settle wait, fill continues in background
- Fix workload CrashLoopBackOff: derive duration from cold start
  calibration (max(10, coldStart×5)) instead of pool-size-based formula
- Move cold start measurement before template creation so longevity can
  compute workload duration from it
- Replace CSV under_1s column with RFC3339 timestamp matching controller
  zap format for log correlation
- Clean up interBatchDelay: Duration math instead of float64(time.Second)
- Replace sort with slices (depguard lint)

Signed-off-by: vvoronko <vvoronko@redhat.com>
…ficient

Add ShutdownPolicy: Delete with TTLSecondsAfterFinished: 0 to every
claim creation site (lifecycle, burst, baseline, benchmarks) via a
shared package-level claimLifecycle var. Without this, the API default
(Retain or nil lifecycle) leaves finished claims, sandboxes, and VMs
alive indefinitely — a resource leak and security gap documented in
kubernetes-sigs#1306.

Add SANDBOX_TTL env var (default 0) to control the TTL for testing
Retain-like behavior with delayed cleanup.

Reduce the longevity batch coefficient from 0.5 to 0.3, giving 3x
refill capacity over drain rate. The previous 0.5 factor depleted
runc pool-40 within 2 minutes on a 4-worker cluster.

Update README: add SANDBOX_TTL to env table, fix workload formula to
max(10, coldStart×5), update coefficient references, document claim
auto-cleanup across all tests, add Design Decisions section linking
to kubernetes-sigs#1306.

Signed-off-by: vvoronko <vvoronko@redhat.com>
@vvoronko vvoronko closed this Jul 29, 2026
vvoronko added a commit that referenced this pull request Aug 26, 2026
- Clarify that claimDefaults handles completed-workload cleanup, not
  running-pod crash recovery
- Broaden cold-start triggers beyond pool exhaustion
- Fix version skew: nil lifecycle is not Retain, it is no management
- Correct Alternative #1 rationale: CRD defaulting only applies when
  lifecycle is explicitly set

Signed-off-by: vvoronko <vvoronko@redhat.com>
vvoronko added a commit that referenced this pull request Oct 5, 2026
- Clarify that claimDefaults handles completed-workload cleanup, not
  running-pod crash recovery
- Broaden cold-start triggers beyond pool exhaustion
- Fix version skew: nil lifecycle is not Retain, it is no management
- Correct Alternative #1 rationale: CRD defaulting only applies when
  lifecycle is explicitly set

Signed-off-by: vvoronko <vvoronko@redhat.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants