Skip to content

fix(shutdown): cancel in-flight chain-event work before retirement drain #2361

Description

@branarakic

Outcome

Cancel in-flight chain-event poll work promptly when shutdown begins, so a slow RPC or callback cannot hold the agent retirement drain until its outer timeout.

This extracts the live shutdown requirement from PR #2044 onto the current lifecycle, chain-event poller, and RPC cancellation boundaries.

Current state

Current shutdown correctly fences new work, stops the chain poller, waits for physical agent work, and quarantines backing-store teardown on a retirement timeout. DKGNode.stopSignal also cancels protocol reads.

The remaining seam is narrower: ChainEventPoller.stop() waits for the active poll, while event dispatch callbacks such as the KA-to-CG nudge do not receive a poll-generation shutdown signal. A slow chain read inside that callback can therefore consume most or all of the retirement budget.

Related current work:

This issue must reuse those ownership boundaries rather than introduce a parallel cancellation model.

Requirements

  • Close new poll/event admission synchronously before the first shutdown await.
  • Give the active poll generation an abort signal composed with existing RPC/read cancellation.
  • Thread cancellation through event dispatch only as far as required; do not make every domain callback own infrastructure policy.
  • An aborted or partially dispatched page must not advance a durable chain-event cursor past uncompleted work.
  • Preserve atomic verified store operations once their existing non-cancellable commit boundary has begun.
  • Keep timeout quarantine as the final fail-stop guard for adapters that ignore cancellation.

Acceptance criteria

  • Shutdown during event scan, callback chain read, admission wait, and protocol wait aborts promptly and releases the poller drain.
  • No event cursor advances past an aborted callback/page.
  • A late settlement cannot restore lifecycle, recovery, or admission state after shutdown.
  • Existing exact-recovery cancellation and background-retirement behavior remains unchanged.
  • Daemon teardown ordering is covered end to end, not only by a direct callback unit test.
  • No new worker, polling loop, persistent cancellation state, or second retirement coordinator is added.

Provenance

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions