Skip to content

fix(workflows): recover exhausted autonomous work after crashes - #112

Merged
duc15052006-dotcom merged 17 commits into
verify/release-candidatefrom
fix/note-21-2-workflow-exhausted-crash-recovery
Sep 20, 2026
Merged

duc15052006-dotcom merged 17 commits into
verify/release-candidatefrom
fix/note-21-2-workflow-exhausted-crash-recovery

Conversation

@duc15052006-dotcom

Copy link
Copy Markdown
Owner

Closes the NOTE 21-2 crash window where a worker can die on its final queue attempt after the durable workflow state already moved forward.

What changes:

  • add a work-queue exhausted cleanup claim lane that leases attempts >= maxAttempts without incrementing attempts;
  • keep normal execution capped: exhausted items are cleanup/reconciliation only, never dispatched again as ordinary side effects;
  • add exact ready-state failure by ready timestamp + attempt;
  • add exact autonomous running-state failure by durable dispatch stamp + attempt;
  • reconcile exhausted ready items that crashed before or after ready->running CAS;
  • reconcile exhausted wait items that crashed before or after waiting->running CAS;
  • finish stale cleanup items harmlessly and release transient cleanup failures for another cleanup pass;
  • prove an old wait stamp cannot fail a newer resumed state even when the workflow attempt number is unchanged;
  • add queue/store/bridge regression coverage, docs and release-preflight invariants.

No second scheduler, credential path, browser bypass or host filesystem path is added.

Base: exact verify/release-candidate HEAD 985229b.
Do not merge aggregate PR #30.

@duc15052006-dotcom
duc15052006-dotcom merged commit 4da0bae into verify/release-candidate Sep 20, 2026
17 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant