Skip to content

Re-recruit an old epoch's backup worker when it dies (ParallelRestoreNewBackupCorrectnessAtomicOp failure) - #14115

Closed
saintstack wants to merge 1 commit into
apple:release-7.4from
saintstack:rerecruit-old-epoch-backup-worker
Closed

saintstack wants to merge 1 commit into
apple:release-7.4from
saintstack:rerecruit-old-epoch-backup-worker

Conversation

@saintstack

@saintstack saintstack commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

ParallelRestoreNewBackupCorrectnessAtomicOp failed 7.4 nightly.

A stopWhenDone partitioned-log backup can hang forever, and the simulation run dies on
TracedTooManyLines rather than on the actual defect.

BackupLogsDispatchTask only finishes when stopWhenDone && restorableVersion.present()
(fdbclient/FileBackupAgent.actor.cpp:3676). For MutationLogType::PARTITIONED_LOG the restorable
version comes from latestBackupWorkerSavedVersion (BackupAgent.actor.h:954), which
monitorBackupProgress only advances while recruitedEpoch == oldestBackupEpoch
(fdbserver/BackupWorker.actor.cpp:614, re-checked in setBackupKeys at :574). So if
oldestBackupEpoch is pinned below the live epoch, the key freezes, the backup never becomes
restorable, and waitBackup never returns.

Root cause

ae0706e016 (#13939) removed the waitFailureClient monitoring of old-epoch backup workers from
TagPartitionedLogSystem::monitorLogSystem, so that an old-epoch backup worker failure no longer
forces a transaction-system recovery. That intent is sound — repeated recoveries hurt availability.

The unintended side effect is that recovery was also the only thing that ever re-recruited
unfinished old-epoch work: recruitBackupWorkers() (fdbserver/ClusterRecovery.actor.cpp) runs once
per recovery and is the sole consumer of BackupProgress::getUnfinishedBackup(). With the failure
signal gone and no later recovery, a dead old-epoch worker's entry stays in LogSet::backupWorkers
forever and pins oldestBackupEpoch. That also defers TLog pops indefinitely.

Observed on tests/slow/ParallelRestoreNewBackupCorrectnessAtomicOp.toml seed 1003233285: the
epoch-8 tag -2:0 worker ac570a326c42885d was killed by machine failure at t=287.38 before
saveProgress; its two peers reported done; zero recoveries followed. oldestBackupEpoch stayed 8,
BackupWorkerSetVersion fired once at t=276 and never again, RestorableVersion stayed -1 across
185 FileBackupLogDispatch events, and the run traced 1,000,001 lines and aborted at t=12586.7.
Collateral: BackupWorkerPopDeferred x1263 and DiskNearCapacity at AvailableSpaceRatio 0.117.

Solution

monitorOldEpochBackupWorker runs one actor per old-epoch recruit. On waitFailureClient it re-reads
durable progress and either re-recruits from max(startVersion, savedVersion + 1), or — if progress
already covers endVersion — releases the slot, covering the case where the worker died after
finishing but before reporting done.

  • ILogSystem::replaceBackupWorker swaps the interface in place, so the epoch's entry count is
    unchanged and oldestBackupEpoch is never recomputed while work is outstanding.
  • Errors are caught and retried, never thrown: an escaping error would fail cluster recovery, which is
    precisely what Avoid recovery on old backup worker failure #13939 set out to avoid.
  • Gated by CC_RERECRUIT_BACKUP_WORKER_ENABLED (default true), mirroring
    CC_RERECRUIT_LOG_ROUTER_ENABLED on main.

Testing

  • Seed 1003233285 now passes in 592.8 simsec (was 12586.7 + abort), 0 SevError, traces 33 MB vs
    560 MB. oldestBackupEpoch advances 8 → 10 → 12; BackupWorkerPopDeferred 1263 → 3;
    BackupWorkerSetVersion 1 → 4; RestorableVersion reaches 654432137.
  • The repair is observed, not inferred: BackupWorkerReplacement / ReplaceBackupWorker fire for the
    same dead worker UID ac570a326c42885d, and the new CODE_PROBEs report Covered="1" on that
    seed and Covered="0" on seeds that never kill an old-epoch worker.
  • Two joshua 100k ensembles, 199,999 passes: one full-suite (267 test files), one rebundled to the 8
    mutationLogType = 1 tests. The single failure was an unrelated fast-restore applier assertion
    (batchData.isValid(), RestoreApplier.actor.cpp:740).
  • Draw-neutral: elapsed simsec is byte-identical across the knob, the probes, and the simplification
    passes on seeds 1003233285 / 471182304 / 12345678.

Scope limit

This fixes one of at least two independent causes of a pinned oldestBackupEpoch. It does not
address a live old-epoch worker that never completes its range — there is no dead worker to replace,
so failure detection cannot see it; that needs progress-stall detection. Reproducer:
tests/slow/BackupNewAndOldRestore.toml seed 2702875306, where the epoch-10 workers
79ad10c0a37972c7 / 3c414ce27aa08121 stay healthy to t=17507 without finishing.

This needs a forward-port...

@saintstack
saintstack requested a review from spraza as a code owner September 23, 2026 15:52
@saintstack saintstack changed the title Re-recruit an old epoch's backup worker when it dies Re-recruit an old epoch's backup worker when it dies (ParallelRestoreNewBackupCorrectnessAtomicOp failure) Sep 23, 2026
@saintstack saintstack added nightly correctness bugs nightlies Issues to address failures in the nighty runs. labels Sep 23, 2026
ae0706e (apple#13939) stopped an old-epoch backup worker failure from forcing a
transaction-system recovery. Recovery was also the only caller of
recruitBackupWorkers(), the sole consumer of BackupProgress::getUnfinishedBackup(),
so nothing re-recruited the unfinished work and the dead worker's slot pinned
oldestBackupEpoch indefinitely. That defers TLog pops and freezes
latestBackupWorkerSavedVersion, which is the restorable version for
PARTITIONED_LOG, so BackupLogsDispatchTask never satisfies
stopWhenDone && restorableVersion.present() and such a backup never completes.

Seen on tests/slow/ParallelRestoreNewBackupCorrectnessAtomicOp.toml seed
1003233285: epoch-8 tag -2:0 worker ac570a326c42885d died at t=287.4 before
saveProgress, no further recovery followed, and RestorableVersion stayed -1 for
12,300 simsec until TracedTooManyLines aborted the run. The seed now passes in
593 simsec with oldestBackupEpoch advancing 8 -> 10 -> 12.

replaceBackupWorker swaps the interface in place so the epoch keeps its hold on
oldestBackupEpoch while its work is outstanding. Knob
CC_RERECRUIT_BACKUP_WORKER_ENABLED (default true) disables the monitor.

This does not address a second, independent cause of the same pin: a live
old-epoch worker that never completes its range, reproducible on
tests/slow/BackupNewAndOldRestore.toml seed 2702875306.
@foundationdb-ci

This comment has been minimized.

@foundationdb-ci

This comment has been minimized.

@foundationdb-ci

This comment has been minimized.

@foundationdb-ci

This comment has been minimized.

@foundationdb-ci

This comment has been minimized.

@foundationdb-ci

This comment has been minimized.

@saintstack
saintstack force-pushed the rerecruit-old-epoch-backup-worker branch from ca2e440 to cadd1ac Compare September 24, 2026 00:43
@foundationdb-ci

Copy link
Copy Markdown
Contributor

Result of foundationdb-pr-clang-arm on Linux RHEL 9

  • Commit ID: cadd1ac
  • Duration 0:45:56
  • Result: ✅ SUCCEEDED
  • Error: N/A
  • Build Log terminal output (available for 30 days)
  • Build Workspace zip file of the working directory (available for 30 days)

@foundationdb-ci

Copy link
Copy Markdown
Contributor

Result of foundationdb-pr-clang on Linux RHEL 9

  • Commit ID: cadd1ac
  • Duration 0:49:02
  • Result: ❌ FAILED
  • Error: Error while executing command: if python3 -m joshua.joshua list --stopped | grep ${ENSEMBLE_ID} | grep -q 'pass=10[0-9][0-9][0-9]'; then echo PASS; else echo FAIL && exit 1; fi. Reason: exit status 1
  • Build Log terminal output (available for 30 days)
  • Build Workspace zip file of the working directory (available for 30 days)

@foundationdb-ci

Copy link
Copy Markdown
Contributor

Result of foundationdb-pr on Linux RHEL 9

  • Commit ID: cadd1ac
  • Duration 0:57:16
  • Result: ✅ SUCCEEDED
  • Error: N/A
  • Build Log terminal output (available for 30 days)
  • Build Workspace zip file of the working directory (available for 30 days)

@foundationdb-ci

Copy link
Copy Markdown
Contributor

Result of foundationdb-pr-macos on macOS 14.x

  • Commit ID: cadd1ac
  • Duration 1:49:09
  • Result: ✅ SUCCEEDED
  • Error: N/A
  • Build Log terminal output (available for 30 days)
  • Build Workspace zip file of the working directory (available for 30 days)

@foundationdb-ci

Copy link
Copy Markdown
Contributor

Result of foundationdb-pr-cluster-tests on Linux RHEL 9

  • Commit ID: cadd1ac
  • Duration 1:49:33
  • Result: ✅ SUCCEEDED
  • Error: N/A
  • Build Log terminal output (available for 30 days)
  • Build Workspace zip file of the working directory (available for 30 days)
  • Cluster Test Logs zip file of the test logs (available for 30 days)

@foundationdb-ci

Copy link
Copy Markdown
Contributor

Result of foundationdb-pr-macos-m1 on macOS 14.x

  • Commit ID: cadd1ac
  • Duration 1:50:37
  • Result: ✅ SUCCEEDED
  • Error: N/A
  • Build Log terminal output (available for 30 days)
  • Build Workspace zip file of the working directory (available for 30 days)

@saintstack

Copy link
Copy Markdown
Contributor Author

Closing in favor of #14240, a more comprehensive version of this approach against main first; will backport after it lands.

@saintstack saintstack closed this Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

nightlies Issues to address failures in the nighty runs. nightly correctness bugs

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants