Repository navigation
Re-recruit an old epoch's backup worker when it dies (ParallelRestoreNewBackupCorrectnessAtomicOp failure) - #14115
Closed
saintstack wants to merge 1 commit into
Closed
saintstack wants to merge 1 commit into
saintstack wants to merge 1 commit into
Conversation
ae0706e (apple#13939) stopped an old-epoch backup worker failure from forcing a transaction-system recovery. Recovery was also the only caller of recruitBackupWorkers(), the sole consumer of BackupProgress::getUnfinishedBackup(), so nothing re-recruited the unfinished work and the dead worker's slot pinned oldestBackupEpoch indefinitely. That defers TLog pops and freezes latestBackupWorkerSavedVersion, which is the restorable version for PARTITIONED_LOG, so BackupLogsDispatchTask never satisfies stopWhenDone && restorableVersion.present() and such a backup never completes. Seen on tests/slow/ParallelRestoreNewBackupCorrectnessAtomicOp.toml seed 1003233285: epoch-8 tag -2:0 worker ac570a326c42885d died at t=287.4 before saveProgress, no further recovery followed, and RestorableVersion stayed -1 for 12,300 simsec until TracedTooManyLines aborted the run. The seed now passes in 593 simsec with oldestBackupEpoch advancing 8 -> 10 -> 12. replaceBackupWorker swaps the interface in place so the epoch keeps its hold on oldestBackupEpoch while its work is outstanding. Knob CC_RERECRUIT_BACKUP_WORKER_ENABLED (default true) disables the monitor. This does not address a second, independent cause of the same pin: a live old-epoch worker that never completes its range, reproducible on tests/slow/BackupNewAndOldRestore.toml seed 2702875306.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
saintstack
force-pushed
the
rerecruit-old-epoch-backup-worker
branch
from
September 24, 2026 00:43
ca2e440 to
cadd1ac
Compare
Contributor
Result of foundationdb-pr-clang-arm on Linux RHEL 9
|
Contributor
Result of foundationdb-pr-clang on Linux RHEL 9
|
Contributor
Result of foundationdb-pr on Linux RHEL 9
|
Contributor
Result of foundationdb-pr-macos on macOS 14.x
|
Contributor
Result of foundationdb-pr-cluster-tests on Linux RHEL 9
|
Contributor
Result of foundationdb-pr-macos-m1 on macOS 14.x
|
Contributor
Author
|
Closing in favor of #14240, a more comprehensive version of this approach against main first; will backport after it lands. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ParallelRestoreNewBackupCorrectnessAtomicOp failed 7.4 nightly.
A
stopWhenDonepartitioned-log backup can hang forever, and the simulation run dies onTracedTooManyLinesrather than on the actual defect.BackupLogsDispatchTaskonly finishes whenstopWhenDone && restorableVersion.present()(
fdbclient/FileBackupAgent.actor.cpp:3676). ForMutationLogType::PARTITIONED_LOGthe restorableversion comes from
latestBackupWorkerSavedVersion(BackupAgent.actor.h:954), whichmonitorBackupProgressonly advances whilerecruitedEpoch == oldestBackupEpoch(
fdbserver/BackupWorker.actor.cpp:614, re-checked insetBackupKeysat:574). So ifoldestBackupEpochis pinned below the live epoch, the key freezes, the backup never becomesrestorable, and
waitBackupnever returns.Root cause
ae0706e016(#13939) removed thewaitFailureClientmonitoring of old-epoch backup workers fromTagPartitionedLogSystem::monitorLogSystem, so that an old-epoch backup worker failure no longerforces a transaction-system recovery. That intent is sound — repeated recoveries hurt availability.
The unintended side effect is that recovery was also the only thing that ever re-recruited
unfinished old-epoch work:
recruitBackupWorkers()(fdbserver/ClusterRecovery.actor.cpp) runs onceper recovery and is the sole consumer of
BackupProgress::getUnfinishedBackup(). With the failuresignal gone and no later recovery, a dead old-epoch worker's entry stays in
LogSet::backupWorkersforever and pins
oldestBackupEpoch. That also defers TLog pops indefinitely.Observed on
tests/slow/ParallelRestoreNewBackupCorrectnessAtomicOp.tomlseed1003233285: theepoch-8 tag
-2:0workerac570a326c42885dwas killed by machine failure at t=287.38 beforesaveProgress; its two peers reported done; zero recoveries followed.oldestBackupEpochstayed 8,BackupWorkerSetVersionfired once at t=276 and never again,RestorableVersionstayed-1across185
FileBackupLogDispatchevents, and the run traced 1,000,001 lines and aborted at t=12586.7.Collateral:
BackupWorkerPopDeferredx1263 andDiskNearCapacityatAvailableSpaceRatio 0.117.Solution
monitorOldEpochBackupWorkerruns one actor per old-epoch recruit. OnwaitFailureClientit re-readsdurable progress and either re-recruits from
max(startVersion, savedVersion + 1), or — if progressalready covers
endVersion— releases the slot, covering the case where the worker died afterfinishing but before reporting done.
ILogSystem::replaceBackupWorkerswaps the interface in place, so the epoch's entry count isunchanged and
oldestBackupEpochis never recomputed while work is outstanding.precisely what Avoid recovery on old backup worker failure #13939 set out to avoid.
CC_RERECRUIT_BACKUP_WORKER_ENABLED(defaulttrue), mirroringCC_RERECRUIT_LOG_ROUTER_ENABLEDonmain.Testing
1003233285now passes in 592.8 simsec (was 12586.7 + abort), 0SevError, traces 33 MB vs560 MB.
oldestBackupEpochadvances 8 → 10 → 12;BackupWorkerPopDeferred1263 → 3;BackupWorkerSetVersion1 → 4;RestorableVersionreaches 654432137.BackupWorkerReplacement/ReplaceBackupWorkerfire for thesame dead worker UID
ac570a326c42885d, and the newCODE_PROBEs reportCovered="1"on thatseed and
Covered="0"on seeds that never kill an old-epoch worker.mutationLogType = 1tests. The single failure was an unrelated fast-restore applier assertion(
batchData.isValid(),RestoreApplier.actor.cpp:740).passes on seeds
1003233285/471182304/12345678.Scope limit
This fixes one of at least two independent causes of a pinned
oldestBackupEpoch. It does notaddress a live old-epoch worker that never completes its range — there is no dead worker to replace,
so failure detection cannot see it; that needs progress-stall detection. Reproducer:
tests/slow/BackupNewAndOldRestore.tomlseed2702875306, where the epoch-10 workers79ad10c0a37972c7/3c414ce27aa08121stay healthy to t=17507 without finishing.This needs a forward-port...