Skip to content

Retained RocksDB WAL growth is unbounded within a session — the #1400 mitigation re-arms at ~27 GB #2395

Description

@branarakic

Split out of PR #2346 (issue #1400), which sizes the Oxigraph readiness deadline from the retained write-ahead log pending replay. That PR mitigates the boot-kill ratchet; this issue tracks the half it cannot fix.

The residual

Retained WAL grows linearly with one session's write volume and resets only at a successful read-write open:

  • a column family does not flush until write_buffer_size (128 MiB) × min_write_buffer_number_to_merge (2),
  • the live OPTIONS file has max_total_wal_size=0 and WAL_ttl_seconds=0, so nothing bounds retention at runtime,
  • oxigraph serve (pinned 0.5.8) exposes only --location/--bind/--cors/--union-default-graph/--timeout-s — there is no RocksDB knob to pass.

The #2346 deadline is capped at 15 min (deliberately: a failed boot is retried up to 5× by the supervisor, so the single-attempt deadline is amplified — an uncapped deadline means hours of apparent hang). The cap starts binding at ~3.35 GiB retained WAL; past that the safety margin shrinks linearly. At ~27 GB retained WAL (at the slowest throughput ever measured, 30 MB/s; ~9 GB on a 10 MB/s disk) replay no longer fits the capped deadline and the kill-loop from #1400 re-arms — 5 supervised 15-min attempts, then the node is down. That is ~7× the worst retention observed in the field (3.7 GiB), so it is a distant cliff, but a single heavy publish campaign in one long session is the profile that approaches it.

Candidate fixes

  1. Daemon-side bound (no upstream dependency): a successful open truncates the WAL, so a scheduled clean stop/reopen of the managed Oxigraph during idle periods bounds retention by construction. The supervisor already owns restart mechanics; this is a policy timer on top (e.g. reopen when measured retained WAL exceeds N GiB, using measureRetainedWalBytes from fix(cli): size the Oxigraph readiness deadline from the WAL pending replay (#1400) #2346).
  2. Upstream: ask oxigraph to expose max_total_wal_size (or flush-on-SIGTERM). Then the whole class disappears.
  3. Escape hatch already present: explicit readyTimeoutMs in store options is honored verbatim and never capped — a stuck node can be recovered by config today.

Also tracked here

The \d+\.log WAL-layout heuristic in packages/cli/src/daemon/oxigraph-wal.ts is calibrated against the pinned oxigraph 0.5.8 / RocksDB layout. A future OXIGRAPH_VERSION bump carries a revalidation obligation — if the layout changes, the measurement silently degrades to the fixed 30s base (fail-safe direction, but the protection evaporates).

Refs: #1400, #2346.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions