Skip to content

Commit ffc4ac1

Browse files
authored
perf(alloc): compile jemalloc's never-purge policy into the binary (#996)
* perf(alloc): compile jemalloc's never-purge policy into the binary jemalloc is unchanged and stays: it is here because the platform allocator keeps freed arena chunks resident, and the recursion campaign measured the same proves reading up to 13 GiB higher under glibc. What changes is its decay policy. Each jemalloc `#[global_allocator]` site now exports jemalloc's compile-time configuration string beside it: dirty_decay_ms:-1,muzzy_decay_ms:-1 The decay timers hand freed pages back to the OS. The prover allocates and frees multi-hundred-MiB host buffers continuously, so those pages are re-faulted almost immediately, and the fault lands on the worker threads doing the proving. The default policy cost about 13 M extra minor faults and about 15 s of extra system time per run on the arms below. Measured on an RTX 5090 box on the recursion campaign's branches, ABBA in each, every arm at one commit and one set of knobs: * WHIR prover (keccak, whir/lfm @ 64393da) — 39.69-39.88 s a block with the setting against 43.94-44.07 s without. * Per-table STARK tree (0e4f461) — 187 / 163 / 166 / 171 s, both never-purge arms under both default arms, 5-24 s a block, proof bytes unmoved. The cost is peak RSS: +1.4 GiB and +2.9-3.3 GiB respectively, an order below the 13 GiB the allocator choice itself is worth, which is why the lever is the decay setting and not the allocator. `background_thread:true` recovers none of it: the cost is the re-touch, not the `madvise` call. This branch's own pipeline has not been measured under the setting. `MALLOC_CONF` / `_RJEM_MALLOC_CONF` in the environment still override the compiled-in default, which is how a benchmark arm puts the old policy back. Nothing about the export is compiler-checked: a misspelled symbol, a wrong value type, or a jemalloc built without the `_rjem_` prefix each leave a binary that links, runs, and quietly purges. `prover/tests/jemalloc_conf.rs` reads `opt.dirty_decay_ms` and `opt.muzzy_decay_ms` back out of the allocator serving the test process and asserts both are -1, after asserting neither environment variable is set so it cannot pass for the wrong reason. It is its own test binary because the check needs a jemalloc process of its own; that the shipped binary carries the symbol is a link-time property, read with `nm`. No proof byte moves: an allocator is not an input to any transcript. * Correct four comment claims in the never-purge export The escape hatch named a variable this binary does not read. `tikv-jemalloc-sys` builds with `--with-jemalloc-prefix=_rjem_` under default features, and jemalloc picks exactly one env name at configure time (`obtain_malloc_conf`, source 3), so only `_RJEM_MALLOC_CONF` is read here — plain `MALLOC_CONF` is inert. An arm that followed the comment to "put the default policy back" would have measured never-purge on both sides, with no warning, and read the lever as dead. The prefixed file source `/etc/_rjem_malloc.conf` is now named too, since it is the one remaining way to set `opt.*` from outside the binary. `jemalloc_conf.rs` was described in two places as what fails if the export stops being read. It carries its own copy of the block and reads its own process, so deleting either production export leaves it green: it pins the pattern, not the two sites that ship it. Both comments now say that, and the test's module doc says it of itself. "An order below the 13 GiB" holds for the +1.4 GiB figure (9x) but not for +2.9-3.3 GiB (4x). Now "a ninth to a quarter of". "+13 M minor faults and +15 s of system time per arm" sits under the WHIR bullet and is that arm's figure; the body says "arms". Scoped to the arm it belongs to. The only non-comment change is the order of the two names in the env guard, so the prefixed one — the one that can actually affect the reading — is checked first. Both still have to be unset; plain `MALLOC_CONF` stays guarded so a future unprefixed build cannot pass here silently.
1 parent c2ac5d5 commit ffc4ac1

3 files changed

Lines changed: 146 additions & 0 deletions

File tree

‎bin/cli/src/main.rs‎

Lines changed: 49 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,55 @@ use clap::{Parser, Subcommand, ValueHint};
1010

1111
#[global_allocator]
1212
static ALLOC: tikv_jemallocator::Jemalloc = tikv_jemallocator::Jemalloc;
13+
14+
// jemalloc, never purging.
15+
//
16+
// The allocator itself is unchanged: jemalloc is here because the platform
17+
// allocator keeps freed arena chunks resident, and on the recursion campaign's
18+
// branch the same proves read up to 13 GiB higher under glibc. What this sets
19+
// is jemalloc's *decay* timers, which hand freed pages back to the OS. The
20+
// prover allocates and frees multi-hundred-MiB host buffers continuously, so
21+
// those pages come straight back as minor faults on the worker threads.
22+
//
23+
// Measured on an RTX 5090 box on the recursion campaign's branches, ABBA in
24+
// each, every arm at one commit and one set of knobs:
25+
// * the WHIR prover (keccak, `whir/lfm` @ 64393da9) — 39.69-39.88 s a block
26+
// with this setting against 43.94-44.07 s without, the default costing
27+
// +13 M minor faults and +15 s of system time per run on this arm;
28+
// * the per-table STARK tree (0e4f4610) — 187 / 163 / 166 / 171 s, both
29+
// never-purge arms under both default arms, 5-24 s a block, with the proof
30+
// bytes unmoved (30 identical lines, 0 differing).
31+
// The cost is peak RSS: +1.4 GiB and +2.9-3.3 GiB respectively, a ninth to a
32+
// quarter of the 13 GiB the allocator choice itself is worth — which is why the
33+
// lever is the decay setting and not the allocator. `background_thread:true` recovers
34+
// none of it: the cost is the re-touch, not the `madvise` call.
35+
//
36+
// This binary's own pipeline has not been measured under the setting; the
37+
// numbers above are from the campaign's branches, where the prover's
38+
// allocation pattern is the same.
39+
//
40+
// `_RJEM_MALLOC_CONF` in the environment still overrides this, which is how a
41+
// measurement arm puts the default policy back. It has to be that spelling:
42+
// `tikv-jemalloc-sys` builds with `--with-jemalloc-prefix=_rjem_` under default
43+
// features, and jemalloc then reads one env name chosen at configure time
44+
// (`jemalloc.c`, `obtain_malloc_conf` source 3) — so plain `MALLOC_CONF` is read
45+
// by nothing here and sets an arm to the default policy without saying it did
46+
// not. The file source is prefixed too: `/etc/_rjem_malloc.conf`.
47+
//
48+
// jemalloc reads this symbol as a `const char *` before `main` is entered, so
49+
// the value has to be in the initializer, and the name is the prefixed one
50+
// `tikv-jemalloc-sys` declares (`#[cfg_attr(prefixed, link_name =
51+
// "_rjem_malloc_conf")]`, its `src/lib.rs`). None of that is compiler-checked.
52+
// `prover/tests/jemalloc_conf.rs` reads both options back out of jemalloc, but
53+
// it carries its own copy of this block and reads its own process — it pins the
54+
// pattern, not this export. Deleting the lines below turns nothing red.
55+
const NEVER_PURGE: &[u8] = b"dirty_decay_ms:-1,muzzy_decay_ms:-1\0";
56+
57+
#[allow(non_upper_case_globals)]
58+
#[unsafe(export_name = "_rjem_malloc_conf")]
59+
pub static malloc_conf: Option<&'static core::ffi::c_char> =
60+
Some(unsafe { &*(NEVER_PURGE.as_ptr() as *const core::ffi::c_char) });
61+
1362
use executor::vm::instruction::decoding::Instruction;
1463
use executor::vm::instruction::execution::{Accelerator, SyscallNumbers};
1564
use executor::{elf::Elf, flamegraph::FlamegraphGenerator, vm::execution::Executor};

‎prover/tests/calibration.rs‎

Lines changed: 16 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -21,6 +21,22 @@ use tikv_jemalloc_ctl::{epoch, stats};
2121
#[global_allocator]
2222
static ALLOC: tikv_jemallocator::Jemalloc = tikv_jemallocator::Jemalloc;
2323

24+
// ...with the shipped binary's purge policy, so this binary is the production
25+
// allocator *configuration* and not just the production allocator. The reason
26+
// and the numbers are at `bin/cli/src/main.rs`. `prover/tests/jemalloc_conf.rs`
27+
// asserts that this export pattern is read, but it does so against its own copy
28+
// in its own process — nothing checks the copy below.
29+
//
30+
// It moves nothing this file asserts — `stats::allocated` is live bytes, which
31+
// the decay timers do not touch; a resident-memory assertion added here later
32+
// would read the wrong configuration without it.
33+
const NEVER_PURGE: &[u8] = b"dirty_decay_ms:-1,muzzy_decay_ms:-1\0";
34+
35+
#[allow(non_upper_case_globals)]
36+
#[unsafe(export_name = "_rjem_malloc_conf")]
37+
pub static malloc_conf: Option<&'static core::ffi::c_char> =
38+
Some(unsafe { &*(NEVER_PURGE.as_ptr() as *const core::ffi::c_char) });
39+
2440
fn allocated_bytes() -> usize {
2541
epoch::advance().ok();
2642
stats::allocated::read().unwrap_or(0)

‎prover/tests/jemalloc_conf.rs‎

Lines changed: 81 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,81 @@
1+
//! jemalloc's purge policy is compiled into the binary — checked by reading it
2+
//! back out of the allocator serving this process.
3+
//!
4+
//! `bin/cli/src/main.rs` and `calibration.rs` each export
5+
//! `_rjem_malloc_conf = "dirty_decay_ms:-1,muzzy_decay_ms:-1"` beside their
6+
//! `#[global_allocator]`, so the never-purge policy travels in the binary
7+
//! rather than in a launcher's environment. Nothing about that export is
8+
//! checked by the compiler: a misspelled symbol, a wrong value type, or a
9+
//! jemalloc built without the `_rjem_` prefix each leave a binary that
10+
//! compiles, links, runs — and quietly purges.
11+
//!
12+
//! This is its own test binary because the check needs a jemalloc process of
13+
//! its own: the prover's lib tests run under the platform allocator, where a
14+
//! `mallctl` read would say nothing, and `calibration.rs` is behind
15+
//! `disk-spill` and pays for a full proof. What it pins is the export pattern —
16+
//! symbol, type, initializer, edition spelling — in the copy below, which is
17+
//! byte-identical to the two production sites but not mechanically tied to
18+
//! them: delete either of those and this still passes. It is a self-test of the
19+
//! pattern, not a regression guard on the two sites that ship it. That the
20+
//! shipped `cli` binary carries the symbol is a link-time property, read with
21+
//! `nm` rather than asserted here.
22+
23+
use tikv_jemalloc_ctl::raw;
24+
25+
#[global_allocator]
26+
static ALLOC: tikv_jemallocator::Jemalloc = tikv_jemallocator::Jemalloc;
27+
28+
/// The same string the two production sites export, character for character.
29+
const NEVER_PURGE: &[u8] = b"dirty_decay_ms:-1,muzzy_decay_ms:-1\0";
30+
31+
#[allow(non_upper_case_globals)]
32+
#[unsafe(export_name = "_rjem_malloc_conf")]
33+
pub static malloc_conf: Option<&'static core::ffi::c_char> =
34+
Some(unsafe { &*(NEVER_PURGE.as_ptr() as *const core::ffi::c_char) });
35+
36+
/// jemalloc's default `opt.dirty_decay_ms`, quoted in the failure message so a
37+
/// red test says which value it found and where that value comes from.
38+
const DEFAULT_DIRTY_DECAY_MS: isize = 10_000;
39+
40+
#[test]
41+
fn jemalloc_never_purge_is_compiled_in() {
42+
// `_RJEM_MALLOC_CONF` sets these same options from the environment, and
43+
// benchmark runs do set it; with it set, reading `-1` back would say nothing
44+
// about the compiled-in export, so refuse to run rather than pass for the
45+
// wrong reason. Plain `MALLOC_CONF` is inert in this prefixed build — guarded
46+
// anyway, so that a future unprefixed build does not silently pass here.
47+
// Not covered: `/etc/_rjem_malloc.conf`, the one remaining source that could
48+
// set `opt.*` from outside this binary.
49+
for var in ["_RJEM_MALLOC_CONF", "MALLOC_CONF"] {
50+
assert!(
51+
std::env::var_os(var).is_none(),
52+
"{var} is set in this process's environment. jemalloc reads \
53+
`_RJEM_MALLOC_CONF` (this build is prefixed), which sets `opt.*` on \
54+
its own, so this test could not tell the compiled-in export from the \
55+
environment; unset it and re-run."
56+
);
57+
}
58+
59+
// `opt.dirty_decay_ms` and `opt.muzzy_decay_ms` are jemalloc `ssize_t`s.
60+
// `raw::read` asserts the mallctl's width equals `size_of::<T>()`, so a
61+
// wrong Rust width fails here rather than reading a truncated value.
62+
let dirty: isize =
63+
unsafe { raw::read(b"opt.dirty_decay_ms\0") }.expect("opt.dirty_decay_ms is readable");
64+
let muzzy: isize =
65+
unsafe { raw::read(b"opt.muzzy_decay_ms\0") }.expect("opt.muzzy_decay_ms is readable");
66+
67+
assert_eq!(
68+
dirty, -1,
69+
"opt.dirty_decay_ms is {dirty}, not -1 (jemalloc's default is \
70+
{DEFAULT_DIRTY_DECAY_MS}): the `_rjem_malloc_conf` export beside this \
71+
file's `#[global_allocator]` is missing, misspelled, or was not read, \
72+
and a binary built this way returns dirty pages to the OS on a timer"
73+
);
74+
assert_eq!(
75+
muzzy, -1,
76+
"opt.muzzy_decay_ms is {muzzy}, not -1 (jemalloc's default is 0): the \
77+
`_rjem_malloc_conf` export beside this file's `#[global_allocator]` is \
78+
missing, misspelled, or was not read, and a binary built this way \
79+
unmaps muzzy pages immediately"
80+
);
81+
}

0 commit comments

Comments
 (0)