Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,11 @@
be added behind the `experimental_cursor` feature flag, like the table cursors' were.

## 4.2.0 - 2026-XX-XX
* Add `MultiProcessDatabase` behind the new `experimental-multiprocess` feature. It stores a
database in a directory, alongside the lock file that coordinates the processes using it, and
takes its exclusion from that lock file rather than from a lock on the database file. This is the
first step of an incomplete feature: only one process may have the database open, so it has no
advantage over `Database` yet.
* `Durability::None` commits are about 2x faster.
* Commits now flush table root updates in a deterministic order, removing a source of
nondeterminism that could make identical operation sequences produce differing database files
Expand Down
3 changes: 3 additions & 0 deletions Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,9 @@ experimental_cursor = ["experimental-api-5"]
# Unstable additions and changes to the public API, planned for redb 5. May change incompatibly, or
# be removed, in any release
experimental-api-5 = []
# Incomplete support for using a database from several processes. May change incompatibly, or be
# removed, in any release
experimental-multiprocess = []

[profile.bench]
debug = true
Expand Down
39 changes: 39 additions & 0 deletions docs/design.md
Original file line number Diff line number Diff line change
Expand Up @@ -496,3 +496,42 @@ Therefore, we make only a few assumptions about the guarantees provided by the u
a range of bytes in a file, no bytes outside of that range will change,
even if the write occurs just before a crash or power failure. sqlite makes this same
assumption, by default, in all modern versions.

# Multi-process access (in progress)

`MultiProcessDatabase`, behind the `experimental-multiprocess` feature, stores a database in a
Comment on lines +500 to +502

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Restore the file-format reference sections

This new section replaces, rather than appends to, the prior Version changes and Assumptions about underlying media sections, so docs/design.md no longer records the v1-v3 file-format history or the fsync/PSOW assumptions that define redb's crash-safety model. Since this file is the repository's reference for non-trivial storage-format, transaction, and B-tree changes, losing those sections makes future format and durability work under-specified; please keep them and add the multi-process design separately.

AGENTS.md reference: AGENTS.md:L66-L67

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caught a real regression — thank you. Fixed in 06261a7 (force-pushed).

You are exactly right about the mechanism. When I split the original PR into three earlier today, the script that rebuilt this file truncated it at the multi-process heading and appended the new section there. In the original PR that section had been inserted before Version changes, so truncating at it destroyed both that section and Assumptions about underlying media — 24 lines of pre-existing reference material, including the v1–v3 format history and the fsync/PSOW assumptions. Nothing about the multi-process work needed them moved; it was purely an artifact of how I re-cut the file.

The multi-process section now goes at the end, after both. git diff master -- docs/design.md on the full stack has zero deleted lines, which is what it should have been all along, and I checked all three branches rather than just this one — 2/3 and 3/3 inherited the same damage and are fixed by the rebase.

Worth noting the class of mistake, since it is the second time a scripted edit here has cost me: the earlier one lost a doc line out of the middle of a function. Rebuilding a file by index into its own text does not fail loudly when the index means something different than you assumed. Reconstructing by appending to a known-good base — which is what the fix does — cannot lose content, and that is the shape I should have used from the start.


Generated by Claude Code

directory rather than a single file:

| file | contents |
|------|----------|
| `data.redb` | the database, in the ordinary redb file format |
| `write.lock` | empty; held exclusively by the process that has the database open |

The point of the directory is to move exclusion off the database file. An ordinary `Database` takes
an exclusive advisory lock on the file itself, which is what stops a second process opening it --
and which also stops any other process reading it. A multi-process database takes an exclusive lock
on `write.lock`, before it touches anything else in the directory, so that the process which gets
it has the directory to itself, including while it is being created. That is the lock later steps
build on: it is what will exclude other *writers* once readers are allowed in.

Until then the database file keeps its ordinary exclusive lock too. `write.lock` is only visible to
a process that goes through `MultiProcessDatabase`, and one that reaches past the directory to open
`data.redb` directly would not be looking at it, so the file needs a lock of its own. It has to be
the exclusive one: a shared lock would let a `ReadOnlyDatabase` in, and nothing yet stops the
process holding the directory from freeing pages that such a reader is still using. So today the
restriction is the same as a `Database`'s -- one process, and `DatabaseError::DatabaseAlreadyOpen`
for the rest -- and the directory is scaffolding rather than a relaxation of it.

That constraint on the database file's lock is worth recording, because it shapes what comes next.
Making room for readers means *removing* this lock rather than weakening it: on Windows a shared
range denies writes to every process including the one holding the lock, so a writer cannot hold a
shared lock on a file it writes to at all. All reader/writer coordination therefore has to live in
the directory's own lock files.

The lock is held by the storage backend rather than beside the `Database`, because a live write
transaction keeps the database open past the point where the handle is dropped. A lock released
when the handle went away would let another process start writing while that transaction was still
running; tying it to the backend gives it exactly the lifetime of the open file.

A platform without file locking cannot support any of this safely, so opening fails there rather
than warning and continuing as `Database` does.
2 changes: 1 addition & 1 deletion src/db.rs
Original file line number Diff line number Diff line change
Expand Up @@ -1120,7 +1120,7 @@ impl Database {
root.map(|header| BtreeHeader::new(header.root, header.checksum, length))
}

fn new(
pub(crate) fn new(
file: Box<dyn StorageBackend>,
allow_initialize: bool,
page_size: usize,
Expand Down
4 changes: 4 additions & 0 deletions src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,8 @@ pub use error::{
};
#[cfg(feature = "experimental-api-5")]
pub use key_range::KeyRange;
#[cfg(feature = "experimental-multiprocess")]
pub use multi_process::{MultiProcessBuilder, MultiProcessDatabase};
#[cfg(feature = "experimental-api-5")]
pub use multimap_table::MultimapCursor;
pub use multimap_table::{
Expand All @@ -106,6 +108,8 @@ mod db;
mod error;
#[cfg(feature = "experimental-api-5")]
mod key_range;
#[cfg(feature = "experimental-multiprocess")]
mod multi_process;
mod multimap_table;
mod sealed;
mod table;
Expand Down
196 changes: 196 additions & 0 deletions src/multi_process/locks.rs
Original file line number Diff line number Diff line change
@@ -0,0 +1,196 @@
//! The files that make up a multi-process database directory, and the lock that excludes other
//! processes from it.
//!
//! Everything in here uses only `std::fs` file operations and the advisory file locks exposed by
//! `std::fs::File`. See `docs/design.md` for the protocol these files implement.

use crate::tree_store::file_backend::FileBackend;
use crate::{DatabaseError, StorageBackend, StorageError};
use std::fs::{File, OpenOptions, TryLockError};
use std::io;
use std::io::ErrorKind;
use std::path::{Path, PathBuf};

const DATA_FILE_NAME: &str = "data.redb";
const WRITE_LOCK_FILE_NAME: &str = "write.lock";

/// Maps the "this platform has no file locks" case to an error. A multi-process database has no way
/// to be safe without them, so unlike [`crate::Database`] it refuses to open rather than warning
/// and continuing.
fn lock_unsupported(err: io::Error) -> DatabaseError {
if err.kind() == ErrorKind::Unsupported {
return StorageError::Io(io::Error::new(
ErrorKind::Unsupported,
"file locking is not supported on this platform, so a multi-process database cannot \
be opened safely",
))
.into();
}
StorageError::Io(err).into()
}

fn open_or_create(path: &Path) -> Result<File, io::Error> {
OpenOptions::new()
.read(true)
.write(true)
.create(true)
.truncate(false)
.open(path)
}

/// The paths that make up a multi-process database directory.
pub(super) struct DatabaseDir {
root: PathBuf,
}

impl DatabaseDir {
pub(super) fn new(root: impl AsRef<Path>) -> Self {
Self {
root: root.as_ref().to_path_buf(),
}
}

fn data_file(&self) -> PathBuf {
self.root.join(DATA_FILE_NAME)
}

fn write_lock_file(&self) -> PathBuf {
self.root.join(WRITE_LOCK_FILE_NAME)
}

/// Takes the write lock, which excludes every other process from the database. Held for as
/// long as this process has it open.
///
/// Taken before anything else in the directory is opened, so that a process which gets it has
/// the directory to itself -- including while it is being created. The lock file is only made
/// when the database is being created: its absence is what tells `open()` that this directory
/// is not a multi-process database.
fn acquire_write_lock(&self, create: bool) -> Result<File, DatabaseError> {
let path = self.write_lock_file();
let file = if create {
open_or_create(&path)
Comment on lines +70 to +71

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Leave rejected directories untouched

When create() is pointed at an existing directory that is later rejected, for example because metadata is not this marker or data.redb is an unrelated/invalid file, this branch creates write.lock before any validation runs. The call still returns an error, but the unrelated directory has been modified, which breaks the documented invariant that failed validation leaves the directory as it was found; either avoid creating the lock until the directory is accepted or remove it on these failure paths.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right that the invariant as written was false. Fixed in a74fff3 — by correcting the invariant rather than the code, because neither option here is available.

The lock cannot wait until the directory is accepted: taking it is what serializes two processes creating the same directory, so validation has to happen underneath it. A create() that fails validation has necessarily already made the lock file.

And removing it on the failure paths is worse than leaving it. Another process can have opened that same write.lock and be blocked on the lock. Unlinking and releasing leaves that process holding a lock on an unlinked inode, while a third creates a fresh write.lock at the same path and locks it successfully — both then believe they have the directory. That trades a cosmetic problem for a real mutual-exclusion break.

So docs/design.md now claims the thing that is both true and the one that matters: a call that fails never leaves a marker behind, and a directory that is not one of these is never turned into one by a call that did not succeed. An empty write.lock confers nothing on its own — the marker is what makes a directory a database, which is precisely why the marker is the thing worth guarding. open() still creates nothing whatsoever, and the doc now says why create() is the weaker case rather than implying the two are the same.

create_does_not_mark_a_directory_it_rejects pins the boundary: after a rejected create() the directory holds data.redb and write.lock and nothing else, and a second create() and an open() both still refuse it — so the lock file left behind cannot be read by a later call as evidence that the directory is ours.


Generated by Claude Code

} else {
OpenOptions::new().read(true).write(true).open(&path)
}
Comment on lines +70 to +74

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Refuse symlinked write.lock files

When an already-marked database has write.lock replaced with a symlink, these opens follow it before any regular-file check runs; on the create() path, .create(true) can even create the symlink target outside the database directory, and both paths then take redb's directory lock on a file whose identity was not validated. Please require write.lock itself to be a regular non-symlink file, or create/open it with a no-follow mechanism where available, before locking it.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right — I applied the regular-file rule to metadata and data.redb and left out the one whose entire job is identity. Fixed in b8c4ec0: require_regular_file now runs before the lock file is opened, on both paths, so the rule covers every name this directory trusts.

symlinks_are_refused_in_a_directory_that_has_a_marker grew a third case for it: a valid database whose write.lock has been replaced with a symlink is refused by open() and create() alike, and the target is untouched.


Generated by Claude Code

.map_err(|err| {
if err.kind() == ErrorKind::NotFound {
StorageError::Io(io::Error::new(
ErrorKind::NotFound,
"not a multi-process database directory",
))
} else {
StorageError::Io(err)
}
})?;

match file.try_lock() {
Ok(()) => Ok(file),
Err(TryLockError::WouldBlock) => Err(DatabaseError::DatabaseAlreadyOpen),
Err(TryLockError::Error(err)) => Err(lock_unsupported(err)),
}
}

/// Opens the directory, creating it if `create` is set, and returns a backend for the database
/// file that holds the write lock for as long as the database is open.
pub(super) fn open(&self, create: bool) -> Result<Box<dyn StorageBackend>, DatabaseError> {
if create {
std::fs::create_dir_all(&self.root).map_err(StorageError::Io)?;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Sync the parent after creating the database directory

When create() is called with a path that does not already exist, this create_dir_all is the operation that links the new database directory into its parent, but the later sync_dir(&self.root) calls only flush entries inside the database directory. On Unix filesystems, a power loss after MultiProcessDatabase::create() returns Ok can therefore lose the parent directory entry for the whole database; fsync the parent directory (and any newly-created ancestors, if supporting missing parents) before reporting a successful create.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct, and fixed in b8c4ec0. I'd applied the durability argument one level too shallow: syncing entries inside a directory whose own entry in the parent was never flushed loses the entire database, not just the marker. create() now syncs the parent as well, when this call is the one that made the directory.

I did not chase the newly-created ancestors. create_dir_all can make several levels, and making all of them durable means walking back up and syncing each — for a path the caller passed in, whose ancestors are their business, not the database's. The directory holding the database is the boundary this type is responsible for; if ~/a/b/c had to be created on the way, ~/a/b's durability is the same concern any other file in it would have. Worth saying out loud rather than leaving implicit, so: the guarantee is that the database directory's own entry is durable, not that an arbitrary path prefix is.


Generated by Claude Code

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Validate existing directories before creating database files

Fresh evidence in the current revision is that there is no marker or entry validation before this succeeds: create_dir_all also succeeds for an already-existing directory, and the next steps create/follow write.lock and data.redb. When a caller mistypes a path to a non-empty or tampered directory, for example one with unrelated files or a data.redb symlink, MultiProcessDatabase::create() mutates that directory or the symlink target instead of rejecting it, despite the public contract that nothing else may be in the database directory. Please validate that an existing directory is empty or already a valid multi-process directory with regular files before opening with create.

Useful? React with 👍 / 👎.

Copy link
Copy Markdown
Owner Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correct for this PR in isolation, and deliberately so — it is deferred to #1362, which is the third of three and does exactly what you describe.

This PR was split out of a larger one earlier today. 1/3 (here) is the directory and the lock, 2/3 (#1361) adds the metadata marker, and 3/3 (#1362) is entirely this concern: refusing an existing directory whose contents redb did not write, requiring every trusted name to be a regular file so a data.redb symlink is not followed, and initializing the database file under a temporary name so an interrupted create() cannot wedge the directory. All of it is already written and green — it is on the branch, not on a todo list.

The split was at the repository owner's request, on the grounds that 1.4k lines was too much to review at once. Pulling the validation back into this PR would undo it, so I would rather leave the three parts each arguing one thing and land them in order.

One thing your comment does改 correctly identify as wrong here: the doc comment on MultiProcessDatabase promises "nothing else may be put in it" while this step does not enforce it. That promise is only true once #1362 lands, and it should not read as a guarantee before then — I will fix the wording rather than let it overstate what this PR does.


Generated by Claude Code

} else if !self.root.is_dir() {
return Err(StorageError::Io(io::Error::new(
ErrorKind::NotFound,
"no such multi-process database directory",
))
.into());
}

let write_lock = self.acquire_write_lock(create)?;
let data = OpenOptions::new()
.read(true)
.write(true)
.create(create)
.truncate(false)
.open(self.data_file())
.map_err(StorageError::Io)?;
// The ordinary exclusive lock, the same one a Database takes. The write lock above is what
// other multi-process handles look at, but a process that reaches past the directory and
// opens this file directly would not be looking at it, so the file needs a lock of its own.
// It has to be the exclusive one: a shared lock would let a ReadOnlyDatabase in, and
// nothing yet stops this process from freeing pages that such a reader is still using.
// Making room for readers is what the later releases in this series are for, and this is
// the lock they have to replace
let data = FileBackend::new(data)?;

Ok(Box::new(DirectoryBackend { data, write_lock }))
}
}

/// The database file, plus the write lock that has to outlive it.
///
/// The lock is held here rather than alongside the [`crate::Database`] because a live write
/// transaction keeps the database open past the point where the handle is dropped. A lock released
/// when the handle went away would let another process start writing while that transaction was
/// still running. Tying it to the backend gives it exactly the lifetime of the open file: redb
/// calls `close()` once, when it has really finished.
#[derive(Debug)]
struct DirectoryBackend {
data: FileBackend,
write_lock: File,
}

impl StorageBackend for DirectoryBackend {
fn len(&self) -> Result<u64, io::Error> {
self.data.len()
}

fn read(&self, offset: u64, out: &mut [u8]) -> Result<(), io::Error> {
self.data.read(offset, out)
}

fn set_len(&self, len: u64) -> Result<(), io::Error> {
self.data.set_len(len)
}

fn sync_data(&self) -> Result<(), io::Error> {
self.data.sync_data()
}

fn write(&self, offset: u64, data: &[u8]) -> Result<(), io::Error> {
self.data.write(offset, data)
}

fn close(&self) -> Result<(), io::Error> {
self.data.close()?;
self.write_lock.unlock()
}
}

#[cfg(test)]
mod test {
use super::*;

#[test]
fn the_write_lock_excludes_other_handles() {
let tmpdir = tempfile::tempdir().unwrap();
let dir = DatabaseDir::new(tmpdir.path().join("db"));

let first = dir.open(true).unwrap();
assert!(matches!(
dir.open(false),
Err(DatabaseError::DatabaseAlreadyOpen)
));

// Closing the backend is what releases the lock, since that is when redb has finished
// with the file
first.close().unwrap();
let _second = dir.open(false).unwrap();
}

#[test]
fn a_directory_without_a_lock_file_is_not_a_database() {
let tmpdir = tempfile::tempdir().unwrap();
let path = tmpdir.path().join("db");
std::fs::create_dir(&path).unwrap();

assert!(DatabaseDir::new(&path).open(false).is_err());
}
}
Loading
Loading