Repository navigation
fix(merge): rebuild signal table at a uniform batch stride - #432
Merged
Merged
Conversation
merge concatenated each input's own Arrow signal batches as-is, offsetting
only byte positions. Almost every POD5's own last batch is short, so
merging N files put a short batch from file k before a full one from file
k+1 at every file boundary but the last. escapepod's own reader walks real
cumulative row counts and is unaffected, but dorado and the official pod5
library assume a constant stride and mis-resolve every read after the
break -- reproduced with dorado 2.1 ("Too few samples in input samples
array") on real multi-file POD5 merges.
merge now flattens every surviving read's compressed signal chunks across
inputs and rebuilds the table at one uniform stride via
write_raw_signal_table, moved from operations::filter into
utils::table_builders (alongside the SignalRow/build_signal_batch
primitives it's built on) and now shared by filter/subset and merge alike.
Compressed bytes are still copied without decompression/recompression.
Verified on a real 16-file, 59,418-read merge: zero signal mismatches
against the pod5 library (ground truth, independent of this crate), and
`escpod inspect summary` no longer reports the batch as NOT PORTABLE.
Pre-existing on main (unrelated to the merge fix); fixing here since it blocks this PR's CI.
5 tasks
jayhesselberth
added a commit
that referenced
this pull request
Oct 5, 2026
…ilders (#434) ## Why Follow-up to #432's merge batch-stride fix: an architecture review afterward surveyed the rest of the workspace for the same cross-module layering smell (a function stranded in one module that a sibling reaches into, or near-duplicate logic copy-pasted across modules). Two findings, filed as #433. ## What - Delete `utils/run_info.rs` — `add_run_infos_deduplicated`/ `map_run_info_index` were never wired into `utils/mod.rs` and have zero call sites; superseded by `pod5_assembler::deduplicate_run_infos`, which `merge`/`filter` already share. - Unify `build_reads_table` (merge, owned `(ReadData, Vec<u64>)` pairs) and `build_reads_table_remapped` (filter, borrowed `FlatReadRef`) behind one generic `build_reads_table_generic<R: PartitionRow + Sync>`. Both were ~200-line copies of the same dictionary-collection/partition/concat/ batch-write skeleton, differing only in input shape — already abstracted at the per-row level by the existing `PartitionRow` trait and `build_partition_inner`. Both public names stay as thin wrappers so callers in `merge.rs`/`operations/filter.rs` are untouched. No behavior or output change — pure extraction, verified by the existing `test_merge_integration.rs` / filter / subset test suite (unchanged, all passing) plus a full `cargo test -p escapepod-pod5` run. ## Testing - `cargo test -p escapepod-pod5`: all green (incl. doctests) - `cargo fmt --all --check`: clean - `RUSTFLAGS=-Dwarnings cargo clippy --workspace --all-targets`: clean Closes #433
Merged
jayhesselberth
added a commit
that referenced
this pull request
Oct 6, 2026
Patch release. - Fixed: `merge` no longer produces non-portable output when inputs don't share one batch stride (#432) - Internal: dead run_info dedup removed, reads-table builders unified (#434) Version bump in root Cargo.toml (workspace + 5 path deps), CHANGELOG rolled, Cargo.lock regenerated; `cargo check --workspace` clean. Tag `v0.31.1` on the merge commit after merge (publishes to GitHub Releases and PyPI).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
mergeconcatenated each input file's own Arrow signal batches as-is, offsetting only byte positions. Almost every POD5 file's own last batch is short (read count rarely divides evenly by batch size), so merging N files put a short batch from filekimmediately before a full one from filek+1at every file boundary but the last — breaking the constant stride dorado and the officialpod5library assume between batches (escapepod's own reader walks real cumulative row counts and is unaffected — same bug class as #195, writer-side).Reproduced on real production data: merging 20 real multi-file flexizyme POD5 directories and basecalling the output with dorado 2.1 gave
Failed to get read signal - 'Invalid: Too few samples in input samples array'on every one of them.escpod inspect summaryindependently confirms:Signal batches: NOT PORTABLE — signal batch 46 has 31 rows, expected 100.What
mergenow flattens every surviving (non-duplicate) read's compressed signal chunks across all inputs and rebuilds the signal table at one uniform stride viawrite_raw_signal_table— the same correctness-first approachfilter/subsetalready used. That writer moved fromoperations::filterintoutils::table_builders, next to theSignalRow/build_signal_batchprimitives it's built from, since it's now shared by three callers instead of one (layering fix, not just the bug fix). Compressed bytes are still copied without decompression/recompression — only the Arrow batch grouping is rebuilt.MergeOptionsgainssignal_batch_size(default1_000, matchingFilterOptions).Side effect, not the point of the change: a duplicate read's signal bytes are no longer carried into the output, since extraction now happens per-surviving-read instead of per-input-file (previously every input file's full signal table was copied regardless of later dedup).
Testing
merge_output_signal_batches_stay_uniform_even_with_short_trailing_inputs): merges two 150-read fixtures (each batches as[100, 50]under defaultWriterOptions) and assertsreader.nonuniform_signal_batch().is_none()on the output — this is exactly the shape that broke before the fix.cargo nextest run -p escapepod-pod5: 257/257 pass, including the existingmerge_preserves_signal_bytewise(no recompression) and the fulltest_read_batch_geometry/test_merge_integrationsuites.cargo clippy -p escapepod-pod5 --all-targets -- -D warnings: clean.cargo build -p escapepod-cli: builds against the newMergeOptionsfield.lysflexizyme merge with this branch's release binary, then compared every read's signal array between the merged output and the original raw files using the officialpod5Python library (independent of this crate): zero mismatches, zero missing, across all 59,418 reads.escpod inspect summaryno longer reportsNOT PORTABLEon the new output (did on the old one, reproduced above).--emit-movesbasecalling of both the raw directory and the fixed merged file no longer errors on either. (Basecalled read yield between the two differs — see follow-up note below; this is a separate, pre-existing issue, not a data-correctness regression from this change.)Follow-up (not in this PR)
While validating, basecalling the full
lyscorpus via dorado 2.1 yielded 59,411/59,418 reads from the raw directory but only 34,186/59,418 from the merged file — despite the signal being byte-identical (proven above). The likely cause:collect_pod5_inputs(escapepod-cli/src/util.rs) sorts input files with plainVec::sort(), which is lexicographic over filenames like..._0.pod5,..._1.pod5,..._10.pod5, ... — scrambling the numeric/chronological file order MinKNOW wrote. If dorado's read-splitter relies on channel/time continuity across adjacent reads in the stream, concatenating files out of chronological order would explain a large, order-sensitive split-rate change with identical bytes. Worth a natural-sort fix and its own dorado-based verification as a separate change.