Warning
leech is alpha quality and under active development. APIs, CLI flags, and output formats may change without notice, and bugs are expected. Validate results before relying on it for anything important.
Learning Enhanced Electrical Classifiers from Hanopore signals
Leech classifies aminoacylation state and amino acid identity from Oxford
Nanopore tRNA sequencing data. It extracts dwell time features from move
tables (the BAM mv tag) and feeds them alongside raw signal and sequence
context into a multi-branch neural network, giving it information that
signal-only tools like Remora
discard.
Requires Python 3.12+
uv add "leech[rust]" # or: pip install "leech[rust]"The rust extra pulls leech-core, the compiled accelerator for data
preparation and inference. leech runs without it — every accelerated path has
a pure-Python fallback — so plain uv add leech is fine if no wheel matches
your platform (wheels are built for manylinux x86_64 and aarch64).
To work on leech itself:
# Install uv if you don't have it
curl -LsSf https://astral.sh/uv/install.sh | sh
# Clone and install
git clone https://github.com/rnabioco/leech.git
cd leech
uv syncuv run leech data prepare \
--pod5 reads.pod5 \
--bam alignments.bam \
--output-dir chunks/ \
--motif CCAGGC --motif-offset 2 \
--label 1 --workers 8uv run leech model train \
--train-data chunks/train.json \
--val-data chunks/val.json \
--model ConvLSTMDwell \
--output-dir models/uv run leech eval test \
--model models/model_best.pt \
--test-data chunks/test.json \
--output metrics.jsonuv run leech predict \
--model models/ \
--pod5 new_reads.pod5 \
--bam new_alignments.bam \
--output predictions.bamPackage multiple pairwise models into a single file and run aggregated inference:
# Bundle all pairwise models
uv run leech model bundle \
--model-dir results/models/pairwise/ \
--output bundle.pt --version 1.0.0
# Inspect bundle contents
uv run leech model bundle-info --bundle bundle.pt
# Run all models (aggregated amino acid prediction)
uv run leech predict \
--bundle bundle.pt --all \
--pod5 reads.pod5 --bam alignments.bam \
--output predictions.bamModels export to ONNX as well as TorchScript, so a runtime that is not PyTorch can load them:
uv run leech model export --model-dir models/ --format onnx -o model.onnxEvery export writes a contract beside the graph, carrying what the graph cannot: which input is which, and that a leech classifier emits a single BCE logit rather than a two-class softmax — reading it as the latter makes every call wrong without erroring.
| Group | Commands | Purpose |
|---|---|---|
leech data |
prepare, merge |
Extract features, merge and split datasets |
leech model |
train, train-crf, optimize, bundle, bundle-info, calibrate, export |
Train, tune, calibrate, and package models |
leech eval |
test, compare, importance, ablation |
Evaluate and analyze models |
leech predict |
Run inference (single model or bundle) |
29 architectures across 5 families, all supporting multi-channel signal input (signal_in_channels):
| Family | Models | Description |
|---|---|---|
| ConvLSTM | ConvLSTMDwell (recommended), ConvLSTMBase | Conv-LSTM with 3 branches (signal, sequence, dwell/level features) |
| ConvLSTM variants | +BN, +Attn, +BNAttn, +GNAttn, +LNAttn | Batch/group/layer normalization and attention pooling |
| Remora-compat | ConvLSTMRemora, ConvLSTMRemoraBase | Remora-compatible architecture for direct comparison |
| Transformer | TransformerDwell, TransformerDwellResidual | Multi-head self-attention; Residual variant uses 2-channel signal (raw + kmer residual) |
| TCN | TCNDwell, +GN, +LN, +Residual | Temporal Convolutional Network with dilated convolutions |
| Other | ResNetDwell, ConvOnly | Residual network; pure CNN with multi-scale convolutions |
Alongside the classifiers, leech.crf trains sequence models: a CTC-CRF
over n_base ** state_len states whose Viterbi traceback emits one base per
move, for reading a sequence out of raw signal rather than assigning it a label.
The formulation is Oxford Nanopore's, introduced in bonito; the architecture is
SeqTagger's published parameters (Genome Res 35:956).
from leech.crf import CrfEncoder, CtcCrfLoss, decode_batch, encoder_config_from_toml, load_config
cfg = encoder_config_from_toml(load_config())
model, criterion = CrfEncoder(cfg), CtcCrfLoss(cfg.n_base, cfg.state_len)
scores = model(signal) # (N, 1, chunk) -> (T, N, n_score)
sequences = decode_batch(scores, cfg.n_base, cfg.state_len)Corpora are described by a manifest
— one row per read naming the signal window and its target — so the vocabulary
of a given assay stays with whatever produced it. plan_corpus/build_corpus
cut a corpus from one, leech model train-crf trains on it, and
leech.crf.evaluate decodes and scores against a reference set by edit
distance:
uv run leech model train-crf --corpus corpus/ldx16 --output-dir models/crf/ \
--epochs 32 --batch-size 256Note the emission rule: a CRF with state_len cannot emit the first
state_len bases of its target, so a target_len target decodes to
target_len - state_len bases at any window width. See the
CRF API reference.
- Loss functions: BCE, focal loss (for class imbalance), and cross-entropy
- Regularization: weight decay, gradient clipping, dropout
- LR scheduling: reduce-on-plateau, cosine annealing with warmup
- Data augmentation: mixup (signal jitter + random scaling)
- Mixed precision: FP16 training on CUDA; TF32 matmul on Ampere+
- Performance:
torch.compilesupport, Rust-accelerated signal statistics (217x) - Class balancing: automatic class weight computation
- Balance-groups sampling: equal contribution per source group per epoch
- K-fold cross-validation: stratified read-level k-fold splits
- Platt calibration: post-hoc Platt scaling for probability calibration
- Signal map refinement: Viterbi-based kmer level table refinement (matches Remora)
- Kmer residual features: expected level, signed/unsigned deviation from kmer table
- Multi-channel signal: 2-channel input (raw + kmer residual) for Residual model variants
- Aggregation: naive, confidence-weighted, and tournament pairwise aggregation
- Composable config: dataclass-based configuration shared between prep and inference
- TorchScript export: standalone model export for deployment without leech
For production workloads, leech includes a Snakemake pipeline supporting:
- Charged vs. uncharged classification
- Pairwise amino acid discrimination
- Grid search optimization
- Multi-architecture comparison
- HPC clusters (SLURM/LSF)
See pipeline/ for configuration and usage.
uv sync --all-extras # Install with dev tools
uv run pytest # Run tests
uv run ruff check . # Lint
uv run ruff format . # Format
uv run ty check src/leech/ # Type checkIf you use leech, please cite:
- This work (publication pending)
- Remora (underlying training framework)
MIT License - see LICENSE for details.