seq2edit
Model the deaminase sequence bias by learning a map from DNA sequence to
per-base editing signal. A deaminase does not edit every accessible cytosine
equally — it has a sequence preference (e.g. DddA’s TC context). seq2edit
trains a convolutional neural network to predict, for a window of one-hot DNA,
the editing each base would receive on the basis of sequence alone. Trained on
a naked/deproteinised-DNA control, this gives an expected track that
downstream footprint and occupancy analyses can divide out to separate enzyme
bias from genuine protein protection.
The design follows the ACCESS-ATAC
cnn_bias_model.
The module has three stages — train (implemented here), predict, and
interpret (planned). This page documents seq2edit train.
Note
seq2edit needs the optional PyTorch dependency, which is not installed by
default:
pip install 'deamtools[seq2edit]'
The data-preparation helpers work without it; only model training/inference
require torch.
Synopsis
deamtools seq2edit train --fasta FILE --bigwig FILE --train_regions FILE \
--out_dir DIR --out_name NAME [options]
Required inputs
Argument |
Description |
|---|---|
|
Reference FASTA indexed with |
|
Per-base editing-signal BigWig (the regression target), e.g. from |
|
BED of regions tiled into training windows. |
|
Output directory. Created automatically if it does not exist. |
|
Base name for the checkpoint. Writes |
Optional arguments
Argument |
Default |
Description |
|---|---|---|
|
(hold-out split) |
BED of validation regions. If omitted, a random |
|
|
Hold-out fraction used when |
|
|
Window width in bp — the model’s input and output length. |
|
|
Spacing between window starts. Use a smaller value to oversample with overlapping windows. |
|
|
Convolutional filters per conv block. |
|
|
Convolution kernel width. |
|
|
Number of training epochs. |
|
|
Mini-batch size. |
|
|
Adam learning rate. |
|
|
Adam weight decay (L2). |
|
|
DataLoader worker processes. |
|
|
RNG seed for reproducibility. |
|
(auto) |
Compute device ( |
|
|
Global flag (before the subcommand): |
How it works
Windowing. Each interval in --train_regions is tiled into non-overlapping
--seq_len-wide windows (a trailing stretch shorter than seq_len is dropped).
For every window:
the reference sequence is one-hot encoded as the input
xof shape(seq_len, 4)with column orderA, C, G, T(non-ACGTbases such asNbecome an all-zero row); andthe per-base
--bigwigsignal over the same coordinates is the targetyof shape(seq_len,)(uncoveredNaNbases are read as0).
Model (EditNet). Two Conv1d → ReLU → MaxPool → Dropout blocks feed a
fully-connected head that regresses the per-base editing vector for the window,
with a final Softplus so the output is a strictly positive rate λ — the mean
of a Poisson over per-base edit counts:
one-hot (seq_len, 4)
→ conv block 1 (n_filters, kernel_size)
→ conv block 2 (n_filters, kernel_size)
→ flatten → FC(1024) → FC(1024) → FC(seq_len) → Softplus
→ predicted per-base editing rate λ (seq_len,)
The flattened size feeding the first dense layer is derived from the
architecture, so non-default --seq_len/--n_filters/--kernel_size work
without code changes (the upstream reference hard-codes it for the 128 / 32 / 5
defaults).
Optimisation. Editing signal is count-like, so the model is fit with a
Poisson negative-log-likelihood loss (PoissonNLLLoss, on the predicted
rate λ) rather than MSE — it maximises the Poisson likelihood of the observed
per-base counts. Adam (--lr 3e-4, --weight_decay 1e-4) with a
ReduceLROnPlateau schedule (patience 10, floor 1e-5) on the validation loss.
Each epoch reports train/validation loss and the current learning rate; the
checkpoint is rewritten whenever the validation loss improves.
Outputs
File |
Description |
|---|---|
|
Best-validation checkpoint: model weights ( |
Examples
# Train on naked-DNA editing signal over a set of regions
deamtools seq2edit train \
--fasta hg38.fa --bigwig naked.bw \
--train_regions regions.bed \
--out_dir models --out_name bias
# Explicit validation regions, larger windows, fewer epochs, on GPU
deamtools seq2edit train \
--fasta hg38.fa --bigwig naked.bw \
--train_regions train.bed --valid_regions valid.bed \
--seq_len 128 --epochs 50 --device cuda \
--out_dir models --out_name bias
Notes
The
--bigwigtarget should come from a control where editing reflects sequence preference, not chromatin state — typically naked or deproteinised DNA — otherwise the model conflates enzyme bias with accessibility.Windows are tiled, so making
--train_regionsbroad (e.g. many peaks or a whole chromosome) gives the model more examples;--stepsmaller than--seq_lenadds overlapping windows for additional coverage.The saved
configlets laterpredict/interpretstages rebuild the exact network from the checkpoint.