Pipeline workflow

Public datasets arrive pre-processed with different basecallers, genome versions and tools, which makes them unfair to compare directly. NanoBaseLib therefore starts from the raw signal and runs one pipeline whose outputs become the inputs of each benchmark task. Each step can be implemented with several tools; the ones we chose follow common practice and double as the baselines in the benchmarks.

Flowchart of the six-step NanoBaseLib pipeline from raw fast5 to modification probabilities
The unified NanoBaseLib preprocessing workflow. Steps 4 and 6 are optional and RNA-specific.

Step 0: get the demo dataset

git clone https://github.com/nanobaselib/NanoBaseLib.git
cd NanoBaseLib
curl -L -o demo_dataset.tar.gz "https://zenodo.org/records/10889896/files/demo_dataset.tar.gz?download=1"
tar -xzf demo_dataset.tar.gz        # creates demo_dataset/ with 0_reference and 1_raw_signal
cd demo_dataset

Step 1: data standardization

Raw data come as single-fast5, multi-fast5 or pod5. Everything is converted to multi-fast5, the format all downstream tools accept. If you want to use Dorado, also produce pod5.

# single-fast5 → multi-fast5 (not needed for the demo dataset)
single_to_multi_fast5 --input_path 1_raw_signal/single_fast5 \
    --save_path 1_raw_signal/multi_fast5 \
    --filename_base demo --batch_size 4000 --recursive

# multi-fast5 → pod5 (only if Dorado is used)
pod5 convert fast5 1_raw_signal/multi_fast5/*.fast5 \
    --output 1_raw_signal/multi_pod5 \
    --one-to-one 1_raw_signal/multi_fast5

Step 2: base calling

Guppy 6.0.1 (high-accuracy models) is the reference basecaller for all datasets; Dorado is run in addition. Configuration files per chemistry are listed in the base calling benchmark.

guppy_basecaller -c rna_r9.4.1_70bps_hac.cfg --num_callers 20 --cpu_threads_per_caller 20 \
    -i 1_raw_signal/multi_fast5 -s 2_base_called/guppy --fast5_out

dorado basecaller rna002_70bps_hac@v3 1_raw_signal/multi_pod5 \
    --estimate-poly-a > 2_base_called/dorado/dorado.bam

Step 3: mapping to the reference

Reads are indexed for Nanopolish, aligned with minimap2 in map-ont mode, then sorted and filtered (-F 2324 drops unmapped, secondary, supplementary and reverse-strand records for direct RNA).

cd 2_base_called
find guppy -type f -name "*.fastq" -exec cat {} + > demo_dataset.fastq
nanopolish index --directory=../1_raw_signal/multi_fast5 demo_dataset.fastq
minimap2 -ax map-ont --MD -t 8 --secondary=no ../0_reference/ref.fa demo_dataset.fastq \
    | samtools sort -o reads-ref.sorted.bam -T reads.tmp

samtools view -b -F 2324 reads-ref.sorted.bam > reads-ref.sorted.filter.bam
samtools index reads-ref.sorted.filter.bam
samtools quickcheck reads-ref.sorted.filter.bam
samtools view -h -o reads-ref.sorted.filter.sam reads-ref.sorted.filter.bam

Step 4 (optional): polyA detection

RNA only. Reads without a detectable poly(A) tail are low quality and can be filtered out. NanoBaseLib ships results from both Nanopolish polya and Tailfindr.

nanopolish polya --threads=32 --reads=demo_dataset.fastq --bam=reads-ref.sorted.filter.bam \
    --genome=../0_reference/ref.fa > ../4_nanopolish/polya.tsv
grep -E 'PASS|readname' ../4_nanopolish/polya.tsv > ../4_nanopolish/polya-pass-only-with-head.tsv
library(tailfindr)
df <- find_tails(fast5_dir = '2_base_called/guppy/workspace',
                 save_dir = '3_tailfindr',
                 csv_filename = 'tails.csv',
                 num_cores = 20)

Step 5: segmentation and event alignment

The raw signal is segmented and each segment aligned to a reference k-mer. Three tools are run: Nanopolish eventalign (with --scale-events and --signal-index), Tombo resquiggle (which needs single-fast5 input), and SegPore eventalign.

nanopolish eventalign --reads demo_dataset.fastq \
    --bam reads-ref.sorted.filter.bam \
    --genome ../0_reference/ref.fa \
    --signal-index --scale-events \
    --summary ../4_nanopolish/summary.txt \
    --threads 32 > ../4_nanopolish/eventalign.txt

cd ..
mkdir -p 5_tombo/single_reads
for i in $(ls 2_base_called/guppy/workspace); do
  multi_to_single_fast5 --input_path 2_base_called/guppy/workspace/${i} \
      --save_path 5_tombo/single_reads --recursive -t 20
done
tombo resquiggle 5_tombo/single_reads 0_reference/ref.fa --overwrite --processes 20

Step 6 (optional): modification detection

Estimates the probability that a nucleotide (or k-mer) at a genomic site is modified. Each tool handles one modification type at a time; see the modification detection benchmark for the tools and how to run them on the eventalign output.

Batch effects and normalization

Sequencing devices are an important source of batch effects. The pipeline mitigates them in two ways: device-specific parameters convert the integer signal to picoamperes, and downstream tools further normalize per read (Tombo: median/MAD; Nanopolish: per-read scaling; SegPore: polyA-based standardization). Details are on the signal processing page.

Processing tools, versions and limitations

SoftwareVersionPipeline stepFunctionLimitation
ont-fast5-api / pod54.0.21 · StandardizationRaw data format conversion (single-fast5, multi-fast5, pod5)–
Bonito0.7.32 · Base callingBase calling–
Causalcall-2 · Base callingBase callingDNA only
Dorado0.5.32 · Base callingBase calling, polyA detection, modification detection–
Guppy6.0.12 · Base callingBase callingRequires a free ONT community account to download
Rodan-2 · Base callingBase callingRNA only
bedtools2.30.33 · MappingGenomic file format conversion–
minimap22.243 · MappingAlign reads to the reference genome / transcriptome–
samtools / htslib-3 · MappingBAM sorting, filtering and indexing–
Nanopolish0.14.04 · PolyA detectionPolyA detection, event alignment, modification detection–
Tailfindr1.44 · PolyA detectionPolyA tail detection and length estimation–
SegPore1.05 · Segmentation & alignmentEvent alignment and modification detectionRNA only
Tombo1.5.15 · Segmentation & alignmentRe-squiggle (segmentation and alignment), modification detection–
CHEUI-6 · Modification detectionRNA modification detectionm6A and m5C only
Epinano1.2.06 · Modification detectionRNA modification detection–
m6Anet1.06 · Modification detectionRNA modification detectionm6A only
MINES-6 · Modification detectionRNA modification detectionm6A only
Nanom6A2.06 · Modification detectionRNA modification detectionm6A only
h5py1.8.18UtilityPython interface to HDF5 (fast5) files–

Guppy and Dorado downloads require a free Oxford Nanopore community account.