Base calling
A seq2seq problem analogous to speech recognition: translate the raw current signal into a nucleotide sequence. Best baseline reaches ~94% match rate; Illumina reaches 99.9%.
NanoBaseLib integrates 16 public Nanopore datasets with over 30 million reads, re-processed with a single unified pipeline and stored in uniform formats, so machine learning researchers can work on base calling, polyA detection, signal segmentation and modification detection without months of bioinformatics preprocessing.
git clone https://github.com/nanobaselib/NanoBaseLib.git
cd NanoBaseLib
# 319 MB demo dataset with every intermediate result
curl -L -o demo_dataset.tar.gz \
"https://zenodo.org/records/10889896/files/demo_dataset.tar.gz?download=1"
tar -xzf demo_dataset.tar.gzEach task takes the output of the previous one as input. NanoBaseLib stores every intermediate result, so you can start at any step with ground truth, baselines and evaluation code already in place.
A seq2seq problem analogous to speech recognition: translate the raw current signal into a nucleotide sequence. Best baseline reaches ~94% match rate; Illumina reaches 99.9%.
Locate the poly(A) tail segment in the signal and estimate its length. Doubles as a quality-control filter: reads without a tail are discarded downstream.
Split the signal into segments and align each to a k-mer of the reference sequence. An aligned k-mer is an “event”, the unit every downstream method reasons about.
Decide whether the central base of a k-mer at a genomic site is modified (m6A, m5C, …) from the pooled signal segments aligned to it. Ground truth comes from orthogonal short-read assays.
All 16 datasets start from raw fast5 files and run through the same tools with the same parameters, so benchmark comparisons are fair and no task depends on a dataset's original, heterogeneous processing.
Convert single-fast5 and pod5 into multi-fast5.
ont-fast5-api · pod5Raw signal → FASTQ reads.
Guppy 6.0.1 · DoradoAlign reads, sort and filter the BAM.
minimap2 · samtoolsTail flag, borders and length per read.
Nanopolish polya · TailfindrMatch signal segments to reference k-mers.
Nanopolish · Tombo · SegPorePer-site modification probability.
m6Anet · CHEUI · Epinano · …Nanopore sequencing measures the ionic current as a DNA or RNA molecule translocates through a protein pore in a membrane. The shape, size and chemistry of the roughly five nucleotides inside the pore (a k-mer) determine the current level, and computational methods decode those levels back into sequence.
Compared with second-generation sequencing this yields ultra-long reads (10–100 kb), simpler library preparation, and direct access to epigenetic marks on native molecules. It powered the Telomere-to-Telomere human genome, rapid SARS-CoV-2 and Ebola surveillance, and direct RNA sequencing for mRNA vaccine development.
The cost is accuracy: RNA base calling sits near 90–94% while Illumina reaches 99.9%, and raw-signal alignment remains unsolved. The full training data of the vendor is proprietary and public data are scattered across repositories in incompatible forms. NanoBaseLib removes that barrier for the machine-learning community.
If NanoBaseLib helps your research, please cite the paper. The processed dataset is released under CC BY 4.0 and the software under Apache 2.0.
NanoBaseLib: A Multi-Task Benchmark Dataset for Nanopore Sequencing.
Guangzhao Cheng, Chengbo Fu, Lu Cheng.
Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Datasets and Benchmarks Track, pp. 76319–76331.
@inproceedings{cheng2024nanobaselib,
title = {NanoBaseLib: A Multi-Task Benchmark Dataset for Nanopore Sequencing},
author = {Cheng, Guangzhao and Fu, Chengbo and Cheng, Lu},
booktitle = {Advances in Neural Information Processing Systems},
volume = {37},
pages = {76319--76331},
year = {2024},
doi = {10.52202/079017-2430}
}