NeurIPS 2024 · Datasets and Benchmarks Track

One dataset,
multiple tasks for Nanopore sequencing

NanoBaseLib integrates 16 public Nanopore datasets with over 30 million reads, re-processed with a single unified pipeline and stored in uniform formats, so machine learning researchers can work on base calling, polyA detection, signal segmentation and modification detection without months of bioinformatics preprocessing.

Get started
git clone https://github.com/nanobaselib/NanoBaseLib.git
cd NanoBaseLib
# 319 MB demo dataset with every intermediate result
curl -L -o demo_dataset.tar.gz \
  "https://zenodo.org/records/10889896/files/demo_dataset.tar.gz?download=1"
tar -xzf demo_dataset.tar.gz
16
public datasets, re-processed with one pipeline
33.7M
reads with raw signal and base calls
4
benchmark tasks with baselines
6
species, DNA and RNA
3.2 TB
raw fast5 data indexed
5
RNA modification types (m6A, m5C, hm5C, inosine, Ψ)
Four benchmark tasks

The transcriptome analysis workflow, task by task

Each task takes the output of the previous one as input. NanoBaseLib stores every intermediate result, so you can start at any step with ground truth, baselines and evaluation code already in place.

Task 1 · BC

Base calling

A seq2seq problem analogous to speech recognition: translate the raw current signal into a nucleotide sequence. Best baseline reaches ~94% match rate; Illumina reaches 99.9%.

Inputraw current signalOutputnucleotide sequenceMetricmatch / mismatch / indel rates
Base calling benchmark
Task 2 · PD

PolyA detection

Locate the poly(A) tail segment in the signal and estimate its length. Doubles as a quality-control filter: reads without a tail are discarded downstream.

Inputraw current signalOutputtail flag, borders, lengthMetricdetection rate, length MSE
PolyA detection benchmark
Task 3 · SA

Segmentation & event alignment

Split the signal into segments and align each to a k-mer of the reference sequence. An aligned k-mer is an “event”, the unit every downstream method reasons about.

Inputraw signal + referenceOutputevent alignment tableMetricavg. std, avg. log-likelihood
Segmentation benchmark
Task 4 · MD

Modification detection

Decide whether the central base of a k-mer at a genomic site is modified (m6A, m5C, …) from the pooled signal segments aligned to it. Ground truth comes from orthogonal short-read assays.

Inputevent alignment resultsOutputmodification probability per siteMetricROC AUC, PR AUC
Modification detection benchmark
Unified preprocessing

Six steps, one pipeline, every intermediate file kept

All 16 datasets start from raw fast5 files and run through the same tools with the same parameters, so benchmark comparisons are fair and no task depends on a dataset's original, heterogeneous processing.

Standardize raw data

Convert single-fast5 and pod5 into multi-fast5.

ont-fast5-api · pod5

Base call

Raw signal → FASTQ reads.

Guppy 6.0.1 · Dorado

Map to reference

Align reads, sort and filter the BAM.

minimap2 · samtools

Detect polyA (RNA)

Tail flag, borders and length per read.

Nanopolish polya · Tailfindr

Segment & align events

Match signal segments to reference k-mers.

Nanopolish · Tombo · SegPore

Detect modifications

Per-site modification probability.

m6Anet · CHEUI · Epinano · …

Read the full pipeline with commands →

Why Nanopore, why now

Long reads, direct modification readout, and a 10-point accuracy gap

Nanopore sequencing measures the ionic current as a DNA or RNA molecule translocates through a protein pore in a membrane. The shape, size and chemistry of the roughly five nucleotides inside the pore (a k-mer) determine the current level, and computational methods decode those levels back into sequence.

Compared with second-generation sequencing this yields ultra-long reads (10–100 kb), simpler library preparation, and direct access to epigenetic marks on native molecules. It powered the Telomere-to-Telomere human genome, rapid SARS-CoV-2 and Ebola surveillance, and direct RNA sequencing for mRNA vaccine development.

The cost is accuracy: RNA base calling sits near 90–94% while Illumina reaches 99.9%, and raw-signal alignment remains unsolved. The full training data of the vendor is proprietary and public data are scattered across repositories in incompatible forms. NanoBaseLib removes that barrier for the machine-learning community.

Schematic of nanopore sequencing: a motor protein feeds a DNA/RNA strand through a nanopore in a membrane while ionic current is measured
Principle of Nanopore sequencing. Adapted from Wang et al., Nature Biotechnology 39, 1348–1365 (2021).
Cite

NanoBaseLib at NeurIPS 2024

If NanoBaseLib helps your research, please cite the paper. The processed dataset is released under CC BY 4.0 and the software under Apache 2.0.

NanoBaseLib: A Multi-Task Benchmark Dataset for Nanopore Sequencing.
Guangzhao Cheng, Chengbo Fu, Lu Cheng.
Advances in Neural Information Processing Systems 37 (NeurIPS 2024), Datasets and Benchmarks Track, pp. 76319–76331.

@inproceedings{cheng2024nanobaselib,
  title     = {NanoBaseLib: A Multi-Task Benchmark Dataset for Nanopore Sequencing},
  author    = {Cheng, Guangzhao and Fu, Chengbo and Cheng, Lu},
  booktitle = {Advances in Neural Information Processing Systems},
  volume    = {37},
  pages     = {76319--76331},
  year      = {2024},
  doi       = {10.52202/079017-2430}
}