Top: polyadenylation of an mRNA transcript. Bottom: the poly(A) tail appears as a long flat segment in the nanopore current between adapter and transcript
Polyadenylation and its signature in the Nanopore signal. Top panel adapted from Weil (2015).

The poly(A) tail is used in library preparation, so a read without a detectable tail is flagged as low quality and removed from downstream analysis. The task outputs a detection flag, the start and end of the tail on the raw signal, and the tail length in nucleotides. Only the tail length has ground truth; the borders cannot be measured independently with current technology.

Detection rate benchmark

Seven test datasets. Tailfindr detects a tail in the largest fraction of reads on every dataset, while Nanopolish polya varies strongly because it relies on alignment to the reference genome; Tailfindr and Dorado are alignment-free.

Grouped bar chart of polyA detection rate for Dorado 0.5.3, Nanopolish 0.14.0 and Tailfindr 1.4 across seven test datasets
PolyA detection rate (fraction of reads with a detected tail) per test dataset.

Tail length estimation benchmark

On ont_ployA_standard the six samples carry synthetic tails of 10, 15, 30, 60, 80 and 100 nucleotides. All three methods give similar distributions; the mean squared error against the known length is the metric.

Box plots of estimated polyA tail length per sample for three tools, with the true lengths 10 to 100 marked
Estimated tail length per sample. Black lines mark the ground truth (10, 15, 30, 60, 80, 100 nt).

Download ont_ployA_standard benchmark

Ground truth

Three synthetic datasets use tails of known length: ont_ployA_standard (10–100 nt, six samples) and eGFP_polyA_DNA / eGFP_polyA_RNA (10, 30, 40, 60, 100 and 150 nt, six samples each). For the ONT standard the sample name is the label. For the eGFP datasets samples are organised by kit and replicate, and the label is recovered from a barcode: expected barcode sequences are aligned against the read sequence preceding the eGFP alignment.

Table of barcode sequences assigned to each polyA length in the eGFP datasets
Barcode-to-tail-length assignment in the eGFP datasets.
Example label table mapping read names to barcode and polyA length
Resulting per-read label file.

Download all label files (tailfindr_label.tar.gz)

Baseline models

ModelVersionApproachCategory
Nanopolish polyav0.14.0Hidden Markov model, alignment-basedStatistical / HMM
Tailfindrv1.4Signal-based R tool, alignment-freeStatistical / HMM
Doradov0.5.3Deep learning basecaller with tail estimation, alignment-freeDeep learning

Nanopolish polya

A hidden Markov model segments each raw signal into five regions: start, leader, adapter, poly(A) tail and transcript, then converts the tail duration into a length estimate.

A squiggle segmented by the HMM into start, leader, adapter, polyA tail and transcript regions, with two cliffs marked in the tail
HMM segmentation of a read: start (cyan), leader (yellow), adapter (red), poly(A) tail (green), transcript (purple). From Workman et al. (2019).

Tailfindr

An R package that estimates poly(A) length from raw fast5 data without alignment, for both RNA and DNA reads including reverse-complement reads with poly(T) stretches.

Dorado

ONT's current basecaller estimates tail length as part of base calling when run with --estimate-poly-a; it is the only deep-learning method among the baselines.

References

  1. Weil, T.T. Post-transcriptional regulation of early embryogenesis. F1000Prime Reports 7 (2015).
  2. Workman, R.E., Tang, A.D., Tang, P.S. et al. Nanopore native RNA sequencing of a human poly(A) transcriptome. Nature Methods 16, 1297–1305 (2019).
  3. Krause, M., Niazi, A.M., Labun, K. et al. tailfindr: alignment-free poly(A) length measurement for Oxford Nanopore RNA and DNA sequencing. RNA 25, 1229–1241 (2019).