PolyA detection benchmark
Polyadenylation adds a poly(A) tail to messenger RNA. Finding that tail in the raw signal is a quality-control step after base calling, and three synthetic datasets with known tail lengths make it a measurable task.

The poly(A) tail is used in library preparation, so a read without a detectable tail is flagged as low quality and removed from downstream analysis. The task outputs a detection flag, the start and end of the tail on the raw signal, and the tail length in nucleotides. Only the tail length has ground truth; the borders cannot be measured independently with current technology.
Detection rate benchmark
Seven test datasets. Tailfindr detects a tail in the largest fraction of reads on every dataset, while Nanopolish polya varies strongly because it relies on alignment to the reference genome; Tailfindr and Dorado are alignment-free.
Tail length estimation benchmark
On ont_ployA_standard the six samples carry synthetic tails of 10, 15, 30, 60, 80 and 100 nucleotides. All three methods give similar distributions; the mean squared error against the known length is the metric.
Download ont_ployA_standard benchmark
Ground truth
Three synthetic datasets use tails of known length: ont_ployA_standard (10–100 nt, six samples) and eGFP_polyA_DNA / eGFP_polyA_RNA (10, 30, 40, 60, 100 and 150 nt, six samples each). For the ONT standard the sample name is the label. For the eGFP datasets samples are organised by kit and replicate, and the label is recovered from a barcode: expected barcode sequences are aligned against the read sequence preceding the eGFP alignment.


Download all label files (tailfindr_label.tar.gz)
Baseline models
| Model | Version | Approach | Category |
|---|---|---|---|
| Nanopolish polya | v0.14.0 | Hidden Markov model, alignment-based | Statistical / HMM |
| Tailfindr | v1.4 | Signal-based R tool, alignment-free | Statistical / HMM |
| Dorado | v0.5.3 | Deep learning basecaller with tail estimation, alignment-free | Deep learning |
Nanopolish polya
A hidden Markov model segments each raw signal into five regions: start, leader, adapter, poly(A) tail and transcript, then converts the tail duration into a length estimate.

Tailfindr
An R package that estimates poly(A) length from raw fast5 data without alignment, for both RNA and DNA reads including reverse-complement reads with poly(T) stretches.
Dorado
ONT's current basecaller estimates tail length as part of base calling when run with --estimate-poly-a; it is the only deep-learning method among the baselines.
References
- Weil, T.T. Post-transcriptional regulation of early embryogenesis. F1000Prime Reports 7 (2015).
- Workman, R.E., Tang, A.D., Tang, P.S. et al. Nanopore native RNA sequencing of a human poly(A) transcriptome. Nature Methods 16, 1297–1305 (2019).
- Krause, M., Niazi, A.M., Labun, K. et al. tailfindr: alignment-free poly(A) length measurement for Oxford Nanopore RNA and DNA sequencing. RNA 25, 1229–1241 (2019).