journal · BMC Bioinformatics · 2011

Efficient alignment of pyrosequencing reads for re-sequencing applications

Francisco Fernandes, Paulo GS da Fonseca, Luís M. S.​Russo, Arlindo L. Oliveira, Ana T. Freitas · 11 citations

View original publication →

See where this sits in the topic map →

Summary AI-generated

TL;DR
Massively parallel DNA sequencing technologies reduce costs and generate huge volumes of data, but require new computational methods to efficiently map reads to reference genomes.
Problem
While modern DNA sequencing technologies drastically lower costs, they also create major computational bottlenecks due to the massive volume of data produced, unique read lengths, and sequencing errors. A critical challenge is mapping these reads efficiently and accurately to a reference genome in re-sequencing projects.
Method
We introduce an efficient local alignment method for pyrosequencing reads from the GS FLX (454) system. The approach leverages data characteristics and combines state-of-the-art indexing techniques with a flexible seed-based approach to create a fast, accurate algorithm requiring minimal user parameterization.
Results
Evaluations using both real and simulated data show that our method outperforms several mainstream tools in alignment quality, quantity, and execution speed.
Contributions
We developed and released TAPyR (Tool for the Alignment of Pyrosequencing Reads), a publicly available software tool implementing the proposed methodology.
Limitations
Not specified in the abstract.
Takeaways
The proposed methodology is available to researchers as an open-source software tool called TAPyR at http://www.tapyr.net.
Applications
Genomic re-sequencing projects using pyrosequencing data.
Topics
Bioinformatics, DNA Sequence Alignment, Computational Biology
For industry
Not specified in the abstract.
Why it matters
Not specified in the abstract.

Abstract

BACKGROUND: Over the past few years, new massively parallel DNA sequencing technologies have emerged. These platforms generate massive amounts of data per run, greatly reducing the cost of DNA sequencing. However, these techniques also raise important computational difficulties mostly due to the huge volume of data produced, but also because of some of their specific characteristics such as read length and sequencing errors. Among the most critical problems is that of efficiently and accurately mapping reads to a reference genome in the context of re-sequencing projects. RESULTS: We present an efficient method for the local alignment of pyrosequencing reads produced by the GS FLX (454) system against a reference sequence. Our approach explores the characteristics of the data in these re-sequencing applications and uses state of the art indexing techniques combined with a flexible seed-based approach, leading to a fast and accurate algorithm which needs very little user parameterization. An evaluation performed using real and simulated data shows that our proposed method outperforms a number of mainstream tools on the quantity and quality of successful alignments, as well as on the execution time. CONCLUSIONS: The proposed methodology was implemented in a software tool called TAPyR--Tool for the Alignment of Pyrosequencing Reads--which is publicly available from http://www.tapyr.net.

References within the group

← All publications