An Efficient Algorithm for the Identification of Structured Motifs in DNA Promoter Sequences
See where this sits in the topic map →Summary AI-generated
- TL;DR
- We propose RISO, a new algorithm for identifying cis-regulatory modules and structured motifs in genomic sequences.
- Problem
- Identifying structured motifs—conserved regions that occur in a well-ordered and regularly spaced manner—is challenging because existing exact algorithms lack efficiency when dealing with spacings between binding sites.
- Method
- The RISO algorithm uses a new data structure called box-link to store information about conserved regions in a well-ordered and regularly spaced manner across sequences.
- Results
- Experimental results demonstrate that the algorithm is much faster than existing approaches, achieving time and space gains that are exponential in the spacings between binding sites, sometimes by more than four orders of magnitude.
- Contributions
- We introduce the RISO algorithm and its underlying box-link data structure, and provide a full implementation made available online along with complexity analysis and experimental validation.
- Limitations
- Not specified in the abstract.
- Takeaways
- RISO efficiently extracts relevant consensus patterns and represents promoter models effectively when applied to biological data sets.
- Applications
- Analyzing biological data sets to extract relevant consensus patterns and study gene regulatory mechanisms.
- Topics
- Computational biology, bioinformatics, genomic sequences, structured motifs, cis-regulatory modules.
- For industry
- Not specified in the abstract.
- Why it matters
- Advances computational methods for researching gene regulatory mechanisms and analyzing DNA promoter sequences.
Abstract
We propose a new algorithm for identifying cis-regulatory modules in genomic sequences. The proposed algorithm, named RISO, uses a new data structure, called box-link, to store the information about conserved regions that occur in a well-ordered and regularly spaced manner in the data set sequences. This type of conserved regions, called structured motifs, is extremely relevant in the research of gene regulatory mechanisms since it can effectively represent promoter models. The complexity analysis shows a time and space gain over the best known exact algorithms that is exponential in the spacings between binding sites. A full implementation of the algorithm was developed and made available online. Experimental results show that the algorithm is much faster than existing ones, sometimes by more than four orders of magnitude. The application of the method to biological data sets shows its ability to extract relevant consensi.