journal · IEEE/ACM Transactions on Computational Biology and Bioinformatics · 2004

Biclustering algorithms for biological data analysis: a survey

Sara C. Madeira, Arlindo L. Oliveira · 2065 citations

View original publication →

See where this sits in the topic map →

Summary AI-generated

TL;DR
This paper presents a comprehensive survey of biclustering algorithms for analyzing gene expression data from microarray experiments.
Problem
Standard clustering methods are limited because there are many experimental conditions under which gene activity is uncorrelated, whether clustering genes or conditions.
Method
The paper reviews biclustering (also known as coclustering or direct clustering), an approach that performs simultaneous clustering on both the row and column dimensions of a data matrix to find subgroups of genes and conditions with correlated activities.
Results
The survey analyzes a wide range of existing biclustering approaches and classifies them according to the types and patterns of biclusters they discover, their search methods, evaluation approaches, and target applications.
Contributions
Not specified in the abstract.
Limitations
Standard clustering is limited by experimental conditions where gene activity remains uncorrelated.
Takeaways
Biclustering offers a powerful alternative to traditional clustering by finding submatrices where genes exhibit highly correlated activities across specific conditions.
Applications
Biological data analysis and information retrieval.
Topics
Computational biology, bioinformatics, and data mining.
For industry
Biotechnology, healthcare, and information retrieval.
Why it matters
Provides a structured classification of biclustering techniques, helping researchers select and evaluate methods for complex data analysis across biological and data mining domains.

Abstract

A large number of clustering approaches have been proposed for the analysis of gene expression data obtained from microarray experiments. However, the results from the application of standard clustering methods to genes are limited. This limitation is imposed by the existence of a number of experimental conditions where the activity of genes is uncorrelated. A similar limitation exists when clustering of conditions is performed. For this reason, a number of algorithms that perform simultaneous clustering on the row and column dimensions of the data matrix has been proposed. The goal is to find submatrices, that is, subgroups of genes and subgroups of conditions, where the genes exhibit highly correlated activities for every condition. In this paper, we refer to this class of algorithms as biclustering. Biclustering is also referred in the literature as coclustering and direct clustering, among others names, and has also been used in fields such as information retrieval and data mining. In this comprehensive survey, we analyze a large number of existing approaches to biclustering, and classify them in accordance with the type of biclusters they can find, the patterns of biclusters that are discovered, the methods used to perform the search, the approaches used to evaluate the solution, and the target applications.

Cited by (group publications)

← All publications