Unmixing mass spectra using machine learning for species identification
MALDI-TOF mass spectrometry is a fast and powerful tool for identifying biological species based on molecular fingerprints. However, when samples contain mixtures, such as multiple species or strains, the resulting spectra overlap, making it hard to tell them apart. To solve this, we can apply machine learning techniques that “unmix” these complex signals into their original components. This concept, known as spectral deconvolution, is similar to separating instruments in a song or resolving mixed cell types in genomics. With access to a high-quality reference spectral database from the EU-funded MALDIBANK project, we can explore statistical and ML models that enhance identification accuracy, potentially transforming how we analyze complex biological samples. This project will develop and test spectral deconvolution models, such as ICA, NMF, and neural networks, to separate overlapping MALDI-TOF signals. This will evaluate how different input formats (raw spectra vs. spectral peaks) and modeling approaches (including Bayesian techniques) affect performance. The goal is to build a prototype pipeline that improves species identification from mixed samples.
Requisites
It is required that the student has a background in Data Science, Machine Learning.
Notes: This project is part of a recently EU-funded project MALDIBANK ( https://cordis.europa.eu/project/id/101188201 ). A fellowship may be offered depending on progress and results. The selected student will have access to the facilities of INESC-ID and the MLKD group ( https://mlkd.idss.inesc-id.pt/cluster ), including computing facilities that include four DELL PowerEdge C41402 servers, eight NVIDIA 32GB Tesla V100, four NVIDIA 48GB A40 and four NVIDIA 64GB Tesla A100, among other computing servers ( https://mlkd.idss.inesc-id.pt/cluster ).