CNMAC-2025

banner - CNMAC-2025
Back

Plenária de Abertura: Alberto Paccanaro, FGV, Brasil

Type:

Plenária

Category:

Palestras

Place:

Auditório Centro Cultural

Date and time:

20:00 to 21:00 on 09/15/2025

Título: Machine Learning on Biological Networks: From Protein Function to Drug Repurposing and Enzyme Design

Resumo: Living cells operate through a multitude of interconnected molecular networks, where proteins, nucleic acids, and other biomolecules interact in precise and coordinated ways. In this talk, I will present machine learning approaches that we developed to solve problems in biology and pharmacology that can be framed in terms of inference in such large-scale networks.

I will first describe S2F (Sequence to Function), a semi-supervised learning method we developed for predicting protein function in newly sequenced organisms, where only sequence information is available. S2F creates an initial set of functional “seeds” that are propagated over networks of predicted functional associations. A key innovation is a novel label propagation algorithm that models overlapping functional communities, improving prediction accuracy. In extensive tests on bacterial genomes, S2F consistently outperformed the best sequence-based methods, achieving substantial gains. The method can be applied to any newly sequenced organism and is available as an easy-to-use, open-access tool.

Next, I will present a new approach that combines ideas from matrix factorization and network medicine to predict which existing drugs can be re-used against specific viral diseases – this problem is known as drug repurposing and our method is the first that can predict host centric antivirals. Our algorithm learns embeddings for the different entities involved (drugs, viruses, and proteins) in a low-dimensional space, making explicit some of their features that are relevant for the problem. I will show that these representations can be interpreted, thus leading to explanations for the predictions that may shed some light on the biological phenomena underlying these problems. Our predictions do not rely on known drug-virus associations and could be applied to new viral diseases and drugs.

Finally, I will discuss a recent work in which we have applied deep learning transformer-based models to learn complex distributions of sequences of amino acids – this is analogous to how Large Language Models model the distribution of word sequences in the context of Natural Language Processing. By fine-tuning these models on specific enzyme families, we can generate enzymes that have virtually the same structure as natural enzymes but with very different amino acid compositions. These new enzymes, while retaining the enzymatic function of their natural counterparts, could have different physical-chemical properties, thus enabling their use in industrial processes where natural enzymes would be unsuitable (joint work with Prof Giorgio Valentini’s Lab at the University of Milan).

CNMAC-2025 Galoá

XLIV Congresso Nacional de Matemática Aplicada e Computacional uses Galoá to painless manage and increase the impact of the event.

Need help planning or organizing your conference? Schedule a call