Proteins are essential molecules that perform a vast array of functions within living organisms. They are made up of long chains of amino acids, which fold into specific three-dimensional structures that determine their function. The relationship between sequence, structure, and function is a central theme in biology. Understanding how a specific sequence leads to a particular function is a complex challenge because even small changes in the sequence (mutations) can significantly alter the protein's behavior. Indeed, mutations in proteins can lead to disease through different mechanisms such as destabilization of the protein fold, by affecting the specific function of the protein, or by causing the protein to aggregate. Predicting the effects of these mutations, or variants, is crucial for understanding their potential impact on health. Accurate predictions can aid in diagnosing genetic disorders, developing treatments, and understanding the underlying mechanisms of disease.
In order to improve our understanding of how mutations lead to genetic diseases we set the following objectives:
- Generate a large dataset of how mutations in human proteins involved in genetic diseases affect the stability of their three-dimensional folds
- Identify which mutations cause genetic diseases through protein destabilization, and analyze the importance of destabilization as a disease mechanism across different diseases and proteins
- Use the data to develop predictive models of how mutations affect protein stability to cover a larger number of pathogenic mutations