Research

The Biophysics of Protein Design

 

Our lab focuses on the development and implementation of methods to understand complex molecular phenomena where pure physics-based or deep learning approaches are insufficient. By synergizing our best understanding of statistical mechanics with modern deep learning architectures, we address problems of conformational change,  entropically-driven binding interactions and lipid-protein mechanics.

Deep Learning Guided Models to Predict and Design Protein Dynamics

Proteins serve as the primary means of an organism’s ability to sense and respond to its internal and external environment. To achieve this, many proteins undergo thermodynamically accessible allosteric changes that initiate biological processes downstream. Mutations and external stimuli can alter the thermodynamic landscapes of proteins. While some of the mechanisms of these conformational changes are understood through studies such as deep mutational scanning, it is not yet possible to predict, engineer or therapeutically alter these changes beyond a subset of well-studied systems. We integrate deep learning with enhanced-sampling guided physics-based simulation to investigate the dynamics and allosteric modulation of proteins. In addition, we use de novo protein design and protein engineering to recreate allosterically responsive proteins from our most basic understanding of protein biophysics. These methods are used to build generalizable models of protein function to determine how allosteric stimuli alter the behavior of proteins. We seek to use this understanding in conjunction with modern computational design methods to build and test established mechanisms of allosteric networks using streamlined de novo protein assemblies that would serve as model systems.

Molecular Mechanisms of Lipid-Protein Interactions

Lipids are intimately involved in nearly every biological process within cells. The loss of membrane homeostasis has been increasingly implicated in many diseases. Dysregulation of sterols, glycosphingolipids and phospholipids is a key feature of late onset neurodegeneration, oncogenic signaling and cardiovascular disease. Many of these lipids interact with transmembrane or membrane-associated proteins to facilitate conformational changes or signaling processes that modulate downstream pathways. The manner in which these lipids affect the behavior of these pathways is multifaceted and challenging to characterize. This is partially due to the complex, heterogeneous nature of lipid-protein interactions, which requires both spatial and temporal resolution to rigorously dissect. We fill this gap by addressing the problem in two directions: 1) Employ deep-learning and physics-based simulation to computationally distill the key interactions that broadly drive protein-lipid headgroup binding and specificity and 2) Use existing methods to investigate the function and dysregulation of glycosphingolipids, an understudied subtype of lipids critical for cell-cell communication. Together, these foci will elucidate our understanding of lipid-protein interactions and lipid behavior within the context of cellular function and disease.

Understanding the Patterns Learned by Protein Language Models

Proteins are encoded as linear sequences of amino acids, and these sequences contain information that shapes protein structure and function. In computational biology, a growing class of methods has adapted concepts from natural language processing to analyze protein sequences, analogous to how language models analyze words and sentences. In this framework, amino acids or short sequence patterns are treated as the basic units of a biological sequence, and the model learns statistical relationships among them from large collections of proteins. These models, commonly referred to as pLMs, have become widely used tools for learning sequence-based protein representations and have been applied to structure prediction, function prediction, and protein design. However, it remains unclear how effectively existing pLMs extract structure-relevant information from sequence, and how this ability differs across models. We systematically compare representative pLMs and quantitatively evaluate their ability to recover structural information from protein sequences. Additionally, we explore protein language models trained on structure in addition to sequence to understanding how these models integrate this information for more complex applications.

Facilities

The Bethel Lab maintains its own private cluster of hundreds of CPUs and GPUs for simulation, inference and training. We also have allocations of the Triton Shared Computing Cluster and the Expanse Supercomputer hosted at UCSD. Additionally, we are users of the Cryo-EM facility, the Thermo Fisher Sandbox, and the Janelia Cores hosted by Howard Hughes Medical Institute.