To build our computational model of RNAPII and RECQ5, we examined each domain to determine how it interacts with the others. Comparing the relative sizes and self-interactions of individual proteins can be done using previously developed methods. But the novel insight we had access to was the internal organization of the condensates from cryogenic electron tomography. An accurate model would preserve individual features and recover the same internal organization at high concentrations of both proteins. Once this was achieved, we could use the dynamic, higher-resolution model to answer questions about condensate formation. For example, we could hypothesize that the number of RNAPII in each condensate might have been limited by the number of free RECQ5 available to attach to it. The long IDR of RECQ5 helped induce phase separation by itself and also might ‘reach’ out to recruit RNAPII. However, it’s important to remember that these results were obtained in highly controlled systems and may not replicate in cells. While our model couldn’t determine how RECQ5 slows transcription, cryogenic electron microscopy images revealed a domain that binds RNAPII near where DNA enters. The domain introduces a small amount of contact, akin to friction, slowing DNA transcription.
The work and methodology from this study have been shared publicly, and we are concluding further experiments in response to the reviewers' comments. Anyone can access the preprint manuscript or use the same computational tools used in the study.
By experimenting with the parameters of the RECQ5-RNAPII system, we observed instances in which mixtures of two proteins were visibly inhomogeneous. Small clusters of a single protein would form, but current techniques were unable to quantify the extent of this phenomenon. We set out to develop a methodology to determine if what we were seeing was significant or biased interpretations.
To develop a new analysis technique, it’s crucial to ensure that it works on more than just one system of two proteins. As a test set, we simulated over 2000 binary mixtures of disordered proteins. Our metric would analyze small environments within the simulation and generate a density distribution to examine the populations of both proteins. If protein A were more inclined to form contacts with itself, rather than protein B, we could measure that preference. When we analyzed proteins involved in transcription and phase separation, we observed the same preferences reported in other experimental studies. Several other studies have reported simulated results from mixtures of IDRs, but ours is the most comprehensive to date due to the diverse range of protein lengths and chemistries. Our work is currently in revision for publication, but the preprint is publicly available in the meantime, along with any code needed to replicate this study. Additionally, we’ve provided a dataset of our mixtures, which we hope will support researchers who want to expand this technique, accelerate simulations, or learn more about how proteins interact in the human body.
While this new technique will help us look inside computational models of transcriptional condensates, we also need a metric to characterize how unique each IDR is relative to the others. If protein A and B have similar chemical groups, then it’s less surprising that they form a homogeneous mixture. However, we again found no agreed-upon methodology to tackle this problem. In analyzing large-scale simulation datasets, we found that, unlike structured proteins, disordered proteins often separate their ‘structure’ from their function. Here, because these proteins are characteristically unstructured, we use structure to refer to the ensemble conformations that are responsible for the protein size and ability to phase separate. Research to understand this complex relationship is ongoing, but our goal is to provide researchers with a simple tool to generate IDRs with specific functions and ensemble properties.