1. A multimodal dataset
We sequenced >750K nuclei in the snRNAseq dataset after thorough QC. All nuclei were annotated using 2 different approaches. First, a classical manual annotation was persued, using canonical marker genes, retrieved both from own experience and the human lung cell atlas. Seconly, we used Celltypist, an automated annotation tool which use annotations from publicly available datasets to predict celltype identities in the new dataset. Integrating these two approaches, we were able to identify 57 cell type indenties over 4 lineages (epithelial, stromal, endothelial, immune cells), as depicted in figure C.
2. The identification of disease stage-specific cell states
Milopy is a tool to identify regions on the neighborhood graph with specific differential abundances. First, it identifies neighborhoods to which nuclei in its proximity with very high similarity in terms of gene expression, contribute. Secondly, it calculates differential abundance of metadata variables in each neighborhood. Hence, we were able to identify neighborhoods with high abundances of control or IPF nuclei on the one hand, and with IPF nuclei originating from mild disease or severe disease on the other hand.
Interestingly, we found neighborhoods with significant enrichment in both dimensions (ctrl vs IPF, mild vs disease) in all four lineages and in almost all cell types. Hence, in almost all cell types, cell states exist which are enriched in IPF, and more specifically in mild IPF and in more severe IPF. As an example, neighborhoods formed by nuclei with the recently described aberrant basaloid cell identity as well as CTHRC1-positive myofibroblast nuclei showed very high enrichment in IPF vs CTRL but did not show enrichment of a specific disease stage, meaning these identities are found in mild disease already but their presence is not restricted to a specific disease stage.
By plotting median milopy-derived differential abundances of neighborhoods enriched for specific cell types, we identified cell type-specific clusters enriched for nuclei of a specific disease stage within the spectrum. Based on the milopy data, cell states were plotted in a two-dimensional space providing an overview to which extent a cell state is enriched for IPF nuclei (x-axis) and for early disease (y-axis), depicted in figure D.
3. The altering epithelial-mesenchymal dialogue throughout disease progression
In a next step, we evaluated whether the dialogue between epithelial and mesenchymal cell states differ. We assessed disease stage-specific ligand-receptor interactions (LRIs) using Nichenet to clusters occurring within the same disease stage. Whereas some LRIs clearly overlap, some are unique to a specific disease stage.
To validate the ligand receptor interactions and further characterize the niches in which these cell states reside, we have setup a spatially resolved proteomics approach. Using a multiplex IF imaging approach (4i), niches are identified after which these niches are selectively laser captured and further processed for mass spectrometry-based spatial proteomics. As this Marie Curie postdoctoral fellowship was terminated early, this will be pursued by the applicant and his supervisor in a follow-up collaboration.