Genomes contain information about the past, but extracting the historical signal from large numbers of genomes remains a major challenge in the biological sciences. In particular, a problem arises when different species of taxa evolve at distinct rates and in addition their genes have heterogeneous signals due to biological process of selection and mutation. Yet, an accurate reconstruction of historical processes using genomes is bound to will bring substantial benefits to the biological sciences.
Among the many approaches that have been proposed for testing your ability to extract historical signals from genomes, the use of simulations has proven particularly promising due to its resemblance to an experimental setting. In addition, methods in machine learning, and in particular methods of unsupervised learning, provide the opportunity to extract the dominant signals in a broad range of data types. This project aimed to perform a detailed simulations study with molecular data evolving under a broad range of conditions, and across large numbers of genes, to examine their possible behaviour of genomic data sets under various statistical analysis frameworks (work package 1). In the second instance, the project aimed to build a software package that was easily accessible to researchers in biology, using methods of unsupervised learning as well as incorporating classical statistical statistical tests for finding the dominant signals of evolutionary rates in genomic data sets (work package 2). Using this novel framework of analysis, additional set of simulations will demonstrate the limitations of the proposed methods, as well as their power and usefulness for analysis of data of different sizes, and across the diversity of evolutionary scenarios (work package 3).
An important objective of the project was to join forces with the bird 10,000 genomes consortium, assessing their data efficiently under the proposed framework described above (work package 4). The framework was used for identifying lineages of birds with unusually fast or slow evolutionary processes, as well as the genes that have been most consequential for their evolutionary success. Overall, the project led to methodological advances in the analysis of molecular genomic data, as well as biological insights within one of the major genome sequencing consortia being led at the host institution. An additional outcome of the project is a long term collaboration between the hosts and the recipient on the development of novel methodological approaches, and the efficient usage of ever-increasing biological data resources. Briefly, the hosts provided world leading knowledge on large molecular genomic data resources while the recipient provided expertise on statistical methods development and analysis.