During its first two years of existence, the LostMA project reached several important milestones in the understanding of textual transmission and survival, in terms of modelling and empirical validation, offering a solution to the century-old ``Bédier's paradox''. This has led to uncover several new challenges, that the project will adress in the coming years.
A null-model for manuscript transmission: In terms of modelling textual transmisision, researchers of the LostMA team have successfully designed and tested a null-model for manuscript transmission, based on a birth-death (BD) process. This process simulates manuscript traditions, allowing to explore parameter spaces for copy/loss rates ( λ/μ ). Results show that drift alone can explain key properties observed in empirical data (in particular, in collections of text genealogical trees, known as stemmata (e.g. 60–65% bifidity in simulations vs. 77% in real data). This provides an answer to the so-called ``Bédier's paradox'', which states that most stemmata show a root bifurcation. In contradiction to Bédier's hypothesis, that this would arise from a bias in the philological method, the model results show that chance (drift) can explain at least a large part of this observation.
A database of 2000 medieval manuscripts: to achieve this result, a important process of data collection has proven necessary. For this, in collaboration with an international team of experts, based in the universities of Antwerp, Heidelberg and Paris Sciences & Lettres, has been established and has led to the organisation of several workshops, in Paris, Brussels and Antwerp. This resulted in a dataset of ~2,000 medieval manuscripts (covering French, Dutch and English medieval chivalric narratives) with metadata (dates, stemmata, witnesses). The data was released as part of OpenStemmata (open-source repository of graph-encoded stemmata) and a Heurist database of textual traditions, and has been made publicly available on Zenodo.
Empirical Validation: collected historical data has been confronted to model outputs, comparing also stochastic processes to the results of other methods, such as Shared Species richness. This has led to the identification of consistet loss rates estimates (40–50% for works, ~1% for manuscripts) for Medieval French chivalric narratives. In addition, LostMA's team was able to confront these results with those resulting from the application of this method to very different corpora, such as the works of the Iberic Church Fathers.
All this work has led to the publication of 9 articles, 2 datasets and a software package, called simMAtree, as well as the organisation of three international workshops.
New challenges: these results have also led to the identification of several new challenges, that the project team has started to adress. They concern the variations in time and space of parameters, the role of intrinsic (taste and reception) or extrinsic factors (e.g. plagues or wars) in textual transmission, the inclusion of multilingual data, and the extraction of more refined features from historical textual traditions.