The PI has conducted empirical studies mainly of Swedish patient organisations to characterize the role and trajectories of patient organizations in the context of healthcare policy and infrastructure and medical research and practice, in order to produce an outline of their structural development, typology and function.
Selection of relevant source segments and analysis of expressions of medical ontologies and modes of thought and conduct in patient organizations’ periodicals and archival documents, mainly done on medical classification internationally and diabetes illness concepts in the Swedish context.
Leading the research engineers (senior research assistants) and the junior research assistant in building a solid technical structure for the project, including developing a database and frontend for the source corpus. This entailed pre-processing already scanned material (applying OCR, converting into txt and xml formats, correcting for errors and ‘cleaning’ the data, lemmatization, POS-tagging etc) and building a database structure, but was significantly delayed for the inclusion of the British material due to the cyber attack on the British library. By the end of 24 months, all material had been collected but work still remained on the final curation of the database. The frontend is up and running.
PI, SRA, JRA and the postdocs have collaborated in running initial computer-based analyses on the material using and adapting available digital tools. Besides established word count- and collocation-based text mining tools and methodologies, we have also conducted significant methodological development of NLP techniques for visualization and segmentation of the data and, using these new tools, performed analyses of the source material that have yielded valuable empirical results. This goes further than we had originally anticipated for the first 24 months.
Furthermore, the two historian postdocs have performed qualitative, empirical studies, primarily on the British patient organisations so far, outlining trends and trajectories and identifying important contextual events and developments in the respective countries (France and UK).
The project uses a high degree of novel methodologies, interdisciplinary work, and knowledge transfer. During our first two years, we have adapted existing NLP visualisation tools to fit historical analysis, including the WordRain visualisation and the Topic Timeline visualisation. Furthermore, the postdoc in digital humanities, Vera Danilova, has developed a framework for genre analysis borrowing from NLP but adapted for historical print media. This development process is intended to help segment historical source materials in a way that makes them more accessible for text mining, a major challenge which excludes many interesting sources from the methods already available in the digital humanities. With the genre annotation technique, different types of content can be separated from each other automatically and analysed, hence reducing noise and producing more fruitful results.
During its first 24 months, the project has generated three scientific publications, and a further eight publications are in various stages of the review process. The papers published so far are focused on the technical aspects of building the database and methods for performing computer-based analyses on the sources. Furthermore, the project team members have presented their work at numerous conferences internationally. In addition, the project team has worked on developing a cross-cutting theory for the use of big datasets in the humanities and biomedicine. We have presented this theory at two conferences and are planning for publication as a research paper.