In the project, we focus on several objectives, which are:
1. Data collection, unification and management
2. Prognostic models incorporating local detail and global context
3. Development of multi-task models for pan-cancer feature discovery
4. Explainability for prognostic models
5. Validation of outcome prediction
The first objective focuses on collection of vast, pan-cancer datasets of 1000s of patients with both digizited histopathology data and outcome data. The steps to collect these data involve developing automated data preprocesisgn tools, such as for quality control, tissue identification and anonymization. Furthermore, after completion of the project all data will be made publicly available.
In objective 2, we will further develop and refine our weakly supervised learning method, Streaming Stochastic Gradient Descent, to be easier to use and apply and more computationally efficient and flexible. This will allow other researchers in the oncology field to use it. The software will also be made open-source. Last, we will provide initial proof that this method is capable of predicting patient outcomes for individual cancer types, such as prostate cancer.
Objective 3 extends this approach to pan-cancer models, AI systems that can incorporate information across different diseases and identify common and distinct prognostic features. By enforcing the models to learn specific, human understandable concepts, these features can later be used to drive, for example, research into drugs targeting specific disease properties. Last, these pan-cancer features could be leveraged as prognostic biomarkers in rare cancer, where not enough data is available to train complex AI systems.
Objective 4 aims to leverage recent developments in large language models (LLMs) to push the explainability of AI systems to the next level, and make it more similar to human communication. Specifically, by combining vision and language, we aim to have AI systems be able to articulate, in writing, their decisions. In this objective we will aim to assess the quality of these 'captions' and determine the risk of hallucinations, thus enabling us to prove the usefulness of these approaches.
Last, Objective 5 validates the full pipeline across three different cancerous entities, specifically prostate, breast, and pancreatic cancer. By prospectively collection new data from existing European initiatives, such as BigPicture and PANCAIM, we will be able to perform this validation at an unprecedented scale, with thousands of patients across many different European countries.