Despite AI models achieving human-level performance in medical tasks, their translation to routine clinical practice remains limited. The core issue is trustworthiness: when deployed clinically, models encounter scenarios fundamentally different from training conditions - equipment varies between hospitals, patient demographics shift, measurement technologies evolve. These "data drift" effects cause even accurate models to fail silently, making incorrect predictions with high confidence.
This is particularly critical in oncology, where treatment decisions have profound consequences. Clinicians need AI tools that not only provide accurate predictions but also reliably communicate uncertainty - tools that "know when they don't know."
TAIPO (Trustworthy AI in Personalized Oncology) addresses this gap through two objectives:
Objective 1: Develop trustworthy AI tools for cancer diagnosis and patient stratification. We assess and enhance model reliability across all clinical scenarios through novel auditing methods (ModelAuditor), model transformation techniques (ModelTransformer), and transparent risk communication (ModelMonitor).
Objective 2: Establish frameworks for robust modeling of therapy decisions and outcomes, extending trustworthiness principles to survival prediction and therapy recommendation with meaningful uncertainty estimates.
We demonstrate broad applicability across different cancer types (melanoma, blood cancers, ...) and diverse data modalities.
We expect impact on two different axes:
Translational Impact: By enabling reliable uncertainty communication with open source tools, we empower efficient human-AI collaboration where clinicians rely on AI for routine cases while focusing expertise on complex cases.
Methodological Impact: Our auditing framework and theoretical advances (bias-variance-covariance decomposition) establish new standards for AI trustworthiness.
Our progress demonstrates these goals are achievable: successful framework development, novel insights into uncertainty quantification of generative models and foundation model calibration, and clinical validation in lymphoma stratification show a path of how trustworthy AI can become clinical reality.