The initial work in the project included assessments of requirements from various stakeholder communities targeted by the project: clinical trial scientists, medical device manufacturers, independent assessors and regulatory bodies, along with foundational research that was completed and reported by the project partners. This included a review of ML methods for producing synthetic datasets, a literature review and classification of digital health applications, a review of data types used in various digital health use cases, as well as an initial standards gap analysis. The requirements assessments and foundational research are driving the technology developments launched during the reporting period for deliverables that will be completed in the next reporting period.
An important technology milestone was achieved with the development of the Data Integration Pipeline Tools, which enable users to define workflows as a series of interconnected tasks, each encapsulated within a Docker container. Its modular architecture and model-based design simplify the management of complex data handling processes. The tool supports integration with various data sources, such as files and databases, and offers flexibility in how these sources are connected. The deliverable includes a comprehensive review of current data pipeline orchestration technologies evaluated against 12 key characteristics derived from the project’s requirements. The analysis concludes that no existing solution fully meets all these criteria. Key features include visual workflow design, automated execution, and built-in monitoring. It has been developed to be user-friendly and adaptable to various synthetic data generation scenarios.