Integrated Data Analysis Pipelines for Large-Scale Data Management, HPC, and Machine Learning

Project Information

DAPHNE

Grant agreement ID: 957407

DOI

10.3030/957407

Project closed

EC signature date 14 July 2020

Start date 1 December 2020

End date 30 November 2024

Funded under

INDUSTRIAL LEADERSHIP - Leadership in enabling and industrial technologies - Information and Communication Technologies (ICT)

Total cost

€ 6 609 665,00

EU contribution

€ 6 609 665,00

6 609 665,00

Coordinated by

KNOW CENTER RESEARCH GMBH
Austria

Project description

New systems for today’s data-driven applications

The infrastructure for data management is growing fast. For accurate predictions, modern data-driven applications leverage large, heterogeneous data collections to uncover interesting patterns. They also build robust machine-learning models for accurate predictions. As a result, new systems have been developed with traditional, high-performance computing and the architecture of underlying hardware clusters. There is also a trend toward complex data analysis pipelines that combine different systems. The EU-funded DAPHNE project will define an open and extensible systems infrastructure for integrated data analysis pipelines. It will build a reference implementation of language abstractions (APIs and a domain-specific language) and an intermediate representation, as well as compilation and runtime techniques.

Objective

Modern data-driven applications leverage large, heterogeneous data collections to find interesting patterns, and build robust machine learning (ML) models for accurate predictions. Large data sizes and advanced analytics spurred the development and adoption of data-parallel computation frameworks like Apache Spark or Flink as well as distributed ML systems like MLlib, TensorFlow, or PyTorch. A key observation is that these new systems share many techniques with traditional high-performance computing (HPC), and the architecture of underlying HW clusters converges. Yet, the programming paradigms, cluster resource management, as well as data formats and representations differ substantially across data management, HPC, and ML software stacks. There is a trend though, toward complex data analysis pipelines that combine these different systems. Examples are workflows of distributed data pre-processing, tuned HPC libraries, and dedicated ML systems, but also HPC applications that leverage ML models for more cost-effective simulation. Major obstacles are (1) limited development productivity for integrated analysis pipelines due to different programming models, and separated cluster environments, (2) unnecessary data movement overhead and underutilization due to separate, statically provisioned clusters, and (3) lack of a common system infrastructure with good interoperability. For these reasons, DAPHNE’s overall objective is the definition of an open and extensible systems infrastructure for integrated data analysis pipelines. We aim at building a reference implementation of language abstractions (i.e. APIs and a domain-specific language), an intermediate representation, as well as compilation and runtime techniques with support for integrating and scheduling heterogeneous accelerator and storage devices. A variety of real-world, high-impact use cases, datasets, and a new benchmark will be used for qualitative and quantitative analysis compared to state-of-the-art.

Fields of science (EuroSciVoc)

CORDIS classifies projects with EuroSciVoc, a multilingual taxonomy of fields of science, through a semi-automatic process based on NLP techniques. See: The European Science Vocabulary.

Keywords

Project’s keywords as indicated by the project coordinator. Not to be confused with the EuroSciVoc taxonomy (Fields of science)

Programme(s)

Multi-annual funding programmes that define the EU’s priorities for research and innovation.

H2020-EU.2.1.1. - INDUSTRIAL LEADERSHIP - Leadership in enabling and industrial technologies - Information and Communication Technologies (ICT) MAIN PROGRAMME
See all projects funded under this programme

Topic(s)

Calls for proposals are divided into topics. A topic defines a specific subject or area for which applicants can submit proposals. The description of a topic comprises its specific scope and the expected impact of the funded project.

ICT-51-2020 - Big Data technologies and extreme-scale analytics
See all projects funded under this topic

Funding Scheme

Funding scheme (or “Type of Action”) inside a programme with common features. It specifies: the scope of what is funded; the reimbursement rate; specific evaluation criteria to qualify for funding; and the use of simplified forms of costs like lump sums.

RIA - Research and Innovation action

See all projects funded under this funding scheme

Call for proposal

Procedure for inviting applicants to submit project proposals, with the aim of receiving EU funding.

(opens in new window) H2020-ICT-2018-20

See all projects funded under this call

Coordinator

KNOW CENTER RESEARCH GMBH

Net EU contribution

€ 737 732,50

Address

SANDGASSE 34
8010 Graz
Austria

Region

Südösterreich Steiermark Graz

Activity type

Research Organisations

Links

Contact the organisation

Website

Participation in EU R&I programmes

HORIZON collaboration network

Total cost

€ 737 732,50

Participants (13)

AVL LIST GMBH

Austria

Net EU contribution

€ 419 175,00

DEUTSCHES ZENTRUM FUR LUFT - UND RAUMFAHRT EV

Germany

Net EU contribution

€ 849 830,00

EIDGENOESSISCHE TECHNISCHE HOCHSCHULE ZUERICH

Switzerland

Net EU contribution

€ 448 032,50

HASSO-PLATTNER-INSTITUT FUR DIGITAL ENGINEERING GGMBH

Germany

Net EU contribution

€ 458 750,00

EREVNITIKO PANEPISTIMIAKO INSTITOUTO SYSTIMATON EPIKOINONION KAI YPOLOGISTON

Greece

Net EU contribution

€ 415 000,00

INFINEON TECHNOLOGIES AUSTRIA AG

Austria

Net EU contribution

€ 436 015,00

INTEL TECHNOLOGY POLAND SPOLKA Z OGRANICZONA ODPOWIEDZIALNOSCIA

Poland

Net EU contribution

€ 271 375,00

IT-UNIVERSITETET I KOBENHAVN

Denmark

Net EU contribution

€ 523 100,00

KAI KOMPETENZZENTRUM AUTOMOBIL - UND INDUSTRIEELEKTRONIK GMBH

Austria

Net EU contribution

€ 470 875,00

TECHNISCHE UNIVERSITAET DRESDEN

Germany

Net EU contribution

€ 409 500,00

UNIVERZA V MARIBORU

Slovenia

Net EU contribution

€ 244 975,00

UNIVERSITAT BASEL

Switzerland

Net EU contribution

€ 700 148,75

TECHNISCHE UNIVERSITAT BERLIN

Germany

Net EU contribution

€ 225 156,25

Project description

New systems for today’s data-driven applications

Objective

Fields of science (EuroSciVoc) CORDIS classifies projects with EuroSciVoc, a multilingual taxonomy of fields of science, through a semi-automatic process based on NLP techniques. See: The European Science Vocabulary.

Keywords Project’s keywords as indicated by the project coordinator. Not to be confused with the EuroSciVoc taxonomy (Fields of science)

Programme(s) Multi-annual funding programmes that define the EU’s priorities for research and innovation.

Topic(s) Calls for proposals are divided into topics. A topic defines a specific subject or area for which applicants can submit proposals. The description of a topic comprises its specific scope and the expected impact of the funded project.

Funding Scheme Funding scheme (or “Type of Action”) inside a programme with common features. It specifies: the scope of what is funded; the reimbursement rate; specific evaluation criteria to qualify for funding; and the use of simplified forms of costs like lump sums.

Call for proposal Procedure for inviting applicants to submit project proposals, with the aim of receiving EU funding.

Coordinator

Participants (13)

Download Download the content of the page

Fields of science (EuroSciVoc)

CORDIS classifies projects with EuroSciVoc, a multilingual taxonomy of fields of science, through a semi-automatic process based on NLP techniques. See: The European Science Vocabulary.

Keywords

Project’s keywords as indicated by the project coordinator. Not to be confused with the EuroSciVoc taxonomy (Fields of science)

Programme(s)

Multi-annual funding programmes that define the EU’s priorities for research and innovation.

Topic(s)

Calls for proposals are divided into topics. A topic defines a specific subject or area for which applicants can submit proposals. The description of a topic comprises its specific scope and the expected impact of the funded project.

Funding Scheme

Funding scheme (or “Type of Action”) inside a programme with common features. It specifies: the scope of what is funded; the reimbursement rate; specific evaluation criteria to qualify for funding; and the use of simplified forms of costs like lump sums.

Call for proposal

Procedure for inviting applicants to submit project proposals, with the aim of receiving EU funding.