Machine learning is a recent scientific field at the crossing of statistics,
computer science, and applied mathematics. It may be seen as the science of
data analysis and potentially impacts every domain where data has become a
potential source of knowledge or economic activity. It has become a key part
of scientific fields that produce a massive amount of data and that are in dire
need of scalable tools to automatically make sense of it. Unfortunately,
classical statistical modeling has often become impractical due to recent
shifts in the amount of data to process, and in the high complexity and large
size of models that are able to take advantage of massive data. The promise of
SOLARIS is to invent a new generation of machine learning models that fulfill
the current needs of large-scale data analysis: high scalability, ability to
deal with huge-dimensional models, fast learning, easiness of use, and
adaptivity to various data structures.
These problems are important for society. Besides a potential impact across
different disciplines, big data is shifting classical paradigms and developing
scalable technology to make sense of massive data has become a strategic issue
for society.
The project has four scientific objectives: (i) Providing more
scalability to nonlinear models in machine learning, and developing learning
schemes that can deal with large model sizes and huge amounts of training data.
(ii) Gaining theoretical insight and principled methodology for deep learning
by marrying two schools of thought that have been considered so far to have
little overlap: kernel methods and deep learning. The former is associated
with a well-understood theory and methodology but lacks scalability, whereas
the latter has obtained significant success on large-scale prediction problems,
notably in computer vision. (iii) Building scalable machine learning techniques
for structured data such as sequences or graphs. (iv) Pushing the frontiers of
image and video modeling.
By focusing on the previous four objectives, the project has led to many
contributions, which range from theory (e.g. for better understanding of deep
neural networks), to algorithms and methods (in machine learning, computer
vision, optimization), software developments in open-source toolboxes, and
real-world applications in various scientific fields (with a strong emphasis on
visual modeling).