The academic output has been substantial, with 59 publications in major peer-reviewed venues in machine learning and computer vision, dissemination in numerous international workshops and summer schools, organisation of scientific meetings and benchmark challenges and the release of several new datasets and open source code implementing the new algorithms. According to impact metrics such as citation counts in the several thousands, this output has been highly impactful in the research community. One of our papers received the best paper award from the Conference on Computer Vision and Pattern Recognition, our largest international meeting -- this paper is awarded yearly to one out of roughly 5000 papers. As part of the project, several postdoctoral researchers and PhD students have been trained, and have since obtained prestigious research position in academia and industry, including becoming professors and joining research labs such as Google Research, Deep Mind, Facebook AI Research and others.
The project has also achieved significant technical progress in all the key challenges.
First, we have developed new methods to understand images in detail with less or no supervision. For example, we have built the first system that can learn about the parts of objects (e.g. that a human’s body is composed of several limbs connected together) in a completely unsupervised manner, by looking for this information on the Web, using Google or a similar search engine automatically. Second, we have invented a new approach to unsupervised learning, which allows a computer to learn about the structure of visual objects all without a single manual annotation other than the images themselves. The learning principle which we discovered, which we call factor learning, is powerful and general and has been demonstrated in numerous applications and examples in the project.
Second, we have made strides in integrated image understanding demonstrating that it is possible, by using certain technical innovations, to build single deep neural network models that can understand very diverse image types, from handwritten digits to images of sport or animals, with a very small incremental cost compared to learning about one of such domains at a time. The resulting models can be one or two order of magnitude smaller than learning models individually, and the overall performance, due to sharing of acquired visual capabilities between domains, is actually improved.
Third, we have developed new methods, including theory, to better understand the outcome of complex learning processes such as deep learning; in fact, current models are, similar to the human brain, black boxes that are induced automatically from empirical experiences. Thus how these models work in practice remains unknown; our new techniques allow to visualize what happens inside a deep network, to better understand how it works and what are its limitations.
Significantly, some of this technology is already been tested for real-world deployment in various research and business contexts, including in orthogonal research areas such as bibliography, material science, and zoology. Furthermore, we have initiated a partnership with Continental corporation on autonomous and assisted driving, where our methods turned out to be extremely useful in the context of learning from large quantities of car-collected data while requiring little manual intervention.