The visual world around us is a source of rich semantic information that guides our higher-level cognitive processes and actions. To tap into this resource, the brain’s visual system engages in complex, intertwined computations to actively sample, extract, and integrate information across space and time. Surprisingly however, the integrative, dynamic, and active nature of vision hardly plays a role in the way we approach it in experimentation and computational modelling. Hence, despite significant advances in our general understanding of the neural mechanisms underlying vision, the brain’s natural modus operandi remains poorly understood. A new approach is needed.
TIME uses more natural viewing paradigms in both experimentation and modelling to close this gap. First, we simultaneously record high-resolution brain activity and eye-movements of participants while they explore and summarise natural visual scenes. The resulting large-scale resource will enable us to better understand the intricate mechanisms at play. In particular, it will enable us to study when and how the brain decides to move on from a given fixation location to the next, and how information from one fixation may affect the processing of the next, i.e. how information is integrated across fixations. In parallel to these analyses, we develop computational models that mirror the overall process. Here, we make use of recent developments in artificial intelligence to derive image-computable, large-scale models. A famous quote by Richard Feynman states that “What I cannot create, I do not understand”. This statement highlights the importance of computational models in science. Observing a system in experimentation and theorising about its inner mechanisms requires models to be built that put our hypotheses to the test. Therefore, in addition to building an AI-based model system, a final work package will test the alignment between brain data and the computational model.
In summary, using an interdisciplinary approach, TIME will establish when, where, and how visual semantic understanding emerges in the brain as it actively samples and integrates information from the world in a continuously updating and dynamic decision process. This approach marks a paradigm shift in how we study the visual system, and promises implications not only for neuroscience, but also for computer vision and artificial intelligence.