In the first scientific reporting period (initial 2 years), we started developing the multimodal GLW architecture and evaluating its abilities related to semantic grounding. With vision and language as input modules, the GLW system can exploit unsupervised training objectives to learn a shared multimodal representation with 5 to 10 times less labelled data than a standard supervised system (Devillers, Maytié & VanRullen, IEEE Transactions in Neural Networks & Learning Systems, 2024). In addition, the resulting multimodal representation can be effectively leveraged for downstream classification tasks, with improved performance compared with state-of-the-art multimodal AI systems like CLIP. We also demonstrated the ability of the GLW architecture to capture sensorimotor affordances for a simulated navigation agent (Kuske & VanRullen, Cognitive Computational Neuroscience 2025). Finally, our work revealed that reinforcement learning (RL) systems using a GLW architecture as an observation space enjoy numerous advantages over standard methods (even those relying on multimodal inputs such as CLIP). Our trained RL agents are capable of zero-shot transfer, i.e. applying a RL policy trained in one modality to inputs recorded in another modality (Maytié, Devillers, Arnold & VanRullen, Reinforcement Learning Journal, 2024). When equipped with a “world model” trained to predict the consequences of planned actions, the RL agents learn even more efficiently, by “dreaming” entire episodes (Maytié, Bertin-Johannet & VanRullen, arXiv 2025); the resulting RL systems outperform state-of-the-art model-based RL methods, like DreamerV3.
In parallel, we have worked on designing a dedicated attention mechanism for module selection in the GLW. We confirmed the mechanism’s usefulness for robust multimodal integration (Bertin-Johannet, Scipio & VanRullen, arXiv 2025), and for sequential deployment of operation modules in the context of an arithmetic addition task (Chateau-Laurent & VanRullen, arXiv 2025).
Our project also aims at testing the relevance of the GLW cognitive architecture as a model of brain processing, with possible applications in brain decoding. To that end, we collected a large-scale multimodal fMRI dataset (SemReps-8K, with 6 human subjects, each viewing more than 8,000 images or captions). We leveraged this dataset to demonstrate the feasibility of modality-agnostic brain decoding, and to identify modality-invariant brain regions (Nikolaus et al, Cognitive Computational Neuroscience, 2025; eLife 2025). These regions constitute a pool of candidates for the neuronal substrates of the Global Workspace. We released the dataset on the open-science platform OpenNeuro:
https://openneuro.org/datasets/ds006798/versions/1.0.0(si apre in una nuova finestra)Finally, our project considers the practical and ethical implications of implementing consciousness theories, like the Global Workspace, in modern AI systems. In collaborative work with renowned experts from other disciplines, we discussed the plausibility of artificial consciousness and its practical evaluation in current and future AI systems (Butlin et al, arXiv 2023; TiCS in-press).