We have collected, annotated and prepared for public release the ECOLANG corpus: a large multimodal corpus of the verbal and non-verbal (speech, gestures, eye-gaze) behaviours by a speaker in interaction with a child or adult partner. Conversations across all dyads are comparable and centre around life-like scenarios. To our knowledge, this will be (public release is imminent) the first large (more than 50 hrs of recording) corpus available annotated for all these different behaviours and we expect it to be useful to researchers from a variety of disciplines.
Analyses of the corpus data have allowed us to make significant progress on understanding when and why speakers use non-verbal behaviours in addition to speech. Overall, speakers use specific non-verbal cues when they are most useful to their listeners. For example, they use iconic non-verbal behaviours (iconic gestures or vocal iconicity) when talking about objects that are not in view, when they talk about novel objects, and when they are about to say a word that is less predictable in context. We were then also able to link these behaviours to the learning of new words and concepts by the listener (child or adult), finding that a small set of non-verbal behaviours (especially points and iconic gestures) support learning in interaction with other linguistic variables, while different non-verbal behaviours (manipulation of objects) instead hinder learning. These findings provide a clear set of “do’s and don’t’s” for improving successful teaching.
In behavioural and electrophysiological studies, we assessed the impact of multimodal cues on word and discourse processing. We found that informative (i.e. more iconic) gestures always speed up processing, while seeing the speaker’s mouth movements only helps when processing is difficult (either because there is noise or the listener is non-native). In electrophysiological studies, we found that all the multimodal cues we investigated affect word processing, indicating that they are central to language processing. We also found that their impact dynamically changes depending upon their informativeness. These studies provide a first snapshot into how the brain dynamically weights audiovisual cues in language comprehension. Finally, in work with people with aphasia, we identified the neural regions involved in integrating speech, iconic gestures and mouth movements. These results are clinically relevant as they provide insight into who can benefit from audio-visual treatment.