The language of native speakers is idiosyncratic. Thus, we know that the early bird gets the worm, we check for the best before date of the grocery we purchase, we jump the line, we face the challenge, and we are sometimes deeply sorry. The MaPPLexiC project focuses on one type of idiosyncratic constructions, which stands out in terms of its frequency of use and as a measure of language proficiency of a learner: (lexical) collocations. The above cited jump [the] line, face [the] challenge, and deeply sorry are examples of them. According to the definition common in lexicography, a collocation is a semi-compositional idiosyncratic co-occurrence of two lexical items that form a syntactico-semantic pattern, where the occurrence of one of them (the “base”) is freely chosen by the Speaker, while the other (the “collocate”) is predetermined by the base and conditioned by the meaning that the Speaker intends to express by the collocate. In face [the] challenge, challenge is the base, face the collocate, ‘deal with’ the intended meaning, and the base is the Direct Grammatical Object of the collocate; in deeply sorry, sorry is the base, the intended meaning is ‘intense’ and deeply is the collocate, which is a modifier of the base.
Collocations are often referred to as “prefabricated phrases”, but the fact that they are prefabricated does not mean that their construction is ad hoc, i.e. not guided by covert
semantic, cognitive, contextual, stylistic and other criteria. Small scale studies show that bases that share primary semantic features also tend to share collocates. There is also evidence from cognitive and socio-cultural studies that collocation creation follows covert, but well-defined criteria. However, the current state of the art still lacks a systematic account of how collocations are constructed across languages. This has far-reaching consequences for the theory of language and for its application alike.
The theoretical consequence is that the claim, advanced by a substantial number of modern linguistic theories, that there is no principled boundary between grammar and the lexicon has so far been supported mainly by evidence from regular syntactic structures involving collocations. It has not yet been complemented by a dynamic grammar of collocation production.The practical consequence is that language learners are still largely required to memorize collocations or look them up in dictionaries, without access to tools that would help them understand how such expressions are constructed.
MaPPLexiC sets out to derive and make explicit the production principles underlying the creation of collocations on a large scale, drawing upon neural large language model (LLM)
technologies. “Collocation production principles” are in this context inherent rules of the language system that determine which semantic, lexical, contextual, and other features of the base words and the surrounding discourse prompt the selection of specific collocates to convey the intended meaning.
In more specific terms, MaPPLexiC targets the following objectives:
O1: Retrieve from the internal neural representations of the LLM human interpretable features that are decisive for the collocation status and collocation type of a given word
combination.
O2: Develop a model for distilling collocation production principles for common collocation patterns from the obtained human interpretable features.
O3: Carry out language-contrastive studies and explore cross-language transfer of collocation production principles, to account for less resourced languages.
O4: Build a test workbench and test extensively the collocation production principles and their cross-language transfer.
The collocation production principles should eventually cover all collocation patterns. However, due to feasibility considerations, MaPPLexiC starts with the 20 most recurrent
patterns in the available language data. To facilitate cross-language studies, MaPPLexiC works with pairs of Germanic (English, German), Romance (French and Spanish), Finno-Ugric (Finnish and Hungarian), and Slavic (Czech and Russian) languages.