Selected Publications
2025
MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
Transactions of the Association for Computational Linguistics (TACL)
We introduce MultiBLiMP 1.0, a massively multilingual benchmark of linguistic minimal pairs, covering 101 languages and 2 types of subject-verb agreement, containing more than 128,000 minimal pairs. Our minimal pairs are created using a fully automated pipeline, leveraging the large-scale linguistic resources of Universal Dependencies and UniMorph. MultiBLiMP evaluates abilities of LLMs at an unprecedented multilingual scale, and highlights the shortcomings of the current state-of-the-art in modelling low-resource languages.
2025
Hybrid Human-LLM Corpus Construction and LLM Evaluation for the Caused-Motion Construction
Northern European Journal of Language Technology (NEJLT)
The caused-motion construction (CMC, "She sneezed the foam off her cappuccino") is one of the most well-studied constructions in Construction Grammar (CxG). It is a prime example for describing how constructions must carry meaning, as otherwise the fact that "sneeze" in this context takes two arguments and causes motion cannot be explained. We form the hypothesis that this remains challenging even for state-of-the-art Large Language Models (LLMs), for which we devise a test based on substituting the verb with a prototypical motion verb. To be able to perform this test at a statistically significant scale, in the absence of adequate CxG corpora, we develop a novel pipeline of NLP-assisted collection of linguistically annotated text. We show how dependency parsing and LLMs can be used to significantly reduce annotation cost and thus enable the annotation of rare phenomena at scale. We then evaluate OpenAI, Gemma3, Llama3, OLMo2, Mistral and Aya models for their understanding of the CMC using the newly collected corpus. We find that most models struggle with understanding the motion component that the CMC adds to a sentence.
2024
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons
LREC-COLING 2024
In this paper, we make a contribution that can be understood from two perspectives: from an NLP perspective, we introduce a small challenge dataset for NLI with large lexical overlap, which minimises the possibility of models discerning entailment solely based on token distinctions, and show that GPT-4 and Llama 2 fail it with strong bias. We then create further challenging sub-tasks in an effort to explain this failure. From a Computational Linguistics perspective, we identify a group of constructions with three classes of adjectives which cannot be distinguished by surface features. This enables us to probe for LLM's understanding of these constructions in various ways, and we find that they fail in a variety of ways to distinguish between them, suggesting that they don't adequately represent their meaning or capture the lexical properties of phrasal heads.
2023
Counting the Bugs in ChatGPT's Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model
EMNLP 2023
Large language models (LLMs) have recently reached an impressive level of linguistic capability, prompting comparisons with human language skills. However, there have been relatively few systematic inquiries into the linguistic capabilities of the latest generation of LLMs, and those studies that do exist (i) ignore the remarkable ability of humans to generalize, (ii) focus only on English, and (iii) investigate syntax or semantics and overlook other capabilities that lie at the heart of human language, like morphology. Here, we close these gaps by conducting the first rigorous analysis of the morphological capabilities of ChatGPT in four typologically varied languages (specifically, English, German, Tamil, and Turkish). We apply a version of Berko's (1958) wug test to ChatGPT, using novel, uncontaminated datasets for the four examined languages. We find that ChatGPT massively underperforms purpose-built systems, particularly in English. Overall, our results — through the lens of morphology — cast a new light on the linguistic capabilities of ChatGPT, suggesting that claims of human-like language skills are premature and misleading.
2023
Construction Grammar Provides Unique Insight into Neural Language Models
First International Workshop on Construction Grammars and NLP (CxGs+NLP, GURT/SyntaxFest 2023)
Construction Grammar (CxG) has recently been used as the basis for probing studies that have investigated the performance of large pretrained language models (PLMs) with respect to the structure and meaning of constructions. In this position paper, we make suggestions for the continuation and augmentation of this line of research. We look at probing methodology that was not designed with CxG in mind, as well as probing methodology that was designed for specific constructions. We analyse selected previous work in detail, and provide our view of the most important challenges and research questions that this promising new field faces.
2022
The Better Your Syntax, the Better Your Semantics? Probing Pretrained Language Models for the English Comparative Correlative
EMNLP 2022
Construction Grammar (CxG) is a paradigm from cognitive linguistics emphasising the connection between syntax and semantics. Rather than rules that operate on lexical items, it posits constructions as the central building blocks of language, i.e., linguistic units of different granularity that combine syntax and semantics. As a first step towards assessing the compatibility of CxG with the syntactic and semantic knowledge demonstrated by state-of-the-art pretrained language models (PLMs), we present an investigation of their capability to classify and understand one of the most commonly studied constructions, the English comparative correlative (CC). We conduct experiments examining the classification accuracy of a syntactic probe on the one hand and the models' behaviour in a semantic application task on the other, with BERT, RoBERTa, and DeBERTa as the example PLMs. Our results show that all three investigated PLMs are able to recognise the structure of the CC but fail to use its meaning. While human-like performance of PLMs on many NLP tasks has been alleged, this indicates that PLMs still suffer from substantial shortcomings in central domains of linguistic knowledge.
All Publications
Synced from my Google Scholar profile.
2026
30th Conference on Computational Natural Language Learning (CoNLL)
2025
Proceedings of the Second International Workshop on Construction Grammars and NLP
Proceedings of the Second International Workshop on Construction Grammars and NLP
2025
Findings of the UniDive 2025 shared task on multilingual Morpho-Syntactic Parsing
Proceedings of The UniDive 2025 Shared Task on Multilingual Morpho-Syntactic Parsing
2025
BabyLM's First Constructions: Causal interventions provide a signal of learning
Empirical Methods in Natural Language Processing (EMNLP)
2025
MultiBLiMP 1.0: A massively multilingual benchmark of linguistic minimal pairs
Transactions of the Association for Computational Linguistics (TACL)
2025
Both direct and indirect evidence contribute to dative alternation preferences in language models
Conference on Language Modeling (COLM)
2025
Constructions are revealed in word distributions
Empirical Methods in Natural Language Processing (EMNLP)
2025
Linguistic generalizations are not rules: Impacts on evaluation of LMs
Second International Workshop on Construction Grammars and NLP (CxG+NLP)
2024
Derivational Morphology Reveals Analogical Generalization in Large Language Models
Proceedings of the National Academy of Sciences (PNAS)
2024
SynthEval: Hybrid Behavioral Testing of NLP Models with Synthetic Evaluation
Findings of Empirical Methods in Natural Language Processing (EMNLP)
2024
Models Can and Should Embrace the Communicative Nature of Human-Generated Math
Workshop on Mathematical Reasoning and AI, NEURIPS
2024
Proceedings of the Sixth Workshop on Teaching NLP
Proceedings of the Sixth Workshop on Teaching NLP
2024
Sixth Workshop on Teaching NLP
2024
2024
Constructions Are So Difficult That Even Large Language Models Get Them Right for the Wrong Reasons
Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING)
2024
Verbing Weirds Language (Models): Evaluation of English Zero-Derivation in Five LLMs
Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING)
2024
UCxn: Typologically Informed Annotation of Constructions Atop Universal Dependencies
Joint International Conference on Computational Linguistics, Language Resources and Evaluation
2024
Hybrid Human-LLM Corpus Construction and LLM Evaluation for the Caused-Motion Construction
Northern European Journal of Language Technology (NEJLT)
2023
Empirical Methods in Natural Language Processing (EMNLP)
2023
Frontiers in Artificial Intelligence
2023
Association for Computational Linguistics (ACL)
2023
Findings of Empirical Methods in Natural Language Processing (EMNLP)
2023
A Crosslingual Investigation of Conceptualization in 1335 Languages
Association for Computational Linguistics (ACL)
2023
4.4 WG4: Finding Idiosyncrasy in Corpora
Report from Dagstuhl Seminar 23191: Universals of Linguistic Idiosyncrasy in Multilingual Computational Linguistics
2023
Construction Grammar Provides Unique Insight into Neural Language Models
CxGs + NLP Workshop, Georgetown University Round Table (GURT)
2022
Empirical Methods in Natural Language Processing (EMNLP)
2022
CaMEL: Case Marker Extraction without Labels
Association for Computational Linguistics (ACL)
2017
Developing a stemmer for German based on a comparative analysis of publicly available stemmers
German Society for Computational Linguistics (GSCL)