Header

Search

WWW-SArDEEn

Project Summary

One of the most fundamental tasks language users face when interpreting utterances is determining the role of arguments in the clause, such as identifying who (subject) did what to whom (object). To do so, we make use of contextual and semantic information, but also four main morphosyntactic strategies: (i) noun inflection (case marking), (ii) verb inflection (agreement), (iii) fixed word order, and (iv) prepositions. In addition to featuring extensively in research across all linguistic fields, the latter are also the subject of one of the most prominent textbook claims about the history of English, specifically the turn from Old English (OE; 650-1150) to Middle English (ME; 1150-1500). The assumption is that English underwent a trade-off between disambiguation strategies in that (i) case marking and (ii) verb inflection were originally more prominent and decreased, (iii) word order used to be more flexible and became fixed, and (iv) prepositions increased. However, despite the popularity of this claim, we know surprisingly little about how these changes are in fact connected, and studies on the history of English are strikingly disconnected from current relevant research in cognitive linguistics, psycholinguistics, typology, and related fields. In addition, research in this area showcases substantial shortcomings of resources and methods in historical linguistics, in particular challenges related to data sparsity and textual features such as extensive spelling variation. These issues have presented great obstacles for data preparation, classification, and analysis, and have also meant that the immense recent progress in the development of artificial intelligence (AI) models which are able to handle language-related questions, and especially the use of large language models (LLMs) for linguistic tasks has not yet reached historical (early English) linguistics.  

The proposed project seeks to bridge these two gaps. On the one hand, it tackles the gap between historical linguistics and recent cognitive, psycholinguistic, and typological insights on disambiguation strategies by extending relevant concepts to the history of English. On the other hand, it brings together historical linguistics and current methodological advances, by developing highly data-driven, LLM-based methods for OE and ME data, informed by profound linguistic, philological, and sociocultural knowledge of these periods.

To accomplish its aims, the project rests on three pillars: Pillar (1) is the empirical, quantitative, statistical substantiation of claims about disambiguation strategies in the history of English, in particular the assumption of a causal trade-off between strategies. Pillar (2) connects changes observed in the history of English to recent theoretical (and methodological) insights on trade-offs vs redundancy, and information-theoretic and cognitive notions such as efficiency or robustness of information transfer. Pillar (3) concerns resource and tool development, using LLMs for lemmatisation, parsing, and semantic classification of historical English data. These novel tools will also be made accessible to the (historical) linguistics community to facilitate future research, and can in turn inform work in natural language processing (NLP).

In sum, the project addresses the crucial question of how we distinguish the different players in an event through linguistic strategies and how this changed in the history of English due to reasons such as efficiency. To do so, it combines careful, historical linguistic research into early English with current findings in other fields and methodological advances in NLP and AI, and aids scientific progress through open resource sharing.