Part of Computational Humanities Research Group and King’s Linguistics Network
Register here: https://forms.gle/D8Kb5xAN5GUTa55t9
Hybrid (King’s College London, Room TBD + online via MS Teams)
19 November 2026, 4-5pm GMT
Roksana Goworek (Queen Mary University of London), Meaning What, Exactly? Measuring and Learning Lexical Meaning Across Time and Languages
Abstract
How can we analyse changes in word meaning across diachronic corpora? Changes in how words are used can reflect broader shifts in language and culture, making their detection relevant to historical linguistics as well as the study of cultural and social change. Contextualised language models have enabled large-scale analysis of semantic change by representing individual word usages in a shared vector space and comparing their distributions across corpora from different time periods. This raises several methodological challenges, from deciding how sets of representations should be compared, to obtaining representations that capture sense distinctions when target-language data are scarce, to creating the high-quality annotated data needed for training and evaluation.
Different approaches to comparing distributions of word usages can produce different estimates of semantic change, and their robustness may depend on the quality and structure of the underlying representations. These issues become particularly important for languages for which suitable models, corpora, or annotated data are limited. When target-language resources are scarce, the choice of training data can influence how well models generalise to those languages.
Constructing suitable training and evaluation data presents its own difficulties. Manual word sense annotation is costly, while rare or emerging usages, which can be particularly informative for detecting changes in meaning, can be difficult to find through conventional corpus sampling. Targeted sampling can make this process more efficient, helping to construct the high-quality human-annotated resources needed to establish whether computational methods work reliably for low-resource languages.
This talk explores how choices in measurement, model training, and data annotation shape our ability to detect semantic change, and what is required to extend such research reliably to a wider range of languages. More broadly, it argues that the methods and resources underlying semantic change detection need to be scrutinised carefully if we are to make reliable claims about changes in meaning.
Bio
Roksana Goworek is a final-year PhD student at Queen Mary University of London. Her research focuses on how language models represent meaning, with interests spanning lexical semantics, cross-lingual transfer and information retrieval. More recently, she has also worked on mechanistic interpretability, investigating how contextual information is processed and propagated through language models, including during a research internship at the Alan Turing Institute.

