Conveners
Parallel sessions 1 (Arnold hall)
- Tanara Zingano Kuhn
Parallel sessions 1 (Arnold hall)
- Jaka ฤibej
Parallel sessions 1 (Arnold hall)
- Kristina Kocijan
Parallel sessions 1 (Arnold hall)
- Kristina Kocijan
Parallel sessions 1 (Arnold hall)
- Ivana Filipoviฤ Petroviฤ
Parallel sessions 1 (Arnold hall)
- Kristina Koppel
Parallel sessions 1 (Arnold hall)
- Carole Tiberius
Parallel sessions 1 (Arnold hall)
- Jelena Kallas
-
Mr Robert Lew, Bartosz Ptasznik11/18/25, 11:00โฏAM
The public release of ChatGPT in late 2022 made an impact on many professional domains. Notwithstanding the many controversies surrounding Generative Artificial Intelligence (GenAI), such as ethics, copyright, accountability, or ecology, we need to acknowledge an important and relevant feature of Large Language Models and chatbot systems built around them: their ability to produce mostly...
Go to contribution page -
Enikล Hรฉja, Lรกszlรณ Simon, Veronika Lipp11/18/25, 11:30โฏAM
Recent findings indicate that current large language models (LLMs) face difficulties in generating clear-cut, well-motivated definitions in a consistent way. This shortcoming is the consequence of their reliance on opaque data sources and their inherently unstable, non-deterministic outputs. In response, this research aims to develop an LLM-based methodology for producing adjectival...
Go to contribution page -
Dominik Schlechtweg, Emma Skรถldberg, Shafqat Mumtaz Virk, James White, Simon Hengchen11/18/25, 12:00โฏPM
Finding non-recorded senses is important for dictionary maintenance, where using automatic methods helps reduce manual efforts. We use automatic Word Sense Induction (WSI) to compare recorded sense numbers among a sample of headwords in a comprehensive Swedish monolingual dictionary with induced sense numbers for the same words in a Swedish corpus. We propose this as a simple technique to find...
Go to contribution page -
Ondลej Herman11/18/25, 12:30โฏPM
Language evolves continuously, rendering static dictionaries quickly outdated. While previous research has addressed the automatic detection of new words, identifying subtler semantic changes in existing words remains a challenge. In this work, we propose a robust, language-independent methodology for the automatic detection of word sense shifts using diachronic corpus data. Our approach...
Go to contribution page -
30. The Dictionary of Contemporary Serbian Language (RSSJ): Advanced Automation and Other ChallengesRanka Stankoviฤ, Rada Stijoviฤ, Mihailo ล koriฤ, Cvetana Krstev11/18/25, 2:30โฏPM
This paper introduces the Dictionary of Contemporary Serbian Language (RSSJ), an ongoing large-scale digital lexicographic project designed to serve both human users via web and mobile applications and machines through APIs. Coordinated by the diaspora association โGathered around the Languageโ and the Society for Language Resources and Technologies (JeRTeh), RSSJ aims to produce a dictionary...
Go to contribution page -
Nathalie Norman, Sanni Nimb, Sussi Olsen, Nina Schneidermann, Bolette S. Pedersen11/18/25, 3:00โฏPM
Large Language Models (LLMs) tend to expose severe language and cultural biases when working in medium- and low-resourced languages. In this paper, we present our work on Danish benchmarking and evaluation of LLMs to more precisely diagnose and potentially remedy such bias. To this aim, we apply available lexical-semantic resources to compile a set of Natural Language Understanding (NLU) tasks...
Go to contribution page -
Maria Tuulik, Ene Vainik, Margit Langemets, Eleri Aedmaa, Lydia Risberg, Esta Prangel, Kristina Koppel, Sirli Zupping11/18/25, 3:30โฏPM
The use of corpora is well established in lexicography, also in Estonia, but since the analysis of corpus data and the post-editing of automatically generated data from the corpus is labour-intensive, the use of large language models (LLMs) has led to growing interest in lexicography (e.g., Evert et al. 2024; Kosem, Gantar et al. 2024; Tiberius et al. 2024). In 2024, the Institute of the...
Go to contribution page -
Polona Gantar, Cyprian Laskowski, Simon Krek11/19/25, 10:05โฏAM
In lexicography, one of the long-standing issues is understanding the nature of its core element of description commonly referred to as the headword (in DMLex and traditional lexicography), canonical form (in OntoLex and the Lexical Markup Framework โ LMF), orthographic form (in the Text Encoding Initiative โ TEI Lex0), lemma (in Wikidata), or lexical unit. With the transition from paper to...
Go to contribution page -
Elinor Hawkes, Phoebe Nicholson, Will Rogers11/19/25, 11:00โฏAM
This paper presents the Oxford English Dictionaryโs (OED) current exploration into the application of artificial intelligence to historical Word Sense Disambiguation (WSD), a fundamental aspect of OEDโs core research. Building on a longstanding tradition of technological innovation, the OED is investigating how Large Language Models (LLMs) can support the identification and retrieval of...
Go to contribution page -
Ivรกn Arias-Arias, Elena Martรญn-Cancela11/19/25, 11:30โฏAM
Generic nouns such as Sache and Ding pose a challenge for semantic annotation due to their referential underspecification and context-dependent meaning. Although frequently classified under categories like {artefact} or {object}, their actual referents often belong to abstract or cognitive domains, as in Der Placeboeffekt ist eines der faszinierendsten Dinge in der Welt der Medizin. Drawing on...
Go to contribution page -
Luise Kรถhler, Gregor Middell, Alexander Geyken11/19/25, 2:30โฏPM
Collocations are a well-covered research area in lexicography. With the advent of evidence-based lexicography and the availability of large text corpora, computational methods of extracting typical co-occurrences from such corpora and supporting lexicographers in identifying collocations among them became a research focus. Especially the statistical properties of collocations (i.e. application...
Go to contribution page -
Lydia Risberg, Eleri Aedmaa, Maria Tuulik, Margit Langemets, Ene Vainik, Esta Prangel, Kristina Koppel, Hanna Pook11/19/25, 3:00โฏPM
Language corpora have long been used in linguistics and lexicography, but recent developments now allow large language models (LLMs) to support or even transform these fields. This study investigates the potential of LLMs for annotating informal language use in Estonian โ a language underrepresented in LLM training data yet supported by a large corpus. Focusing on the informal register label...
Go to contribution page -
Aleksandra Markoviฤ, Ranka Stankoviฤ11/19/25, 3:30โฏPM
Automation has revolutionised lexicography, introducing the โpost-editing lexicographyโ model, where the role of the lexicographer involves refining automatically generated dictionary drafts. Since the launch of ChatGPT in November 2022, numerous papers have explored the potential applications of LLMs in dictionary production. The rapid evolution of LLMs necessitates a re-evaluation of...
Go to contribution page -
Urลกka Vranjek Oลกlak11/19/25, 4:30โฏPM
This paper explores the applicability of generative artificial intelligence in the field of language consulting, focusing on ChatGPT-4 and the Slovenian language. The analysis is based on an experiment involving 30 real user questions submitted to the Language Consulting Service (LCS) of the Fran Ramovลก Institute of the Slovenian Language. The questions cover a range of linguistic categories...
Go to contribution page -
Olena Synchak, Vasyl Starko, Mariana Burak, Mykhaylo Svystun11/19/25, 5:00โฏPM
While CEFR-aligned vocabulary profiles have been developed for many languages (e.g., English, German, and Swedish), Ukrainian as a foreign language (UFL) still lacks an empirically grounded lexical profile. A foundational issue in creating such profiles is combining lexical frequency data with expert knowledge to assign CEFR-level labels. Existing UFL word lists rely primarily on professional...
Go to contribution page -
Matej Meterc, Nataลกa Jakop11/19/25, 5:30โฏPM
In preparing phraseological units for the third edition of the Standard Slovenian Dictionary (eSSKJ), the authors aimed to identify the most relevant comparative phrasemes in the contemporary standard language using objective corpus-based criteria. A key goal was to determine how representative specific phrasemes and their variants are in actual use. Two lists of the hundred most frequent...
Go to contribution page -
Ana Frankenberg-Garcia11/20/25, 10:05โฏAM
The use of LLMs in lexicography is a hot topic and indeed the focus of eLex 2025. In the past couple of years, several papers have emerged comparing existing dictionary entries with zero-shot chatbot queries (e.g. Nichols 2023) or with dictionary-like content obtained through the dynamic interaction between experts and chatbots (e.g. Lew 2023, Jakubรญฤek & Rundell 2023). However, studies so far...
Go to contribution page -
Simon Krek, Primoลพ Ponikvar, Andraลพ Repar, Iztok Kosem, David Lindemann11/20/25, 11:00โฏAM
This paper presents an experimental workflow for converting legacy digitized dictionaries into the DMLex standard and subsequently importing them into a Wikibase instance. DMLex, a serialization-independent model developed by the OASIS LEXIDMA Technical Committee, aims to provide a universal and modular representation of lexicographic data. The study tested whether dictionaries from...
Go to contribution page -
Andrej Perdih, Dejan Gabrovลกek, Janoลก Jeลพovnik11/20/25, 11:30โฏAM
This paper evaluates the results of using GPT-4o mini language model batch processing with image recognition capability to align 1,572 images of 398 polysemous nouns in the Dictionary of the Slovenian Standard Language (second edition) to their specific dictionary senses, and it compares them to the results of the manual image-to-sense alignment process. The images were manually assigned to...
Go to contribution page -
Kris Heylen, Vincent Prins, Katrien Depuydt, Jesse de Does, Laura van Eerten, Thomas Haga11/20/25, 12:00โฏPM
Representative monitor corpora with detailed metadata offer a solid empirical basis for documenting lexical innovation and change (Kosem et al. 2021). However, continuously updated time-stamped textual data presents challenges for data management, lexicographic analysis, and visualization. Building on its existing corpus infrastructure, the Dutch Language Institute (INT) has developed...
Go to contribution page