Nov 17โ€‰โ€“โ€‰20, 2025
Bled, Slovenia
Europe/Ljubljana timezone

Session

Parallel sessions 1 (Arnold hall)

Nov 18, 2025, 11:00โ€ฏAM
Arnold hall

Arnold hall

Conveners

Parallel sessions 1 (Arnold hall)

  • Tanara Zingano Kuhn

Parallel sessions 1 (Arnold hall)

  • Jaka ฤŒibej

Parallel sessions 1 (Arnold hall)

  • Kristina Kocijan

Parallel sessions 1 (Arnold hall)

  • Kristina Kocijan

Parallel sessions 1 (Arnold hall)

  • Ivana Filipoviฤ‡ Petroviฤ‡

Parallel sessions 1 (Arnold hall)

  • Kristina Koppel

Parallel sessions 1 (Arnold hall)

  • Carole Tiberius

Parallel sessions 1 (Arnold hall)

  • Jelena Kallas

Presentation materials

There are no materials yet.

  1. Mr Robert Lew, Bartosz Ptasznik
    11/18/25, 11:00โ€ฏAM

    The public release of ChatGPT in late 2022 made an impact on many professional domains. Notwithstanding the many controversies surrounding Generative Artificial Intelligence (GenAI), such as ethics, copyright, accountability, or ecology, we need to acknowledge an important and relevant feature of Large Language Models and chatbot systems built around them: their ability to produce mostly...

    Go to contribution page
  2. Enikล‘ Hรฉja, Lรกszlรณ Simon, Veronika Lipp
    11/18/25, 11:30โ€ฏAM

    Recent findings indicate that current large language models (LLMs) face difficulties in generating clear-cut, well-motivated definitions in a consistent way. This shortcoming is the consequence of their reliance on opaque data sources and their inherently unstable, non-deterministic outputs. In response, this research aims to develop an LLM-based methodology for producing adjectival...

    Go to contribution page
  3. Dominik Schlechtweg, Emma Skรถldberg, Shafqat Mumtaz Virk, James White, Simon Hengchen
    11/18/25, 12:00โ€ฏPM

    Finding non-recorded senses is important for dictionary maintenance, where using automatic methods helps reduce manual efforts. We use automatic Word Sense Induction (WSI) to compare recorded sense numbers among a sample of headwords in a comprehensive Swedish monolingual dictionary with induced sense numbers for the same words in a Swedish corpus. We propose this as a simple technique to find...

    Go to contribution page
  4. Ondล™ej Herman
    11/18/25, 12:30โ€ฏPM

    Language evolves continuously, rendering static dictionaries quickly outdated. While previous research has addressed the automatic detection of new words, identifying subtler semantic changes in existing words remains a challenge. In this work, we propose a robust, language-independent methodology for the automatic detection of word sense shifts using diachronic corpus data. Our approach...

    Go to contribution page
  5. Ranka Stankoviฤ‡, Rada Stijoviฤ‡, Mihailo ล koriฤ‡, Cvetana Krstev
    11/18/25, 2:30โ€ฏPM

    This paper introduces the Dictionary of Contemporary Serbian Language (RSSJ), an ongoing large-scale digital lexicographic project designed to serve both human users via web and mobile applications and machines through APIs. Coordinated by the diaspora association โ€œGathered around the Languageโ€ and the Society for Language Resources and Technologies (JeRTeh), RSSJ aims to produce a dictionary...

    Go to contribution page
  6. Nathalie Norman, Sanni Nimb, Sussi Olsen, Nina Schneidermann, Bolette S. Pedersen
    11/18/25, 3:00โ€ฏPM

    Large Language Models (LLMs) tend to expose severe language and cultural biases when working in medium- and low-resourced languages. In this paper, we present our work on Danish benchmarking and evaluation of LLMs to more precisely diagnose and potentially remedy such bias. To this aim, we apply available lexical-semantic resources to compile a set of Natural Language Understanding (NLU) tasks...

    Go to contribution page
  7. Maria Tuulik, Ene Vainik, Margit Langemets, Eleri Aedmaa, Lydia Risberg, Esta Prangel, Kristina Koppel, Sirli Zupping
    11/18/25, 3:30โ€ฏPM

    The use of corpora is well established in lexicography, also in Estonia, but since the analysis of corpus data and the post-editing of automatically generated data from the corpus is labour-intensive, the use of large language models (LLMs) has led to growing interest in lexicography (e.g., Evert et al. 2024; Kosem, Gantar et al. 2024; Tiberius et al. 2024). In 2024, the Institute of the...

    Go to contribution page
  8. Polona Gantar, Cyprian Laskowski, Simon Krek
    11/19/25, 10:05โ€ฏAM

    In lexicography, one of the long-standing issues is understanding the nature of its core element of description commonly referred to as the headword (in DMLex and traditional lexicography), canonical form (in OntoLex and the Lexical Markup Framework โ€“ LMF), orthographic form (in the Text Encoding Initiative โ€“ TEI Lex0), lemma (in Wikidata), or lexical unit. With the transition from paper to...

    Go to contribution page
  9. Elinor Hawkes, Phoebe Nicholson, Will Rogers
    11/19/25, 11:00โ€ฏAM

    This paper presents the Oxford English Dictionaryโ€™s (OED) current exploration into the application of artificial intelligence to historical Word Sense Disambiguation (WSD), a fundamental aspect of OEDโ€™s core research. Building on a longstanding tradition of technological innovation, the OED is investigating how Large Language Models (LLMs) can support the identification and retrieval of...

    Go to contribution page
  10. Ivรกn Arias-Arias, Elena Martรญn-Cancela
    11/19/25, 11:30โ€ฏAM

    Generic nouns such as Sache and Ding pose a challenge for semantic annotation due to their referential underspecification and context-dependent meaning. Although frequently classified under categories like {artefact} or {object}, their actual referents often belong to abstract or cognitive domains, as in Der Placeboeffekt ist eines der faszinierendsten Dinge in der Welt der Medizin. Drawing on...

    Go to contribution page
  11. Luise Kรถhler, Gregor Middell, Alexander Geyken
    11/19/25, 2:30โ€ฏPM

    Collocations are a well-covered research area in lexicography. With the advent of evidence-based lexicography and the availability of large text corpora, computational methods of extracting typical co-occurrences from such corpora and supporting lexicographers in identifying collocations among them became a research focus. Especially the statistical properties of collocations (i.e. application...

    Go to contribution page
  12. Lydia Risberg, Eleri Aedmaa, Maria Tuulik, Margit Langemets, Ene Vainik, Esta Prangel, Kristina Koppel, Hanna Pook
    11/19/25, 3:00โ€ฏPM

    Language corpora have long been used in linguistics and lexicography, but recent developments now allow large language models (LLMs) to support or even transform these fields. This study investigates the potential of LLMs for annotating informal language use in Estonian โ€“ a language underrepresented in LLM training data yet supported by a large corpus. Focusing on the informal register label...

    Go to contribution page
  13. Aleksandra Markoviฤ‡, Ranka Stankoviฤ‡
    11/19/25, 3:30โ€ฏPM

    Automation has revolutionised lexicography, introducing the โ€™post-editing lexicographyโ€™ model, where the role of the lexicographer involves refining automatically generated dictionary drafts. Since the launch of ChatGPT in November 2022, numerous papers have explored the potential applications of LLMs in dictionary production. The rapid evolution of LLMs necessitates a re-evaluation of...

    Go to contribution page
  14. Urลกka Vranjek Oลกlak
    11/19/25, 4:30โ€ฏPM

    This paper explores the applicability of generative artificial intelligence in the field of language consulting, focusing on ChatGPT-4 and the Slovenian language. The analysis is based on an experiment involving 30 real user questions submitted to the Language Consulting Service (LCS) of the Fran Ramovลก Institute of the Slovenian Language. The questions cover a range of linguistic categories...

    Go to contribution page
  15. Olena Synchak, Vasyl Starko, Mariana Burak, Mykhaylo Svystun
    11/19/25, 5:00โ€ฏPM

    While CEFR-aligned vocabulary profiles have been developed for many languages (e.g., English, German, and Swedish), Ukrainian as a foreign language (UFL) still lacks an empirically grounded lexical profile. A foundational issue in creating such profiles is combining lexical frequency data with expert knowledge to assign CEFR-level labels. Existing UFL word lists rely primarily on professional...

    Go to contribution page
  16. Matej Meterc, Nataลกa Jakop
    11/19/25, 5:30โ€ฏPM

    In preparing phraseological units for the third edition of the Standard Slovenian Dictionary (eSSKJ), the authors aimed to identify the most relevant comparative phrasemes in the contemporary standard language using objective corpus-based criteria. A key goal was to determine how representative specific phrasemes and their variants are in actual use. Two lists of the hundred most frequent...

    Go to contribution page
  17. Ana Frankenberg-Garcia
    11/20/25, 10:05โ€ฏAM

    The use of LLMs in lexicography is a hot topic and indeed the focus of eLex 2025. In the past couple of years, several papers have emerged comparing existing dictionary entries with zero-shot chatbot queries (e.g. Nichols 2023) or with dictionary-like content obtained through the dynamic interaction between experts and chatbots (e.g. Lew 2023, Jakubรญฤek & Rundell 2023). However, studies so far...

    Go to contribution page
  18. Simon Krek, Primoลพ Ponikvar, Andraลพ Repar, Iztok Kosem, David Lindemann
    11/20/25, 11:00โ€ฏAM

    This paper presents an experimental workflow for converting legacy digitized dictionaries into the DMLex standard and subsequently importing them into a Wikibase instance. DMLex, a serialization-independent model developed by the OASIS LEXIDMA Technical Committee, aims to provide a universal and modular representation of lexicographic data. The study tested whether dictionaries from...

    Go to contribution page
  19. Andrej Perdih, Dejan Gabrovลกek, Janoลก Jeลพovnik
    11/20/25, 11:30โ€ฏAM

    This paper evaluates the results of using GPT-4o mini language model batch processing with image recognition capability to align 1,572 images of 398 polysemous nouns in the Dictionary of the Slovenian Standard Language (second edition) to their specific dictionary senses, and it compares them to the results of the manual image-to-sense alignment process. The images were manually assigned to...

    Go to contribution page
  20. Kris Heylen, Vincent Prins, Katrien Depuydt, Jesse de Does, Laura van Eerten, Thomas Haga
    11/20/25, 12:00โ€ฏPM

    Representative monitor corpora with detailed metadata offer a solid empirical basis for documenting lexical innovation and change (Kosem et al. 2021). However, continuously updated time-stamped textual data presents challenges for data management, lexicographic analysis, and visualization. Building on its existing corpus infrastructure, the Dutch Language Institute (INT) has developed...

    Go to contribution page
Building timetable...