Conveners
Parallel sessions 2 (Sonce hall)
- Mojca Stritar Kuฤuk
Parallel sessions 2 (Sonce hall)
- Valeria Caruso
Parallel sessions 2 (Sonce hall)
- Philipp Stรถckle
Parallel sessions 2 (Sonce hall)
- Philipp Stรถckle
Parallel sessions 2 (Sonce hall)
- Slobodan Beliga
Parallel sessions 2 (Sonce hall)
- David Lindemann
Parallel sessions 2 (Sonce hall)
- Janoลก Jeลพovnik
Parallel sessions 2 (Sonce hall)
- Ana Frankenberg
-
Annalisa Greco, Matteo Delsanto, Andrea Di Fabio, Lorenzo Mori, Cristina Onesti, Daniele Paolo Radicioni, Calogero Jerik Scozzaro11/18/25, 11:00โฏAM
The present research explores the use of large language models (LLMs) in digital lexicography, specifically for translating Italian multiword expressions (MWEs) into English and French.
The study aims to assess the capability of contemporary LLMs in providing accurate and reliable translation equivalents, examples and definitions of Italian MWEs into English and French, while also...
Go to contribution page -
Jelena Kallas, Kristina Koppel, Kris Heylen, Ilan Kernerman, Ana Ostroลกki Aniฤ, Federica Vezzani, ล pela Arhar Holdt11/18/25, 11:30โฏAM
The COST Action โEuropean Network on Lexical Innovationโ (ENEOLI) has conducted a comprehensive survey in October-November 2024 regarding the methods, practices, tools, and resources used in the study and documentation of lexical innovations, including neologisms and novel senses. The 249 respondents from 50 countries represented linguists, lexicographers, terminologists, translators, software...
Go to contribution page -
Isidora Despotidou, Zoe Gavriilidou11/18/25, 12:30โฏPM
The purpose of the presentation is to explore the design and development of an innovative online pedagogical dictionary of Greek Sign Language, specifically tailored to the linguistic and educational needs of Deaf and Hard-of-Hearing (DHH) learners in Greece. Emphasizing accessibility and pedagogical usability, the dictionary integrates Artificial Intelligence (AI) technologies to support...
Go to contribution page -
Geraint Paul Rees11/18/25, 2:30โฏPM
Online dictionaries have many advantages over their physical counterparts. However, the ephemeral nature of web content means that they are often changed without notice and no ostensible record of what came before remains. This makes research on historical online dictionaries difficult and perhaps explains why, while the history of printed monolingual English learnersโ dictionaries (MELDs) has...
Go to contribution page -
Mojca Stritar Kuฤuk11/18/25, 3:00โฏPM
This paper examines how a learner corpus can support lexicographic work by classifying learner vocabulary according to the CEFR scale. Using a corpus-driven methodology, I explore the potential of AI to complement traditional analysis. The study focuses on a selection of texts from the Slovene learner corpus KOST, balanced according to the pragmatically assigned levels of learnersโ language...
Go to contribution page -
Slobodan Beliga, Ivana Filipoviฤ Petroviฤ11/18/25, 3:30โฏPM
The paper introduces a hybrid methodology for cross-linguistic identification of phraseme constructions, developed within the scope of a pilot study on Croatian repetitive constructions. The study explores how artificial intelligence and corpus technologies can be systematically combined to uncover functionally equivalent patterns across languages. The proposed strategy rests on three...
Go to contribution page -
Theo J.D. Bothma, Rufus H. Gouws11/19/25, 10:05โฏAM
The focus of this paper is on Generative Artificial Intelligence (GenAI), chatbots and some implications for lexicography and dictionary use. It has been well documented that chatbots originally tended to โhallucinateโ if they did not have an answer to the prompt put to them. Much larger training databases have, however, been developed and chatbots have become more accurate. Multiple...
Go to contribution page -
Lena De Pourcq, Marie Grรฉgoire, Leonardo Zilio11/19/25, 11:00โฏAM
This study explores the use of several chatbots based on recent generative large language models for automatic term extraction (ATE) from smaller text samples. The samples were selected from three domains: board games, ice hockey, and kitesurfing; and they cover three languages: English, French, and Portuguese. We used four prompting strategies: zero shot, one shot, few shots, and few shots...
Go to contribution page -
Michael Rundell, Miloลก Jakubรญฤek, Vojtฤch Kovรกล, Ondลej Matuลกka, Michal Cukr11/19/25, 11:30โฏAM
In this paper we show how the academic content and computational tools featured in Lexicom form a parallel history of the last 25 years of innovation in lexicography. Lexicom is a 5-day intensive workshop offering handson training in corpus-based dictionary creation, from collecting and annotating language data to publishing the final product. Since it was launched in 2001, by Sue Atkins, Adam...
Go to contribution page -
Iztok Kosem, ล pela Arhar Holdt11/19/25, 2:30โฏPM
This paper presents two tasks involving large language models (LLMs)โGemini-2.0-flash and GPT-4oโused to generate distractors (i.e., incorrect options) for synonym and collocation questions in a language game. The lexical data for both tasks was sourced from the Digital Dictionary Database of Slovene (DDDS). Prompts were initially tested on a sample dataset with both models, and the...
Go to contribution page -
Markus Kunzmann11/19/25, 3:00โฏPM
The project Dictionary of Bavarian Dialects in Austria "Wรถrterbuch der bairischen Mundarten in รsterreich"(WBร) project maintains an archive of approximately 3.6 million handwritten dialectal paper slips documenting dialectal evidence. While 2.4 million entries have been manually digitized and converted to TEI format, the remaining 1.2 million paper slips from sections A-C require automated...
Go to contribution page -
Tomasz Michta, Ana Frankenberg-Garcia11/19/25, 3:30โฏPM
Studies comparing dictionary entries generated with AI with those of well-established dictionaries edited by lexicographers show that LLMs tend to perform better in some tasks (e.g. writing definitions) than in others (e.g. word-sense disambiguation (e.g. Nichols 2023, Lew 2023, Jakubรญฤek & Rundell 2023, Rees & Lew 2024). One of the problems resulting from the latter is that of โfalse...
Go to contribution page -
Antonio San Martรญn11/19/25, 4:30โฏPM
This paper presents contextonym analysis as a hybrid method combining corpus-based techniques and generative artificial (GenAI) tools to support the writing of precise, context-sensitive terminological definitions. Grounded in the Flexible Terminological Definition Approach, this method is based on the premise that definitions should reflect the most relevant conceptual content activated in...
Go to contribution page -
Jesรบs Torres del Rey, Marรญa Garcรญa Garmendia11/19/25, 5:00โฏPM
While the move to the digital design of lexical resources has, in principle, enhanced the physical and sensory accessibility of dictionaries, a lack of adherence to accessibility standards such as WCAG 2 (Web Content Accessibility Guidelines) (Campbell et all 2023) can introduce significant barriers (NCD 2006; Botelho 2021). These barriers often hinder access to the information and...
Go to contribution page -
Ondลej Herman, Miloลก Jakubรญฤek, Jan Kraus, Vรญt Suchomel11/19/25, 5:30โฏPM
This paper presents a long-term privately-funded programme focusing on collecting of timestamped monitor corpora in a wide range of (currently 25) languages. These corpora are primarily designed for researching linguistic trends (including neology) and language change over time. They are available through the Sketch Engine platform and vary significantly in size โ from 3 million tokens for...
Go to contribution page -
Philipp Stรถckle, Daniel Elsner, Wolfgang Koppensteiner, Katharina Korecky-Krรถll11/20/25, 10:05โฏAM
This paper investigates the potential of LLMs in supporting lexicographic work on non-standard linguistic varieties using data from the Dictionary of Bavarian Dialects in Austria (WBร). Based on approx. 2.4 million digitized and TEI-encoded dialect paper slips published via the Lexical Information System Austria (LIร), we construct a domain-specific corpus and evaluate LLMs in semantic...
Go to contribution page -
Loryn Isaacs, Santiago Chambรณ, Pilar Leรณn-Araรบz11/20/25, 11:00โฏAM
Corpus-based conceptual analysis for the Humanitarian Encyclopedia (HE) grapples with vast amounts of lexical data to describe the meaning of key humanitarian notions and detect conceptual variation among actors (Odlum & Chambรณ, 2022). By building on Frame-based Terminology (Faber, 2015, 2022), the HE is incorporating qualitative methods necessary to subsume lexical data into manageable...
Go to contribution page -
Marek Blahuลก, Miloลก Jakubรญฤek, Vojtฤch Kovรกล, Frantiลกek Kovaลรญk11/20/25, 11:30โฏAM
This paper explores the theory of measuring vocabulary size, including the various methods that can be used and the parameters that have to be set. We have examined the experiments carried out on English and Dutch. Gouldenet al. (1990) claims the average native speaker knows about 17,000 English base words (non-derived words). Keuleers et al. (2015) and Brysbaert et al. (2016) claim the...
Go to contribution page -
Marek Blahuลก, Ota Mikuลกek11/20/25, 12:00โฏPM
We present a collection of monolingual text corpora derived from the steno protocols of 30 parliamentary chambers across 22 EU member states, covering 20 languages. The corpora are continuously and automatically updated, enabling intralingual and cross-lingual analysis of parliamentary discussions. Each chamberโs protocols are regularly downloaded, processed, and transformed into a unified...
Go to contribution page -
Matej Meterc, Nataลกa Jakop
In preparing phraseological units for the third edition of the Standard Slovenian Dictionary (eSSKJ), the authors aimed to identify the most relevant comparative phrasemes in the contemporary standard language using objective corpus-based criteria. A key goal was to determine how representative specific phrasemes and their variants are in actual use. Two lists of the hundred most frequent...
Go to contribution page