-
11/17/25, 2:00 PM
-
Slobodan Beliga11/17/25, 2:15 PM
-
Polona Gantar11/17/25, 2:40 PM
-
Ivana Filipović Petrović (Croatian Academy of Sciences and Arts)11/17/25, 3:05 PM
-
11/17/25, 4:00 PM
-
11/17/25, 5:00 PM
-
11/17/25, 6:00 PM
-
Marko Robnik-Šikonja11/18/25, 9:30 AM
Currently, large language models (LLMs) are redefining methodological approaches in many scientific areas, including linguistics and lexicography. LLMs are pretrained on huge text corpora by predicting the next tokens and adapted for human interaction with the instruction following datasets. This does not make them immune to hallucinations and biases, requiring a human-in-the-loop approach. In...
Go to contribution page -
Annalisa Greco, Matteo Delsanto, Andrea Di Fabio, Lorenzo Mori, Cristina Onesti, Daniele Paolo Radicioni, Calogero Jerik Scozzaro11/18/25, 11:00 AM
The present research explores the use of large language models (LLMs) in digital lexicography, specifically for translating Italian multiword expressions (MWEs) into English and French.
The study aims to assess the capability of contemporary LLMs in providing accurate and reliable translation equivalents, examples and definitions of Italian MWEs into English and French, while also...
Go to contribution page -
Iryna Ostapova, Yevhen Kupriianov, Mykyta Yablochkov11/18/25, 11:00 AM
The paper outlines technological and methodological ways to arrange the dictionary parsing process. The Spanish Dictionary (Diccionario de la lengua Española 23 ed. – DLE 23) website (https://dle.rae.es/) serves as a basis for the research. First of all, asthe most complex multi-parameter lexicographic frameworks, explanatory dictionaries of national languages are of the most interest because...
Go to contribution page -
Mr Robert Lew, Bartosz Ptasznik11/18/25, 11:00 AM
The public release of ChatGPT in late 2022 made an impact on many professional domains. Notwithstanding the many controversies surrounding Generative Artificial Intelligence (GenAI), such as ethics, copyright, accountability, or ecology, we need to acknowledge an important and relevant feature of Large Language Models and chatbot systems built around them: their ability to produce mostly...
Go to contribution page -
Enikő Héja, László Simon, Veronika Lipp11/18/25, 11:30 AM
Recent findings indicate that current large language models (LLMs) face difficulties in generating clear-cut, well-motivated definitions in a consistent way. This shortcoming is the consequence of their reliance on opaque data sources and their inherently unstable, non-deterministic outputs. In response, this research aims to develop an LLM-based methodology for producing adjectival...
Go to contribution page -
Jelena Kallas, Kristina Koppel, Kris Heylen, Ilan Kernerman, Ana Ostroški Anić, Federica Vezzani, Špela Arhar Holdt11/18/25, 11:30 AM
The COST Action ‘European Network on Lexical Innovation’ (ENEOLI) has conducted a comprehensive survey in October-November 2024 regarding the methods, practices, tools, and resources used in the study and documentation of lexical innovations, including neologisms and novel senses. The 249 respondents from 50 countries represented linguists, lexicographers, terminologists, translators, software...
Go to contribution page -
Madis Jürviste, Tiina Paet11/18/25, 11:30 AM
Traditionally, historical texts’ optical character recognition (OCR) has primarily been conducted using specialised software such as Transkribus, eScriptorium, Kraken, and similar tools. To achieve accurate character recognition, these systems require extensive pre-training and the creation of a refined "ground truth" dataset. The comprehensiveness of model pre-training directly correlates...
Go to contribution page -
Dominik Schlechtweg, Emma Sköldberg, Shafqat Mumtaz Virk, James White, Simon Hengchen11/18/25, 12:00 PM
Finding non-recorded senses is important for dictionary maintenance, where using automatic methods helps reduce manual efforts. We use automatic Word Sense Induction (WSI) to compare recorded sense numbers among a sample of headwords in a comprehensive Swedish monolingual dictionary with induced sense numbers for the same words in a Swedish corpus. We propose this as a simple technique to find...
Go to contribution page -
Dominika Kovarikova11/18/25, 12:00 PM
GramatiKat is a freely accessible online application designed to support lexicographic and grammatical work on morphologically rich languages. It provides grammatical profiles, a frequency distribution of lemmas inflected forms, for thousands of Czech nouns, adjectives, and verbs based on large annotated corpora. The concept of grammatical profiling is rooted in the work of Janda and...
Go to contribution page -
Ondřej Herman11/18/25, 12:30 PM
Language evolves continuously, rendering static dictionaries quickly outdated. While previous research has addressed the automatic detection of new words, identifying subtler semantic changes in existing words remains a challenge. In this work, we propose a robust, language-independent methodology for the automatic detection of word sense shifts using diachronic corpus data. Our approach...
Go to contribution page -
Isidora Despotidou, Zoe Gavriilidou11/18/25, 12:30 PM
The purpose of the presentation is to explore the design and development of an innovative online pedagogical dictionary of Greek Sign Language, specifically tailored to the linguistic and educational needs of Deaf and Hard-of-Hearing (DHH) learners in Greece. Emphasizing accessibility and pedagogical usability, the dictionary integrates Artificial Intelligence (AI) technologies to support...
Go to contribution page -
Lian Chen11/18/25, 12:30 PM
ONLINE PRESENTATION
The creation of ontologies—traditionally the domain of linguists and knowledge engineers—is undergoing a significant transformation thanks to advances in artificial intelligence and natural language processing (NLP). These developments open new avenues for phraseology, a field where multi-word expressions (MWEs)—often opaque and non-compositional—must be identified,...
Go to contribution page -
Geraint Paul Rees11/18/25, 2:30 PM
Online dictionaries have many advantages over their physical counterparts. However, the ephemeral nature of web content means that they are often changed without notice and no ostensible record of what came before remains. This makes research on historical online dictionaries difficult and perhaps explains why, while the history of printed monolingual English learners’ dictionaries (MELDs) has...
Go to contribution page -
Heete Sahkai, Geda Paulsen, Ene Vainik, Jelena Kallas, Ahto Kiil, Katrin Tsepelina, Kertu Saul, Arvi Tavast11/18/25, 2:30 PM
Constructicography, or the description of grammatical constructions in a lexicographic format, is an emerging field currently in the stage of developing and automating methods for treating large numbers of (semi-)schematic constructions. This study explores how existing lexicographic data and language models can be used to facilitate the constructicographic workflow. Our results suggest that...
Go to contribution page -
30. The Dictionary of Contemporary Serbian Language (RSSJ): Advanced Automation and Other ChallengesRanka Stanković, Rada Stijović, Mihailo Škorić, Cvetana Krstev11/18/25, 2:30 PM
This paper introduces the Dictionary of Contemporary Serbian Language (RSSJ), an ongoing large-scale digital lexicographic project designed to serve both human users via web and mobile applications and machines through APIs. Coordinated by the diaspora association “Gathered around the Language” and the Society for Language Resources and Technologies (JeRTeh), RSSJ aims to produce a dictionary...
Go to contribution page -
Nathalie Norman, Sanni Nimb, Sussi Olsen, Nina Schneidermann, Bolette S. Pedersen11/18/25, 3:00 PM
Large Language Models (LLMs) tend to expose severe language and cultural biases when working in medium- and low-resourced languages. In this paper, we present our work on Danish benchmarking and evaluation of LLMs to more precisely diagnose and potentially remedy such bias. To this aim, we apply available lexical-semantic resources to compile a set of Natural Language Understanding (NLU) tasks...
Go to contribution page -
Mojca Stritar Kučuk11/18/25, 3:00 PM
This paper examines how a learner corpus can support lexicographic work by classifying learner vocabulary according to the CEFR scale. Using a corpus-driven methodology, I explore the potential of AI to complement traditional analysis. The study focuses on a selection of texts from the Slovene learner corpus KOST, balanced according to the pragmatically assigned levels of learners’ language...
Go to contribution page -
37. The Mangalam Dictionary of Buddhist Sanskrit: automating lexicographic data with generative LLMsLigeia Lugli11/18/25, 3:00 PM
This paper reports on recent advancements in the development of the Mangalam Dictionary of Buddhist Sanskrit, the first corpus-driven dictionary dedicated to Buddhist Sanskrit. This is a low-resource, historical, and domain-specific language variety instantiated in South Asian Buddhist literature dating from approximately the first millennium CE. The paper focusses on advances in the...
Go to contribution page -
Slobodan Beliga, Ivana Filipović Petrović11/18/25, 3:30 PM
The paper introduces a hybrid methodology for cross-linguistic identification of phraseme constructions, developed within the scope of a pilot study on Croatian repetitive constructions. The study explores how artificial intelligence and corpus technologies can be systematically combined to uncover functionally equivalent patterns across languages. The proposed strategy rests on three...
Go to contribution page -
Maria Tuulik, Ene Vainik, Margit Langemets, Eleri Aedmaa, Lydia Risberg, Esta Prangel, Kristina Koppel, Sirli Zupping11/18/25, 3:30 PM
The use of corpora is well established in lexicography, also in Estonia, but since the analysis of corpus data and the post-editing of automatically generated data from the corpus is labour-intensive, the use of large language models (LLMs) has led to growing interest in lexicography (e.g., Evert et al. 2024; Kosem, Gantar et al. 2024; Tiberius et al. 2024). In 2024, the Institute of the...
Go to contribution page -
Lorna Morris11/18/25, 3:30 PM
South Africa is in a literacy crisis, with learners not progressing in school because they are being taught in a second language when they are not functionally literate in their first language. Fewer than 10% of South Africans have English as a home language, but 90% of learners are being taught in English. Many South African schools are under resourced and are not able to give learners the...
Go to contribution page -
Carole Tiberius, Jesse de Does11/19/25, 9:00 AM
The Dutch Language Institute (INT) has a long tradition compiling historic and contemporary dictionaries and other types of lexicographic databases, mainly for Dutch but also for some other languages with a relation to Dutch. Lexicographic work at the institute is computer-supported but there is still a great deal of manual work involved. Therefore, INT is exploring how new technologies...
Go to contribution page -
Vasyl Starko, Andriy Rysin11/19/25, 10:05 AM
Due to the policy of Russification in the 20th century, the Ukrainian language underwent an influx of Russianisms, among other forms of interference with its structure. Today, many Ukrainians require guidance regarding non-Russified usage, and a Large Electronic Dictionary of Ukrainian (VESUM, vesum.nlp.net.ua) is designed to meet this need. With a register of over 430,000 lemmas, it is the...
Go to contribution page -
Theo J.D. Bothma, Rufus H. Gouws11/19/25, 10:05 AM
The focus of this paper is on Generative Artificial Intelligence (GenAI), chatbots and some implications for lexicography and dictionary use. It has been well documented that chatbots originally tended to “hallucinate” if they did not have an answer to the prompt put to them. Much larger training databases have, however, been developed and chatbots have become more accurate. Multiple...
Go to contribution page -
Polona Gantar, Cyprian Laskowski, Simon Krek11/19/25, 10:05 AM
In lexicography, one of the long-standing issues is understanding the nature of its core element of description commonly referred to as the headword (in DMLex and traditional lexicography), canonical form (in OntoLex and the Lexical Markup Framework – LMF), orthographic form (in the Text Encoding Initiative – TEI Lex0), lemma (in Wikidata), or lexical unit. With the transition from paper to...
Go to contribution page -
Lena De Pourcq, Marie Grégoire, Leonardo Zilio11/19/25, 11:00 AM
This study explores the use of several chatbots based on recent generative large language models for automatic term extraction (ATE) from smaller text samples. The samples were selected from three domains: board games, ice hockey, and kitesurfing; and they cover three languages: English, French, and Portuguese. We used four prompting strategies: zero shot, one shot, few shots, and few shots...
Go to contribution page -
Elinor Hawkes, Phoebe Nicholson, Will Rogers11/19/25, 11:00 AM
This paper presents the Oxford English Dictionary’s (OED) current exploration into the application of artificial intelligence to historical Word Sense Disambiguation (WSD), a fundamental aspect of OED’s core research. Building on a longstanding tradition of technological innovation, the OED is investigating how Large Language Models (LLMs) can support the identification and retrieval of...
Go to contribution page -
47. Navigating linguistic diversity: modelling diatopic and bibliographic information with TEI Lex-0Veronika Engler, Karlheinz Mörth, Stephan Procházka, Michaela Rausch-Supola, Daniel Schopper11/19/25, 11:00 AM
The Vienna Corpus of Arabic Varieties (VICAV) is a digital research infrastructure for the documentation and analysis of the linguistic diversity of Arabic varieties^. Integrating methods from language technology and the digital humanities, VICAV provides a modular, sustainable platform for the creation, management, and publication of heterogeneous language resources within a shared data...
Go to contribution page -
Mojca Kompara Lukančič11/19/25, 11:30 AM
The article describes the use of artificial intelligence in compiling English dictionary entries for a dictionary of abbreviations (Slovar krajšav), published in 2025 and financed by the Slovenian Research and Innovation Agency (ARIS). Together with the Slovenian dictionary of abbreviations (Slovenski slovar krajšav) published in 2023, the mentioned dictionary adopted a pioneering approach to...
Go to contribution page -
Iván Arias-Arias, Elena Martín-Cancela11/19/25, 11:30 AM
Generic nouns such as Sache and Ding pose a challenge for semantic annotation due to their referential underspecification and context-dependent meaning. Although frequently classified under categories like {artefact} or {object}, their actual referents often belong to abstract or cognitive domains, as in Der Placeboeffekt ist eines der faszinierendsten Dinge in der Welt der Medizin. Drawing on...
Go to contribution page -
Michael Rundell, Miloš Jakubíček, Vojtěch Kovář, Ondřej Matuška, Michal Cukr11/19/25, 11:30 AM
In this paper we show how the academic content and computational tools featured in Lexicom form a parallel history of the last 25 years of innovation in lexicography. Lexicom is a 5-day intensive workshop offering handson training in corpus-based dictionary creation, from collecting and annotating language data to publishing the final product. Since it was launched in 2001, by Sue Atkins, Adam...
Go to contribution page -
Gregor Middell11/19/25, 12:00 PM
POSTER
Eric Raymond’s influential essay (Raymond 1999) about the community-based software development as practiced in the Open Source movement vs. the previously dominant, closed, top-down approach mostly preferred in the commercial realm proved also instructive for the Wikiverse. Its flagship project Wikipedia with a comparable approach to knowledge production and dissemination disrupted...
Go to contribution page -
Nathalie Norman, Nicolai Hartvig Sørensen, Jonas Jensen, Kirsten Appel, Sanni Nimb11/19/25, 12:00 PM
POSTER
Writing dictionary entries is not only time-consuming but also an expensive process due to the highly specialized knowledge and experience required of the lexicographer. To facilitate the task of compiling the Danish monolingual dictionary DDO (ordnet.dk/ddo), we aim to establish an automatic assistant based on applied language technology (e.g. n-gram analysis, word embeddings, etc.)...
Go to contribution page -
Špela Arhar Holdt, Iztok Kosem11/19/25, 12:00 PM
DEMO
CJVT igre (https://igre.cjvt.si/) is a new digital platform offering word games designed to foster lexical awareness and engagement with standard Slovene. Developed by the Centre for Language Resources and Technologies at the University of Ljubljana, the portal currently hosts three games—Cvetka, Besedolov, and Vezalka—with two more in development. Each game utilizes curated lexical...
Go to contribution page -
Martina Pavić, Daša Farkaš11/19/25, 12:00 PM
POSTER
The representation of medical adjectives in Croatian general dictionaries reveals significant inconsistencies, reflected in uneven lemma inclusion, ambigous or absent domain labels, and limited definitional precision. This paper analyzes the 80 most frequent adjectives, based on corpus data from the Croatian Medical Corpus (CMC) (Kocijan, Kurolt & Mijić, 2020), in the three major...
Go to contribution page -
Mititelu Catalin11/19/25, 12:00 PM
POSTER
Recently, the digitization of resources of any type has become an increasingly discussed topic. In the linguistic field, lexicography is among the most influenced by this process, with digital dictionaries playing an essential role both for online consultation by specialists and for the automatic development of useful resources in natural language processing, as well as downstream...
Go to contribution page -
Krešimir Šojat, Kristina Kocijan11/19/25, 12:00 PM
POSTER
This paper presents a novel approach to exploring derivational families within the framework of Intelligent Lexicography, using the ŠKOLARAC corpus: a collection of Croatian school essays written by L1 learners (native-speaking students) in grades 5 through 8 and enriched with metadata such as gender, grade level, and region. By combining rule-based linguistic processing in NooJ, a...
Go to contribution page -
Bálint Sass, Éva Dömötör, Balázs Indig, Mátyás Lagos Cortes, Veronika Lipp, Márton Makrai, Gergely Pethő11/19/25, 12:00 PM
POSTER
Taking seriously the common construction grammar statement that “it’s constructions all the way down” (Goldberg, 2006: 18), the Hungarian Constructicon aims to encompass the widest possible range of constructions. As it is a dictionary-based constructicon, it naturally contains what a dictionary can provide — from morphemes to words, and to partially schematic multiword constructions...
Go to contribution page -
Laura Rebosio11/19/25, 12:00 PM
POSTER
This paper explores the differences between the Phrase-based Active Dictionary (PAD) and FrameNet in their approaches to meaning representation, focusing on the verbs agree and follow. The PAD, a component of the PhraseBase project, adopts a splitting-friendly methodology that emphasizes granularity and ontological consistency, ensuring a more comprehensive coverage of polysemy. In...
Go to contribution page -
Jasmina Jelčić Čolakovac11/19/25, 12:00 PM
POSTER
The lack of normative resources for the Croatian language has incited the development of a novel resource that would not only compile normative data for Croatian but also focus on an underrepresented group of linguistic units – figurative multi-word (MWE) expressions. Thus, the creation of a normative database for figurative MWEs in Croatian is a significant step in the right...
Go to contribution page -
Ligeia Lugli11/19/25, 12:00 PM
DEMO
This demo introduces lexicographR (citation withheld for anonymization), a prototype computer application aimed at facilitating the creation of digital dictionaries for scholars working in low-tech environments, where access to programming skills is severely hindered by lack of funding, institutional support and technical training. Based on recent user-surveys (Lugli 2024b), these...
Go to contribution page -
Sarah Piepkorn11/19/25, 12:00 PM
POSTER
A system of lexicographic presentational devices for data on verbal aspect has been developed that is aimed at providing advanced foreign language learners of English, German or Italian with data for individual verbs and their different readings. It is part of a monolingual, production-oriented electronic dictionary, the Phrase-based Active Dictionary (DiMuccio-Failla, 2025;...
Go to contribution page -
Mykyta Yablochkov, Alona Dorozhynska, Iryna Ostapova, Iuliia Verbynenko11/19/25, 12:00 PM
POSTER
The objective of the research is to develop a technology for converting specialized dictionary text into a website with a developed user interface.
The object of the study was “Dictionary of Ukrainian biological terminology” (7,342 entries and about 26,000 terms in Ukrainian, Russian and English), that contains definitions, terms polysemy, synonymy, stresses for Slavic languages,...
Go to contribution page -
Ana Ostroški Anić, Jaka Čibej, Ivana Filipović Petrović, Martina Pavić, Siniša Runjaić, Robert Sviben11/19/25, 12:00 PM
POSTER
As part of the COST Action CA21167 Universality, Diversity and Idiosyncrasy in Language Technology (UniDive), the ELEXIS-WSD Parallel Sense-Annotated Corpus (Martelli et al., 2021; Čibej et al., 2025) is being expanded to include subcorpora in additional languages—among them, Croatian—as well as new annotation layers. Each language subcorpus of ELEXIS-WSD contains the same 2,024...
Go to contribution page -
Verginica Barbu Mititelu, Voula Giouli, Gražina Korvel, Chaya Liebeskind, Irina Lobzhanidze, Rusudan Makhachashvili, Stella Markantonatou, Alexandra Markovic, Ivelina Stoyanova11/19/25, 12:00 PM
POSTER
In this paper, we provide a comprehensive overview of the way in which the morpho-syntactic properties of multiword expressions are represented in lexical resources to support Natural Language Processing downstream applications. Starting from an up-to-date and comprehensive overview of the existing lexica dedicated to multiword expressions and containing their syntactic description,...
Go to contribution page -
Tanara Zingano Kuhn11/19/25, 12:00 PM
POSTER
The dictionary of pluricentric Portuguese project, which is at its initial stage at (University of Coimbra), aims at providing a free, online dictionary that describes Portuguese as it is used in several territories around the globe. The purpose of this poster is to present theoretical questions that need to be answered to guide the methodological decisions for the creation of this...
Go to contribution page -
Dace Šostaka, Inguna Skadiņa11/19/25, 12:00 PM
POSTER
The Information and Communication Technologies (ICT) field has evolved rapidly in recent decades. Thus, to describe new devices, activities, and concepts that appear yearly, a vast number of terms are created primarily in English, while other languages rely on secondary term formation (STF) for ICT end-users (ETSI Guide, 2022). Systematic secondary rendering and dissemination...
Go to contribution page -
Jaka Čibej11/19/25, 12:00 PM
POSTER
Lexicons of taboo language are useful language resources that can serve multiple purposes. In addition to their direct use to either automatically censor words deemed inappropriate for a given context (e.g. to help mitigate the problem of online hate speech), they can also help filter out materials not suitable for educational purposes (see Zingano Kuhn et al., 2022), games with a...
Go to contribution page -
Luise Köhler, Gregor Middell, Alexander Geyken11/19/25, 2:30 PM
Collocations are a well-covered research area in lexicography. With the advent of evidence-based lexicography and the availability of large text corpora, computational methods of extracting typical co-occurrences from such corpora and supporting lexicographers in identifying collocations among them became a research focus. Especially the statistical properties of collocations (i.e. application...
Go to contribution page -
Anna Dziemianko, Mojca M. Hočevar11/19/25, 2:30 PM
ONLINE PRESENTATION
Technology has largely affected the way language learners seek information. Digital formats virtually superseded the paper dictionary (Ptasznik, Wolfer and Lew, 2024), online translators gained much importance (O’Neill, 2019), and web browsers became the first port of call (Kosem et al., 2019). Obviously, generative AI systems imitating human-like communication mark...
Go to contribution page -
Iztok Kosem, Špela Arhar Holdt11/19/25, 2:30 PM
This paper presents two tasks involving large language models (LLMs)—Gemini-2.0-flash and GPT-4o—used to generate distractors (i.e., incorrect options) for synonym and collocation questions in a language game. The lexical data for both tasks was sourced from the Digital Dictionary Database of Slovene (DDDS). Prompts were initially tested on a sample dataset with both models, and the...
Go to contribution page -
Markus Kunzmann11/19/25, 3:00 PM
The project Dictionary of Bavarian Dialects in Austria "Wörterbuch der bairischen Mundarten in Österreich"(WBÖ) project maintains an archive of approximately 3.6 million handwritten dialectal paper slips documenting dialectal evidence. While 2.4 million entries have been manually digitized and converted to TEI format, the remaining 1.2 million paper slips from sections A-C require automated...
Go to contribution page -
Sylwia Wojciechowska11/19/25, 3:00 PM
ONLINE PRESENTATION
A major change in dictionary exemplification was brought about by the arrival of corpus data, which replaced lexicographer-made examples with authentic ones from real spoken and written discourse. Monolingual English learners’ dictionaries (MELDs) prefer a third type of examples, corpus-based ones, with unnecessarily complex vocab and structure, and unclear content...
Go to contribution page -
Lydia Risberg, Eleri Aedmaa, Maria Tuulik, Margit Langemets, Ene Vainik, Esta Prangel, Kristina Koppel, Hanna Pook11/19/25, 3:00 PM
Language corpora have long been used in linguistics and lexicography, but recent developments now allow large language models (LLMs) to support or even transform these fields. This study investigates the potential of LLMs for annotating informal language use in Estonian – a language underrepresented in LLM training data yet supported by a large corpus. Focusing on the informal register label...
Go to contribution page -
Tomasz Michta, Ana Frankenberg-Garcia11/19/25, 3:30 PM
Studies comparing dictionary entries generated with AI with those of well-established dictionaries edited by lexicographers show that LLMs tend to perform better in some tasks (e.g. writing definitions) than in others (e.g. word-sense disambiguation (e.g. Nichols 2023, Lew 2023, Jakubíček & Rundell 2023, Rees & Lew 2024). One of the problems resulting from the latter is that of “false...
Go to contribution page -
Aleksandra Marković, Ranka Stanković11/19/25, 3:30 PM
Automation has revolutionised lexicography, introducing the ’post-editing lexicography’ model, where the role of the lexicographer involves refining automatically generated dictionary drafts. Since the launch of ChatGPT in November 2022, numerous papers have explored the potential applications of LLMs in dictionary production. The rapid evolution of LLMs necessitates a re-evaluation of...
Go to contribution page -
Stefania Spina, Fabio Zanda, Irene Fioravanti, Luciana Forti, Damiano Perri, Osvaldo Gervasi11/19/25, 3:30 PM
ONLINE PRESENTATION
In this presentation we describe the DICI-A (Dizionario delle collocazioni italiane per apprendenti), a new learner dictionary of Italian collocations.
The DICI-A includes ca. 11,000 collocations belonging to six syntactic relations: i. Verb + Direct object (mantenere una promessa, ‘to keep a promise’); ii. Adjective + Noun/Noun + Adjective, where the adjective is a...
Go to contribution page -
Monique Rabé, Martin J. Puttkammer, Gerhard B. van Huyssteen11/19/25, 4:30 PM
ONLINE PRESENTATION
Taboo-language resources remain scarce for under-resourced languages like Afrikaans – despite their clear relevance for natural language processing (NLP) and applications in artificial intelligence (AI). Although Afrikaans has a long-standing lexicographic tradition, it still lacks an open-access reusable lexical database for the taboo language. One of the most crucial...
Go to contribution page -
Antonio San Martín11/19/25, 4:30 PM
This paper presents contextonym analysis as a hybrid method combining corpus-based techniques and generative artificial (GenAI) tools to support the writing of precise, context-sensitive terminological definitions. Grounded in the Flexible Terminological Definition Approach, this method is based on the premise that definitions should reflect the most relevant conceptual content activated in...
Go to contribution page -
Urška Vranjek Ošlak11/19/25, 4:30 PM
This paper explores the applicability of generative artificial intelligence in the field of language consulting, focusing on ChatGPT-4 and the Slovenian language. The analysis is based on an experiment involving 30 real user questions submitted to the Language Consulting Service (LCS) of the Fran Ramovš Institute of the Slovenian Language. The questions cover a range of linguistic categories...
Go to contribution page -
Olena Synchak, Vasyl Starko, Mariana Burak, Mykhaylo Svystun11/19/25, 5:00 PM
While CEFR-aligned vocabulary profiles have been developed for many languages (e.g., English, German, and Swedish), Ukrainian as a foreign language (UFL) still lacks an empirically grounded lexical profile. A foundational issue in creating such profiles is combining lexical frequency data with expert knowledge to assign CEFR-level labels. Existing UFL word lists rely primarily on professional...
Go to contribution page -
Jesús Torres del Rey, María García Garmendia11/19/25, 5:00 PM
While the move to the digital design of lexical resources has, in principle, enhanced the physical and sensory accessibility of dictionaries, a lack of adherence to accessibility standards such as WCAG 2 (Web Content Accessibility Guidelines) (Campbell et all 2023) can introduce significant barriers (NCD 2006; Botelho 2021). These barriers often hinder access to the information and...
Go to contribution page -
Esra Abdelzaher, Ágoston Tóth11/19/25, 5:00 PM
ONLINE PRESENTATION
Taboo words present a challenge for a lexicographer to include and describe in a language resource, as they are forms of verbal violence. However, discarding offensive words from general-purpose lexicographic wordlists disregards the representation of an integral part of the mental lexicon. The present study aims at using lexicographic scenarios to jailbreak four GPT...
Go to contribution page -
Irina Lobzhanidze, Rusudan Gersamia11/19/25, 5:30 PM
ONLINE PRESENTATION
This paper presents a corpus-based approach to compiling a bilingual Megrelian-English online dictionary. The Megrelian language belongs to the UNESCO Atlas of the World’s Languages in Danger group of “increasingly endangered” languages, and faces a number of critical challenges, among them a lack of standardised resources, intergenerational transmission, and minimal...
Go to contribution page -
Ondřej Herman, Miloš Jakubíček, Jan Kraus, Vít Suchomel11/19/25, 5:30 PM
This paper presents a long-term privately-funded programme focusing on collecting of timestamped monitor corpora in a wide range of (currently 25) languages. These corpora are primarily designed for researching linguistic trends (including neology) and language change over time. They are available through the Sketch Engine platform and vary significantly in size — from 3 million tokens for...
Go to contribution page -
Matej Meterc, Nataša Jakop11/19/25, 5:30 PM
In preparing phraseological units for the third edition of the Standard Slovenian Dictionary (eSSKJ), the authors aimed to identify the most relevant comparative phrasemes in the contemporary standard language using objective corpus-based criteria. A key goal was to determine how representative specific phrasemes and their variants are in actual use. Two lists of the hundred most frequent...
Go to contribution page -
Michal Měchura11/20/25, 9:00 AM
It has been almost half a century since we started “doing” lexicography on computers. Let’s stop for a minute now and take a critical look at the data models we have been using to represent the structure of dictionaries in dictionary writing systems and other software.
In this talk, I will trace the history of lexicographic data modelling from its beginnings as text markup for...
Go to contribution page -
Ana Frankenberg-Garcia11/20/25, 10:05 AM
The use of LLMs in lexicography is a hot topic and indeed the focus of eLex 2025. In the past couple of years, several papers have emerged comparing existing dictionary entries with zero-shot chatbot queries (e.g. Nichols 2023) or with dictionary-like content obtained through the dynamic interaction between experts and chatbots (e.g. Lew 2023, Jakubíček & Rundell 2023). However, studies so far...
Go to contribution page -
Philipp Stöckle, Daniel Elsner, Wolfgang Koppensteiner, Katharina Korecky-Kröll11/20/25, 10:05 AM
This paper investigates the potential of LLMs in supporting lexicographic work on non-standard linguistic varieties using data from the Dictionary of Bavarian Dialects in Austria (WBÖ). Based on approx. 2.4 million digitized and TEI-encoded dialect paper slips published via the Lexical Information System Austria (LIÖ), we construct a domain-specific corpus and evaluate LLMs in semantic...
Go to contribution page -
María Iglesias Vázquez, Charlotte Venema, Marie Steffens11/20/25, 10:05 AM
This contribution focuses on the methodological aspects of the ICoMuTe project aiming to design a corpus-based multilingual terminology database for Intercultural Communication (ICC). The project seeks to explore how ICC terms relate to each other within six European languages (Dutch, English, German, French, Italian, Spanish), how these terms are connected to their scientific and cultural...
Go to contribution page -
Thomas Widmann11/20/25, 11:00 AM
This paper presents a modular pipeline for automated dictionary creation using large language models (LLMs). It addresses the well-known limitations of prompting systems such as ChatGPT to produce entire entries in a single step – outputs that may read fluently but often lack structural consistency, transparency, originality and verifiability. The proposed system overcomes these weaknesses by...
Go to contribution page -
Simon Krek, Primož Ponikvar, Andraž Repar, Iztok Kosem, David Lindemann11/20/25, 11:00 AM
This paper presents an experimental workflow for converting legacy digitized dictionaries into the DMLex standard and subsequently importing them into a Wikibase instance. DMLex, a serialization-independent model developed by the OASIS LEXIDMA Technical Committee, aims to provide a universal and modular representation of lexicographic data. The study tested whether dictionaries from...
Go to contribution page -
Loryn Isaacs, Santiago Chambó, Pilar León-Araúz11/20/25, 11:00 AM
Corpus-based conceptual analysis for the Humanitarian Encyclopedia (HE) grapples with vast amounts of lexical data to describe the meaning of key humanitarian notions and detect conceptual variation among actors (Odlum & Chambó, 2022). By building on Frame-based Terminology (Faber, 2015, 2022), the HE is incorporating qualitative methods necessary to subsume lexical data into manageable...
Go to contribution page -
Nikola Bakarić11/20/25, 11:30 AM
The task of automatic detection of idiomatic expressions such as proverbs is an established problem in natural language processing. Before the advent of large language models, attempts were made to describe proverbs by modelling their syntactic structure (Rassi et al., 2014). Later, others employed contextual embeddings and neural networks to identify idioms (Škvorc et al., 2022) which is a...
Go to contribution page -
Andrej Perdih, Dejan Gabrovšek, Janoš Ježovnik11/20/25, 11:30 AM
This paper evaluates the results of using GPT-4o mini language model batch processing with image recognition capability to align 1,572 images of 398 polysemous nouns in the Dictionary of the Slovenian Standard Language (second edition) to their specific dictionary senses, and it compares them to the results of the manual image-to-sense alignment process. The images were manually assigned to...
Go to contribution page -
Marek Blahuš, Miloš Jakubíček, Vojtěch Kovář, František Kovařík11/20/25, 11:30 AM
This paper explores the theory of measuring vocabulary size, including the various methods that can be used and the parameters that have to be set. We have examined the experiments carried out on English and Dutch. Gouldenet al. (1990) claims the average native speaker knows about 17,000 English base words (non-derived words). Keuleers et al. (2015) and Brysbaert et al. (2016) claim the...
Go to contribution page -
Marek Blahuš, Ota Mikušek11/20/25, 12:00 PM
We present a collection of monolingual text corpora derived from the steno protocols of 30 parliamentary chambers across 22 EU member states, covering 20 languages. The corpora are continuously and automatically updated, enabling intralingual and cross-lingual analysis of parliamentary discussions. Each chamber’s protocols are regularly downloaded, processed, and transformed into a unified...
Go to contribution page -
Kris Heylen, Vincent Prins, Katrien Depuydt, Jesse de Does, Laura van Eerten, Thomas Haga11/20/25, 12:00 PM
Representative monitor corpora with detailed metadata offer a solid empirical basis for documenting lexical innovation and change (Kosem et al. 2021). However, continuously updated time-stamped textual data presents challenges for data management, lexicographic analysis, and visualization. Building on its existing corpus infrastructure, the Dutch Language Institute (INT) has developed...
Go to contribution page -
11/20/25, 2:00 PM
-
Nikos Mathioudakis11/20/25, 2:10 PM
-
Jinhong Huang11/20/25, 2:35 PM
-
Jiang Li, Wang Yi11/20/25, 3:00 PM
-
Alexander Geyken11/20/25, 3:25 PM
-
24. ENEOLI Wikibase: A collaborative working platform for the European Network on Lexical InnovationAna Salgado, David Lindemann11/20/25, 4:20 PM
-
Marius Glebus11/20/25, 4:45 PM
-
Cécile Poix, Natalya Shevchenko11/20/25, 5:10 PM
-
11/20/25, 5:35 PM
-
Matej Meterc, Nataša Jakop
In preparing phraseological units for the third edition of the Standard Slovenian Dictionary (eSSKJ), the authors aimed to identify the most relevant comparative phrasemes in the contemporary standard language using objective corpus-based criteria. A key goal was to determine how representative specific phrasemes and their variants are in actual use. Two lists of the hundred most frequent...
Go to contribution page
Choose timezone
Your profile timezone: