Conveners
Parallel sessions 3 (Zrak hall)
- Miloลก Jakubรญฤek
Parallel sessions 3 (Zrak hall)
- Margit Langemets
Parallel sessions 3 (Zrak hall)
- Markus Kunzmann
Parallel sessions 3 (Zrak hall)
- Markus Kunzmann
Parallel sessions 3 (Zrak hall)
- Kris Heylen
Parallel sessions 3 (Zrak hall)
- Geraint Rees
Parallel sessions 3 (Zrak hall)
- Bรกlint Sass
Parallel sessions 3 (Zrak hall)
- Tomasz Michta
-
Iryna Ostapova, Yevhen Kupriianov, Mykyta Yablochkov11/18/25, 11:00โฏAM
The paper outlines technological and methodological ways to arrange the dictionary parsing process. The Spanish Dictionary (Diccionario de la lengua Espaรฑola 23 ed. โ DLE 23) website (https://dle.rae.es/) serves as a basis for the research. First of all, asthe most complex multi-parameter lexicographic frameworks, explanatory dictionaries of national languages are of the most interest because...
Go to contribution page -
Madis Jรผrviste, Tiina Paet11/18/25, 11:30โฏAM
Traditionally, historical textsโ optical character recognition (OCR) has primarily been conducted using specialised software such as Transkribus, eScriptorium, Kraken, and similar tools. To achieve accurate character recognition, these systems require extensive pre-training and the creation of a refined "ground truth" dataset. The comprehensiveness of model pre-training directly correlates...
Go to contribution page -
Dominika Kovarikova11/18/25, 12:00โฏPM
GramatiKat is a freely accessible online application designed to support lexicographic and grammatical work on morphologically rich languages. It provides grammatical profiles, a frequency distribution of lemmas inflected forms, for thousands of Czech nouns, adjectives, and verbs based on large annotated corpora. The concept of grammatical profiling is rooted in the work of Janda and...
Go to contribution page -
Lian Chen11/18/25, 12:30โฏPM
ONLINE PRESENTATION
The creation of ontologiesโtraditionally the domain of linguists and knowledge engineersโis undergoing a significant transformation thanks to advances in artificial intelligence and natural language processing (NLP). These developments open new avenues for phraseology, a field where multi-word expressions (MWEs)โoften opaque and non-compositionalโmust be identified,...
Go to contribution page -
Heete Sahkai, Geda Paulsen, Ene Vainik, Jelena Kallas, Ahto Kiil, Katrin Tsepelina, Kertu Saul, Arvi Tavast11/18/25, 2:30โฏPM
Constructicography, or the description of grammatical constructions in a lexicographic format, is an emerging field currently in the stage of developing and automating methods for treating large numbers of (semi-)schematic constructions. This study explores how existing lexicographic data and language models can be used to facilitate the constructicographic workflow. Our results suggest that...
Go to contribution page -
37. The Mangalam Dictionary of Buddhist Sanskrit: automating lexicographic data with generative LLMsLigeia Lugli11/18/25, 3:00โฏPM
This paper reports on recent advancements in the development of the Mangalam Dictionary of Buddhist Sanskrit, the first corpus-driven dictionary dedicated to Buddhist Sanskrit. This is a low-resource, historical, and domain-specific language variety instantiated in South Asian Buddhist literature dating from approximately the first millennium CE. The paper focusses on advances in the...
Go to contribution page -
Lorna Morris11/18/25, 3:30โฏPM
South Africa is in a literacy crisis, with learners not progressing in school because they are being taught in a second language when they are not functionally literate in their first language. Fewer than 10% of South Africans have English as a home language, but 90% of learners are being taught in English. Many South African schools are under resourced and are not able to give learners the...
Go to contribution page -
Vasyl Starko, Andriy Rysin11/19/25, 10:05โฏAM
Due to the policy of Russification in the 20th century, the Ukrainian language underwent an influx of Russianisms, among other forms of interference with its structure. Today, many Ukrainians require guidance regarding non-Russified usage, and a Large Electronic Dictionary of Ukrainian (VESUM, vesum.nlp.net.ua) is designed to meet this need. With a register of over 430,000 lemmas, it is the...
Go to contribution page -
47. Navigating linguistic diversity: modelling diatopic and bibliographic information with TEI Lex-0Veronika Engler, Karlheinz Mรถrth, Stephan Prochรกzka, Michaela Rausch-Supola, Daniel Schopper11/19/25, 11:00โฏAM
The Vienna Corpus of Arabic Varieties (VICAV) is a digital research infrastructure for the documentation and analysis of the linguistic diversity of Arabic varieties^. Integrating methods from language technology and the digital humanities, VICAV provides a modular, sustainable platform for the creation, management, and publication of heterogeneous language resources within a shared data...
Go to contribution page -
Mojca Kompara Lukanฤiฤ11/19/25, 11:30โฏAM
The article describes the use of artificial intelligence in compiling English dictionary entries for a dictionary of abbreviations (Slovar krajลกav), published in 2025 and financed by the Slovenian Research and Innovation Agency (ARIS). Together with the Slovenian dictionary of abbreviations (Slovenski slovar krajลกav) published in 2023, the mentioned dictionary adopted a pioneering approach to...
Go to contribution page -
Anna Dziemianko, Mojca M. Hoฤevar11/19/25, 2:30โฏPM
ONLINE PRESENTATION
Technology has largely affected the way language learners seek information. Digital formats virtually superseded the paper dictionary (Ptasznik, Wolfer and Lew, 2024), online translators gained much importance (OโNeill, 2019), and web browsers became the first port of call (Kosem et al., 2019). Obviously, generative AI systems imitating human-like communication mark...
Go to contribution page -
Sylwia Wojciechowska11/19/25, 3:00โฏPM
ONLINE PRESENTATION
A major change in dictionary exemplification was brought about by the arrival of corpus data, which replaced lexicographer-made examples with authentic ones from real spoken and written discourse. Monolingual English learnersโ dictionaries (MELDs) prefer a third type of examples, corpus-based ones, with unnecessarily complex vocab and structure, and unclear content...
Go to contribution page -
Stefania Spina, Fabio Zanda, Irene Fioravanti, Luciana Forti, Damiano Perri, Osvaldo Gervasi11/19/25, 3:30โฏPM
ONLINE PRESENTATION
In this presentation we describe the DICI-A (Dizionario delle collocazioni italiane per apprendenti), a new learner dictionary of Italian collocations.
The DICI-A includes ca. 11,000 collocations belonging to six syntactic relations: i. Verb + Direct object (mantenere una promessa, โto keep a promiseโ); ii. Adjective + Noun/Noun + Adjective, where the adjective is a...
Go to contribution page -
Monique Rabรฉ, Martin J. Puttkammer, Gerhard B. van Huyssteen11/19/25, 4:30โฏPM
ONLINE PRESENTATION
Taboo-language resources remain scarce for under-resourced languages like Afrikaans โ despite their clear relevance for natural language processing (NLP) and applications in artificial intelligence (AI). Although Afrikaans has a long-standing lexicographic tradition, it still lacks an open-access reusable lexical database for the taboo language. One of the most crucial...
Go to contribution page -
Esra Abdelzaher, รgoston Tรณth11/19/25, 5:00โฏPM
ONLINE PRESENTATION
Taboo words present a challenge for a lexicographer to include and describe in a language resource, as they are forms of verbal violence. However, discarding offensive words from general-purpose lexicographic wordlists disregards the representation of an integral part of the mental lexicon. The present study aims at using lexicographic scenarios to jailbreak four GPT...
Go to contribution page -
Irina Lobzhanidze, Rusudan Gersamia11/19/25, 5:30โฏPM
ONLINE PRESENTATION
This paper presents a corpus-based approach to compiling a bilingual Megrelian-English online dictionary. The Megrelian language belongs to the UNESCO Atlas of the Worldโs Languages in Danger group of โincreasingly endangeredโ languages, and faces a number of critical challenges, among them a lack of standardised resources, intergenerational transmission, and minimal...
Go to contribution page -
Marรญa Iglesias Vรกzquez, Charlotte Venema, Marie Steffens11/20/25, 10:05โฏAM
This contribution focuses on the methodological aspects of the ICoMuTe project aiming to design a corpus-based multilingual terminology database for Intercultural Communication (ICC). The project seeks to explore how ICC terms relate to each other within six European languages (Dutch, English, German, French, Italian, Spanish), how these terms are connected to their scientific and cultural...
Go to contribution page -
Thomas Widmann11/20/25, 11:00โฏAM
This paper presents a modular pipeline for automated dictionary creation using large language models (LLMs). It addresses the well-known limitations of prompting systems such as ChatGPT to produce entire entries in a single step โ outputs that may read fluently but often lack structural consistency, transparency, originality and verifiability. The proposed system overcomes these weaknesses by...
Go to contribution page -
Nikola Bakariฤ11/20/25, 11:30โฏAM
The task of automatic detection of idiomatic expressions such as proverbs is an established problem in natural language processing. Before the advent of large language models, attempts were made to describe proverbs by modelling their syntactic structure (Rassi et al., 2014). Later, others employed contextual embeddings and neural networks to identify idioms (ล kvorc et al., 2022) which is a...
Go to contribution page