Building a multilingual dataset for Romansh and Ladin Large Language Models

Angela Heldstab1, Samuel Frontull2, Ignacio Pérez Prat3, Jannis Vamvas1, Paul Videsott4

  1. University of Zurich, Zurich, Switzerland
  2. University of Innsbruck, Innsbruck, Austria
  3. Lia Rumantscha, Chur, Switzerland
  4. Free University of Bozen-Bolzano, Bolzano, Italy

Background

Large Language Models (LLMs) have gained considerable attention over the past few years, with their capabilities continuously being researched and evaluated. Due to the necessity of immense data being available, LLMs have improved substantially through training on increasingly large and diverse datasets and continue to do so with more data being made available. LLMs are largely able to apply learned rules from one language to another. However, their performance remains limited for many low-resource minority languages due to the scarcity of high-quality training data. This may often force native speakers to switch to dominant and thus often non-native languages if they wish to use an LLM.

Motivation

Previous experiences with native speakers of Rhaeto-Romance varieties indicate that the interest in using a technology in their native language is high. We aim to contribute to solving this problem for two Rhaeto-Romance language continuums (Romansh and Ladin) spoken within Switzerland and South Tyrol (Northern Italy) by creating a high-quality dataset. Creating high-quality datasets is a prerequisite for many NLP applications and thus remains an important research contribution in itself.

Our dataset may then be used to fine-tune LLMs for all six Romansh and three Ladin varieties. This includes Rumantsch Grischun, Sursilvan, Sutsilvan, Surmiran, Puter, Vallader, and Ladin of Val Badia, Gherdëina, and Ladin Dolomitan.

Methods

To create our dataset, we build on previous research: Two recently developed translation systems, ALAS for Romansh and a neural machine translation system for Ladin (Val Badia and Gherdëina), will be used to translate source text from German into the minorized languages. Our source text in German stems from OASST2, a previously built dataset containing human-written prompts, and answers from an LLM. The automatically generated text will then be post-edited by native speakers and professional translators, ensuring a high-quality dataset for all nine varieties. To this end, we will also be making annotation guidelines available to all annotators. The existing machine translation system for Ladin does not currently support Ladin Dolomitan. However, the close syntactic proximity to the Ladin of Val Badia allows for this data to be produced via targeted lexical and orthographic adaptations of the Val Badia translations.

Outlook

The final dataset will contain prompts and LLM-generated answers in German, all prompts and answers machine-translated to all nine varieties, and finally, human-post-edited answers for all nine varieties. The dataset will be made available publicly, which we expect to support the continued development of LLMs. This would further allow native speakers to use those in their own language as well.

From a socio-linguistic standpoint, this has a high added value for the own perception of the language. The dataset enables several downstream applications, especially through its intended usage of fine-tuning LLMs; for example, this may allow native speakers to polish their own texts, contribute to the retention of the languages, or aid learners in their language acquisition. An introduction of a dataset such as this will further allow more foundational research for these minorized languages.

We further intend to continue contributing to the dataset even after its first publication, leading to the possibility of continuous downstream improvement.