Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Resource-Lean Lexicon Induction for German Dialects

This repository contains the code, data and evaluation scripts to reproduce the results of the paper Resource-Lean Lexicon Induction for German Dialects, which has been presented at LREC 2026.

Overview

  1. Dialect Variation Dictionaries
  2. Reproduce Results
    1. DiaLemma BLI (Table 2, Figure 1)
    2. WikiDIR BLI (Tables 3-5)
    3. Cross-Dialect IR (Table 6)
  3. Citation

📖 Dialect Variation Dictionaries

We provide automatically generated dictionaries for five German dialects, which we evaluated extrinsically on the task of cross-dialect retrieval (query expansion). The files can be downloaded from the folder data/multilemma/.

Dialect Lemmas Variants V/L File
als 38,129 88,114 2.31 als_dictionary.jsonl
bar 27,974 51,392 1.86 bar_dictionary.jsonl
ksh 6,889 9,384 1.36 ksh_dictionary.jsonl
pfl 9,127 13,050 1.43 pfl_dictionary.jsonl
nds 21,974 39,547 1.80 nds_dictionary.jsonl

The dictionaries above have been automatically induced following the DiaLemma annotation framework . We trained a classifier (annotation model) on human-annotated Bavarian word pairs and classified (annotated) for German reference terms their ten lexical nearest neighbors.

Example (als):

{
  "term": "Ortschaft",
  "variants": [
    "ortschaft",
    "Ortschàft",
    "Oortschaft"
  ],
  "inflected_variants": [
    "Ortschaftä",
    "Ortschofte",
    "Ortschafta",
    "Ortschaftè"
  ],
  "unrelated": [
    "Wûrtschaft",
    "Wértschaft",
    "Wrtschaft"
  ]
}

Format:

  • "term": German reference term.
  • "variants" / "inflected_variants": Dialect terms that were classified as (inflected) translations.
  • "unrelated": Dialect terms that were classified as being unrelated to the reference term.

Reproduce Results

Create a python environment and install the required packages:

conda create --name multilemma python=3.10
conda activate multilemma
pip install -r requirements.txt
chmod +x scripts/*

Run scripts/reproduce.sh to download input files and generate result files:

  • Input files can be found in data/.
  • Result files can be found in results/.

Below are the steps to reproduce specific results. Note: Scripts need to be run in the project root folder.

🤔 DiaLemma BLI

  • Download DiaLemma files and create splits: scripts/setup_dialemma.sh
  • Reproduce results in Table 2: scripts/run.sh dialemma main
  • Reproduce results in Figure 1: scripts/run.sh dialemma ablation and python src/plot.py

Results are written to dialemma_main.csv, dialemma_ablation.csv, and Figure-1.pdf.

🌐 WikiDIR BLI

  • Download WikiDIR files and create splits: scripts/setup_wikidir.sh
  • Reproduce results in Tables 3-5: scripts/run.sh wikidir main

Results are written to wikidir_main_precision.csv, wikidir_main_recall.csv, and wikidir_main_f1.csv.

🔍 Cross-Dialect IR

  • Set the JAVA_HOME environment variable (e.g., export JAVA_HOME=/path/to/java/jdk-21.0.4).
  • Download retrieval data and index dialect corpora: scripts/setup_cdir.sh
  • Reproduce results in Table 6: python src/run_cdir.py

Results are written to bm25_results.csv.

Citation

Please consider citing our paper if you use resources from this repository:

@inproceedings{litschko-etal-2026-resource,
  title = {Resource-Lean Lexicon Induction for German Dialects},
  author = {Litschko, Robert and Plank, Barbara and Frassinelli, Diego},
  booktitle = {Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026)},
  month = {May},
  year = {2026},
  pages = {9044--9050},
  address = {Palma, Mallorca, Spain},
  publisher = {European Language Resources Association (ELRA)},
  editor = {Piperidis, Stelios and Bel, Núria and van den Heuvel, Henk and Ide, Nancy and Krek, Simon and Toral, Antonio},
  doi = {10.63317/2feouaji2rxe},
}

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages