Publications

Schacht, Carmen; Landwehr, Isabell

CoBra: A Compound Branching Resource for Nominal Triconstituent Compounds in English and German Inproceedings

Proceedings of the Ninth Workshop on Universal Dependencies (UDW, LREC 2026), pp. 128-141, Palma de Mallorca, Spain, 2026.

We present CoBra, a resource containing triconstituent nominal compounds in English and German. This addresses an understudied aspect of compound processing, since research and resources in psycholinguistics and NLP have mostly focused on two-constituent compounds. In addition, our resource covers both general and scientific language, allowing for a register-informed perspective on compounds. It provides syntactic and semantic annotation of compound structure, in particular of the branching direction (i.e. the internal embedding structure, the Compound Branching) and the semantic relationship between constituents. Annotations are implemented using extensions of Universal Dependencies (UD) labels. To explore applications of our new resource, we also conduct a pilot study investigating the relationship between semantic transparency and branching direction. Our results indicate that there is indeed a correlation. Overall, our resource contributes to gaining a more detailed understanding of the structure and processing of morphologically complex words within the UD framework.

@inproceedings{Schacht_etal_2026:Cobra,
title = {CoBra: A Compound Branching Resource for Nominal Triconstituent Compounds in English and German},
author = {Carmen Schacht and Isabell Landwehr},
url = {http://lrec-conf.org/proceedings/lrec2026/workshops/udw/2026.udw-1.0.pdf},
year = {2026},
date = {2026},
booktitle = {Proceedings of the Ninth Workshop on Universal Dependencies (UDW, LREC 2026)},
pages = {128-141},
address = {Palma de Mallorca, Spain},
abstract = {We present CoBra, a resource containing triconstituent nominal compounds in English and German. This addresses an understudied aspect of compound processing, since research and resources in psycholinguistics and NLP have mostly focused on two-constituent compounds. In addition, our resource covers both general and scientific language, allowing for a register-informed perspective on compounds. It provides syntactic and semantic annotation of compound structure, in particular of the branching direction (i.e. the internal embedding structure, the Compound Branching) and the semantic relationship between constituents. Annotations are implemented using extensions of Universal Dependencies (UD) labels. To explore applications of our new resource, we also conduct a pilot study investigating the relationship between semantic transparency and branching direction. Our results indicate that there is indeed a correlation. Overall, our resource contributes to gaining a more detailed understanding of the structure and processing of morphologically complex words within the UD framework.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Projects:   B1 C6

Bagdasarov, Sergei; Alves, Diego; Fischer, Stefan; Teich, Elke

Using LLMs for Automatic Discipline Annotation in a Diachronic Corpus of English Scientific Papers Inproceedings

Piperidis, Stelios; Bel, Núria; van den Heuvel, Henk; Ide, Nancy; Krek, Simon; Toral, Antonio (Ed.): Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), European Language Resources Association (ELRA), pp. 2376--2386, Palma, Mallorca, Spain, 2026.

This study investigates the potential of generative large language models (LLMs) to automatically identify the disciplines of scientific papers in the Royal Society Corpus (RSC) – an extensive collection of English scientific publications spanning more than three centuries. We evaluated eight open-source, state-of-the-art LLMs from four model families on a manually annotated subset and further validated the three best-performing models on a corpus of modern scientific texts. These models were subsequently used for large-scale annotation of the RSC. The models exhibited robust and consistent performance, with at least two LLMs agreeing on the same label for 98.3% of the documents. We then conducted an error analysis of papers assigned divergent labels and a diachronic case study of disciplinary trends within the corpus. The error analysis revealed that most discrepancies occurred in twentieth-century texts, reflecting the growing interdisciplinarity of research. The diachronic analysis showed a gradual decline in disciplinary diversity over time as well as fluctuations corresponding to major paradigm shifts such as the Chemical Revolution and key twentieth-century developments in Physics. The discipline labels generated by the three models will be made publicly available.

@inproceedings{bagdasarov-etal-2026-llms,
title = {Using LLMs for Automatic Discipline Annotation in a Diachronic Corpus of English Scientific Papers},
author = {Sergei Bagdasarov and Diego Alves and Stefan Fischer and Elke Teich},
editor = {Stelios Piperidis and Núria Bel and Henk van den Heuvel and Nancy Ide and Simon Krek and Antonio Toral},
url = {https://lrec.elra.info/lrec2026-main-187},
doi = {https://doi.org/10.63317/3j9wvu86v48t},
year = {2026},
date = {2026},
booktitle = {Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026)},
pages = {2376--2386},
publisher = {European Language Resources Association (ELRA)},
address = {Palma, Mallorca, Spain},
abstract = {

This study investigates the potential of generative large language models (LLMs) to automatically identify the disciplines of scientific papers in the Royal Society Corpus (RSC) – an extensive collection of English scientific publications spanning more than three centuries. We evaluated eight open-source, state-of-the-art LLMs from four model families on a manually annotated subset and further validated the three best-performing models on a corpus of modern scientific texts. These models were subsequently used for large-scale annotation of the RSC. The models exhibited robust and consistent performance, with at least two LLMs agreeing on the same label for 98.3% of the documents. We then conducted an error analysis of papers assigned divergent labels and a diachronic case study of disciplinary trends within the corpus. The error analysis revealed that most discrepancies occurred in twentieth-century texts, reflecting the growing interdisciplinarity of research. The diachronic analysis showed a gradual decline in disciplinary diversity over time as well as fluctuations corresponding to major paradigm shifts such as the Chemical Revolution and key twentieth-century developments in Physics. The discipline labels generated by the three models will be made publicly available.

},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Alves, Diego; Bagdasarov, Sergei; Teich, Elke

Cognitive Signatures of Multi-Word Expressions: Reading-Time and Surprisal Inproceedings

Kr. Ojha, Atul; Barbu Mititelu, Verginica; Constant, Mathieu; Stoyanova, Ivelina; Seza Doğruöz, A.; Rademaker, Alexandre (Ed.): Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026), Association for Computational Linguistics, pp. 48-53, Rabat, Marocco, 2026, ISBN 979-8-89176-363-0.

This study investigates whether eye-tracking measures predict if a word is the final token of a multi-word expression (MWE), focusing on two understudied MWE types: fixed expressions (e.g., due to) and phrasal verbs (e.g., turn out). Using mixed-effects logistic regression, we compared tokens in MWE contexts with the same tokens in non-MWE contexts. Results reveal a clear difference in processing. For fixed expressions, reading-time measures significantly predict MWEhood. In contrast, phrasal verbs show no consistent predictive effects. Additionally, we compared the reading-time models to models that included GPT-2 surprisal as a predictor. While surprisal does predict MWEhood, it fails to capture the distinction between types. These findings highlight the need to consider MWE typology in models of formulaic language processing.

@inproceedings{alves-etal-2026-cognitive,
title = {Cognitive Signatures of Multi-Word Expressions: Reading-Time and Surprisal},
author = {Diego Alves and Sergei Bagdasarov and Elke Teich},
editor = {Atul Kr. Ojha and Verginica Barbu Mititelu and Mathieu Constant and Ivelina Stoyanova and A. Seza Doğru{\"o}z and Alexandre Rademaker},
url = {https://aclanthology.org/2026.mwe-1.5/},
doi = {https://doi.org/10.18653/v1/2026.mwe-1.5},
year = {2026},
date = {2026},
booktitle = {Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026)},
isbn = {979-8-89176-363-0},
pages = {48-53},
publisher = {Association for Computational Linguistics},
address = {Rabat, Marocco},
abstract = {This study investigates whether eye-tracking measures predict if a word is the final token of a multi-word expression (MWE), focusing on two understudied MWE types: fixed expressions (e.g., due to) and phrasal verbs (e.g., turn out). Using mixed-effects logistic regression, we compared tokens in MWE contexts with the same tokens in non-MWE contexts. Results reveal a clear difference in processing. For fixed expressions, reading-time measures significantly predict MWEhood. In contrast, phrasal verbs show no consistent predictive effects. Additionally, we compared the reading-time models to models that included GPT-2 surprisal as a predictor. While surprisal does predict MWEhood, it fails to capture the distinction between types. These findings highlight the need to consider MWE typology in models of formulaic language processing.

},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Steuer, Julius; Krielke, Marie-Pauline; Degaetano-Ortlieb, Stefania; Teich, Elke; Klakow, Dietrich

Modeling the Memory-Surprisal Trade-Off over Time: Communicative Efficiency Decreases with Lexico-Grammatical Change in Scientific English Inproceedings

Piperidis, Stelios; Bel, Núria; van den Heuvel, Henk; Ide, Nancy; Krek, Simon; Toral, Antonio (Ed.): Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), European Language Resources Association (ELRA), pp. 11309-11319, Palma, Mallorca, Spain, 2026.

The memory-surprisal trade-off (MST) has been shown to hold cross-linguistically as a general principle of communicative efficiency: languages that exhibit information locality tend to have word orders that allow for efficient memory use, i.e., lower surprisal at a fixed memory budget. In this paper, we explore the influence of diachronic variation on the MST. We compare scientific English in the Royal Society Corpus (RSC, 18thc. – 20thc.) to „general language“ in the Corpus of Historical American English (COHA) to assess the impact of intra-linguistic variation (register). We find that both time and register influence the shape of the tradeoff: Over time, vocabulary expansion raises minimal surprisal, while the shape of the MST curves changes. Decreasing distances between syntactic dependencies due to more local nominal encodings change how predictive information is distributed across memory scales. The effects are stronger for the RSC than for COHA.

@inproceedings{steuer-etal-2026-modeling,
title = {Modeling the Memory-Surprisal Trade-Off over Time: Communicative Efficiency Decreases with Lexico-Grammatical Change in Scientific English},
author = {Julius Steuer and Marie-Pauline Krielke and Stefania Degaetano-Ortlieb and Elke Teich and Dietrich Klakow},
editor = {Stelios Piperidis and Núria Bel and Henk van den Heuvel and Nancy Ide and Simon Krek and Antonio Toral},
url = {https://lrec.elra.info/lrec2026-main-884},
doi = {https://doi.org/10.63317/4txotk7fwkhp},
year = {2026},
date = {2026},
booktitle = {Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026)},
pages = {11309-11319},
publisher = {European Language Resources Association (ELRA)},
address = {Palma, Mallorca, Spain},
abstract = {The memory-surprisal trade-off (MST) has been shown to hold cross-linguistically as a general principle of communicative efficiency: languages that exhibit information locality tend to have word orders that allow for efficient memory use, i.e., lower surprisal at a fixed memory budget. In this paper, we explore the influence of diachronic variation on the MST. We compare scientific English in the Royal Society Corpus (RSC, 18thc. – 20thc.) to "general language" in the Corpus of Historical American English (COHA) to assess the impact of intra-linguistic variation (register). We find that both time and register influence the shape of the tradeoff: Over time, vocabulary expansion raises minimal surprisal, while the shape of the MST curves changes. Decreasing distances between syntactic dependencies due to more local nominal encodings change how predictive information is distributed across memory scales. The effects are stronger for the RSC than for COHA.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Projects:   B1 B4

Alves, Diego; Bagdasarov, Sergei; Teich, Elke

Cognitive Signatures of Multi-Word Expressions: Reading-Time and Surprisal Inproceedings

Kr. Ojha, Atul; Barbu Mititelu, Verginica; Constant, Mathieu; Stoyanova, Ivelina; Seza Doğruöz, A.; Rademaker, Alexandre (Ed.): Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026), Association for Computational Linguistics, pp. 48-53, Rabat, Marocco, 2026, ISBN 979-8-89176-363-0.

This study investigates whether eye-tracking measures predict if a word is the final token of a multi-word expression (MWE), focusing on two understudied MWE types: fixed expressions (e.g., \textit{due to}) and phrasal verbs (e.g., \textit{turn out}). Using mixed-effects logistic regression, we compared tokens in MWE contexts with the same tokens in non-MWE contexts. Results reveal a clear difference in processing. For fixed expressions, reading-time measures significantly predict MWEhood. In contrast, phrasal verbs show no consistent predictive effects. Additionally, we compared the reading-time models to models that included GPT-2 surprisal as a predictor. While surprisal does predict MWEhood, it fails to capture the distinction between types. These findings highlight the need to consider MWE typology in models of formulaic language processing.

@inproceedings{alves-etal-2026-cognitive,
title = {Cognitive Signatures of Multi-Word Expressions: Reading-Time and Surprisal},
author = {Diego Alves and Sergei Bagdasarov and Elke Teich},
editor = {Atul Kr. Ojha and Verginica Barbu Mititelu and Mathieu Constant and Ivelina Stoyanova and A. Seza Doğru{\"o}z and Alexandre Rademaker},
url = {https://aclanthology.org/2026.mwe-1.5/},
doi = {https://doi.org/10.18653/v1/2026.mwe-1.5},
year = {2026},
date = {2026},
booktitle = {Proceedings of the 22nd Workshop on Multiword Expressions (MWE 2026)},
isbn = {979-8-89176-363-0},
pages = {48-53},
publisher = {Association for Computational Linguistics},
address = {Rabat, Marocco},
abstract = {This study investigates whether eye-tracking measures predict if a word is the final token of a multi-word expression (MWE), focusing on two understudied MWE types: fixed expressions (e.g., \textit{due to}) and phrasal verbs (e.g., \textit{turn out}). Using mixed-effects logistic regression, we compared tokens in MWE contexts with the same tokens in non-MWE contexts. Results reveal a clear difference in processing. For fixed expressions, reading-time measures significantly predict MWEhood. In contrast, phrasal verbs show no consistent predictive effects. Additionally, we compared the reading-time models to models that included GPT-2 surprisal as a predictor. While surprisal does predict MWEhood, it fails to capture the distinction between types. These findings highlight the need to consider MWE typology in models of formulaic language processing.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Menzel, Katrin

Eine korpusbasierte diachrone Untersuchung zu übersetzten Wissenschaftsartikeln aus den Zeitschriften der Royal Society of London Journal Article

chronotopos – A Journal of Translation History, 6, pp. 31--57, 2025.

Dieser Beitrag beschreibt eine Korpusstudie zu den englischen Übersetzungen von naturwissenschaftlichen Texten, die in Zeitschriften der Royal Society of London seit dem 17. Jhd. veröffentlicht wurden. Als Datengrundlage dient das Royal Society Corpus (RSC), welches vor allem originalsprachliche englische Fachartikel, aber auch eine beachtliche Anzahl von übersetzten englischen Beiträgen aus Zeitschriften wie den Philosophical Transactions und den Proceedings der Royal Society beinhaltet. In einem ersten Schritt werden die übersetzten Fachartikel in den Daten identifiziert und in einem zusammenfassenden Überblick im Hinblick auf ihre genauen Entstehungszeiten, Themen, Ausgangssprachen und Übersetzer analysiert. Dabei stellt sich u. a. heraus, dass die meisten Übersetzungen im RSC aus dem 18. Jhd. stammen. Daher werden in einem nächsten Schritt speziell diese Texte in Bezug auf ausgewählte linguistische Merkmale untersucht, welche geeignet sind, um eine übersetzungswissenschaftliche Universalienhypothese, und zwar die ‚Normalisierungshypothese‘, in einem historischen Kontext zu überprüfen. Hierbei soll geklärt werden, ob die übersetzten Texte durch sprachlich weniger innovative Merkmale geprägt sind als nicht-übersetzte englische Vergleichstexte aus dem RSC. Insgesamt zeigen die Ergebnisse, dass Normalisierung und eine stärkere Nutzung von konventionelleren sprachlichen Strukturen keine auf die historischen Wissenschaftsübersetzungen zutreffende Übersetzungspraktiken waren. Anschließend wird ein Ausblick auf den Aufbau eines multilingualen Parallelkorpus mit den übersetzten Fachartikeln und ihren jeweiligen Ausgangstexten gegeben, um weitere Untersuchungen zu prototypischen Übersetzungseigenschaften zu ermöglichen, bei denen auch der Einfluss der Ausgangstexte berücksichtigt werden kann.

@article{Menzel_2025,
title = {Eine korpusbasierte diachrone Untersuchung zu {\"u}bersetzten Wissenschaftsartikeln aus den Zeitschriften der Royal Society of London},
author = {Katrin Menzel},
url = {https://chronotopos.eu/cts/article/view/130},
doi = {https://doi.org/10.70596/cts130},
year = {2025},
date = {2025},
journal = {chronotopos – A Journal of Translation History},
pages = {31--57},
volume = {6},
number = {2},
abstract = {Dieser Beitrag beschreibt eine Korpusstudie zu den englischen {\"U}bersetzungen von naturwissenschaftlichen Texten, die in Zeitschriften der Royal Society of London seit dem 17. Jhd. ver{\"o}ffentlicht wurden. Als Datengrundlage dient das Royal Society Corpus (RSC), welches vor allem originalsprachliche englische Fachartikel, aber auch eine beachtliche Anzahl von {\"u}bersetzten englischen Beitr{\"a}gen aus Zeitschriften wie den Philosophical Transactions und den Proceedings der Royal Society beinhaltet. In einem ersten Schritt werden die {\"u}bersetzten Fachartikel in den Daten identifiziert und in einem zusammenfassenden {\"U}berblick im Hinblick auf ihre genauen Entstehungszeiten, Themen, Ausgangssprachen und {\"U}bersetzer analysiert. Dabei stellt sich u. a. heraus, dass die meisten {\"U}bersetzungen im RSC aus dem 18. Jhd. stammen. Daher werden in einem n{\"a}chsten Schritt speziell diese Texte in Bezug auf ausgew{\"a}hlte linguistische Merkmale untersucht, welche geeignet sind, um eine {\"u}bersetzungswissenschaftliche Universalienhypothese, und zwar die ‚Normalisierungshypothese‘, in einem historischen Kontext zu {\"u}berpr{\"u}fen. Hierbei soll gekl{\"a}rt werden, ob die {\"u}bersetzten Texte durch sprachlich weniger innovative Merkmale gepr{\"a}gt sind als nicht-{\"u}bersetzte englische Vergleichstexte aus dem RSC. Insgesamt zeigen die Ergebnisse, dass Normalisierung und eine st{\"a}rkere Nutzung von konventionelleren sprachlichen Strukturen keine auf die historischen Wissenschafts{\"u}bersetzungen zutreffende {\"U}bersetzungspraktiken waren. Anschlie{\ss}end wird ein Ausblick auf den Aufbau eines multilingualen Parallelkorpus mit den {\"u}bersetzten Fachartikeln und ihren jeweiligen Ausgangstexten gegeben, um weitere Untersuchungen zu prototypischen {\"U}bersetzungseigenschaften zu erm{\"o}glichen, bei denen auch der Einfluss der Ausgangstexte ber{\"u}cksichtigt werden kann.},
pubstate = {published},
type = {article}
}

Copy BibTeX to Clipboard

Project:   B1

Alves, Diego; Bagdasarov, Sergei; Teich, Elke

Surprisal Dynamics for the Detection of Multi-Word Expressions in English Inproceedings

Inui, Kentaro; Sakti, Sakriani; Wang, Haofen; F. Wong, Derek; Bhattacharyya, Pushpak; Banerjee, Biplab; Ekbal, Asif; Chakraborty, Tanmoy; Pratap Singh, Dhirendra (Ed.): Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, The Asian Federation of Natural Language Processing and The Association for Computational Linguistics, pp. 1185-1194, Mumbai, India, 2025, ISBN 979-8-89176-303-6.

This work examines the potential of surprisal slope as a feature for identifying multi-word expressions (MWEs) in English, leveraging token-level surprisal estimates from the GPT-2 language model. Evaluations on the DiMSUM and SemEval-2022 datasets reveal that surprisal slope provides moderate yet meaningful discriminative power with a trade-off between specificity and coverage: while high recall indicates that surprisal slope captures many true MWEs, the slightly lower precision reflects false positives, particularly for non-MWEs that follow formulaic patterns (e.g., adjective-noun or verb-pronoun structures). The method performs particularly well for conventionalized expressions, such as idiomatic bigrams in the SemEval-2022 corpus. Both idiomatic and literal usages of these bigrams exhibit negative slopes, with idiomatic instances generally showing a more pronounced decrease.Overall, surprisal slope offers a cognitively motivated and interpretable signal that complements existing MWE identification methods, particularly for conventionalized expressions.

@inproceedings{alves-etal-2025-surprisal,
title = {Surprisal Dynamics for the Detection of Multi-Word Expressions in English},
author = {Diego Alves and Sergei Bagdasarov and Elke Teich},
editor = {Kentaro Inui and Sakriani Sakti and Haofen Wang and Derek F. Wong and Pushpak Bhattacharyya and Biplab Banerjee and Asif Ekbal and Tanmoy Chakraborty and Dhirendra Pratap Singh},
url = {https://aclanthology.org/2025.findings-ijcnlp.72/},
doi = {https://doi.org/10.18653/v1/2025.findings-ijcnlp.72},
year = {2025},
date = {2025},
booktitle = {Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics},
isbn = {979-8-89176-303-6},
pages = {1185-1194},
publisher = {The Asian Federation of Natural Language Processing and The Association for Computational Linguistics},
address = {Mumbai, India},
abstract = {This work examines the potential of surprisal slope as a feature for identifying multi-word expressions (MWEs) in English, leveraging token-level surprisal estimates from the GPT-2 language model. Evaluations on the DiMSUM and SemEval-2022 datasets reveal that surprisal slope provides moderate yet meaningful discriminative power with a trade-off between specificity and coverage: while high recall indicates that surprisal slope captures many true MWEs, the slightly lower precision reflects false positives, particularly for non-MWEs that follow formulaic patterns (e.g., adjective-noun or verb-pronoun structures). The method performs particularly well for conventionalized expressions, such as idiomatic bigrams in the SemEval-2022 corpus. Both idiomatic and literal usages of these bigrams exhibit negative slopes, with idiomatic instances generally showing a more pronounced decrease.Overall, surprisal slope offers a cognitively motivated and interpretable signal that complements existing MWE identification methods, particularly for conventionalized expressions.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Alves, Diego; Bagdasarov, Sergei; Teich, Elke

Surprisal Dynamics for the Detection of Multi-Word Expressions in English Inproceedings

Inui, Kentaro; Sakt, Sakriani; Wang, Haofen; F. Wong, Derek; Bhattacharyya, Pushpak; Banerjee, Biplab; Ekbal, Asif; Chakraborty, Tanmoy; Pratap Singh, Dhirendra (Ed.): Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, The Asian Federation of Natural Language Processing and The Association for Computational Linguistics, pp. 1185-1194, Mumbai, India, 2025, ISBN 979-8-89176-303-6.

This work examines the potential of surprisal slope as a feature for identifying multi-word expressions (MWEs) in English, leveraging token-level surprisal estimates from the GPT-2 language model. Evaluations on the DiMSUM and SemEval-2022 datasets reveal that surprisal slope provides moderate yet meaningful discriminative power with a trade-off between specificity and coverage: while high recall indicates that surprisal slope captures many true MWEs, the slightly lower precision reflects false positives, particularly for non-MWEs that follow formulaic patterns (e.g., adjective-noun or verb-pronoun structures). The method performs particularly well for conventionalized expressions, such as idiomatic bigrams in the SemEval-2022 corpus. Both idiomatic and literal usages of these bigrams exhibit negative slopes, with idiomatic instances generally showing a more pronounced decrease.Overall, surprisal slope offers a cognitively motivated and interpretable signal that complements existing MWE identification methods, particularly for conventionalized expressions.

@inproceedings{alves-etal-2025-surprisal,
title = {Surprisal Dynamics for the Detection of Multi-Word Expressions in English},
author = {Diego Alves and Sergei Bagdasarov and Elke Teich},
editor = {Kentaro Inui and Sakriani Sakt and Haofen Wang and Derek F. Wong and Pushpak Bhattacharyya and Biplab Banerjee and Asif Ekbal and Tanmoy Chakraborty and Dhirendra Pratap Singh},
url = {https://aclanthology.org/2025.findings-ijcnlp.72/},
doi = {https://doi.org/10.18653/v1/2025.findings-ijcnlp.72},
year = {2025},
date = {2025},
booktitle = {Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics},
isbn = {979-8-89176-303-6},
pages = {1185-1194},
publisher = {The Asian Federation of Natural Language Processing and The Association for Computational Linguistics},
address = {Mumbai, India},
abstract = {This work examines the potential of surprisal slope as a feature for identifying multi-word expressions (MWEs) in English, leveraging token-level surprisal estimates from the GPT-2 language model. Evaluations on the DiMSUM and SemEval-2022 datasets reveal that surprisal slope provides moderate yet meaningful discriminative power with a trade-off between specificity and coverage: while high recall indicates that surprisal slope captures many true MWEs, the slightly lower precision reflects false positives, particularly for non-MWEs that follow formulaic patterns (e.g., adjective-noun or verb-pronoun structures). The method performs particularly well for conventionalized expressions, such as idiomatic bigrams in the SemEval-2022 corpus. Both idiomatic and literal usages of these bigrams exhibit negative slopes, with idiomatic instances generally showing a more pronounced decrease.Overall, surprisal slope offers a cognitively motivated and interpretable signal that complements existing MWE identification methods, particularly for conventionalized expressions.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Alves, Diego; Fischer, Stefan; Teich, Elke

Syntagmatic Productivity of MWEs in Scientific English Inproceedings

Kr. Ojha, Atul; Giouli, Voula; Barbu Mititelu, Verginica; Constant, Mathieu; Korvel, Gražina; Seza Doğruöz, A.; Rademaker, Alexandre (Ed.): Proceedings of the 21st Workshop on Multiword Expressions (MWE 2025), Association for Computational Linguistics, pp. 1-6, Albuquerque, New Mexico, U.S.A., 2025, ISBN 979-8-89176-243-5.

This paper presents an analysis of the syntagmatic productivity (SynProd) of different classes of multiword expressions (MWEs) in English scientific writing over time (mid 17th to 20th c.). SynProd refers to the variability of the syntagmatic context in which a word or other kind of linguistic unit is used. To measure SynProd, we use entropy. The study reveals that, similar to single-token units of various parts of speech, MWEs exhibit an increasing trend in syntagmatic productivity over time, particularly after the mid-19th century. Furthermore, when compared to similar parts of speech (PoS), MWEs show a more pronounced increase in SynProd over time.

@inproceedings{alves-etal-2025-syntagmatic,
title = {Syntagmatic Productivity of MWEs in Scientific English},
author = {Diego Alves and Stefan Fischer and Elke Teich},
editor = {Atul Kr. Ojha and Voula Giouli and Verginica Barbu Mititelu and Mathieu Constant and Gra{\v{z}ina Korvel and A. Seza Doğru{\"o}z and Alexandre Rademaker},
url = {https://aclanthology.org/2025.mwe-1.1/},
doi = {https://doi.org/10.18653/v1/2025.mwe-1.1},
year = {2025},
date = {2025},
booktitle = {Proceedings of the 21st Workshop on Multiword Expressions (MWE 2025)},
isbn = {979-8-89176-243-5},
pages = {1-6},
publisher = {Association for Computational Linguistics},
address = {Albuquerque, New Mexico, U.S.A.},
abstract = {This paper presents an analysis of the syntagmatic productivity (SynProd) of different classes of multiword expressions (MWEs) in English scientific writing over time (mid 17th to 20th c.). SynProd refers to the variability of the syntagmatic context in which a word or other kind of linguistic unit is used. To measure SynProd, we use entropy. The study reveals that, similar to single-token units of various parts of speech, MWEs exhibit an increasing trend in syntagmatic productivity over time, particularly after the mid-19th century. Furthermore, when compared to similar parts of speech (PoS), MWEs show a more pronounced increase in SynProd over time.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Landwehr, Isabell

The Interplay of Noun Phrase Complexity and Modification Type in Scientific Writing Inproceedings

Chen, Xinying; Wang, Yaqin (Ed.): Proceedings of the Third Workshop on Quantitative Syntax (QUASY, SyntaxFest 2025), Association for Computational Linguistics, pp. 72-82, Ljubljana, Slovenia, 2025, ISBN 979-8-89176-293-0.

We investigate the interplay of noun phrase (NP) complexity and modification type, namely the choice between pre- and postmodification, using a corpus-based approach. Our dataset is the Royal Society Corpus (RSC, Fischer et al. 2020), a diachronic corpus of English scientific writing. We find that the number of dependents, length of the head noun and distance to the head noun{‚}s own syntactic head (typically the main verb) affect the likelihood of pre- vs. postmodification: NPs with more dependents are more likely to be premodified, NPs with a longer head noun and a head noun closer to its own head are more likely to be postmodified. In addition, we find an effect of syntactic role and definiteness as well as time: The likelihood of premodification over postmodification increases with time and subject NPs as well as indefinite NPs are more likely to be premodified than NPs in other syntactic roles or definite NPs.

@inproceedings{landwehr-2025-interplay,
title = {The Interplay of Noun Phrase Complexity and Modification Type in Scientific Writing},
author = {Isabell Landwehr},
editor = {Xinying Chen and Yaqin Wang},
url = {https://aclanthology.org/2025.quasy-1.10/},
year = {2025},
date = {2025},
booktitle = {Proceedings of the Third Workshop on Quantitative Syntax (QUASY, SyntaxFest 2025)},
isbn = {979-8-89176-293-0},
pages = {72-82},
publisher = {Association for Computational Linguistics},
address = {Ljubljana, Slovenia},
abstract = {We investigate the interplay of noun phrase (NP) complexity and modification type, namely the choice between pre- and postmodification, using a corpus-based approach. Our dataset is the Royal Society Corpus (RSC, Fischer et al. 2020), a diachronic corpus of English scientific writing. We find that the number of dependents, length of the head noun and distance to the head noun{'}s own syntactic head (typically the main verb) affect the likelihood of pre- vs. postmodification: NPs with more dependents are more likely to be premodified, NPs with a longer head noun and a head noun closer to its own head are more likely to be postmodified. In addition, we find an effect of syntactic role and definiteness as well as time: The likelihood of premodification over postmodification increases with time and subject NPs as well as indefinite NPs are more likely to be premodified than NPs in other syntactic roles or definite NPs.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Landwehr, Isabell; Krielke, Marie-Pauline; Degaetano-Ortlieb, Stefania; Zhao, Jin; Wang, Mingyang; Liu, Zhu

Exploring the Effect of Nominal Compound Structure in Scientific Texts on Reading Times of Experts and Novices Inproceedings

Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop), Association for Computational Linguistics, pp. 396-408, Vienna, Austria, 2025, ISBN 979-8-89176-254-1.

We explore how different types of nominal compound complexity in scientific writing, in particular different types of compound structure, affect the reading times of experts and novices. We consider both in-domain and out-of-domain reading and use PoTeC (Jakobi et al. 2024), a corpus containing eye-tracking data of German native speakers reading passages from scientific textbooks. Our results suggest that some compound types are associated with longer reading times and that experts may not only have an advantage while reading in-domain texts, but also while reading out-of-domain.

@inproceedings{landwehr-etal-2025-exploring,
title = {Exploring the Effect of Nominal Compound Structure in Scientific Texts on Reading Times of Experts and Novices},
author = {Isabell Landwehr and Marie-Pauline Krielke and Stefania Degaetano-Ortlieb andJin Zhao and Mingyang Wang and Zhu Liu},
url = {https://aclanthology.org/2025.acl-srw.25/},
doi = {https://doi.org/10.18653/v1/2025.acl-srw.25},
year = {2025},
date = {2025},
booktitle = {Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 4: Student Research Workshop)},
isbn = {979-8-89176-254-1},
pages = {396-408},
publisher = {Association for Computational Linguistics},
address = {Vienna, Austria},
abstract = {We explore how different types of nominal compound complexity in scientific writing, in particular different types of compound structure, affect the reading times of experts and novices. We consider both in-domain and out-of-domain reading and use PoTeC (Jakobi et al. 2024), a corpus containing eye-tracking data of German native speakers reading passages from scientific textbooks. Our results suggest that some compound types are associated with longer reading times and that experts may not only have an advantage while reading in-domain texts, but also while reading out-of-domain.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Alves, Diego

Information Theory and Linguistic Variation: A Study of Brazilian and European Portuguese Inproceedings

Scherrer, Yves; Jauhiainen, Tommi; Ljubešić, Nikola; Nakov, Preslav; Tiedemann, Jorg; Zampieri, Marcos (Ed.): Proceedings of the 12th Workshop on NLP for Similar Languages, Varieties and Dialects, Association for Computational Linguistics, pp. 9-19, Abu Dhabi, UAE, 2025.

We present a general analysis of the lexical and grammatical differences between Brazilian and European Portuguese by applying entropy measures, including Kullback-Leibler divergence and word order entropy, across various linguistic levels. Using a parallel corpus of BP and EP sentences translated from English, we quantified these differences and identified characteristic phenomena underlying the divergences between the two varieties. The highest divergence was observed at the lexical level due to word pairs unique to each variety but also related to grammatical distinctions. Furthermore, the analysis of parts-of-speech (POS), dependency relations, and POS tri-grams provided information concerning distinctive grammatical constructions. Finally, the word order entropy analysis revealed that while most of the syntactic features analysed showed similar patterns across BP and EP, specific word order preferences were still apparent.

@inproceedings{alves-2025-information,
title = {Information Theory and Linguistic Variation: A Study of Brazilian and European Portuguese},
author = {Diego Alves},
editor = {Yves Scherrer and Tommi Jauhiainen and Nikola Ljubešić and Preslav Nakov and Jorg Tiedemann and Marcos Zampieri},
url = {https://aclanthology.org/2025.vardial-1.2/},
year = {2025},
date = {2025},
booktitle = {Proceedings of the 12th Workshop on NLP for Similar Languages, Varieties and Dialects},
pages = {9-19},
publisher = {Association for Computational Linguistics},
address = {Abu Dhabi, UAE},
abstract = {We present a general analysis of the lexical and grammatical differences between Brazilian and European Portuguese by applying entropy measures, including Kullback-Leibler divergence and word order entropy, across various linguistic levels. Using a parallel corpus of BP and EP sentences translated from English, we quantified these differences and identified characteristic phenomena underlying the divergences between the two varieties. The highest divergence was observed at the lexical level due to word pairs unique to each variety but also related to grammatical distinctions. Furthermore, the analysis of parts-of-speech (POS), dependency relations, and POS tri-grams provided information concerning distinctive grammatical constructions. Finally, the word order entropy analysis revealed that while most of the syntactic features analysed showed similar patterns across BP and EP, specific word order preferences were still apparent.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Alves, Diego

Diachronic Analysis of Phrasal Verbs in English Scientific Writing Inproceedings

Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025), University of Tartu Library, Tallinn, Estonia, 2025.
Phrasal verbs (PVs) are a specific type of multi-word expressions and a specific feature of the English language. However, their usage in scientific prose is limited. Our study focuses on the analysis of phrasal verbs in the scientific domain using information theory methods to describe diachronic phenomena such as conventionalization and diversification regarding the usage of PVs. Thus, we analysed their developmental trajectory over time from the mid-17th century to the end of the 20th century by measuring the relative entropy (Kullback-Leibler divergence), predictability in context of the phrasal verbs particles (surprisal), and the paradigmatic variability using word embedding spaces. We were able to identify interesting phenomena such as the process of conventionalization over the 20th century and the peaks of diversification throughout the centuries.

@inproceedings{Alves-2025,
title = {Diachronic Analysis of Phrasal Verbs in English Scientific Writing},
author = {Diego Alves},
url = {https://dspace.ut.ee/items/ef26bd7f-e708-41b3-b5c8-84cf8057ab71},
year = {2025},
date = {2025},
booktitle = {Proceedings of the Joint 25th Nordic Conference on Computational Linguistics and 11th Baltic Conference on Human Language Technologies (NoDaLiDa/Baltic-HLT 2025)},
publisher = {University of Tartu Library},
address = {Tallinn, Estonia},
abstract = {

Phrasal verbs (PVs) are a specific type of multi-word expressions and a specific feature of the English language. However, their usage in scientific prose is limited. Our study focuses on the analysis of phrasal verbs in the scientific domain using information theory methods to describe diachronic phenomena such as conventionalization and diversification regarding the usage of PVs. Thus, we analysed their developmental trajectory over time from the mid-17th century to the end of the 20th century by measuring the relative entropy (Kullback-Leibler divergence), predictability in context of the phrasal verbs particles (surprisal), and the paradigmatic variability using word embedding spaces. We were able to identify interesting phenomena such as the process of conventionalization over the 20th century and the peaks of diversification throughout the centuries.
},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Menzel, Katrin

Noun + noun Compounds and Verbal Complements as Non-normalised Features in Late Modern English Scientific Translations Inproceedings

Proceedings of 7th Translation in Transition Conference, Batumi: Shota Rustaveli State University, 2024.

This paper presents a study on the usage of noun+noun compounds and verbal complement structures in 18th century scientific articles in the Royal Society Corpus (RSC) comparing translated to non-translated English texts. Departing from the hypothesis that the translations will conform stronger to traditional patterns of the English language, the analysis shows that these historical translations and non-translated texts are similarly marked by the ongoing reorganisation of the noun phrase, but translations
contain more innovative complementation patterns. Additionally, a surprisal analysis shows that the analysed patterns tend to occur in more predictable and conventionalised contexts in non-translated texts than in translation.

 

@inproceedings{Menzel2024Noun,
title = {Noun + noun Compounds and Verbal Complements as Non-normalised Features in Late Modern English Scientific Translations},
author = {Katrin Menzel},
url = {https://sites.google.com/view/tt2024/schedule-and-proceedings},
year = {2024},
date = {2024-12-26},
booktitle = {Proceedings of 7th Translation in Transition Conference},
address = {Batumi: Shota Rustaveli State University},
abstract = {This paper presents a study on the usage of noun+noun compounds and verbal complement structures in 18th century scientific articles in the Royal Society Corpus (RSC) comparing translated to non-translated English texts. Departing from the hypothesis that the translations will conform stronger to traditional patterns of the English language, the analysis shows that these historical translations and non-translated texts are similarly marked by the ongoing reorganisation of the noun phrase, but translations
contain more innovative complementation patterns. Additionally, a surprisal analysis shows that the analysed patterns tend to occur in more predictable and conventionalised contexts in non-translated texts than in translation.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Menzel, Katrin

Initialisms in Scientific Writing in the 19th and Early 20th Centuries Journal Article

Zeitschrift für Wortbildung / Journal of Word Formation (ZWJW) (Special issue Historical English Word-Formation), 8, pp. 7-27, 2024.
This paper focusses on the role of initialisms in scientific English articles in the Royal Society Corpus (Fischer et al. 2020; Kermes et al. 2016). The development of scientific initialisms is illustrated with frequency data, a discussion of the evolution of the text topics obtained from topic modelling and an analysis of the development of information-theoretic surprisal values of initialisms in three time spans between 1830 and 1919. The overall frequency and diversity of initialisms for scientific concepts has risen considerably between 1830 and 1919 in the context of the ongoing specialisation of the sciences. Particularly from the 1860s onwards scientific initialisms increasingly become shortcuts for multiword units with wordhood and term status. The surprisal values of scientific initialisms decrease over time as such forms more regularly occur in conventionalised textual contexts and fixed expressions. Overall, the analysis of the RSC texts shows that key developments towards the conventionalisation of scientific initialisms as term formation patterns took place in the transitional period from Late Modern to Present-day English.

@article{Menzel2024,
title = {Initialisms in Scientific Writing in the 19th and Early 20th Centuries},
author = {Katrin Menzel},
url = {https://journals.linguistik.de/zwjw/article/view/108},
doi = {https://doi.org/10.21248/zwjw.2024.2.108},
year = {2024},
date = {2024},
journal = {Zeitschrift f{\"u}r Wortbildung / Journal of Word Formation (ZWJW) (Special issue Historical English Word-Formation)},
pages = {7-27},
volume = {8},
number = {2},
abstract = {

This paper focusses on the role of initialisms in scientific English articles in the Royal Society Corpus (Fischer et al. 2020; Kermes et al. 2016). The development of scientific initialisms is illustrated with frequency data, a discussion of the evolution of the text topics obtained from topic modelling and an analysis of the development of information-theoretic surprisal values of initialisms in three time spans between 1830 and 1919. The overall frequency and diversity of initialisms for scientific concepts has risen considerably between 1830 and 1919 in the context of the ongoing specialisation of the sciences. Particularly from the 1860s onwards scientific initialisms increasingly become shortcuts for multiword units with wordhood and term status. The surprisal values of scientific initialisms decrease over time as such forms more regularly occur in conventionalised textual contexts and fixed expressions. Overall, the analysis of the RSC texts shows that key developments towards the conventionalisation of scientific initialisms as term formation patterns took place in the transitional period from Late Modern to Present-day English.
},
pubstate = {published},
type = {article}
}

Copy BibTeX to Clipboard

Project:   B1

Steuer, Julius; Krielke, Marie-Pauline; Fischer, Stefan; Degaetano-Ortlieb, Stefania; Mosbach, Marius; Klakow, Dietrich

Modeling Diachronic Change in English Scientific Writing over 300+ Years with Transformer-based Language Model Surprisal Inproceedings

Zweigenbaum, Pierre; Rapp, Reinhard; Sharoff, Serge (Ed.): Proceedings of the 17th Workshop on Building and Using Comparable Corpora (BUCC) @ LREC-COLING 2024, ELRA and ICCL, pp. 12-23, Torino, Italia, 2024.

This study presents an analysis of diachronic linguistic changes in English scientific writing, utilizing surprisal from transformer-based language models. Unlike traditional n-gram models, transformer-based models are potentially better at capturing nuanced linguistic changes such as long-range dependencies by considering variable context sizes. However, to create diachronically comparable language models there are several challenges with historical data, notably an exponential increase in no. of texts, tokens per text and vocabulary size over time. We address these by using a shared vocabulary and employing a robust training strategy that includes initial uniform sampling from the corpus and continuing pre-training on specific temporal segments. Our empirical analysis highlights the predictive power of surprisal from transformer-based models, particularly in analyzing complex linguistic structures like relative clauses. The models’ broader contextual awareness and the inclusion of dependency length annotations contribute to a more intricate understanding of communicative efficiency. While our focus is on scientific English, our approach can be applied to other low-resource scenarios.

@inproceedings{steuer-etal-2024-modeling ,
title = {Modeling Diachronic Change in English Scientific Writing over 300+ Years with Transformer-based Language Model Surprisal},
author = {Julius Steuer and Marie-Pauline Krielke and Stefan Fischer and Stefania Degaetano-Ortlieb and Marius Mosbach and Dietrich Klakow},
editor = {Pierre Zweigenbaum and Reinhard Rapp and Serge Sharoff},
url = {https://aclanthology.org/2024.bucc-1.2/},
year = {2024},
date = {2024},
booktitle = {Proceedings of the 17th Workshop on Building and Using Comparable Corpora (BUCC) @ LREC-COLING 2024},
pages = {12-23},
publisher = {ELRA and ICCL},
address = {Torino, Italia},
abstract = {This study presents an analysis of diachronic linguistic changes in English scientific writing, utilizing surprisal from transformer-based language models. Unlike traditional n-gram models, transformer-based models are potentially better at capturing nuanced linguistic changes such as long-range dependencies by considering variable context sizes. However, to create diachronically comparable language models there are several challenges with historical data, notably an exponential increase in no. of texts, tokens per text and vocabulary size over time. We address these by using a shared vocabulary and employing a robust training strategy that includes initial uniform sampling from the corpus and continuing pre-training on specific temporal segments. Our empirical analysis highlights the predictive power of surprisal from transformer-based models, particularly in analyzing complex linguistic structures like relative clauses. The models’ broader contextual awareness and the inclusion of dependency length annotations contribute to a more intricate understanding of communicative efficiency. While our focus is on scientific English, our approach can be applied to other low-resource scenarios.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Projects:   B1 B4

Bagdasarov, Sergei; Teich, Elke

Multi-word expressions in biomedical abstracts and their plain English adaptations Inproceedings

Hämäläinen, Mika; Öhman, Emily; Miyagawa, So; Alnajjar, Khalid; Bizzoni, Yuri (Ed.): Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities, Association for Computational Linguistics, pp. 483-488, Miami, USA, 2024.

This study analyzes the use of multi-word expressions (MWEs), prefabricated sequences of words (e.g. in this case, this means that, healthcare service, follow up) in biomedical abstracts and their plain language adaptations. While English academic writing became highly specialized and complex from the late 19th century onwards, recent decades have seen a rising demand for a lay-friendly language in scientific content, especially in the health domain, to bridge a communication gap between experts and laypersons. Based on previous research showing that MWEs are easier to process than non-formulaic word sequences of comparable length, we hypothesize that they can potentially be used to create a more reader-friendly language. Our preliminary results suggest some significant differences between complex and plain abstracts when it comes to the usage patterns and informational load of MWEs.

@inproceedings{bagdasarov-teich-2024-multi,
title = {Multi-word expressions in biomedical abstracts and their plain English adaptations},
author = {Sergei Bagdasarov and Elke Teich},
editor = {Mika H{\"a}m{\"a}l{\"a}inen and Emily {\"O}hman and So Miyagawa and Khalid Alnajjar and Yuri Bizzoni},
url = {https://aclanthology.org/2024.nlp4dh-1.46/},
doi = {https://doi.org/10.18653/v1/2024.nlp4dh-1.46},
year = {2024},
date = {2024},
booktitle = {Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities},
pages = {483-488},
publisher = {Association for Computational Linguistics},
address = {Miami, USA},
abstract = {This study analyzes the use of multi-word expressions (MWEs), prefabricated sequences of words (e.g. in this case, this means that, healthcare service, follow up) in biomedical abstracts and their plain language adaptations. While English academic writing became highly specialized and complex from the late 19th century onwards, recent decades have seen a rising demand for a lay-friendly language in scientific content, especially in the health domain, to bridge a communication gap between experts and laypersons. Based on previous research showing that MWEs are easier to process than non-formulaic word sequences of comparable length, we hypothesize that they can potentially be used to create a more reader-friendly language. Our preliminary results suggest some significant differences between complex and plain abstracts when it comes to the usage patterns and informational load of MWEs.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Alves, Diego; Gamallo, Pablo; Claro, Daniela; Teixeira, António; Real, Livy; Garcia, Marcos; Gonçalo, Hugo; Amaro, Raquel

An evaluation of Portuguese language models' adaptation to African Portuguese varieties Inproceedings

Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1, Association for Computational Lingustics, pp. 544-550, Santiago de Compostela, Galicia/Spain, 2024.

In this study, we conduct a comparative evaluation of two state-of-the-art language models, Albertina PT-PT and Albertina PT-BR, which are trained on European Portuguese and Brazilian Portuguese, respectively. Our aim is to assess their suitability for African varieties of Portuguese. To evaluate their performance, we create two test sets for each variety, encompassing both spoken and written language. We measure the percentage of sentences in which one model outperforms the other in terms of perplexity. This evaluation seeks to ascertain whether one model shows more adaptability to the African varieties of Portuguese. Our findings reveal that Albertina PT-PT consistently outperforms Albertina PT-BR in scenarios involving spoken language corpora. However, in written registers, the advantage of Albertina PTPT is less pronounced for the Portuguese varieties of Guinea-Bissau, Mozambique, and São Tomé and Principe. These insights contribute to our understanding of the adaptability of existing language models to African Portuguese varieties and emphasize the need for specialized models to address the unique linguistic nuances of this region.

@inproceedings{alves-2024-evaluation,
title = {An evaluation of Portuguese language models' adaptation to African Portuguese varieties},
author = {Diego Alves andPablo Gamallo and Daniela Claro and António Teixeira and Livy Real and Marcos Garcia and Hugo Gonçalo Oliveira and Raquel Amaro},
url = {https://aclanthology.org/2024.propor-1.58/},
year = {2024},
date = {2024},
booktitle = {Proceedings of the 16th International Conference on Computational Processing of Portuguese - Vol. 1},
pages = {544-550},
publisher = {Association for Computational Lingustics},
address = {Santiago de Compostela, Galicia/Spain},
abstract = {In this study, we conduct a comparative evaluation of two state-of-the-art language models, Albertina PT-PT and Albertina PT-BR, which are trained on European Portuguese and Brazilian Portuguese, respectively. Our aim is to assess their suitability for African varieties of Portuguese. To evaluate their performance, we create two test sets for each variety, encompassing both spoken and written language. We measure the percentage of sentences in which one model outperforms the other in terms of perplexity. This evaluation seeks to ascertain whether one model shows more adaptability to the African varieties of Portuguese. Our findings reveal that Albertina PT-PT consistently outperforms Albertina PT-BR in scenarios involving spoken language corpora. However, in written registers, the advantage of Albertina PTPT is less pronounced for the Portuguese varieties of Guinea-Bissau, Mozambique, and São Tom{\'e} and Principe. These insights contribute to our understanding of the adaptability of existing language models to African Portuguese varieties and emphasize the need for specialized models to address the unique linguistic nuances of this region.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Alves, Diego; Degaetano-Ortlieb, Stefania; Schmidt, Elena; Teich, Elke

Diachronic Analysis of Multi-word Expression Functional Categories in Scientific English Inproceedings

Bhatia, Archna; Bouma, Gosse; Seza Dogruoz, A.; Evang, Kilian; Garcia, Marcos; Giouli, Voula; Han, Lifeng; Nivre, Joakim; Rademaker, Alexandre (Ed.): Proceedings of the Joint Workshop on Multiword Expressions and Universal Dependencies (MWE-UD) @ LREC-COLING 2024, ELRA and ICCL, pp. 81-87, Torino, Italia, 2024.

We present a diachronic analysis of multi-word expressions (MWEs) in English based on the Royal Society Corpus, a dataset containing 300+ years of the scientific publications of the Royal Society of London. Specifically, we investigate the functions of MWEs, such as stance markers (“is is interesting”) or discourse organizers (“in this section”), and their development over time. Our approach is multi-disciplinary: to detect MWEs we use Universal Dependencies, to classify them functionally we use an approach from register linguistics, and to assess their role in diachronic development we use an information-theoretic measure, relative entropy.

@inproceedings{alves-etal-2024-diachronic,
title = {Diachronic Analysis of Multi-word Expression Functional Categories in Scientific English},
author = {Diego Alves and Stefania Degaetano-Ortlieb and Elena Schmidt and Elke Teich},
editor = {Archna Bhatia and Gosse Bouma and A. Seza Dogruoz and Kilian Evang and Marcos Garcia and Voula Giouli and Lifeng Han and Joakim Nivre and Alexandre Rademaker},
url = {https://aclanthology.org/2024.mwe-1.12},
year = {2024},
date = {2024},
booktitle = {Proceedings of the Joint Workshop on Multiword Expressions and Universal Dependencies (MWE-UD) @ LREC-COLING 2024},
pages = {81-87},
publisher = {ELRA and ICCL},
address = {Torino, Italia},
abstract = {We present a diachronic analysis of multi-word expressions (MWEs) in English based on the Royal Society Corpus, a dataset containing 300+ years of the scientific publications of the Royal Society of London. Specifically, we investigate the functions of MWEs, such as stance markers (“is is interesting”) or discourse organizers (“in this section”), and their development over time. Our approach is multi-disciplinary: to detect MWEs we use Universal Dependencies, to classify them functionally we use an approach from register linguistics, and to assess their role in diachronic development we use an information-theoretic measure, relative entropy.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Bagdasarov, Sergei; Degaetano-Ortlieb, Stefania

Applying Information-theoretic Notions to Measure Effects of the Plain English Movement on English Law Reports and Scientific Articles Inproceedings

Bizzoni, Yuri; Degaetano-Ortlieb, Stefania; Kazantseva, Anna; Szpakowicz, Stan (Ed.): Proceedings of the 8th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2024), Association for Computational Linguistics, pp. 101-110, St. Julians, Malta, 2024.

We investigate the impact of the Plain English Movement (PEM) on the complexity of legal language in UK law reports from the 1950s-2010s, contrasting it with the evolution of scientific language. The PEM, emerging in the late 20th century, advocated for clear and understandable legal language. We define complexity through the concept of surprisal – an information-theoretic measure correlating with cognitive processing difficulty. Our research contrasts surprisal with traditional readability measures, which often overlook content. We hypothesize that, if the PEM has influenced legal language, there would be a reduction in complexity over time and a shift from a nominal to a more verbal style. We analyze text complexity and lexico-grammatical changes in line with PEM recommendations. Results indicate minimal impact of the PEM on both legal and scientific domains. This finding suggests future research should consider processing effort when advocating for linguistic norms to enhance accessibility.

@inproceedings{bagdasarov-degaetano-ortlieb-2024-applying,
title = {Applying Information-theoretic Notions to Measure Effects of the Plain English Movement on English Law Reports and Scientific Articles},
author = {Sergei Bagdasarov and Stefania Degaetano-Ortlieb},
editor = {Yuri Bizzoni and Stefania Degaetano-Ortlieb and Anna Kazantseva and Stan Szpakowicz},
url = {https://aclanthology.org/2024.latechclfl-1.11},
year = {2024},
date = {2024},
booktitle = {Proceedings of the 8th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (LaTeCH-CLfL 2024)},
pages = {101-110},
publisher = {Association for Computational Linguistics},
address = {St. Julians, Malta},
abstract = {We investigate the impact of the Plain English Movement (PEM) on the complexity of legal language in UK law reports from the 1950s-2010s, contrasting it with the evolution of scientific language. The PEM, emerging in the late 20th century, advocated for clear and understandable legal language. We define complexity through the concept of surprisal - an information-theoretic measure correlating with cognitive processing difficulty. Our research contrasts surprisal with traditional readability measures, which often overlook content. We hypothesize that, if the PEM has influenced legal language, there would be a reduction in complexity over time and a shift from a nominal to a more verbal style. We analyze text complexity and lexico-grammatical changes in line with PEM recommendations. Results indicate minimal impact of the PEM on both legal and scientific domains. This finding suggests future research should consider processing effort when advocating for linguistic norms to enhance accessibility.},
pubstate = {published},
type = {inproceedings}
}

Copy BibTeX to Clipboard

Project:   B1

Successfully