1 of 9

CHALLENGES IN THE DETECTION OF DIALECT FOR HISTORICAL LANGUAGES

THE CASE OF OLD IRISH TEXT RESOURCES

ADRIAN DOYLE

ADRIAN.ODUBHGHAILL@MU.IE

2 of 9

OLD IRISH

Archaic Irish

(400 – 600)

Old Irish

(600 – 900)

Middle Irish

(900 – 1200)

Early Modern Irish

(1200 – 1600)

Modern Dialectal Irish

(1600 – Present)

Modern Standard Irish

(1958 – Present)

Scottish Gaelic

Manx Gaelic

3 of 9

OLD IRISH AND DIALECTS

Early Old Irish

(600 – 700)

Mid. Old Irish

(700 – 800)

Late Old Irish

(800 – 900)

RIA MS 23 E 25 (Leabhar na hUidhre)

(isos.dias.ie)

St. Gallen, Stiftsbibliothek, Cod. Sang. 904, f. 194

(www.e-codices.ch)

  • Thomas F. O’Rahilly. 1932. Irish Dialects Past and Present. Dublin Institute for Advanced Studies, Dublin: 16.

4 of 9

COMPUTATIONAL APPROACHES

The Problem

    • No Old Irish text data labelled for dialect

Neural Models and LLMs

    • Can’t train models
    • Can’t test models’ accuracy

Stylometric Approach

    • Cannot quantify accuracy
    • Clustering allows visual assessment
  • Maged S. Al-Shaibani and Moataz Ahmed. 2025. The Arabic AI Fingerprint: Stylometric Analysis and Detection of Large Language Models Text.
  • Jack Grieve.2023. Register Variation Explains Stylometric Authorship Analysis. Corpus Linguistics and Linguistic Theory, 19(1): 47–77.
  • M.Z. Lahjouji-Seppälä, A. Rabus, and R. von Waldenfels. 2022. Ukrainian Standard Variants in the 20th Century: Stylometry to the Rescue. Russian Linguistics, 46(3): 217–232.

5 of 9

METHODOLOGY

    • Corpus PalaeoHibernicum (CorPH)
    • 12 texts in total

Texts all drawn from one source:

    • Linked to named geographical location
    • Linked to named author
    • Large corpus of glosses

Text selected because:

    • Each sentence labelled by source text

Texts split into sentences

6 of 9

METHODOLOGY

    • Pre-trained model: paraphrase-multilingual-mpnet-base-v2
    • 768-dimensional vector space
    • Cosine Similarity

Sentence Embeddings

    • HDBSCAN
    • Min. cluster size = 15
    • Min. samples = 5

Clustering

    • UMAP

Dimensionality Reduction

7 of 9

RESULTS

.i. adǽ

.i. adæ /�.i. adae

adæ

adǽ /

ádǽ /

adé /

ádæ

8 of 9

FUTURE WORK

  1. Application to modern Gaelic languages

​

  1. Reattempt on Old Irish with more data

9 of 9

GO RAIBH MAITH AGAIBH!���THANK YOU!