1 of 17

IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages?

Aravapalli Akhilesh Marreddy Mounika

Radhika Mamidi Manish Gupta Subba Reddy Oota

2 of 17

Harry never thought he would

Harry never thought he ???

Language models (LMs) are trained to predict missing words

3 of 17

3

BERTology studies focused on investigating the internal workings and linguistic representations of language models (LM)

Hierarchy of linguistic info ⇒ how BERT encodes linguistic properties across layers

Conneau et al. 2018

Jawahar et al. 2019

Rogers et al. 2020

Investigate via a probing tasks

Tasks:

    • Surface – Sentence Length, Word Content
    • Syntactic – Bigram shift, Tree depth, Top constituent
    • Semantic – Tense, Subject Number, Object Number, Coordination Inversion and Semantic Odd Man Out.

4 of 17

4

Surface

Syntactic

Semantic

Jawahar et al. 2019 ACL

BERT composes a hierarchy of linguistic signals ranging from surface to semantic features

5 of 17

5

Multi-lingual language models are pretrained on many languages and learn representations for each language

Devlin et al. 2019

Kakwani et al. 2020

Khanuja et al. 2021

Universal multi-lingual language models

  • Pretrained on 100+ languages
  • larger vocabulary to represent tokens in many languages 

Indic multi-lingual language models

  • Pretrained on 11+ Indic languages
  • Indic vocabulary to represent tokens in Indic languages 

IndicBERT

MuRIL

! No prior work has examined which linguistic properties multi-lingual language models encode for different Indic languages.

https://www.ruder.io/state-of-multilingual-ai/

How multi-lingual Transformer-based language models capture linguistic properties across layers for different Indic languages?

6 of 17

6

IndicSentEval: multi-lingual probing tasks for Indic languages

  • How effectively and robustly are English language properties encoded by these universal and Indic-specific models?
  • IndicBertology: How effectively do universal and Indic multi-lingual models encode the hierarchy of linguistic structure in Indic languages?

7 of 17

7

How effectively and robustly are English language properties encoded by these universal and Indic-specific models?

8 of 17

Probing Results for en

  • For En language across all multi-lingual models:
    • Surface property encoded in early layers
    • Syntactic property encoded in middle layers
    • Semantic property encoded in later layers

Surface

Syntactic

Semantic

9 of 17

9

How effectively do multi-lingual models encode hierarchy of linguistic structure for Indic languages?

10 of 17

Probing Results: Universal vs. Indic multi-lingual language models

  • Indic-specific models: best at capturing language properties within the realm of Indic languages, may be due to their targeted training.
  • Universal multi-lingual models: Both encoder-based and decoder-based models show mixed results, may be due to their broader training across languages.

Surface

Syntactic

Semantic

11 of 17

11

IndicSentEval: text perturbations for Indic languages

! Do multilingual models exhibit greater robustness to specific text perturbations while preserving linguistic hierarchies across Indic languages?

  • Indic-specific models are better at encoding linguistic hierarchy across six Indic languages

12 of 17

12

Which multilingual models are more robust to

perturbations in Indic languages?

13 of 17

Perturbation Results: Universal vs. Indic multi-lingual language models

  • Universal multi-lingual models: show greater resilience to perturbations in at least four languages
  • BERT-specific models: display a more significant accuracy drop across all the Indic languages.

Multilingual models vs. Indic languages

Multilingual models vs. Probing tasks

  • Universal multi-lingual models have greater robustness compared to Indic models and mBERT
  • Surface and syntactic probing tasks are significantly impacted by perturbations compared to semantic properties.

14 of 17

14

Which layers are more affected due to text perturbations for Indic languages?

15 of 17

Perturbation Results: layer-wise robustness analysis

  • Surface property: early layers are impacted across Indic languages
  • Syntactic property: middle layers are affected for all languages except te
  • Semantic property: for ur-early to middle layers, while middle to later layers for remaining languages

Probing tasks vs. Indic languages

16 of 17

Conclusions

  1. For En: probing analysis reveals that almost all multilingual models demonstrate consistent encoding performance and similar language hierarchy

  • For Indic languages, Indic-specific multilingual models capture better language hierarchy while universal models show mixed results.

  • Intriguingly, universal models broadly exhibit better robustness compared to Indic-specific models

17 of 17

IndicSentEval: How Effectively do Multilingual Transformer Models encode Linguistic Properties for Indic Languages? (IJCNLP-AACL 2025)

Manish Gupta

Radhika Mamidi

Mounika Marreddy

Akhilesh Aravapalli

Subba Reddy Oota