1 of 19

Word-level Language Identification in Dravidian Languages (CoLi-Dravidian)

16th Forum for Information Retrieval Evaluation (FIRE 2024)

DA-IICT, Gandhinagar, India

2 of 19

Organizing Team

Asha Hegde

Department of Computer Science, Mangalore University, India

Fazlourrahman Balouchzahi

CIC, IPN, Mexico

Sabur Butt

IFE, Tecnologico de Monterrey, Mexico

Sharal Coelho

Department of Computer Science, Mangalore University, India

Kavya G

Department of Computer Science, Mangalore University, India

Harshitha S Kumar

Department of Computer Science, Mangalore University, India

Sonith D

Department of Computer Science, Mangalore University, India

H.L. Shashirekha

Department of Computer Science, Mangalore University, India

Ameeta Agrawal

Department of Computer Science, Portland State University, USA

3 of 19

Outline

1

Introduction

2

Objective

3

Task Description

4

Evaluation Metrics

5

Baselines

6

Overview of the Submitted Systems

7

Conclusion

4 of 19

Introduction

Code-mixing:

  • A linguistic phenomenon where multiple languages are blended within a single text.
  • It has become increasingly prevalent in multilingual societies, particularly in digital communication.

Word-level Language Identification (LI)

It aims to identify the language of individual words within a given sentence.

Though there are several tools/models for word-level LI for high-resource languages, under-resourced languages like Tulu, Kannada etc., are less explored due to lack of annotated data.

Dravidian language family

It includes languages such as Tamil, Kannada, Tulu, and Malayalam.

  • These languages are characterized by unique morphological structures, and significant linguistic variations that demand specialized computational methodologies beyond traditional NLP frameworks

In the context of Dravidian languages, code-mixing often involves alternating between English and regional languages.

5 of 19

Objective

Research

Inviting researchers to develop models capable of classifying words in code-mixed texts involving Dravidian languages

Address the challenges

To address the challenges of word-level LI in code-mixed Dravidian languages - Tamil, Kannada, Malayalam, and Tulu,

6 of 19

Task Description

  • The CoLi-Dravidian shared task, an extension of the CoLiKanglish [1] and CoLi-Tunglish [3] shared tasks, invites researchers to develop models for word-level LI using the CoLI-Dravidian datasets.
  • This dataset consists of code-mixed YouTube comments in four Dravidian languages: Tamil, Kannada, Malayalam, and Tulu.
  • The comments are pre-processed by removing punctuation and control characters and then tokenized into words.
  • Afterward, the comments are romanized using the libindic library. (Latin Script)
  • Annotation —Native speakers of the respective languages, fluent in English, manually annotate the words for the task.

7 of 19

The CoLi-Dravidian dataset comprises:

  • Language classes (Tamil/Kannada/Malayalam/Tulu and English)
  • ‘Number’ for digits
  • ‘Name’ for person names
  • ‘Location’ for geographical locations
  • ‘Mixed’ for code-mixed words formed by blending Dravidian languages and English at the word or sub-word level
  • ‘Other’ for unclassified terms.

Conventionally, the word-level LI datasets are imbalanced, and Figure 1 shows the classwise distribution of CoLi-Dravidian datasets.

8 of 19

9 of 19

Samples of Tamil data

Samples of Malayalam data

Samples of tulu data

Samples of Different Datasets

Samples of Kannada data

10 of 19

Evaluation Metrics

CoLI-Dravidian datasets exhibit an imbalanced label distribution across their classes, making the evaluation process crucial.

To assess the performance of the submitted models, both macro-averaged and weighted-averaged F1 scores are used, as these metrics are well-suited for imbalanced datasets.

  • The macro F1 score treats all classes equally, while the weighted F1 score accounts for class imbalances by giving more importance to larger classes.
  • These evaluation metrics ensure a fair assessment of the models' performance across all classes, regardless of the class distribution.

11 of 19

Baselines

Support Vector Machines (SVMs)

Effective in high-dimensional spaces, suitable for word embeddings. Computationally expensive for very large datasets.

Decision Trees (DTs)

Offer an interpretable model, easily visualizing the decision-making process. Prone to overfitting, especially with complex linguistic features.

Logistic Regression (LR)

Provides a computationally efficient approach. May not capture intricate data relationships as effectively as SVMs or more sophisticated models.

12 of 19

Table 1. Rank lists of Tamil and Kannada languages

Task Description - Results

13 of 19

Table 2. Rank lists of Malayalam and Tulu languages

14 of 19

Overview of the Submitted Systems

Team PonsubashRaj

Excellent accuracy and efficiency with a novel model.

Team Kaivalya

High accuracy, but with slightly lower efficiency.

Team NLPnorth

Unique model design with promising results.

Team Awasthama

Consistent accuracy and efficiency across metrics.

15 of 19

  • Ten teams from three different countries (India, Iran, and Denmark) and various types of institutions (university, research center, and industry) submitted their predictions to the shared task.

  • Their predictions are based on ML, transfer learning, and DL approaches and the shared task overview paper presents the in-depth evaluations of the submitted approaches.

  • Among 10 teams who submitted their predictions for the shared task
  • 8 teams submitted the working notes of their models

  • The top-performing models achieved macro F1 scores of
    • 0.7656 - Tamil
    • 0.9293 - Kannada
    • 0.8939 - Malayalam
    • 0.8678 -Tulu

  • Remarkably, the best score is achieved by the MuRiL model for Kannada,
  • The MACHAMP model with an additional CRF layer for Malayalam and Tulu languages, and the voting classifier with LR, DT, and SVM for Tamil.

.

Overview of the Submitted Systems

16 of 19

TEAMS

  • MUCS
  • MUCSNLPLab
  • abadian
  • denis_gordeev
  • Srihari V K
  • TextTitans
  • Ram
  • CUFE
  • PNB

Approaches

Overall, the majority of participating teams experimented various traditional ML techniques, while significant number of participants chose DL and transformer models and the same is reflected in Figure.

17 of 19

Conclusion

1

The task is focused on four low-resource Dravidian languages - Tamil, Kannada, Malayalam, and Tulu, intertwined with English.

2

By using the datasets of this shared task, researchers can focus on adding more context and improving transformer models to better understand the unique details of Dravidian languages in real-world tasks like SA,MT, and monitoring social media.

3

The shared task's outcomes emphasize the importance of continued research into code-mixed LI, which is crucial for preserving linguistic diversity in the digital age.

18 of 19

Reference

[1] F Balouchzahi, S Butt, A Hegde, N Ashraf, HL Shashirekha, G Sidorov, and A Gelbukh. Overview of CoLI-Kanglish: Word Level Language Identification in Code-mixed Kannada-English Texts at ICON 2022. Shared Task on Word Level Language Identification in Code-mixed Kannada-English Texts, page 38, 2022.

[2] Asha Hegde, F Balouchzahi, Sharal Coelho, Shashirekha H L, Hamada A Nayel, and Sabur Butt. CoLI@FIRE2023: Findings of Word-level Lan_x0002_guage Identification in Code-mixed Tulu Text. In Proceedings of the

15th Annual Meeting of the Forum for Information Retrieval Evaluation, FIRE ’23, page 25–26. Association for Computing Machinery, 2024.

[3] Asha Hegde, F Balouchzahi, Sharal Coelho, HL Shashirekha, Hamada A Nayel, and Sabur Butt. Overview of CoLI-Tunglish: Word-level Language Identification in Code-mixed Tulu Text at FIRE 2023. In FIRE

(Working Notes), pages 179–190, 2023.

[4] Fazlourrahman Balouchzahi, Hosahalli Lakshmaiah Shashirekha, Grigori Sidorov, and Alexander Gelbukh. A Comparative Study of Syllables and Character Level N-grams for Dravidian Multi-script and Code_x0002_

mixed Offensive Language Identification. Journal of Intelligent & Fuzzy Systems, 43(6):6995–7005, 2022.

[5] Asha Hegde and Hosahalli Lakshmaiah Shashirekha. Syllable-Level Morphological Segmentation of Kannada and Tulu Words. Automatic Speech Recognition and Translation for Low Resource Languages, pages

113–133, 2024.

[6] Asha Hegde, Shubhanker Banerjee, Bharathi Raja Chakravarthi, Ruba Priyadharshini, Hosahalli Shashirekha, John Philip McCrae, et al. Overview of the shared task on machine translation in Dravidian

languages. In Proceedings of the second workshop on speech and language technologies for Dravidian languages, pages 271–278, 2022.

[7] Hegde Asha and Shashirekha Hosahalli Lakshmaiah. Kt2: Kannada tulu parallel corpus construction for neural machine translation. In Proceedings of the 20th International Conference on Natural Language

Processing (ICON), pages 743–753, 2023.

[8] Asha Hegde and Hosahalli Lakshmaiah Shashirekha. Kansan: Kannada sanskrit parallel corpus construction for machine translation. In International Conference on Speech and Language Technologies for Low-resource Languages, pages 3–18. Springer, 2022.

[9] Shashirekha Hosahalli Lakshmaiah, Fazlourrahman Balouchzahi, Mudoor Devadas Anusha, and Grigori Sidorov. Coli-machine learning approaches for code-mixed language identification at the word level in kannada-english texts. Acta Polytechnica Hungarica, 19(10), 2022.

[10] Asha Hegde, Fazlourrahman Balouchzahi, Sabur Butt, Sharal Coelho, Kavya G, Harshitha S Kumar, Sonith D, Shashirekha Hosahalli Lakshmaiah, and Ameeta Agrawal. Overview of CoLI-Dravidian: Word-level Code-mixed Language Identification in Dravidian Languages. In Forum for Information Retrieval Evaluation FIRE - 2024, 2024.

19 of 19

Thank You