Word-level Language Identification in Dravidian Languages (CoLi-Dravidian)
16th Forum for Information Retrieval Evaluation (FIRE 2024)
DA-IICT, Gandhinagar, India
Organizing Team
Asha Hegde
Department of Computer Science, Mangalore University, India
Fazlourrahman Balouchzahi
CIC, IPN, Mexico
Sabur Butt
IFE, Tecnologico de Monterrey, Mexico
Sharal Coelho
Department of Computer Science, Mangalore University, India
Kavya G
Department of Computer Science, Mangalore University, India
Harshitha S Kumar
Department of Computer Science, Mangalore University, India
Sonith D
Department of Computer Science, Mangalore University, India
H.L. Shashirekha
Department of Computer Science, Mangalore University, India
Ameeta Agrawal
Department of Computer Science, Portland State University, USA
Outline
1
Introduction
2
Objective
3
Task Description
4
Evaluation Metrics
5
Baselines
6
Overview of the Submitted Systems
7
Conclusion
Introduction
Code-mixing:
Word-level Language Identification (LI)
It aims to identify the language of individual words within a given sentence.
Though there are several tools/models for word-level LI for high-resource languages, under-resourced languages like Tulu, Kannada etc., are less explored due to lack of annotated data.
Dravidian language family
It includes languages such as Tamil, Kannada, Tulu, and Malayalam.
In the context of Dravidian languages, code-mixing often involves alternating between English and regional languages.
Objective
Research
Inviting researchers to develop models capable of classifying words in code-mixed texts involving Dravidian languages
Address the challenges
To address the challenges of word-level LI in code-mixed Dravidian languages - Tamil, Kannada, Malayalam, and Tulu,
Task Description
The CoLi-Dravidian dataset comprises:
Conventionally, the word-level LI datasets are imbalanced, and Figure 1 shows the classwise distribution of CoLi-Dravidian datasets.
Samples of Tamil data
Samples of Malayalam data
Samples of tulu data
Samples of Different Datasets
Samples of Kannada data
Evaluation Metrics
CoLI-Dravidian datasets exhibit an imbalanced label distribution across their classes, making the evaluation process crucial.
To assess the performance of the submitted models, both macro-averaged and weighted-averaged F1 scores are used, as these metrics are well-suited for imbalanced datasets.
Baselines
Support Vector Machines (SVMs)
Effective in high-dimensional spaces, suitable for word embeddings. Computationally expensive for very large datasets.
Decision Trees (DTs)
Offer an interpretable model, easily visualizing the decision-making process. Prone to overfitting, especially with complex linguistic features.
Logistic Regression (LR)
Provides a computationally efficient approach. May not capture intricate data relationships as effectively as SVMs or more sophisticated models.
Table 1. Rank lists of Tamil and Kannada languages
Task Description - Results
Table 2. Rank lists of Malayalam and Tulu languages
Overview of the Submitted Systems
Team PonsubashRaj
Excellent accuracy and efficiency with a novel model.
Team Kaivalya
High accuracy, but with slightly lower efficiency.
Team NLPnorth
Unique model design with promising results.
Team Awasthama
Consistent accuracy and efficiency across metrics.
.
Overview of the Submitted Systems
TEAMS
Approaches
Overall, the majority of participating teams experimented various traditional ML techniques, while significant number of participants chose DL and transformer models and the same is reflected in Figure.
Conclusion
1
The task is focused on four low-resource Dravidian languages - Tamil, Kannada, Malayalam, and Tulu, intertwined with English.
2
By using the datasets of this shared task, researchers can focus on adding more context and improving transformer models to better understand the unique details of Dravidian languages in real-world tasks like SA,MT, and monitoring social media.
3
The shared task's outcomes emphasize the importance of continued research into code-mixed LI, which is crucial for preserving linguistic diversity in the digital age.
Reference
[1] F Balouchzahi, S Butt, A Hegde, N Ashraf, HL Shashirekha, G Sidorov, and A Gelbukh. Overview of CoLI-Kanglish: Word Level Language Identification in Code-mixed Kannada-English Texts at ICON 2022. Shared Task on Word Level Language Identification in Code-mixed Kannada-English Texts, page 38, 2022.
[2] Asha Hegde, F Balouchzahi, Sharal Coelho, Shashirekha H L, Hamada A Nayel, and Sabur Butt. CoLI@FIRE2023: Findings of Word-level Lan_x0002_guage Identification in Code-mixed Tulu Text. In Proceedings of the
15th Annual Meeting of the Forum for Information Retrieval Evaluation, FIRE ’23, page 25–26. Association for Computing Machinery, 2024.
[3] Asha Hegde, F Balouchzahi, Sharal Coelho, HL Shashirekha, Hamada A Nayel, and Sabur Butt. Overview of CoLI-Tunglish: Word-level Language Identification in Code-mixed Tulu Text at FIRE 2023. In FIRE
(Working Notes), pages 179–190, 2023.
[4] Fazlourrahman Balouchzahi, Hosahalli Lakshmaiah Shashirekha, Grigori Sidorov, and Alexander Gelbukh. A Comparative Study of Syllables and Character Level N-grams for Dravidian Multi-script and Code_x0002_
mixed Offensive Language Identification. Journal of Intelligent & Fuzzy Systems, 43(6):6995–7005, 2022.
[5] Asha Hegde and Hosahalli Lakshmaiah Shashirekha. Syllable-Level Morphological Segmentation of Kannada and Tulu Words. Automatic Speech Recognition and Translation for Low Resource Languages, pages
113–133, 2024.
[6] Asha Hegde, Shubhanker Banerjee, Bharathi Raja Chakravarthi, Ruba Priyadharshini, Hosahalli Shashirekha, John Philip McCrae, et al. Overview of the shared task on machine translation in Dravidian
languages. In Proceedings of the second workshop on speech and language technologies for Dravidian languages, pages 271–278, 2022.
[7] Hegde Asha and Shashirekha Hosahalli Lakshmaiah. Kt2: Kannada tulu parallel corpus construction for neural machine translation. In Proceedings of the 20th International Conference on Natural Language
Processing (ICON), pages 743–753, 2023.
[8] Asha Hegde and Hosahalli Lakshmaiah Shashirekha. Kansan: Kannada sanskrit parallel corpus construction for machine translation. In International Conference on Speech and Language Technologies for Low-resource Languages, pages 3–18. Springer, 2022.
[9] Shashirekha Hosahalli Lakshmaiah, Fazlourrahman Balouchzahi, Mudoor Devadas Anusha, and Grigori Sidorov. Coli-machine learning approaches for code-mixed language identification at the word level in kannada-english texts. Acta Polytechnica Hungarica, 19(10), 2022.
[10] Asha Hegde, Fazlourrahman Balouchzahi, Sabur Butt, Sharal Coelho, Kavya G, Harshitha S Kumar, Sonith D, Shashirekha Hosahalli Lakshmaiah, and Ameeta Agrawal. Overview of CoLI-Dravidian: Word-level Code-mixed Language Identification in Dravidian Languages. In Forum for Information Retrieval Evaluation FIRE - 2024, 2024.
Thank You