1 of 15

BFCAI at CoLI-Tunglish@FIRE 2023: Machine Learning Based Model for Word-level Language Identification in Code-mixed Tulu Texts

AHMED MEGAHED; HAMADA NAYEL

DEPARTMENT OF COMPUTER SCIENCE,

FACULTY OF COMPUTERS AND ARTIFICIAL INTELLIGENCE,

BENHA UNIVERSITY

2 of 15

TABLE OF CONTENT

CoLI-Tunglish@FIRE 2023

2

06

01

02

03

04

05

Introduction

Problem Defination

Dataset

Methodology

Results

Conclusion & Future Work

3 of 15

Language Identification

CoLI-Tunglish@FIRE 2023

3

  • Automatically recognizing the languages present in a document

  • It’s a critical task in various text processing pipelines.

4 of 15

Problem Formulation

CoLI-Tunglish@FIRE 2023

4

  • The primary goal of the shared task is to develop a new approach for LI in

mixed languages.

  • The task entails dealing with tokens from various categories, including English, Kannada, Mixed-language, Tulu, names, locations, symbol, and other.

  • Each word in the Test set has to be assigned with one of these eight categories.

5 of 15

Dataset

CoLI-Tunglish@FIRE 2023

5

The CoLI-Tunglish dataset contains English and Kannada words written in Roman script,

The data is divided into eight categories:

6 of 15

Dataset

CoLI-Tunglish@FIRE 2023

6

The sources of data are extracted from Tulu YouTube video comments to construct Code-mixed Tulu-English Language Identification (CoLI-Tunglish) dataset.

7 of 15

Model Structure

CoLI-Tunglish@FIRE 2023

7

Training Set

Test Set

Feature Engineering

Preprocessing

Training the Model

Model

Output

8 of 15

Methodology:- Preprocessing

  • By augmenting the dataset to incorporate the additional attribute(word length)
  • This feature provides valuable insights into the relationship between word length and language classification.

CoLI-Tunglish@FIRE 2023

8

Words

Language

Word Length

Oo

English

2

Anna

Kannada

4

Ninna

Tulu

5

Pukuli

Tulu

6

naddh

Tulu

5

9 of 15

Methodology:- Why Word Length?

  • Different languages often exhibit distinct characteristics in terms of word length.
  • English, tend to have shorter words on average, while others, like Mixed may have longer words.
  • The variations in average word length indicate differences in the linguistic structures of these languages

CoLI-Tunglish@FIRE 2023

9

10 of 15

Methodology:- Feature Engineering (vectorization)

  • TF-IDF(Term Frequency-Inverse Document Frequency) has been used as a vector space model for feature extraction.
  • The combination of the TF-IDF and word length leads to better understanding of the core patterns in language identification.

CoLI-Tunglish@FIRE 2023

10

11 of 15

Methodology:- Algorithms

  • Support Vector Machine (SVM)
  • Stochastic Gradient Descent (SGD)
  • K-Nearest Neighbors (KNN)
  • Multi-Layer Perceptron (MLP)

CoLI-Tunglish@FIRE 2023

11

12 of 15

Results:

CoLI-Tunglish@FIRE 2023

12

Comparison of machine learning algorithms scores on the development set

13 of 15

Results:

CoLI-Tunglish@FIRE 2023

13

Comparison of macro average scores with top ranked teams.

14 of 15

Conclusion and Future work

CoLI-Tunglish@FIRE 2023

14

  • Different ML algorithms have been implemented.
  • SVM achieved the highest F1-Score in the CoLI-Tunglish dataset.
  • Our approach achieved the second rank among all other submissions.
  • Deep learning structure can be used in future work.
  • Transfer learning can be applied using word embedding.

15 of 15

CoLI-Tunglish@FIRE 2023

15

THANKS!