BFCAI at CoLI-Tunglish@FIRE 2023: Machine Learning Based Model for Word-level Language Identification in Code-mixed Tulu Texts
AHMED MEGAHED; HAMADA NAYEL
DEPARTMENT OF COMPUTER SCIENCE,
FACULTY OF COMPUTERS AND ARTIFICIAL INTELLIGENCE,
BENHA UNIVERSITY
TABLE OF CONTENT
CoLI-Tunglish@FIRE 2023
2
06
01
02
03
04
05
Introduction
Problem Defination
Dataset
Methodology
Results
Conclusion & Future Work
Language Identification
CoLI-Tunglish@FIRE 2023
3
Problem Formulation
CoLI-Tunglish@FIRE 2023
4
mixed languages.
Dataset
CoLI-Tunglish@FIRE 2023
5
The CoLI-Tunglish dataset contains English and Kannada words written in Roman script,
The data is divided into eight categories:
Dataset
CoLI-Tunglish@FIRE 2023
6
The sources of data are extracted from Tulu YouTube video comments to construct Code-mixed Tulu-English Language Identification (CoLI-Tunglish) dataset.
Model Structure
CoLI-Tunglish@FIRE 2023
7
Training Set
Test Set
Feature Engineering
Preprocessing
Training the Model
Model
Output
Methodology:- Preprocessing
CoLI-Tunglish@FIRE 2023
8
Words | Language | Word Length |
Oo | English | 2 |
Anna | Kannada | 4 |
Ninna | Tulu | 5 |
Pukuli | Tulu | 6 |
naddh | Tulu | 5 |
Methodology:- Why Word Length?
CoLI-Tunglish@FIRE 2023
9
Methodology:- Feature Engineering (vectorization)
CoLI-Tunglish@FIRE 2023
10
Methodology:- Algorithms
CoLI-Tunglish@FIRE 2023
11
Results:
CoLI-Tunglish@FIRE 2023
12
Comparison of machine learning algorithms scores on the development set
Results:
CoLI-Tunglish@FIRE 2023
13
Comparison of macro average scores with top ranked teams.
Conclusion and Future work
CoLI-Tunglish@FIRE 2023
14
CoLI-Tunglish@FIRE 2023
15
THANKS!