1 of 39

Ice Day Presentation

2 of 39

Bootstrapping NLP for Hinglish

Presented by:

R214220750 - Neomi Sule

R214220499 - Harshit

R2142201919 - Sarthak Patel

R214220770 - Nitesh Vishwakarma

Mentored By:

Mr. Bikram Pratim Bhuyan

3 of 39

Problem Statement

Nowadays, people decide to get something or go somewhere based on other people's reviews. Primarily when it is written in Hinglish, this mix-coded language makes it difficult to segregate them into their respective sentiments. Even in research and data analysis work, most of the real-time data contains Hinglish text due to the hike in the use of Hinglish language. Hence, our project aims to be helpful in all these cases.

4 of 39

Motivation

  • This project has better expansion scope.

  • Less work done on Hinglish dataset analysis in past.

  • Lack of previous work done in this sector this can be later diverted to research oriented work.

5 of 39

Technical Background of Project

This project uses two different approaches.

1) Self-developed techniques

2) Utilizing transformers

1) Self-developed techniques include letter-by-letter transliterations and the creation of erroneous words for increased robustness.

2) Transformers are pretrained models on massive datasets that undergo countless hours of training to produce a useful set of trained weights for the required purpose.

Diberta-large and Roberta-large:

These were two of the transformers utilised in the study to obtain predictions on translated English texts.

Indic-Trans and Google-Trans:

These were utilised for transliteration and translation purposes.

6 of 39

Technical Concepts and Algorithms Used

English Pipeline:

Traditional Mathematical Model

Pseudo code

pos ← 0

neg ← 0

sentence ← “some input”

for id in sentence do

if id = ‘positive’

pos ← pos +1

if id = ‘negative’

neg←neg +1

result ← pos - neg

If result<0

classify(“negative”)

Else

classify(“positive”)

Training Model

Steps

1→ Raw Data

2→ Creation of embedding (Word2Vec)

3→ Training various models (Logistic Regression)

4→ Sentiment Classification

7 of 39

Hinglish Pipeline:

1) Using Indic-Trans, transliterate - Romanized Hindi into Devnagri.

2) Generating Hinglish texts using our existing English vocabulary and comparing the input texts' similarities to Hinglish text.

3) Using an artificially-induced change in vowels and artificially-induced mistaken alterations for typos will build a model to predict the original word.

Text Pre-processing:

  • Removing punctuations like . , ! $( ) * % @
  • Removing URLs
  • Removing Stop words
  • Lower casing
  • Tokenization
  • Lemmatization

8 of 39

Social media monitoring :

With the help of sentiment analysis software , you can wade through all that data in minutes, to analyze individual emotions and overall public sentiment on every social platform.

Product analysis :

Find out what the public is saying about a new product right after launch, or analyze years of feedback you may have never seen.

Market and competitor research :

Use sentiment analysis for market and competitor research. Find out who’s receiving positive mentions  among your competitors, and how your marketing efforts compare

Area of Application

9 of 39

Dataset and Input Format

We used dataset from kaggle and paper with code and the input format is in excel and csv file.

https://www.kaggle.com/datasets/kazanova/sentiment140

https://zenodo.org/record/3974927

10 of 39

Literature Review

1. In (Biradar et al., 2022), researchers shows that applied sentimental analysis with unsupervised clustering of data into specific domains and supervised machine learning techniques handle large amounts of twitter data in an efficient way. The developed tool 1.5 times faster than that of traditional database to Hadoop cluster and also the accuracy is nearly 80 %, which helps the user in computing, analyzing and interpreting interaction and associations between people, topics and ideas.

2. In (Mir et al., 2022), researchers Nostalgic investigation is frequently performed on text based information to assist organizations with observing brand and item slant in client criticism and comprehend client needs. Nostalgic investigation acts like an incredible asset for clients to separate the needful data just as to total the aggregate assessments of the surveys. Nostalgic investigation is likewise used to extricate information from web-based media stages, for example, twitter and so forth and dissect the content.

3. In (Shrestha et al., 2020), researchers presents a practical dynamic approach on to find the polarity of any sentence and analyse the opinion of the particular sentence. The proposed Sentimental Analysis of Hindi (SAH) script have adopted two different classifier Naïve Bayes Classifier and Decision Tree Classifier is used for the text extraction. The positive, neutral and negative result validation shows a comparative result of sentimental analysis.

11 of 39

4. In (Kumar et al., 2011), researchers discuss the development of an aggression tagset and an annotated corpus of Hindi-English code-mixed data from two of the most popular social networking / social media platforms in India – Twitter and Facebook. The corpus is annotated using a hierarchical tagset of 3 top-level tags and 10 level 2 tags. The final dataset contains approximately 18k tweets and 21k facebook comments and is being released for further research in the field.

5. In (Varma et al., 2021), researchers introduce an annotated dataset for Sentiment Analysis in CMTET. Also, they report an accuracy of 80.22% on this dataset using novel unsupervised data normalization with a Multilayer Perceptron (MLP) model. This proposed data normalization technique can be extended to any NLP task involving CMTET. Further, they report an increase of 2.53% accuracy due to this data normalization approach in our best model.

6. In (Sharma et al., 2022), researchers proposed work focuses on analyzing hate speech in Hindi-English code-switched language. Our method explores transformation techniques to capture precise text representation. To contain the structure of data and yet use it with existing algorithms, they developed ‘MoH’ or (Map Only Hindi), which means ‘Love’ in Hindi. ‘MoH’ pipeline which consists of language identification, Roman to Devanagari Hindi transliteration using a knowledge base of Roman Hindi words, and finally employs the fine-tuned Multilingual Bert, and MuRIL language models.

12 of 39

SWOT Analysis

13 of 39

Objective

Main Objective

Research outcome for Sentiment Analysis on Hinglish

Sub Objective

1. To implement an existing system for English language.

2. Develop and implement a basic algorithm to interpret Hinglish tokens.

3. Integrate Hinglish module in the existing system.

14 of 39

Methodology

Reference Software model

15 of 39

Steps

Our methods can improve upon the older NLP concepts utilized for sentiment analysis. The inputs are processed sequentially by many pipelines that make our technique.

1. Hinglish Pipeline

This pipeline will take the raw inputs and pre-process them for easy processing of the Machine learning models.

As the first step, we will gather and clean up our English dataset. Next, a trained model will analyze the input and assign a language label to each word in the sentence. The raw data is then split into parts and worked on separately in either English or Hinglish.

Hinglish corpora goes through normalization. We have employed three different techniques for normalization. The one with the best results will be selected.

These three methods are:

1) Using Indic-Trans, transliterate - Romanized Hindi into Devanagari

2) Generating Hinglish texts using our existing English vocabulary and comparing the input texts' similarities to Hinglish text.

3) Using an artificially-induced change in vowels and artificially-induced mistaken alterations for typos will build a model to predict the original word.

Transliterate both the Roman Hinglish and the Devnagri Hindi (converted from Hinglish) portions into complete Devnagri Hindi. Last but not least, we combine the altered corpora and send them to the English pipeline.

16 of 39

2. English Pipeline

For better results, this pipeline will take the inputs from the previous Hindi layer and further pre-process them.

We will begin by gathering and cleaning our English dataset, then lemmatize the dataset to get the root words of the input text, convert the meaning of the emoji, and remove URLs, tags and new lines. Finally, we calculate accuracy using a traditional mathematical model. However, we planned to change this model to incorporate neural networks and transformers to achieve better results.

17 of 39

Timeline

18 of 39

Working Model

Requirement analysis (Link of SRS)

https://docs.google.com/document/d/1oqa0pW-DE09NmIxdcSnAFK6rATl6CKLfNi7Z5qfr0QU/edit#

Technical Diagram

English Pipeline:

19 of 39

Hinglish Dataset :

20 of 39

Full Workflow:

21 of 39

22 of 39

Input Driven :

23 of 39

Logistic Regression Model on Word2Vec:

24 of 39

25 of 39

Erroneous Text :

26 of 39

Tests Cases

sentence chosen for test case :

“tum bhai kaise ho all good?"

Results

27 of 39

Method 1 :

In this method, we first check for hinglish vocab from the sentence, and change it with the key of the adjacency value. For the unchanged value, we find levenshtein distance greater than 2 and add it into hinglish vocab. We change English words in Hindi and Hindi to Devnagri and then the entire sentence is transliterated to English.

28 of 39

Method 2 :

In this method, we first check for english vocab from the sentence, and change it with the key of the adjacency value. For the unchanged value, we check with Hinglish vocab and we use phonemes and levenshtein distance <=2 on it and replace it with first satisfied word for both English and Hinglish vocab. Further we use pyenchant for suggestions on still unchanged words.

29 of 39

Outcome

Method 1 One Shot Accuracy Without Training - 49

Method 2 One Shot Accuracy Without Training - 50.01

30 of 39

Training Score :

31 of 39

Comparative Studies :

1st method for one shot prediction on test dataset :

2nd method for one shot prediction on test dataset :

Experiment 1

32 of 39

Experiment 4

Experiment 3

Experiment 2

33 of 39

Deberta on semeval unprocessed dataset with train eval test :

Merged dataset with 1st method translated 14k + 2nd method translated 14k

Training accuracy :

34 of 39

Translated 14k method 1+ translated 14k method 2 + processed hinglish 14k train = 45k dataset

35 of 39

Top best accuracy : 75%

Our accuracy : 64%

36 of 39

  • We wanted to implement an existing system to understand how sentiment analysis actually works on the global language, English.

  • We wanted to develop and implement basic algorithm on a code mixed language, Hinglish in our case, to edge up the sentiment analysis which is done on just one language.

  • We wanted to integrate Hinglish module in the existing system for it to be helpful in review analysis, social media monitoring, getting feedbacks on services to name a few.

Justification Of Objectives

37 of 39

  • Multilingual

  • Improve vocabulary

  • Developing it into a google extension

  • Improving erroneous word generation function like including letter replacement with nearby keyboard configuration.

  • Improve time and space complexity of the algorithm

Future Scope

38 of 39

Reference

  • Biradar, S. H., Gorabal, J. V., & Gupta, G. (2022). Machine learning tool for exploring sentiment analysis on twitter data. Materials Today: Proceedings, 56, 1927–1934. https://doi.org/10.1016/j.matpr.2021.11.199
  • Mir, F. A., Singh, R. P., & Mehra, D. M. (2022). Sentimental Analysis of a Sentence. International Journal of Innovative Research in Computer Science & Technology, 1, 10–14.https://doi.org/10.55524/ijircst.2022.10.1.3
  • Shrestha, H., Dhasarathan, C., Munisamy, S., & Jayavel, A. (2020). Natural Language Processing Based Sentimental Analysis of Hindi (SAH) Script an Optimization Approach.
  • International Journal of Speech Technology, 23(4), 757–766. https://doi.org/10.1007/s10772-020-09730-x
  • Kumar, R., Reganti, A. N., Bhatia, A., & Maheshwari, T. (2011). Aggression-annotated Corpus of Hindi-English Code-mixed Data.
  • Varma, K. S. S., Sathineni, P., & Mamidi, R. (2021). Sentiment Analysis in Code-Mixed Telugu-English Text with Unsupervised Data Normalization. International Conference Recent Advances in Natural Language Processing, RANLP, 753–760. https://doi.org/10.26615/978-954-452-072-4_086
  • Sharma, A., Kabra, A., & Jain, M. (2022). Ceasing hate with MoH: Hate Speech Detection in Hindi–English code-switched language. Information Processing and Management, 59(1). https://doi.org/10.1016/j.ipm.2021.102760
  • Bhat, I. A., Mujadia, V., Tammewar, A., Bhat, R. A., & Shrivastava, M. (2014). IIIT-H system submission for FIRE2014 shared task on transliterated search. ACM International Conference Proceeding Series, 05-07-Dec-2014, 48–53. https://doi.org/10.1145/2824864.2824872
  • Joshi, R. B. (2022). L3Cube-HindBERT and DevBERT : Pre-Trained BERT Transformer models for Devanagari based Hindi and Marathi Languages L3Cube-HindBERT and DevBERT : Pre-Trained BERT Transformer models for Devanagari based Hindi and Marathi Languages . October, 10–15. https://doi.org/10.13140/RG.2.2.14606.84809
  • Mathur, P., Sawhney, R., Ayyar, M., & Shah, R. R. (2018). Did you offend me? Classification of Offensive Tweets in Hinglish Language. 2nd Workshop on Abusive Language Online - Proceedings of the Workshop, Co-Located with EMNLP 2018, 138–148. https://doi.org/10.18653/v1/w18-5118

39 of 39

Thank You