1 of 21

War in Ukraine

Machine Learning Sentiment Prediction

Final Data Analysis Berkeley Bootcamp Project

�Module 20

2 of 21

February 24 2022 was started Russo-Ukrainian War, which is going to have huge impacts around the world, perhaps even ending the globalized era as we know it. It is imperative that we capture and analyze the massive amounts of data being put out as a result of this war.

Project roles:

  • Dataset - Jeyme Liu

  • DataBase - Veronika Lobkina

  • Remote Server - Jesse Fernandez

  • Machine Learning - Olga Podolska

3 of 21

The Question We Are Asking

Can we predict if a certain tweet about Ukrainian War �is negative?

4 of 21

Structure of Sentiment Prediction project

5 of 21

Extract Transform Load

49.74M tweets on Kaggle, 12 GB

6 of 21

Data Cleaning

  1. Twitter users created before 2008
  2. English Tweets
  3. Standards tweets to alphabet characters
  4. Transform data types example change tweetscreated from text to date
  5. Dropped Unneeded Columns

7 of 21

Pre-Machine Learning

Joined Twitter and Events Data Set

8 of 21

RoBERTa Sentiment Analysis

What is RoBERTa?

  • RoBERTa is an AI developed by the Meta Research team.
  • It’s a model trained on more than 124M tweets (from January 2018 to December 2021) for self-supervised natural language processing (NLP)

What is sentiment analysis?

  • Sentiment analysis (also known as opinion mining) is a natural language processing (NLP) algorithm to identify, extract, and quantify the emotional tone behind a body of text.

9 of 21

Server Setup

Why do we need a server?

  • Resource intensive
  • Limited runtimes in Google Colab
  • Limited resources on personal laptops

Setup:

  • Windows Server 2022
  • Git Bash
  • Anaconda
    • Python
    • ML virtual enviro
    • Jupyter notebook

10 of 21

Machine Learning Model - Step 1 : Prep dataset for RoBERTa ingestion

Step 1:

  • Clean Twitter and Events dataset to have just text
    • Convert Upper case letters to lower case
    • Remove punctuations and special characters
    • Remove emoji unicode
    • Remove numbers

Step 2:

  • Create a Word Cloud just for fun!

11 of 21

Machine Learning Model - Step 2 : RoBERTa to Output Sentiment Weights

Step 3:

  • Run the cleaned text from the Twitter data and run it through RoBERTa

12 of 21

RoBERTa to Output Sentiment Weights

Step 4:

  • Get sentiment scores for the tweets
  • Add the scores back in to Twitter dataset
    • Negative
    • Neutral
    • Positive

13 of 21

Post-RoBERTa Machine Learning

Sentiments Data Set Added

14 of 21

Data Exploration and Visualization

15 of 21

ERD

Relationship

Chart

16 of 21

Final Preprocessing

713009 rows x 22 columns => 710355 rows x 14 columns

136805 hashtags labeled

17 of 21

Final Exploratory Analysis

18 of 21

Supervised Machine Learning Model

Linear Regression Model prediction �for sentiment of each tweet

Linear Regression Model prediction �for the average sentiment of the day

19 of 21

Deep Machine Learning Model

ReLU + Linear

Probability Density Function

20 of 21

LightBGM Classifier Model

0.6527 accuracy score�Confusion Matrix

21 of 21

We will glad to answer your questions!