1 of 42

Depression Prediction from Twitter Using LIWC and LEAPS

2 of 42

Supervisor: Tohedul Islam

Assistant Professor of Department of Computer Science(CS)

American International University-Bangladesh(AIUB)

Name

ID

MD. SAIFUR RAHMAN

17-33944-1

MD. NAKIBUL ISLAM HRIDOY

17-33961-1

MD.AL AMIN

17-34214-1

MD. AREFUR RAHMAN

17-33945-1

3 of 42

Outline

Introduction

Motivation

Applications

Research Questions

Related Work

Methodology

Results

Finding

Conclusion & Future Work

4 of 42

Introduction

Depression

  • Depression is a mental disorder.
  • Depression is a serious condition that negatively affects how a person thinks, feels, and behaves.
  • Suffering greatly and function badly at work, at school and in the family.
  • More than 264 million people affected by the depression through worldwide.

Twitter

  • A social media where people share their thoughts and feelings through photo, text and comment.
  • 330 million monthly active user and 145 million daily active users.

5 of 42

Introduction(Continued)

  • Collected data of 200 depressed and 200 not depressed users.
  • Prepare dataset by pre-processing
  • Used LEAPS, LIWC on the dataset.

6 of 42

Motivation

  • Every year up to 800000 people died because of depression.
  • Around 15 to 29 years old people’s death main reason is depression.
  • Women are affected much through depression more than men.
  • Between 76% to 85% in low and middle-income countries people do not receive any treatment for this mental disorder.
  • The motivation behind our thesis is to utilize user's social media contents to find out a depressed person.
  • Taking counseling before some health hazards emerge.

7 of 42

Application

  • By predict depression it will help us through decreasing death rate, not attempt suicide and recommend them for counseling .
  • Preventing depression will improve our mental health as well as physical health, work performance, creativity and quality of life.

8 of 42

Research Question

  • How can we predict depression by tweets?
  • Which type of attribute should we pick to predict depression?

9 of 42

Related Work

  • “Depression detection via harvesting social media: A multimodal dictionary learning solution.”- Shows social media can be used to predict depression and make a multimodal dictionary learning solution for predict depression.

  • Limitations:
  • Detecting Normal and Depressed for data collection was not standard.
  • Did not use any linguistic analyzer for analyzing data.

10 of 42

Related Work(Continued )

  • “Predicting depression levels using social media posts.”- investigate how social network user’s posts can help to classify users mental health levels with the help of their generated contents.

  • Limitations:
  • Did not preprocess collected data for removing unnecessary information.
  • Did not use any linguistic analyzer for analyzing data.

11 of 42

Related Work(Continued )�

  • “Predicting depression tendency based on image, text and behaviour data from Instagram.”- Applied deep learning to collect sample user and consider both image and textual data to make model.

  • Limitations:
  • Depends more on hyperlinks to preprocess data.
  • Did not use any linguistic analyzer for analyzing data.

12 of 42

Related Work(Continued )

  • “A model for Prediction of human depression using Apriori algorithm.”- Collect data by direct interview and applied Apriori and Associate role mining Concept to extract necessary information.

  • Limitations:
  • Did not preprocess data.
  • Did not use any linguistic analyzer for analyzing data.

13 of 42

Why our work is different?�

  • Data collection is rich enough for the manual search.
  • Affluent lexicon is used to analyze data.
  • A smart tool called LEAPS keeps a great impact on summarizing lexicon result.
  • Different classifier algorithm are applied.

14 of 42

Methodology

User

Selection

Data Collection

Data preprocessing

Categorized

Data using LIWC

Leaps Package

Feature

Selection

Model

Construction

By different

Classifier

Algorithms

Predict

Depression

Figure: Workflow

15 of 42

User Selection ( with depression tendency)

Find users using these keywords:

  • Depressed
  • Anxiety
  • Frustration
  • Tired
  • Money problem
  • Breakup
  • Want to be alone

16 of 42

User Selection ( without depression tendency)

Find users using these keywords:

  • Happiness
  • Having fun
  • Feeling grateful
  • Celebrating
  • Blessed
  • Delightful
  • Wonderful

17 of 42

Example tweets

18 of 42

Filtered Twitter User

Filtered

  • Organizations
  • Twitter bot
  • Retweets of others

19 of 42

Data Collection Using Tweepy

Figure: Python Code to collect data from Twitter using Tweepy

20 of 42

Data Preprocessing

Cleaning

  • Urls removed
  • ReTweet removed
  • Hashtags removed

Filtered

  • Inactive users
  • Business accounts
  • Official Accounts

21 of 42

Feature Selection

  • Feature selection algorithms are those types of algorithms where input variables are filtered by reducing irrelevant, redundant, or noisy features from the dataset.
  • By choosing the right subset improve the accuracy of a model
  • Reduce the complexity of a model.

22 of 42

Feature Selection(Workflow)

Figure : Feature Selection Workflow

23 of 42

Feature Selection(Text Analysis)

  • For analysis purpose we use a text analysis tools called LIWC.
  • Linguistic Inquiry and Word Count(LIWC).
  • LIWC is a crystalline text analysis program that counts words in psychologically meaningful categories.
  • Almost 70 categories

24 of 42

Feature Selection(Picking Most Relevant Attributes)

Regsubsets (Regression subset selection model)

  •  Regression is a statistical method. Relationship between one dependent variable and a series of independent variables.
  • Subset selection refers to the task of finding a small subset of the available independent variables that does a good job of predicting the dependent variable. Exhaustive searches are possible for regressions.
  • Best subset of predictors to include in the model, among all possible subsets of predictors, is referred to as variable selection.

25 of 42

Features Selection(Approaches)

  • LIWC(Linguistic Inquiry and Word Count)

  • Leaps(R package)

26 of 42

Feature Selection(LIWC)

Figure : Analysis tweet’s text by LIWC and its output

27 of 42

Feature Selection (regsubsets function , Leaps)

28 of 42

Feature Selection (regsubsets output)

Table : Relevant Features extract from leaps output.

Nvmax = maximum size of subsets to examine

29 of 42

Relevant Feature Explanations

Table : Explanation selected variables

30 of 42

Model Constructing (Workflow)

Figure : Model Construction

31 of 42

Model Constructing (Classifier generation)

Algorithm to be applied :

  • Naive Bayes
  • PART
  • RANDOM FOREST
  • J48

32 of 42

Model Constructing (Naive Bayes )

33 of 42

Model Constructing (J48 )

34 of 42

Model Constructing (PART )

35 of 42

Model Constructing (RANDOM FOREST)

36 of 42

Classifier Generation (Result)

  • Random Forest has correctly classified instance = 72.25% ( approx. 72%)
  • It has higher TPR,ROC value that others.

TPR

FPR

ROC

CLASS

0.730

0.285

0.781

Yes

0.715

0.270

0.781

No

37 of 42

Test Data Set

38 of 42

Discussion

  • Why choose Random Forest :
  • Detects as best classifier which give us most correct instance value.
  • Use of both classification and regression task.
  • Handles missing values
  • Maintaining accuracy

39 of 42

Overview

  • Our main goal was to predict depressed person by analysing tweets.
  • We collected tweets from different countries.
  • We pre-processed our data and for analyse.
  • We used LIWC a text analysis tools.
  • Leaps in R was applied to LIWC.

  • Applying feature selection algorithms we get most relevant features from leaps output.
  • After constructing classifier four types of algorithm were applied: Naive Bayes. ,Random Forest, J48, PART,
  • We tested their accuracy and consider this things ROC , TPR , correctly classified instances.
  • After that we construct a classifier which could give us most accurate result.

40 of 42

Conclusion & Future Work

  • In our research, we construct classifier to detect depressed person by analysing their tweets.
  • We successfully detect that Random forest is the best classifier than other classifier.
  • In future we will use other text analysis tools like LIWC. Similar tools EMPATH.
  • We would like to work with large amount dataset to improve our work.
  • In future we would like to collect data from people by taking their answer in a real time question-answer session

41 of 42

Final Target

  • Our target is to build a system where we can identify a real depressed person. By accessing their public information which they share when they feel depressed. These types of people need well counseling to improve their mental condition. Right time care can help to improve their condition.

42 of 42

Thank You

Any Questions?