1 of 28

Big Data Analytics

CSCI 6502

Big Data Analytics

Product Feature Analysis Based On Amazon Customer Reviews

Aishwarya Satwani, Swapnil Sethi, Sanjeev

2 of 28

Introduction

  • Today, online stores collect a lot of customer feedback in the form of surveys, reviews, and comments.
  • This feedback is categorized, and in some cases responded to, but in general it is underutilized – even though customer satisfaction is essential to the success of their business.

Big Data Analytics

3 of 28

Problem Statement

  • In this project, given reviews for a product, we essentially identify its features/attributes and then extract opinion about those product features and classify them as positive or negative. We will analyze the opinion expressed for each product feature and observe the good and bad features from a customer's perspective.

Big Data Analytics

4 of 28

Dataset

  • Amazon Customer reviews Dataset (~1.5 GB per product)
  • Rich source of information for academic researchers in fields of NLP and ML.
  • Sample view of the dataset:

Big Data Analytics

5 of 28

Tech Stack

  • Python - NLP libraries
  • Spark
  • Visualizations - Tableau

Big Data Analytics

6 of 28

Infrastructure

Set-up

  • Spark
  • AWS
  • PostgreSQL

Big Data Analytics

7 of 28

POC - AWS

Big Data Analytics

8 of 28

Database - PostgreSQL

Big Data Analytics

9 of 28

Data Wrangling

Big Data Analytics

10 of 28

11 of 28

Data Overview/Sampling

Big Data Analytics

12 of 28

Binary Classification

Big Data Analytics

x1

Count the number of positive lexicon

3

x2

Count the number of negative lexicon

2

x3

1 - If the word “no” is present

0 - Otherwise

1

x4

Count 1st and 2nd pronouns

3

x5

1 - If the word “!” is present

0 - Otherwise

0

x6

Log of the word count of the review

 

13 of 28

  • Processing the reviews and extracting feature vectors that are appropriate for use with Logistic regression then finding the set of weights that are effective
  • Implementation of Stochastic Gradient Descent using cross entropy as the loss function

Big Data Analytics

14 of 28

Topic Modeling with LDA

  • The purpose of topic modeling is to extract “topics” from a collection of documents.
  • The goal is to discover a list of meaningful, non-overlapping, exhaustive “topics” that best reflect the product features and aspects that customers commented on.

Big Data Analytics

15 of 28

  • Latent Dirichlet Allocation is a generative probabilistic model for a collection of documents (corpus).
  • Each document is generated from a random mixture of latent topics and each word in a document is generated from multinomial distributions conditioned on topics.
  • In the model output, each topic is represented by a weighted list of words, and each document is assigned a weighted list of topics, where the weights represent multinomial probabilities.

Big Data Analytics

16 of 28

Big Data Analytics

17 of 28

18 of 28

Feature Sentiment Analysis

  • Most times when we consider sentiment analysis, we look at the overall positive/negative scores.
  • There are times when you want your sentiment analysis to be aspect-based, or otherwise called topic-based.

Big Data Analytics

19 of 28

  • We have employed Parts of Speech (POS) tagging for this, along with paying attention to the child tokens, so that we’re able to pick up intensifiers such as “very”, “quite”, and more.
  • For example

Big Data Analytics

20 of 28

  • We then perform sentiment analysis on its description.
  • We have used TextBlob library for this. It has a bag-of-words approach, meaning that it has a list of words such as “good”, “bad”, and “great” that have a sentiment score attached to them. It is also able to pick up modifiers (such as “not”) and intensifiers (such as “very”) that affect the sentiment score.

Big Data Analytics

21 of 28

Visualization

Big Data Analytics

22 of 28

Big Data Analytics

23 of 28

Big Data Analytics

24 of 28

Evaluation

  • We calculated the average sentiment score from our model and compared it with the star-rating given by the user.
  • If the overall sentiment of the review containing the product feature is +ve then we verify that the star-provided by the same user is 4/5 or 5/5 stars.

Big Data Analytics

25 of 28

Results

    • Precision: 0.920
    • Recall: 1.000
    • Accuracy: 0.920
    • F1 Score: 0.959

Big Data Analytics

26 of 28

Future Enhancements

  • Dealing with the review biases
  • Overcome dataset limitations
  • Make the topic modeling more independent
  • Evaluate the model against other existing solutions in the market.

Big Data Analytics

27 of 28

Citations

28 of 28

Thank you!

Big Data Analytics