1 of 26

Online Booksellers

Deep-Dive into ML Classification

Zain Raza

May 13, 2020

2 of 26

Project Summary

Objective

Goals

Solution

Project Outline

3 of 26

Objective

Flipkart vs. Amazon:

  • Flipkart: widely successful e-commerce startup
  • Main category: books
  • acquired for $18B USD in 2018
  • BEAT out Amazon for #1 e-commerce in India
    • What was so different about books sold on Flipkart?
    • Would a machine be able to tell the difference?

📚📚🤖📚📚

4 of 26

Goals

How does Flipkart differentiate?

The goal of this experiment is to uncover differences on how books are sold between Amazon and Flipkart.

This insight can inform the way smart products work, such as recommender-systems that direct consumers to one platform or another.

5 of 26

Solution

Using Logistic Regression, a classification model was built, with an F1-Score of approximately 0.7034.

Input:

Book listed online for sale

Output:

Classification of who the bookseller was - Flipkart or Amazon?

6 of 26

Project Outline

The experiment was divided into performing the following tasks:

  1. Exploration of the Dataset
  2. Data Preprocessing
  3. Implementation of Classification Models
  4. Comparison of Models

7 of 26

Technologies and Tools

The Environment

8 of 26

The Environment

This experiment was made possible using the following software tools:

  • Python 3.7.6: programming language�
  • Jupyter Notebook: executes Python commands in the browser�
  • SciKit-Learn: framework for implementing Machine Learning algorithms

Full list of imported dependencies

9 of 26

Exploratory Data Analysis

The Dataset

Scatter Plots

Feature Distributions

Correlation Heatmap

10 of 26

The Dataset

“Amazon Vs. Flipkart Book Prices” on Kaggle

Compares 1,382 of the same books�sold differently on Flipkart and Amazon

Continuous, numerical features analyzed:

  • Price of book (USD)
  • Average Rating
  • Number of Reviews

Dataset posted by: mandan

11 of 26

Scatter Plots of Book Features, by Company

Books on Amazon

Books on Flipkart

12 of 26

Distribution of Book Price

13 of 26

Distribution of Book Rating

14 of 26

Distribution of Book Reviews Count

15 of 26

Correlation

Heatmap

16 of 26

Methodology

Modification of the Dataset

17 of 26

Modification of the Dataset

  • Dimensionality
    • Features clearly low in variance, so I decided to compare PCA vs. all 3 features in ML models
  • Data Preprocessing
    • Comparing data normalized vs data scaled to a distribution 0 - 1
    • Outliers removed using IQR
    • Classes equally weighted
  • Train-Testing Split
    • 75%/25%
  • Model Evaluation
    • Emphasis on F1-Score
    • 5 Fold Cross Validation to Compare Different Models Types

18 of 26

The Classifier Models

K-Nearest Neighbors

Support Vector Machine

Logistic Regression

Decision Tree

Random Forest

19 of 26

K-Nearest Neighbors

The best results of the KNN models --->

  • F1-Score: 0.8799
  • K = 1
  • non-reduced data (3 dimensions)
  • Normalized to a �Standard Distribution
  • Probably overfitted, based on�visualization of model trained on �principal components�(F1-Score: only 0.865) ------------------------>

20 of 26

Support Vector Machine

Hyperparameters tuned using grid search: C = 1; gamma = 1 (RBF kernel)�Data normalized to a Z-distribution

(least overfitted)

21 of 26

Logistic Regression

Conclusions:�- Best results on PCA components, normal distribution�- Slight decrease in Recall, increase in F1-score from MinMax

22 of 26

Decision Tree

Problems in binarizing the data,�because the ranges were very similar�(lack of variance to begin with)

No Scaling Applied, all categorical values

Feature Importance:

  1. Review Count
  2. Rating
  3. Price

23 of 26

Random Forest

NO modification applied to dataset

  • Scaling, binarizing, PCA, etc.

100 Estimators

Feature Ranking Differences from 1�decision tree

Better results than 1 decision tree!

24 of 26

Final Conclusions

Cross Validation

Conclusions/Future Improvements

25 of 26

Cross Validation

Using 5 fold cross validation,�the Logistic Regression model �showed the least variance in�its F1-Score.

The K-Nearest Neighbors classifier�showed the highest mean�F1-Score.

26 of 26

Final Conclusion/Future Improvements

Most Accurate: Logistic Regression (F1-Score: .7034; Accuracy: 62%)

  • i.e. Approx. 62% of a random guess for the bookseller being correct

Highest Contributing Feature:

  • Number of reviews, based upon Random Forest
  • May be attributed to difference in size of user base

Future Improvements

  • Experiment on a dataset of books exclusively sold on Amazon/Flipkart
  • Feature engineering to improve Random Forest