Online Booksellers
Deep-Dive into ML Classification
Zain Raza
May 13, 2020
Project Summary
Objective
Goals
Solution
Project Outline
Objective
Flipkart vs. Amazon:
📚📚🤖📚📚
Goals
How does Flipkart differentiate?
The goal of this experiment is to uncover differences on how books are sold between Amazon and Flipkart.
This insight can inform the way smart products work, such as recommender-systems that direct consumers to one platform or another.
Solution
Using Logistic Regression, a classification model was built, with an F1-Score of approximately 0.7034.
Input:
Book listed online for sale
Output:
Classification of who the bookseller was - Flipkart or Amazon?
Project Outline
The experiment was divided into performing the following tasks:
Technologies and Tools
The Environment
The Environment
This experiment was made possible using the following software tools:
Full list of imported dependencies
Exploratory Data Analysis
The Dataset
Scatter Plots
Feature Distributions
Correlation Heatmap
The Dataset
“Amazon Vs. Flipkart Book Prices” on Kaggle
Compares 1,382 of the same books�sold differently on Flipkart and Amazon
Continuous, numerical features analyzed:
Dataset posted by: mandan
Scatter Plots of Book Features, by Company
Books on Amazon
Books on Flipkart
Distribution of Book Price
Distribution of Book Rating
Distribution of Book Reviews Count
Correlation
Heatmap
Methodology
Modification of the Dataset
Modification of the Dataset
The Classifier Models
K-Nearest Neighbors
Support Vector Machine
Logistic Regression
Decision Tree
Random Forest
K-Nearest Neighbors
The best results of the KNN models --->
Support Vector Machine
Hyperparameters tuned using grid search: C = 1; gamma = 1 (RBF kernel)�Data normalized to a Z-distribution
(least overfitted)
Logistic Regression
Conclusions:�- Best results on PCA components, normal distribution�- Slight decrease in Recall, increase in F1-score from MinMax
Decision Tree
Problems in binarizing the data,�because the ranges were very similar�(lack of variance to begin with)
No Scaling Applied, all categorical values
Feature Importance:
Random Forest
NO modification applied to dataset
100 Estimators
Feature Ranking Differences from 1�decision tree
Better results than 1 decision tree!
Final Conclusions
Cross Validation
Conclusions/Future Improvements
Cross Validation
Using 5 fold cross validation,�the Logistic Regression model �showed the least variance in�its F1-Score.
The K-Nearest Neighbors classifier�showed the highest mean�F1-Score.
Final Conclusion/Future Improvements
Most Accurate: Logistic Regression (F1-Score: .7034; Accuracy: 62%)
Highest Contributing Feature:
Future Improvements