PYTHON FOR MACHINE LEARNING
Hands-on Walkthrough of Applying Machine Learning to Your Data Science Projects
NUS Statistics Society
Link For Workshop Materials
https://bit.ly/35AS7bk
Prerequisites and Learning Objectives
Learning Objectives:
TABLE OF CONTENTS
INTRO TO DATA SCIENCE AND ML
NUMPY AND PANDAS QUICKSTART
EXPLORATORY DATA ANALYSIS
DATA PREPROCESSING
MODEL TRAINING AND EVALUATION
OVERVIEW OF SCIKIT-LEARN ALGORITHMS
01
03
02
04
05
06
INTRO TO DATA SCIENCE AND ML
01
What is Data Science?
Data science is a multidisciplinary field that uses scientific methods, processes, algorithms and systems to extract knowledge and insights data
Why Learn Data Science?
Data Science Process
What is Machine Learning?
model
Input
Output
What is Machine Learning?
Supervised Learning: Classification vs Regression
NumPy and Pandas Quickstart
02
Basic Numpy Operations
Basic Numpy Operations
Basic Numpy Operations
Basic Numpy Operations
Pandas
Pandas
Pandas
Reading CSV
EXPLORATORY DATA ANALYSIS
03
What is Exploratory Data Analysis?
Exploratory Data Analysis refers to the process of performing initial investigations on data with the help of summary statistics and graphical representations, so as to:
Seaborn
Detecting Outliers
The box plot shows the three quartile values of the distribution along with extreme values. The “whiskers” extend to points that lie within 1.5 IQRs of the lower and upper quartile, and then observations that fall outside this range are displayed independently
Detecting Outliers
Plotting Distribution of a Feature
Plotting Features Against Each Other
Plotting Features Against Each Other
Plotting Features Against Each Other
Plotting Correlation Heatmap
DATA PREPROCESSING
04
Data Preprocessing
Now that we have a good understanding of the dataset, we need to:
Data Cleaning
Handling duplicate entries
Handling missing values
Handling missing values
One-Hot Encoding
One-Hot Encoding
One-Hot Encoding
Feature Engineering
Feature engineering refers to the process of creating new input features from your existing ones to improve model performance.
Coming up with features is difficult, time-consuming, requires expert knowledge. “Applied machine learning” is basically feature engineering. ~ Andrew Ng
Feature Engineering: Indicator Variables
The first type of feature engineering involves using indicator variables to help your algorithm “focus” on important signals in the data.
Feature Engineering: Interaction Features
The next type of feature engineering involves highlighting interactions between two or more features. Specifically, look for opportunities to take the sum, difference, product, or quotient of multiple features.
Feature Engineering: Feature Representation
Your data won’t always come in the ideal format. You should consider if you’d gain information by representing the same feature in a different way.
Feature Scaling
Scikit-Learn
Feature Scaling: Min-Max Scaling
Scikit-Learn provides a transformer called MinMaxScaler for this. It has a feature_range hyperparameter that lets you change the range if, for some reason, you don’t want 0– 1.
Feature Scaling: Standardization
MODEL TRAINING AND EVALUATION
05
Model Training
At this point, we have:
Linear Regression
where X are input features, and β are feature weights
Minimize squared distance between true Y and predicted Y
(Linear Least Squares)
Model Training
Fitting a linear regression model, we obtain the feature weights β. We can then use these weights to infer on new input:
Example:
Final grade = 0.3 x attendance + 1.5 x hours spent studying
Remembering The Goal
model
Input
Output
Underfitting and Overfitting
Underfitting
Overfitting
Underfitting and Overfitting
What do we do when the model doesn’t work?
Problems with Data
“Bad Data” Problems | |||
Problem | Insufficient Data | Non-representative Training Data | Erroneous or Noisy Data |
How to Diagnose? | Simple model underfits, complex model overfits | Good validation performance, poor test performance | Closer inspection of data |
Solution | Collect More Data | Improve Data Collection Methods to Reduce Sampling Bias | Clean Data to Remove Errors/ Outliers |
Choosing an Evaluation Metric
Accuracy, Precision and Recall
= TP / Total Predicted Positive
= TP / Total Actual Positive
Data Splits
Validation Set
K-Fold Cross Validation
Training a Linear Regression Model
Training a Linear Regression Model
Training a Random Forest Model
Hyper-parameter Tuning
Grid and Randomized Search
OVERVIEW OF MACHINE LEARNING ALGORITHMS
06
Commonly used Machine Learning Algorithms
Linear Regression
Logistic Regression
Decision Tree
Random Forest
K-Nearest Neighbours
Choosing the right model
In a competition setting:
Summary
Today we learnt how to:
What’s Next?
FURTHER RESOURCES
07
Further Resources
Projects to Try
Post Workshop Survey
https://bit.ly/2OV69i3
Thank you and we hope you have fun in your machine learning journey!
23 October 2019 (Wed), 7:00 pm – 9:00 pm
NUS Science LT 32
BestTop X NUS Statistics Society Career Talk Event:
How To Prepare Your Career in Data Science Industry
An information and networking session for you and your friends to interact with three honoured speakers, and learn more about the data science industry.
Free refreshments are provided!
https://orgsync.com/
167409/forms/370841
Our Social Media Links:
FaceBook:
https://www.facebook.com/nus.statistics.soc/
LinkedIn: https://www.linkedin.com/company/nusstatssoc/
Instagram:
https://www.instagram.com/nusstatssoc/
Extra Slides
Installing Python
Machine Learning Tools Covered
> pip install jupyter pandas numpy scikit-learn seaborn
Basic Numpy Operations
Basic Numpy Operations
Virtual Environments
Field of Data Science
How do Machine Learning Models Work?
How do Machine Learning Models Work?
How do Machine Learning Models Work?
Bayesian Optimization
Bayesian Optimization
Linear Regression
Algorithm
Sklearn Example
Description
Support Vector Machines
Diagram
Description
Decision Tree
Algorithm
Sklearn Example
Description
K-means Clustering
Algorithm
Applications
Description
Multi-layer Perceptron
Diagram
from sklearn.neural_network import MLPClassifier
Details
Algorithm