Live Lecture 7.1
Introduction to Classification
Summer 2020
DATA 8
Spring 2020
Announcements
Agenda
Predicting an Observation’s Outcome
Machine Learning Algorithm
Classification Example: Spam Emails
Examples (cont.)
Agenda
Training a Classifier
Classifier
Variable(s) of an observation
Predicted label of the observation
Population
Labels
Sample
Training
Set
Test
Set
Model the association between attributes & labels
Estimate the accuracy of the classifier
Training a Classifier
Population
Attributes
Labels
Training
Set
Test
Set
Classifier
Variables of an observation
Predicted label of the observation
Agenda
Sketch of K Nearest Neighbors
Variable 1
Variable 2
Legend:
Class 1:
Class 2:
Sketch of K Nearest Neighbors
Variable 1
Variable 2
Legend:
Class 1:
Class 2:
Sketch of K Nearest Neighbors
Variable 1
Variable 2
Legend:
Class 1:
Class 2:
?
Distance: Measure of Similarity
(x₀, y₀)
(x₁, y₁)
y₀ - y₁
x₀ - x₁
Distance Between Two Points
Standardize if Necessary
Chronic Kidney Disease data set
Finding the K Nearest Neighbors (KNN)
To find the k nearest neighbors of a new observation:
Sketch of K Nearest Neighbors
Variable 1
Variable 2
Legend:
Class 1:
Class 2:
Agenda
Accuracy of a Classifier
The accuracy of a classifier on a labeled data set is the proportion of examples that are labeled correctly
Need to compare classifier predictions to true labels
If the labeled data set is sampled at random from a population, then we can infer accuracy on that population
Sample
Labels
Training
Set
Test
Set
Population
Start with a Representative Sample
Questions?