Offline evaluation and analysis
CV4E August 23, 2023 - Sam Lapp
Offline evaluation and analysis
CV4E August 23, 2023 - Sam Lapp
Probing your model’s performance
Offline evaluation and analysis
CV4E August 23, 2023 - Sam Lapp
Logits: raw output
sigmoid(logits)
An aside: Single-target vs Multi-target Classification
Sigmoid activation
Binary Cross Entropy loss
Multi-target
binary classification of every class independently
Each sample has 1 label
Single target
Softmax activation
Cross Entropy Loss
Alien: 0.9
Will my model detect aliens in this upcoming year’s field data if they appear?
Why does my model miss 50% of aliens?
What non-aliens are being confused for aliens? (and the other way around)
Why does my model work on last year’s data but not this year’s data?
Do I just need more labeled data?
What’s the probability that this sample has an alien?
Is my model even paying attention to the right stuff?
Why does my model suddenly detect lots of aliens? (btw, there aren’t any)
Histograms! Qualitative interpretation
Positive sample scores
Negative sample scores
“logits”: Raw output score of classifier (-inf to inf)
What data to use for model evaluation?
Training data!
Validation data
Test data (labeled)
Target data (unlabeled)
Evaluating model training
WandB plug wandb.ai
Why is my validation performance better than train?
Train
Validation
Why does my validation performance decrease while my training performance continues increasing?
Train
Validation
What’s next?
Should I keep tweaking training hyperparameters to squeeze out a higher Mean Average Precision?
Do I just need more data?
Inspecting samples!!
Getting a “feel” for your model’s performance
without any labels!
Grids of images + scores help build intuition
Most “confident” raccoon images
Most “confusing” raccoon images
Other categories to inspect
Worst performance:
Lowest-scoring positives
Highest-scoring negatives
Other stratifications:
By device / site / location
By date
Specific confusions in a confusion matrix
What is your model looking at?
Object Detection
“Skunk”
“Skunk”
Saliency Maps
Off-the-shelf from FlashTorch, with a CoLab Demo
Simonyan et al., 2014
Saliency Maps
Also implemented in the Pytorch Grad Cam package
Simonyan et al., 2014
Grad Cam
Guided backpropagation
Spectrogram with bird songs
Generalizability & Domain Shift
Validation versus test performance
Data Split
Top-1 Accuracy
Domain shifts: what is shifting?
Domain shifts: what is shifting?
Training data
Test data
Recap: causes of domain-shift performance drop
Generalizing to new data is hard
How generalizable do you need your model to be?
Learning = Bias. You decide what your model should learn and how it should be biased.
What can we do with continuous scores?
Histogram of unlabeled data
Histograms: shifts across sites
Histograms: shifts across sites
Histograms: shifts across sites
Histograms: shifts across sites
Histograms: shifts across sites
Precision ~ 80%
Precision ~ 10%
Calibration
Calibrated probability score:
Score = 0.9 means 90% chance that sample is positive
In other words, 90% of the clips scoring 0.9 should be positives
ML outputs are not calibrated probabilities!
Calibration
Fraction of samples
That are positive
model score
1
0
0
1
Over-confident
model
Account for the error
Calibration
Fraction of samples
That are positive
model score
1
0
0
1
Over-confident
model
Account for the error
Predict with train set ->
Calibrate with validation set ->
Check calibration on test set
Methods:
“Temperature” Calibration (variant of Platt scaling)
Temperature scaling:
p = softmax( logit / T )
T often around 1.5 - 3
Fit a logistic regression: logit(y) ~ a0 + a1p
Alternatives to calibration?
Will my model detect aliens in this upcoming year’s field data if they appear?
Why does my model work on last year’s data but not this year’s data?
Why does my model miss 50% of aliens?
What non-aliens are being confused for aliens? (and the other way around)
Do I just need more labeled data?
What’s the probability that this sample has an alien?
Is my model even paying attention to the right stuff?
Why does my model suddenly detect lots of aliens? (btw, there aren’t any)
Next time:
Is my model good enough? Useful?
What do I do with it?