1 of 46

Offline evaluation and analysis

CV4E August 23, 2023 - Sam Lapp

2 of 46

Offline evaluation and analysis

CV4E August 23, 2023 - Sam Lapp

3 of 46

Probing your model’s performance

Offline evaluation and analysis

CV4E August 23, 2023 - Sam Lapp

4 of 46

Logits: raw output

sigmoid(logits)

5 of 46

An aside: Single-target vs Multi-target Classification

Sigmoid activation

Binary Cross Entropy loss

​

Multi-target

​

binary classification of every class independently

Each sample has 1 label

Single target

​

Softmax activation

Cross Entropy Loss

6 of 46

Alien: 0.9

7 of 46

Will my model detect aliens in this upcoming year’s field data if they appear?

Why does my model miss 50% of aliens?

What non-aliens are being confused for aliens? (and the other way around)

Why does my model work on last year’s data but not this year’s data?

Do I just need more labeled data?

What’s the probability that this sample has an alien?

Is my model even paying attention to the right stuff?

Why does my model suddenly detect lots of aliens? (btw, there aren’t any)

8 of 46

Histograms! Qualitative interpretation

Positive sample scores

Negative sample scores

“logits”: Raw output score of classifier (-inf to inf)

9 of 46

What data to use for model evaluation?

Training data!

​

Validation data

​

Test data (labeled)

​

Target data (unlabeled)

10 of 46

Evaluating model training

11 of 46

WandB plug wandb.ai

12 of 46

Why is my validation performance better than train?

Train

​

Validation

13 of 46

Why does my validation performance decrease while my training performance continues increasing?

Train

​

Validation

14 of 46

What’s next?

​

Should I keep tweaking training hyperparameters to squeeze out a higher Mean Average Precision?

15 of 46

Do I just need more data?

16 of 46

Inspecting samples!!

17 of 46

Getting a “feel” for your model’s performance

​

without any labels!

​

Grids of images + scores help build intuition

​

18 of 46

Most “confident” raccoon images

19 of 46

Most “confusing” raccoon images

20 of 46

Other categories to inspect

Worst performance:

Lowest-scoring positives

Highest-scoring negatives

​

Other stratifications:

By device / site / location

By date

Specific confusions in a confusion matrix

21 of 46

What is your model looking at?

22 of 46

Object Detection

  • More complex training labels ($$)
  • More finicky to train
  • More interpretable outputs

“Skunk”

“Skunk”

23 of 46

Saliency Maps

Off-the-shelf from FlashTorch, with a CoLab Demo

Simonyan et al., 2014

24 of 46

Saliency Maps

Also implemented in the Pytorch Grad Cam package

​

Simonyan et al., 2014

Grad Cam

Guided backpropagation

Spectrogram with bird songs

25 of 46

Generalizability & Domain Shift

26 of 46

Validation versus test performance

Data Split

Top-1 Accuracy

27 of 46

Domain shifts: what is shifting?

28 of 46

Domain shifts: what is shifting?

29 of 46

Training data

30 of 46

Test data

31 of 46

Recap: causes of domain-shift performance drop

Generalizing to new data is hard

  • Lower quality / more difficult / “dirtier” real-world samples
  • New stuff that wasn’t seen during training
  • Background change
  • Class distribution shift
  • (and other unpredictable things…)

32 of 46

How generalizable do you need your model to be?

  • Can you train on a annotated subset of your real data?� A model that’s great for your data beats a model that’s ok everywhere� Investing in project-specific training data may be worth the effort�
  • Should you allow your model to learn from class imbalance?

Learning = Bias. You decide what your model should learn and how it should be biased.

33 of 46

What can we do with continuous scores?

34 of 46

Histogram of unlabeled data

35 of 46

Histograms: shifts across sites

36 of 46

Histograms: shifts across sites

37 of 46

Histograms: shifts across sites

38 of 46

Histograms: shifts across sites

39 of 46

Histograms: shifts across sites

Precision ~ 80%

Precision ~ 10%

40 of 46

Calibration

Calibrated probability score:

Score = 0.9 means 90% chance that sample is positive

In other words, 90% of the clips scoring 0.9 should be positives

ML outputs are not calibrated probabilities!

  • Caveat: based on the previous slides, notice that even a “calibrated” model often isn’t calibrated when applied to a new dataset

41 of 46

Calibration

Fraction of samples

That are positive

model score

1

0

0

1

Over-confident

model

Account for the error

42 of 46

Calibration

Fraction of samples

That are positive

model score

1

0

0

1

Over-confident

model

Account for the error

Predict with train set ->

Calibrate with validation set ->

Check calibration on test set

​

Methods:

  • histogram bins
    • Generalize with isotonic regression
  • Assume a functional relationship
    • Platt scaling (logistic regression)
    • Temperature scaling

43 of 46

“Temperature” Calibration (variant of Platt scaling)

Temperature scaling:

​

p = softmax( logit / T )

​

T often around 1.5 - 3

Fit a logistic regression: logit(y) ~ a0 + a1p

44 of 46

Alternatives to calibration?

45 of 46

Will my model detect aliens in this upcoming year’s field data if they appear?

Why does my model work on last year’s data but not this year’s data?

Why does my model miss 50% of aliens?

What non-aliens are being confused for aliens? (and the other way around)

Do I just need more labeled data?

What’s the probability that this sample has an alien?

Is my model even paying attention to the right stuff?

Why does my model suddenly detect lots of aliens? (btw, there aren’t any)

46 of 46

Next time:

Is my model good enough? Useful?

What do I do with it?