1 of 33

1

Uncertainty in Classification

MSCV capstone project

Jia Shi

Advisor: Deva Ramanan

2 of 33

  1. Why care about uncertainty?

2

Lack of explainability exist in deep learning, decisions process is often unknown

3 of 33

  1. Why care about uncertainty?

3

Spacecraft

Autonomous Vehicle

Surgical Robots

“One miss shot” will cause catastrophe on safety-critical application

Model can’t say, “I am not sure”. When encounter out-of-distribution(OOD)/noisy sample.

4 of 33

  1. What kind of uncertainty do we care?

4

Mean: dice

OOD/rare from train set

“In practice, the categorization of uncertainty can be much more complicated than that

and highly task specific”

ID, but noisy data

Mean: knowledge

Model Uncertainty

Data Uncertainty

5 of 33

Two track of project under uncertainty

Data Uncertainty:

Visual Features Disentanglement into Interpretable Attributes Features using Prompts

Question: How do we describe the distribution of data?

(ICCV 2023 submission)

Model Uncertainty:

Improving Generalization with Semantic Consistency Feature

Question: How do we tell whether a model could generalize better on OOD?

(Aim for Neurips 2023 submission)

5

6 of 33

Visual Features Disentanglement into Interpretable Attributes Features using Prompts

Data Uncertainty

6

Question: How do we describe the distribution of data?

7 of 33

human perform classification from visual concept

7

8 of 33

Problem of Feature Entanglement

8

end to end training —-> feature entanglement —-> not interpretable

9 of 33

Feature entanglement limit trustworthy

9

Animal

watermelon Skin

Rino body

grass background

10 of 33

Visual Feature disentanglement

10

Our goal is to propose a way to disentangle visual feature entanglement to a list of human interpretable attribute

11 of 33

Attribute CLIP - Use prompts to get attributes from original clip embedding

12 of 33

Method: Disentangle Feature

12

  1. Visual Feature encoding with CLIP

13 of 33

Method: Disentangle Feature

13

2. Attribute text encoding with CLIP

14 of 33

Method: Disentangle Feature

14

3. Adopting cosine similarity score as feature!

Visual feature become a weighted combination of interpretable attribute feature

15 of 33

Application: 1. Few shot generalization

15

Using disentangled feature could achieve superior performance on very few shot setting

zero shot evaluation

(time consuming)

(poor performance)

(model ensemble)

16 of 33

Application: 2. Quantifying distribution shift

16

Top attribute from ImageNet-Sketch and ImageNet-Rendition could well reflection the description on those two dataset.

Colorless; Paper; Gray; Light Gray; Lined

ImageNet-Sketch:

17 of 33

17

18 of 33

Application: 2. Quantifying distribution shift

18

Top attribute from ImageNet-Sketch and ImageNet-Rendition could well reflection the description on those two dataset.

Cartoon; Painting; Graffitied; Tattooed;

(ImageNet-R(endition) contains art, cartoons, graffiti, embroidery, graphics, origami, paintings, patterns, plastic objects, plush objects, sculptures, sketches, tattoos, toys, and video game renditions of ImageNet classes.)

ImageNet-Rendition:

19 of 33

Application: 3. Open vocabulary object classification

19

“Swapping” feature weight could adapt to OOD concept without retraining

Our propose dataset:

9 color, 37 fruit, 6050 images

20 of 33

Application: 4. Attribute guided Detection

20

Per-grid classification without accessing the class name.

Potential for Detection Diagnosis

21 of 33

Conclusion

  1. In this work, we propose a simple yet efficient method to disentangle dense visual features to a weighted combination of interpretable attribute feature in language domain.

  • We shows that our interpretable feature is a semantically meaningful yet a competitive feature in adapting visual concept by showing our superior performance on few shot evaluation.

  • We also shows the potential of the attribute feature by showing a skew of downstream application like measuring distribution shrift, solving open-vocabulary classification and detection.

21

22 of 33

Improving Generalization with Semantic Consistency Feature

Model Uncertainty

22

Question: How do we tell whether a model could generalize better on OOD?

23 of 33

How to predict OOD performance?

23

In distribution

OOD distribution

24 of 33

Previous work

24

ID acc -> OOD acc

Miller, etc.

model agreement-> OOD acc

Baek, etc.

https://arxiv.org/abs/2107.04649

https://arxiv.org/abs/2206.13089

25 of 33

Generalization = Semantic Consistency

25

strawberry!

in OOD dataset, for class with same semantic label, there are often shrift in visual appearance

ID vs OOD : Semantic consistency ✅ Visual consistency ❌

example from ImageNet-Rendition

26 of 33

Visual Similarity != Semantic Similarity

26

Supervised method with class-discrimination could not properly preserve semantic consist feature

semantical ideal embedding

(world distribution)

27 of 33

Visual Similarity != Semantic Similarity

27

Supervised method with class-discrimination could not properly preserve semantic consist feature

How to measure the semantic consistency of model embedding?

28 of 33

Semantic Consistency = Mistake Severity

28

Frog

Dog

Cat

Car

Mammal

Animal

Things

What is better mistake?

with target = Cat

Predicting Dog is a better mistake than predicting a Car

29 of 33

Why Semantic Consistency = Mistake Severity

29

GT

Pred 1

Pred 2

Mistake severity is a measurement of how close prediction are semantically close to target

Thus a model with ‘better mistake’ could have more semantically consistent embedding!

30 of 33

Measure (mistake severity using taxonomy loss)

30

+

31 of 33

Current Result

31

Compare to Top 1 ID accuracy (ImageNet acc) , our proposed semantic embedding measurement is a better indicator on performance on OOD dataset ( ImageNet-A acc)

correct trend is monotonically increasing !

Top 1 ID acc =

how separate each class clustering in embedding space

Semantic measurement =

how class cluster embedding are distributed semantically

32 of 33

Conclusion

Future Step

  • In this work, we show that our propose metric of semantic measurement is a better indicator of predicting out of distribution accuracy than the Top1 in-distribution accuracy.

  • We show that model could generalize better by optimizing our proposed metric.
  • We will extend our measurement on more model to testify it’s universality.

32

33 of 33

Thank you!

33