1 of 62

Detecting Anomalies in Object Appearance and Motion Dynamics

�Mazen Alotaibi

1

2 of 62

2

https://loc.gov/pictures/resource/var.1680/

3 of 62

What is Common Sense?

Object Permanence:�“Understanding that items and people �still exist even when you can't see or �hear them.” - webmd

3

https://loc.gov/pictures/resource/var.1680/

4 of 62

4

https://www.youtube.com/watch?v=rVqJacvywAQ

5 of 62

Why Machine Common Sense is important?

  • Self-driving cars

5

6 of 62

Why Machine Common Sense is important?

  • Self-driving cars
  • Machines don’t learn common �sense automatically
  • Teaching all of cases where �common sense knowledge is �applied is hard

6

7 of 62

Benchmark[1]

  • Measures the performance of an agent perceiving the world without any interaction on many common sense reasoning tasks.
    • Object Permanence
    • Gravity
  • Expects from agent to detect anomalies.

7

[1] Ricochet et al. “IntPhys 2019: A Benchmark for Visual Intuitive Physics Understanding”

8 of 62

Prior work

  • Utilize a physical simulation module to predict future object states from current beliefs.
    • Develops beliefs by answering two questions without �update[1].
    • Develops initial beliefs for object states but update when�it is reasonable[2].

8

[1] Battaglia et al. “Simulation as an engine of physical scene understanding”

[2] Smith et al. “Modeling Expectation Violation in Intuitive Physics with Coarse Probabilistic Object Representations”

9 of 62

Table of Content

  • MCS Challenge
  • Our Approach
  • Evaluation
  • Future work

9

10 of 62

MCS Challenge

10

11 of 62

Problem Statement

  • Can a machine predict whether a scene is plausible or implausible based on a set of expectations?

11

Scene

Plausible Y/N

Machine

12 of 62

12

13 of 62

Collisions (COLL)

Focus objects can't change their �motion or appearance without �an explanation.

13

14 of 62

Gravity Support (GRAV)

A dropped focus object follows �gravity.

14

15 of 62

Object Permanence (OP)

Focus objects can't appear �or disappear from the scene �without an explanation.

15

16 of 62

Shape Constancy (SC)

Focus objects can't change �their appearance.

16

replace with video

17 of 62

Spatiotemporal Continuity (STC)

A thrown focus object needs �to have a continuous motion

17

18 of 62

Input

  • RGB images
  • Ground truth segmentation masks
  • Depth images

18

19 of 62

Problem Statement

  • Can a machine predict whether a scene is plausible or implausible based on a set of expectations?

19

SEG

RGB

Depth

Plausible Y/N

Machine

20 of 62

Approach

20

21 of 62

System - Pipeline

21

Mask to 2dbox

Tracker

SEG

RGB

Depth

Learning

component

Reasoning�Agent

Plausible Y/N

Rule-base

component

Role Assigner

3D Amodal Detector

Algorithm

22 of 62

Mask to 2dbbox

  • Given ground truth segmentation masks
    • Converts masks to 2D bounding boxes coordinates

22

Mask to 2dbox

Tracker

Reasoning

Agent

Role Assigner

3D amodal detector

23 of 62

Tracker

  • Given 2D bounding boxes, create tracklets

23

Mask to 2dbox

Tracker

Reasoning

Agent

Role Assigner

3D amodal detector

Tracker

2D �Bounding Boxes

Tracks

RGB

24 of 62

Tracker - Pipeline

24

1st�Stage

Learning

component

Algorithm

Voting

Candidates�pairing

Merging

RGB

Tracks

2D �Bounding �Boxes

2nd Stage

25 of 62

Tracker - 1st Stage

  • Bi-Linear LSTM model[1] (out-of-the-box)
    • Online tracker
    • Relies on motion and appearance cues.
  • Given 2D bounding boxes, outputs tracklets of continuous motion.

25

[1] “Discriminative appearance modeling with multi-track pooling for real-time multi-object tracking”, � Kim et al. - CVPR 2021

1st Stage

Tracklets

2D �Bounding Boxes

RGB

26 of 62

Tracker - 1st Stage

Objects before and after�occlusion have different�object IDs.

26

27 of 62

Tracker - 2nd Stage

  • Components:
    • Candidates pairing
    • Voting
    • Merging

27

2nd Stage

Tracks

Continuous�Tracklets

RGB

28 of 62

Tracker - 2nd Stage (Candidates Pairing)

  • Creates a candidate pair that exists after each other without being visible in any of the frames that they exist on.

28

29 of 62

Tracker - 2nd Stage (Voting)

  • Runs a similarity scoring network on multiple pairs from two different tracklets to generate a set of similarity scores.

  • If the average of the similarity scores is above a threshold, both tracklets be considered as potential merge candidates.

29

30 of 62

Tracker - 2nd Stage (Voting)

30

31 of 62

Tracker - 2nd Stage (Voting)

  • Siamese Network with pre-trained ResNet18 as backbone
  • Relies on appearance cues only
  • Given two images, outputs similarity score of two images.

31

Similarity�Score

Network

32 of 62

Tracker - 2nd Stage (Merging)

  • Merges tracklets based on the voting score.

32

33 of 62

Tracker - Pipeline

33

1st�Stage

Learning

component

Algorithm

Voting

Candidates�pairing

Merging

RGB

Tracks

2D �Bounding �Boxes

2nd Stage

34 of 62

Role Assigner

  • Pre-trained ResNet18
  • Given tracklets, for each, assign a role of
    • Focus object (solid)
    • Environmental object (dashes)

34

Mask to 2dbox

Tracker

Reasoning

Agent

Role Assigner

3D amodal detector

Role�Assigner

Focus Object�Y/N

Tracks

RGB

35 of 62

3D amodal detector

  • CenterTrack[1] (out-of-the-box).
  • Given 2D bounding box of an object and its depth, estimate the 3D bounding cube.

35

[1] “Tracking Objects as Points”, Xingyi Zhou et al. - ECCV 2020

Mask to 2dbox

Tracker

Reasoning

Agent

Role Assigner

3D amodal detector

3D�Amodal�Detector

3D �Tracks

2D Tracks

Depth

36 of 62

Rule-based Reasoning Agent

  • Given tracklets, predict if the scene is plausible or implausible.
  • For each scene type, we have developed rule-based VOE classifiers.

36

Mask to 2dbox

Tracker

Reasoning

Agent

Role Assigner

3D amodal detector

Reasoning�Agent

Plausible Y/N

3D

Tracks

37 of 62

Rule-based Reasoning Agent - GRAV

  • Assumptions for fallen focus �object:
    • Has only gravity force applied �to it.
    • Is supported by the support �object if:
      • Lands on top the support object.
      • Well-balanced.

37

38 of 62

Rule-based Reasoning Agent - GRAV

  • Given the first position of the fallen focus object and the position of the support object.
  • Estimate the last relative position of the fallen object to set an expectation.

38

39 of 62

Rule-based Reasoning Agent - SC

  • Assumptions for focus �objects:
    • Don’t change appearance.
    • Enter the scene from the �edges of the frame only.

39

40 of 62

Rule-based Reasoning Agent - SC

  • Check if a new track was initialized in the middle of the scene.

40

41 of 62

Rule-based Reasoning Agent - OP

  • Assumptions for focus �objects:
    • Enter the scene from the �edges of the frame only.
    • Once they leave the scene, �they can’t re-appear.

41

42 of 62

Rule-based Reasoning Agent - OP

  • Objects can only enter and leave the sense from the sides.
  • When an object leaves, the object shouldn’t re-appear after
  • When an object doesn’t leave and was occluded, we expect it to be behind the occluder wall.

42

43 of 62

Rule-based Reasoning Agent - STC

  • Assumptions for focus �objects:
    • Enter the scene from the �edges of the frame only.
    • Have continuous motion.

43

44 of 62

Rule-based Reasoning Agent - STC

  • When a focus object has a gap in its track
    • Check if the gap has occluder walls placed throughout the entire distance where the object was absent.

44

45 of 62

Rule-based Reasoning Agent - COLL

  • Assumptions for focus �objects:
    • Enter the scene from the �edges of the frame only.
    • Two objects can only interact �if they share depth and �intersect.

45

46 of 62

Rule-based Reasoning Agent - COLL

  • Two objects can collide if they share the same depth and path.

46

47 of 62

Evaluation

47

48 of 62

Evaluation

  • Evaluation set:
    • For each scene type, randomly sample 100 scenes of plausible and implausible.
  • Comparing:
    • Every component against baselines.
    • The system against a baseline system.

48

49 of 62

Tracker

  • Baseline:
    • ResNet18 pre-traiend on ImageNet, not fine-tuned on MCS data
    • Normalization then dot product
  • Metrics:
    • Number of Identity Switch (IDSW)
    • Multiple Object Tracking Accuracy (MOTA)

49

50 of 62

Tracker

50

Baseline

Ours

IDSW↓

MOTA↑

IDSW↓

MOTA↑

Collisions

0.1

0.9944

0.1

0.9944

Object Permanence

0.14

0.9966

0.18

0.9961

Shape Constancy

0.44

0.9867

0.29

0.9891

Spatiotemporal Continuity

0.0

0.9924

0.0

0.9924

Gravity Support

0.0

0.9998

0.0

0.9998

All

0.136

0.9939

0.114

0.9943

51 of 62

Role Assigner

  • Baseline:
    • Decision Tree model to classify roles based on object area and position.
  • Metric:
    • F1-score, positives are Focus objects

51

52 of 62

Role Assigner

52

Baseline

Ours

Precision

Recall

F1-score

Precision

Recall

F1-score

Collisions

0.71

0.82

0.76

1.0

1.0

1.0

Object Permanence

0.6

1.0

0.75

1.0

1.0

1.0

Shape Constancy

0.47

1.0

0.64

1.0

0.98

0.99

Spatiotemporal Continuity

0.42

1.0

0.6

1.0

1.0

1.0

Gravity Support

1.0

0.75

0.86

1.0

0.96

0.98

All

0.72

0.99

53 of 62

Reasoning Agent

  • Baseline:
    • Learned Reasoning Agent trained on plausible scenes.
  • Metric:
    • F1-score, positives are implausible scenes

53

54 of 62

Reasoning Agent - Ground-Truth Tracks

54

Baseline

Ours

Precision

Recall

F1-score

Precision

Recall

F1-score

Collisions

0.54

1.0

0.70

0.29

0.16

0.21

Object Permanence

0.47

0.88

0.61

1.0

0.88

0.94

Shape Constancy

0.0

0.0

0.0

1.0

1.0

1.0

Spatiotemporal Continuity

0.56

1.0

0.72

1.0

0.84

0.91

Gravity Support

0.0

0.0

0.0

1.0

0.92

0.96

All

0.41

0.8

55 of 62

Reasoning Agent - Actual Tracks

55

Baseline

Ours

Precision

Recall

F1-score

Precision

Recall

F1-score

Collisions

0.5

1.0

0.67

0.45

0.64

0.53

Object Permanence

0.5

1.0

0.67

1.0

0.6

0.75

Shape Constancy

0.0

0.0

0.0

1.0

0.86

0.92

Spatiotemporal Continuity

0.55

1.0

0.7

1.0

0.84

0.91

Gravity Support

0.0

0.0

0.0

1.0

0.92

0.95

All

0.41

0.82

56 of 62

Future

56

57 of 62

Improvement

  • Data generation: more variation and better quality.
  • Tracker
    • 1st Stage: introduce a heuristic to split tracklets when object changes its color without occlusion.
    • 2nd Stage: train on hard-examples - similar shapes with same colors.
  • Reasoning Agent: design rules that are more robust to noise in the pipeline.

57

58 of 62

Questions?

58

59 of 62

Thank you!

59

60 of 62

Appendix

60

61 of 62

Number of Identity Switch (IDSW↓)

  • Counts the number of emergences when a ground truth target i is matched to hypothesis j and the last known assignment was k (k != j)

61

Leal-Taixe et al. MOTChallenge 2015: Towards a Benchmark for Multi-Target Tracking

62 of 62

Multiple Object Tracking Accuracy (MOTA↑)

  • Compute a set of measures per frame
    • Perform matching between predictions and ground truth
    • FP = False positives
    • FN = False negatives (missing detections)
    • IDSW: identity switches

62

Leal-Taixe et al. MOTChallenge 2015: Towards a Benchmark for Multi-Target Tracking