1 of 40

Multitask CNN Architecture for Online 3D Human Pose Estimation and Multi-person Tracking

Orestis Zambounis

​

Master Thesis

Supervised by Stefan Leutenegger, Margarita Grinvald, Roland Siegwart 

Orestis Zambounis

1

08.03.2019

|

|

Autonomous Systems Lab

2 of 40

Motivation

  • Navigation of robotic system in dynamic environment
    • Robot must detect dynamic objects and predict their motion

Orestis Zambounis

2

07.03.2019

?

|

|

Autonomous Systems Lab

3 of 40

Motivation

  • RGB-D camera integrated in many Robotics systems
    • Leverage depth information

Orestis Zambounis

3

07.03.2019

RGB Depth

|

|

Autonomous Systems Lab

4 of 40

Motivation

  • Estimate 3D human poses
  • Create human poses dataset
  • Use dataset to learn to predict human motion

Orestis Zambounis

4

07.03.2019

* Images taken from openframeworks.cc/ofBook/chapters/image_processing_computer_vision.html,

Fragkiadaki, Katerina, et al. "Recurrent network models for human dynamics."

|

|

Autonomous Systems Lab

5 of 40

Preliminaries

Orestis Zambounis

5

07.03.2019

Object detection

Human pose estimation

Instance segmentation

Multi-person tracking

|

|

Autonomous Systems Lab

6 of 40

Multi-Stage Prediction

Orestis Zambounis

6

07.03.2019

Object detector

Human pose detector

Multiple person tracker

Human motion predictor

into the future

RGB-D video

|

|

Autonomous Systems Lab

7 of 40

Mask R-CNN1

Orestis Zambounis

7

07.03.2019

* Images taken from [1]

objects, classes & masks

human poses & masks

|

|

Autonomous Systems Lab

8 of 40

Mask R-CNN1

Orestis Zambounis

8

07.03.2019

keypoint branch

|

|

Autonomous Systems Lab

9 of 40

Tracking as a Graph Problem

Orestis Zambounis

9

07.03.2019

* Image taken from [2]

|

|

Autonomous Systems Lab

10 of 40

Combining Multiple Cues

Orestis Zambounis

10

07.03.2019

* Image taken from [3]

|

|

Autonomous Systems Lab

11 of 40

State-of-the-art: KCF4

  • Integrates Single Object Tracking (SOT) method
    • SOT doesn't rely on object detector
    • SOT depends on the model learned at the first frame
  • Fuse information from detections and SOT
  • Recovery of lost targets
  • Complex multi-stage tracking approach

Orestis Zambounis

11

07.03.2019

|

|

Autonomous Systems Lab

12 of 40

MatchNet5

  • Patch based
  • CNN
  • Bottleneck
    • Reduces number dimensions for the feature maps
    • FC layer

​

Orestis Zambounis

12

07.03.2019

|

|

Autonomous Systems Lab

13 of 40

Method

  • Extend Mask R-CNN
  • Pairwise comparison of all object detections across 2 images

​

Orestis Zambounis

13

07.03.2019

1st stage of Mask R-CNN

|

|

Autonomous Systems Lab

14 of 40

Method

  • MatchNet5 affinity metric
  • Output: Similarity matrix
    • Optimization problem

​

Orestis Zambounis

14

07.03.2019

1st stage of Mask R-CNN

MatchNet head

|

|

Autonomous Systems Lab

15 of 40

Back-Tracking

  • Object detector not always reliable
  • Re-identify new objects by comparing to stored lost tracks

Orestis Zambounis

15

07.03.2019

t - 2

t - 1

t

|

|

Autonomous Systems Lab

16 of 40

Method

Orestis Zambounis

16

07.03.2019

from previous image

|

|

Autonomous Systems Lab

17 of 40

Method

Orestis Zambounis

17

07.03.2019

from previous image

|

|

Autonomous Systems Lab

18 of 40

Dataset and Metrics

  • MOT Benchmark6
    • Standardized multi-object tracking framework
    • 14 sequences (7 train / 7 test)
    • ~11’000 frames, ~ 1’300 tracks
    • Variance (static / moving, POV, frame rate, pedestrian density)
  • MOTA (Multi-Object Tracking Accuracy)
    • Metric which coincides best with human judgement

​

Orestis Zambounis

18

07.03.2019

|

|

Autonomous Systems Lab

19 of 40

Training

  •  

Orestis Zambounis

19

07.03.2019

|

|

Autonomous Systems Lab

20 of 40

Training

  • Freeze backbone
  • Train match head
  • Sample images from a 2s time window

Orestis Zambounis

20

07.03.2019

1st stage of Mask R-CNN

|

|

Autonomous Systems Lab

21 of 40

Training

Orestis Zambounis

21

07.03.2019

Training loss

Validation loss

|

|

Autonomous Systems Lab

22 of 40

Back-tracking

Orestis Zambounis

22

07.03.2019

number of frames

number of frames

|

|

Autonomous Systems Lab

23 of 40

Results

  • 03-SDP

Orestis Zambounis

23

07.03.2019

|

|

Autonomous Systems Lab

24 of 40

Results

  • 06-SDP

Orestis Zambounis

24

07.03.2019

|

|

Autonomous Systems Lab

25 of 40

Results

  • 14-SDP

Orestis Zambounis

25

07.03.2019

|

|

Autonomous Systems Lab

26 of 40

Results

Orestis Zambounis

26

07.03.2019

​

MOTA [%] ↑

FP ↓

Avg. Rank* ↓

State-of-the-art

54.7

​

21.5

Ours

37.4

4'415

64.2

Worst

36.4

50'903

71.3

* Average rank taking all metrics into account (MOTA, IDF1, MT, ML, FP, FN, ID Sw., Frag, Hz)

|

|

Autonomous Systems Lab

27 of 40

Results

  • 03-SDP multitask

Orestis Zambounis

27

07.03.2019

|

|

Autonomous Systems Lab

28 of 40

Results

  • 14-SDP multitask

Orestis Zambounis

28

07.03.2019

|

|

Autonomous Systems Lab

29 of 40

Results

  • 3D-map

Orestis Zambounis

29

07.03.2019

|

|

Autonomous Systems Lab

30 of 40

Timings and Memory

Orestis Zambounis

30

07.03.2019

​

Time / image [ms]

GPU memory [GB]

Object detection

220

2.2

Tracking

80

0.2

Mask prediction

30

0.2

Keypoint detection

170

1.65

Total

500

4.2

@Nvidia GTX 1070, 640x480 image 

|

|

Autonomous Systems Lab

31 of 40

Conclusion

  • Multi-task architecture (boxes, masks, poses, tracking)
    • Extended Mask R-CNN with tracking head
    • New approach for multi-person tracking
  • 3D pose estimation showcase
  • Limitations
    • Useful but not state-of-the-art tracking

​

​

​

Orestis Zambounis

31

07.03.2019

|

|

Autonomous Systems Lab

32 of 40

Future Work

  • Include spatial information?
  • End-to-end multi-task training
  • Tracking model improvement
    • Model alternations
    • Hyperparameter optimization
    • Weight regularization
  • 2D → 3D Human pose estimation
  • Human motion prediction

Orestis Zambounis

32

07.03.2019

|

|

Autonomous Systems Lab

33 of 40

​

​

Thank you for your attention!

​

Questions?

​

Orestis Zambounis

33

07.03.2019

|

|

Autonomous Systems Lab

34 of 40

References

  1. He, Kaiming, et al. "Mask r-cnn." Computer Vision (ICCV), 2017 IEEE International Conference on. IEEE, 2017.
  2. Tang, Siyu, et al. "Multi-person tracking by multicut and deep matching." European Conference on Computer Vision. Springer, Cham, 2016.
  3. Sadeghian, Amir, Alexandre Alahi, and Silvio Savarese. "Tracking the untrackable: Learning to track multiple cues with long-term dependencies." Proceedings of the IEEE International Conference on Computer Vision. 2017.
  4. Chu, Peng, et al. "Online Multi-Object Tracking with Instance-Aware Tracker and Dynamic Model Refreshment." arXiv preprint arXiv:1902.08231 (2019).
  5. Han, Xufeng, et al. "Matchnet: Unifying feature and metric learning for patch-based matching." Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2015.
  6. Milan, Anton, et al. "MOT16: A benchmark for multi-object tracking." arXiv preprint arXiv:1603.00831 (2016).

​

Orestis Zambounis

34

07.03.2019

|

|

Autonomous Systems Lab

35 of 40

A1: Mask R-CNN

Orestis Zambounis

35

07.03.2019

|

|

Autonomous Systems Lab

36 of 40

A2: Combining Multiple Cues

Orestis Zambounis

36

07.03.2019

* Image taken from [3]

|

|

Autonomous Systems Lab

37 of 40

A3: Tracking Preliminaries

  • Post Processing vs. Online
    • Online Processing is wanted for robotic applications
  • Tracking-by-detection
    • Tracking algorithm heavily relies on boxes form an object-detector

Orestis Zambounis

37

07.03.2019

|

|

Autonomous Systems Lab

38 of 40

A4: Inference

Orestis Zambounis

38

07.03.2019

|

|

Autonomous Systems Lab

39 of 40

A5: Threshold

Orestis Zambounis

39

07.03.2019

|

|

Autonomous Systems Lab

40 of 40

A6: Results

Orestis Zambounis

40

07.03.2019

*********** Your MOT17 Results ************�Rcll  Prcn |   FP     FN   IDs  Frag|  MOTA

40.6  98.1 | 4415 334950 13924 14261|  37.4

where T is the total number of true detections, and Φ is the total number of fragmentations

|

|

Autonomous Systems Lab