1 of 20

ArgusRoad: Road Activity Detection with Connectionist Spatiotemporal Proposals

Lijun Yu, Yijun Qian, Xiwen Chen, Wenhe Liu

and Alexander G. Hauptmann

10/16/2021

Speaker: Lijun Yu

2 of 20

Introduction – Task 

  • Activity Detection
    • From autonomous driving perspective
  • Sub Tasks
    • Spatial Localization
    • Temporal Localization
    • Action Classification
  • Evaluation Metric
    • Video mAP @ 3D IoU

[1] ROAD: The ROad event Awareness Dataset for Autonomous Driving,arXiv 2021. Gurkirt Singh et. al.

3 of 20

Introduction – Dataset

Annotation

ROAD’s annotated frames cover multiple agents and actions, recorded at different weather conditions (overcast, sun, rain) at different times of the day (morning, afternoon and night)[1].

Challenges

ROAD dataset

  • Weather and time condition
  • Camera view varies
  • Scene varies

[1] ROAD: The ROad event Awareness Dataset for Autonomous Driving,arXiv 2021. Gurkirt Singh et. al.

4 of 20

Introduction – Argus++ Framework

  • Designed for real-time activity recognition in extended videos in NIST ActEV evaluations
    • Best performing system since fall 2020
    • Measure false positive at frame level, but true positive at instance level
    • Do not care about spatial localization or integrity in temporal dimension
    • Real-time processing on consumer-level hardware
  • Adaptations needed for ROAD challenge
    • Video mAP matches each instance with only one prediction, need merging of short predictions
    • 3D IoU measures spatial-temporal alignment of tubes, need precise bounding box in each frame
    • No efficiency requirements

5 of 20

ArgusRoad Framework

  • A special application of Argus++: proposal stride is 1 frame, no proposal filtering
  • Key Concept: Spatiotemporal Cube 
  • Proposal Generation: Generate candidate cubes with overlapping temporal context
  • Temporal Localization: Measure temporal boundaries by connecting neighbor cubes

Proposal Generation

Activity Recognition

Object Detection

Object Tracking

Temporal Localization

Activity Instances

Video Stream

6 of 20

Detection and Tracking

  • Object detection
    • Mask R-CNN with Resnet-101 backbone from detectron2
    • Classes: person, vehicle (car, truck, etc.), traffic light
    • For better tracking and spatial localization, process every frame
    • For better efficiency, skip frames as long as tracking is reasonable
  • Multi-object tracking
    • Towards-Realtime-MOT
    • Based on RoI features from detection backbone

7 of 20

Proposal Generation

  • Proposal Paradigm
    • Previous: spatial-temporal tube proposal
      • Use whole trajectory of each tracked object
      • Still require temporal localization
      • Object’s shape changes when resized for feature extraction
    • New: spatial-temporal cube proposal:
      • A simple six-tuple defining the boundaries in three dimensions

Why Cube is better than Tube for action recognition?​

8 of 20

Proposal Sampling

  •  

9 of 20

Proposal Generation: An Example

10 of 20

11 of 20

Proposal Sampling

  • Non-overlapping

  • Overlapping

  • Overlapping, stride=1�����Equivalent to sampling cubes at every frame with a temporal context

12 of 20

Proposal Evaluation

  • Idea: the upper bound performance if you have a perfect classifier
  • Match proposal cubes and converted ground truth cubes

     Score = Spatial_IoU(detection, cube gt)  *  score(cube gt[Temporal coverage]

  • Label assignment: Faster R-CNN [1]
    • For each ground truth, assign to the detection with the highest score (> low_thres)
    • For each detection, assign as the ground truth with score >= high_thres
    • For each detection, assign as negative if all score <= low_thres
    • low_thres: 0, high_thres: 0.5

[1] Ren, Shaoqing, et al. "Faster r-cnn: Towards real-time object detection with region proposal networks." Advances in neural information processing systems 28 (2015): 91-99.

13 of 20

Proposal Evaluation Results

  • Under the assumption of a perfect classifier
    • Use proposals and assigned labels to generate outputs �(output generation method covered later)

3D IoU Threshold

mAP

mR

0.1

59.38

88.69

0.2

51.20

79.26

0.5

33.49

51.45

Average

48.02

73.13

Performance on ROAD val3

14 of 20

Activity Recognition - Training

  • Multi-label Classification
    • Binary cross entropy loss
    • Weighted by proposal scores
    • Balance activity-wise pos/neg samples 
    • Balance samples of different activities
    • Balance samples of different datasets

15 of 20

Activity Recognition - Model

T x D x D

3D Conv

1 x D x D

T x 1 x 1

(2+1)D Conv

Sports 1M

Kinetics 400

Comparison with the state-of-the-art on Kinetics and Sports 1M. [1]

[1] Tran, Du, et al. "A closer look at spatiotemporal convolutions for action recognition." Proceedings of the IEEE conference on Computer Vision and Pattern Recognition. 2018.

R(2+1)D

R3D

16 of 20

Temporal Localization

  • Process each type of action separately
  • Select consecutive cubes by a threshold, subject to a minimum length
  • Determine thresholds by grid search on validation set
    • For ROAD dataset at 12fps�score >= 0.005 �length >= 20 frames

17 of 20

Leaderboard Results

https://eval.ai/web/challenges/challenge-page/1059/leaderboard/2748

CMU-INF team won the 1st place

18 of 20

Performance on Surveillance Dataset

CMU-INF team won the 1st place

NIST ActEV SDL Leaderboard https://actev.nist.gov/sdl#tab_leaderboard

19 of 20

Takeaways

  • Easier than expected to adapt Argus++ to video mAP @ 3D IoU metric
    • Spatial localization works well on object detection
    • Temporal localization can be acquired with dense sampling
  • Dense sampling consumes more resource
    • Stride 1 vs. stride 16 (an empirical value for MEVA/VIRAT datasets)
    • Harder to run in real-time
  • Still in need of big improvement for autonomous driving

20 of 20

Q&A

Thanks for listening!