1 of 14

Why is the video analytics accuracy fluctuating,�and what can we do about it?

Sibendu Paul, Kunal Rao, Giuseppe Coviello, Murugan Sankaradas, Oliver Po,

Y. Charlie Hu, and Srimat Chakradhar

ECCV AROW 2022

1

2 of 14

2

Large-Scale Video Analytics Pipeline (VAP)

Image Acquisition

Image Transmission

Edge Cameras

Analytics Server (cloud)

Wireless Network(5G)

High-definition

video streams

Human Computer Interaction

Automated

Surveillance

Ambulance

Content based

Video Retrieval

Intelligent Transportation System

Remote Supervision

3 of 14

Importance of VAPs

  • Important applications in many domains such as intelligent transportation, retail solutions, smart cities and healthcare

  • Safety-critical 5G apps like biometrics-based payments, AR/VR, remote healthcare require higher accuracy analytics under all conditions

  • Optimized video analytics apps in 5G slices is a multi-billion-dollar market

3

4 of 14

Observation: Accuracy of VAP fluctuates !

  • Methodology
    • Object detection -- One of the most common tasks
    • Use different object detectors to detect cars and persons on video snippets from Roadway dataset (videos with almost static contents)
  • Observation:
    • Accuracy (detection) on consecutive frames fluctuate while ground-truth stays almost the same
    • Also observed for face-detection models (retinaNet, MTCNN, Neoface-v4)

4

Accuracy =

True-positive object detection count

5 of 14

How to quantify accuracy fluctuation?

  • F2 – Max detection count variation over mean ground-truth detections over consecutive 2 frames

  • E.g., the F2 metric values are 48%, 42%, 20%, 25% for figures on the right

5

Accuracy =

True-positive object detection count

6 of 14

Possible factors for accuracy fluctuation

  • Motion of objects in camera FoV

🡪 Switch to a video of static 3D scenes, using AXIS Q1615 camera

6

6

Object Motion

F2 = 17.4%

Static 3D scene

7 of 14

Possible factors for accuracy fluctuation

  • Motion of objects in camera FoV

🡪 Switch to video of static 3D scenes

  • Lossy Image/Video compression, e.g., MJPEG/H.264 (Video)

🡪 disable compression

7

Static 3D scene

Object Motion

Compression

F2 = 13.0%

8 of 14

Possible factors for accuracy fluctuation

  • Motion of objects in camera FoV

🡪 Switched to video of static 3D scenes

  • Lossy Image/Video compression, MJPEG/H.264 (Video)

🡪 disabled compression

  • Environmental Condition change -- Flickering of light source

🡪 Switch to flicker-free light bulbs

8

Object Motion

Compression

Flicker

F2 = 8.7%

Static 3D scene

9 of 14

Possible factors for accuracy fluctuation

  • Motion of objects in camera FoV

🡪 Switch to video of static 3D scenes

  • Lossy Image/Video compression, e.g., MJPEG/H.264 (Video)

🡪 disable compression

  • Environmental Condition change -- Flickering of light source

🡪 Switch to flicker-free light bulbs

  • Does it only happen to some particular camera model?

🡪 try 2 other cameras

9

Object Motion

Compression

Flicker

Static 3D scene

10 of 14

The Truth Lies Within (the Camera)!

To capture visually pleasing and smooth video, video cameras try to find optimal values of automatically tuned (AUTO) parameters for each frame, which results in segments of 1 high-quality frame followed by some low-quality frames.

Example AUTO parameters: exposure time, shutter speed

10

11 of 14

Validation of Hypothesis

  • Methodology: narrowing the range of an AUTO camera parameter, max exposure time

11

Max exposure time = ¼ second

Max exposure time = 1/120 second

Reducing the max exposure time setting results in lowered accuracy fluctuations

12 of 14

A Quick Solution: Transfer learning using video datasets

  • Idea:
    • Retrain a SOTA detector (trained using image datasets) on video datasets using transfer learning to extract video-specific features

  • Case study:
    • Applying the idea to yolov5 reduces detection fluctuation (F2-metrix) from 42% to 32%
    • Which translates to better object tracking performance: ∼40% fewer mistakes in tracking

12

Roadway Dataset

Yolov5 trained on video content shows less detection fluctuation

3D Static Scene

13 of 14

Key Takeaways

  • The mystery: After eliminating possible factors, popular Deep learning models still show accuracy fluctuation in consecutive frames

  • The explanation: By automatically changing camera parameters to capture visually pleasing videos, the camera acts as an unintentional adversary

  • What can we do about it? Deep learning models trained on video specific characteristics can reduce this unwanted performance fluctuation

13

14 of 14

Thank you !

Any Questions?

14