Shoplifting recognition from video surveillance cameras
02
04
03
01
Table of contents
Dataset
Data Preprocessing
Models
Best Model
“You can have all of the fancy tools, but if [your] data quality is not good, you’re nowhere.” — Veda Bawo, director of data governance, Raymond James
“Without clean data, or clean enough data, your data science is worthless.” — Michael Stonebraker,
Adjunct Professor, MIT
Dataset
01
Data! Data! Data! I can’t make bricks without clay!
Dataset Summary
565
436
129
Non Shoplifter
Shoplifter
The duration of each video is between 10 to 18 seconds with fps = 25
Total Number of Videos
Data Preprocessing
02
If you torture the data long enough, it will confess.
How to Preprocess a video?
A video is?
Pre-�processing
Sampling Frames
Randomly sample N frames from each video
Sample N frames from each video with a uniform distribution
Sample N frames from each video while giving more importance to middle frames
Subsampling in the temporal dimension
Preprocessing the Frames
We need to preprocess our sampled frames to match the input size of the model
Minor Class ?
Preprocessing Steps
Reshaping frame to expected size for most models.
Encode the labels
Temporal Subsampling
Regular Sampling
Now we are ready to start training our models!
Models
03
Used Models
Hybrid 2D approach
Mvit
R3D
A 3D convolutional neural network architecture designed for video understanding tasks.
A video understanding model based on the Multiscale Vision Transformer architecture.
More formally, using 2D convolutional backbone along with Recurrent units to capture temporal relations
Our Strategy
03
01
02
Hybrid 2D Approach
Hybrid 2D Approach
The Whole architecture
Hybrid Model Performance
Precision�0.75 %
Recall�14.28%
Didn’t stand for a lot of time
Accuracy�54.76%
R3d model
R3d Model Key Features
R3D leverages 3D convolutions to extract both spatial (image) and temporal (motion) features from videos, unlike standard CNNs limited to 2D images.
Inspired by ResNet, the R3D model uses residual connections to train deeper networks effectively by mitigating the vanishing gradient problem, boosting performance as the network depth increases.
3D Convolutional Layers
Residual Connections
R3d Model Performance
Precision�100 %
Recall�53%
Accuracy�76%
First place for a short of time
What is going on with the recall ?
We have zero false positives and very high false negatives. The model never says a positive if its not sure of it, and it says a negative easily !
MViT Model
MViT Model Key Features
MViT Model Performance
Precision�100 %
Recall�96%
First Place for most of phase 1
Accuracy�98%
Best Model
04
“The key is not the will to win... everybody has that. It is the will to prepare to win that is important.” —Bobby Knight
Performance On The Last Hidden Set
Precision�94 %
Recall�100%
F1-score�97%
Accuracy�97%
Our Performance : Lorethril
Happy Ending !
Thanks!
Do you have any
questions?
Mohammed
Alaa
Mohamed Ibrahim
Our Team