SUDIPTA SARKAR
(Registration Number: A03-1112-0571-23)
Under the supervision of
Prof. Abir Das
(Assistant Professor, CSE Department, IIT Kharagpur. )
Super Image For Efficient Large Scale Video Action Recognition
May 19, 2025
1/22
Contents
1. Video Action Recognition
2. Motivation
3. Creation of Super Image
4. Related Works
5. Our Approach
6. Learning Framework
7. Experiments and Results
8. Conclusion
2/22
INPUT :
OUTPUT :
<Soccer_Penalty, 0.91>
Action Recognition
Model
Resource Efficient : Achieving high performance (e.g., accuracy) while using less computational power, memory, or energy.
Video Action Recognition is the task of identifying and classifying actions or activities happening in a video, such as "running," "jumping," or "cooking."
Figure: 1.1
Figure: 1.2
2. Motivation
3/22
�
2. Motivation
3/22
�
2. Motivation
3/22
�
Yes.
2. Motivation
3/22
�
Yes.
3. Creation of Super Image
4/22
3 x 3 Super Image
Video
4. Related Works
5/22
Can an Image Classifier suffice for action recognition?
SIFAR | ICLR | 2022
Swin Transformer
Figure : 4.1
5. Our Approach
6/22
Instead of using Swin Transformer, we adopt the Hiera (Hierarchical Vision Transformer), a more recent transformer-based classifier, as the backbone of our model.
5. Our Approach
6/22
Instead of using Swin Transformer, we adopt the Hiera (Hierarchical Vision Transformer), a more recent transformer-based classifier, as the backbone of our model.
�
Why Hiera?
5. Our Approach
6/22
Instead of using Swin Transformer, we adopt the Hiera (Hierarchical Vision Transformer), a more recent transformer-based classifier, as the backbone of our model.
�
Why Hiera?
5. Our Approach
7/22
Our Working Pipeline
Input Video
↓
Frame Extraction
↓
Super Image Creation
↓
Hiera Transformer (as Image Classifier)
↓
Predicted Action Class
5. Our Approach
7/22
Our Working Pipeline
SIFAR: Working Pipeline
Input Video
↓
Frame Extraction
↓
Super Image Creation
↓
Hiera Transformer (as Image Classifier)
↓
Predicted Action Class
Input Video
↓
Frame Extraction
↓
Super Image Creation
↓
Swin Transformer (as Image Classifier)
↓
Predicted Action Class
6. Learning Framework
8/22
Phase 1 : Two varieties of Super Images (input of the transformer)
Phase 2 : Hiera Transformer (as Super Image Classifier)
6. Learning Framework
9/22
Phase 1 : Two varieties of Super Images
Video
4x4 Super Image (896 × 896)
3 x3 Super Image (672 x 672)
6. Learning Framework
10/22
Phase 1 : Two varieties of Super Images
16 video frames arranged in a 4 × 4 grid. Each frame is cropped to 224 × 224 pixels, resulting in a final image resolution of 896 × 896.
8 video frames arranged in a 3 × 3 grid with one padded slot. Each frame is resized to 224 × 224 pixels, resulting in a final image resolution of 672 × 672.
6. Learning Framework
10/22
Phase 1 : Two varieties of Super Images
16 video frames arranged in a 4 × 4 grid. Each frame is cropped to 224 × 224 pixels, resulting in a final image resolution of 896 × 896.
8 video frames arranged in a 3 × 3 grid with one padded slot. Each frame is resized to 224 × 224 pixels, resulting in a final image resolution of 672 × 672.
224 x 224
224 x 224
6. Learning Framework
11/22
Phase 2: Hiera Transformer
<disk throw, 0.90>
Figure: 6.1 Super Image Classification using Hiera
Super Image as input
Output Label
12/22
7. Experiments and Results
13/22
7.1. Dataset Details
Dataset | Total videos | Source | Resolution(p) | Frame per sec | Avg length of clip (sec) | #classes |
SSV2 | 220,847 | Crowdsourced | 240 – 720 | 12–30 | 4 | 174 |
Kinetics-400 | 306,245 | YouTube | – | – | 10 | 400 |
Table 1: Description of Something-Something V2 (SSV2) and Kinetics-400 datasets
13/22
7.1. Dataset Details
Dataset | Total videos | Source | Resolution(p) | Frame per sec | Avg length of clip (sec) | #classes |
SSV2 | 220,847 | Crowdsourced | 240 – 720 | 12–30 | 4 | 174 |
Kinetics-400 | 306,245 | YouTube | – | – | 10 | 400 |
Table 1: Description of Something-Something V2 (SSV2) and Kinetics-400 datasets
“Bowling” in Kinetics-400
“Throwing something” in SSV2
Diversity of Classes Present in Kinetics 400 and SSV2
14/22
7.2. Training Description
Optimizer:
AdamW, LR (e.g. 5e-5, 4e-4), Warmup epochs (e.g. 5), Cosine LR decay
Training Config:
Epochs (e.g. 30, 50), Batch Size (e.g. 48, 128) Weight Decay (e.g. 0.05, 0.1)
Augmentation and Regularization:
Mixup, Cut Mix, Label Smoothing, Drop Path
Hardware:
NVIDIA H100 GPU, NVIDIA L40 GPU, NVIDIA Tesla V100 GPU (MIT-Satori)
15/22
7.3. Evaluation Description
Multi-Crop Testing (3 crops):� Center, left, and right spatial crops → Improves robustness to object location variation.�
Multi-Clip Testing (5 clips):� Five temporal segments per video → Better captures motion and context.�
Why It Matters:� Enhances accuracy and reliability by reducing sensitivity to Spatial and Temporal biases.
15/22
7.3. Evaluation Description
Multi-Crop Testing (3 crops):� Center, left, and right spatial crops → Improves robustness to object location variation.�
Multi-Clip Testing (5 clips):� Five temporal segments per video → Better captures motion and context.�
Why It Matters:� Enhances accuracy and reliability by reducing sensitivity to Spatial and Temporal biases.
Evaluation Metrics: (in accordance with SIFAR)
16/22
Results
Table 2 : Our results on Kinetics-400 and SSV2
16/22
Results
Table 2 : Our results on Kinetics-400 and SSV2
16/22
Results
Table 2 : Our results on Kinetics-400 and SSV2
16/22
Results
Table 2 : Our results on Kinetics-400 and SSV2
16/22
Results
Table 2 : Our results on Kinetics-400 and SSV2
17/22
Results
Table 3: Comparison between SIFAR and Hiera on the Something-Something V2 (SSV2) dataset
17/22
Results
Table 3: Comparison between SIFAR and Hiera on the Something-Something V2 (SSV2) dataset
18/22
Results
Table 4: Comparison between SIFAR and Hiera on the Kinetics-400 dataset.
18/22
Results
Table 4: Comparison between SIFAR and Hiera on the Kinetics-400 dataset.
19/22
Results
Table 5: Compact summary of some significant experiments on SSV2
19/22
Results
Table 5: Compact summary of some significant experiments on SSV2
20/22
Results
Table 6: Compact summary of some significant experiments on Kinetics-400
20/22
Results
Table 6: Compact summary of some significant experiments on Kinetics-400
21/22
7.4. Conclusion
Summary of Contributions�
22/22
7.4. Conclusion
Limitations & Future Directions
References
[1] Fuli Yu, Alexander Kirillov, Trevor Darrell, and Rui Collobert. Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles. NeurIPS, 2023.
[2] Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale Vision Transformers. ICCV, 2021.
[3] Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. ICCV, 2021.
[4] Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, and Cordelia Schmid. SViT: Self-supervised Vision Transformers for Action Recognition. CVPR, 2021.
[5] Quanfu Fan, Richard Chen, and Rameswar Panda. Can an Image Classifier Suffice for Action Recognition? International Conference on Learning Representations (ICLR), 2022.
�[6] Karen Simonyan Brian Zhang Chloe Hillier Sudheendra Vijayanarasimhan Fabio Viola Tim Green Trevor Back Paul Natsev Mustafa Suleyman Andrew Zisserman Will Kay, Joao Carreira. The Kinetics Human Action Video Dataset. Computer Vision and Pattern Recognition, 2017.
�[7] Andrew Zisserman Joao Carreira. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. Computer Vision and Pattern Recognition, 2017.
[8] Fabian Caba Heilbron; Victor Escorcia; Bernard Ghanem; Juan Carlos Niebles. ActivityNet: A large-scale video benchmark for human activity understanding. IEEE, 2015.
�[9] Mubarak Shah. Khurram Soomro, Amir Roshan Zamir. UCF101 - Action Recognition Data Set. Center for Research in Computer Vision, 2013.
[10] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention Is All You Need. Advances in Neural Information Processing Systems (NeurIPS), 2017.
[11] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby.�
Thank You!
Any Questions
Appendix
Positional Embedding Interpolation
Enables model to handle variable input image sizes.
einops Library
Used to rearrange frames into a super image.
Group Augmentation
Applied on frame groups (8 or 16) for better temporal learning.
GPU Limitation
Large-scale experiments required high GPU resources.
ResNet-50 Results (8 Frames)
Kinetics-400: 69.67% (Top-1)
SSV2: 37.89% (Top-1)�