1 of 43

SUDIPTA SARKAR

(Registration Number: A03-1112-0571-23)

​

Under the supervision of

​

Prof. Abir Das

​

(Assistant Professor, CSE Department, IIT Kharagpur. )

Super Image For Efficient Large Scale Video Action Recognition

​

May 19, 2025

2 of 43

1/22

Contents

1. Video Action Recognition

​

2. Motivation

​

3. Creation of Super Image

​

4. Related Works

5. Our Approach

​

6. Learning Framework

​

7. Experiments and Results

8. Conclusion

3 of 43

  1. Video Action Recognition

2/22

INPUT :

OUTPUT :

<Soccer_Penalty, 0.91>

Action Recognition

Model

Resource Efficient : Achieving high performance (e.g., accuracy) while using less computational power, memory, or energy.

Video Action Recognition is the task of identifying and classifying actions or activities happening in a video, such as "running," "jumping," or "cooking."

​

Figure: 1.1

Figure: 1.2

4 of 43

2. Motivation

3/22

  • Video action recognition is computation-heavy.�
  • Traditional models use 3D convolutions or video transformers.�
  • Goal: Improve efficiency without compromising accuracy.�

�

​

5 of 43

2. Motivation

3/22

  • Video action recognition is computation-heavy.�
  • Traditional models use 3D convolutions or video transformers.�
  • Goal: Improve efficiency without compromising accuracy.�
  • Can we use 2D vision transformers on videos via a clever frame-to-image transformation?

​

�

6 of 43

2. Motivation

3/22

  • Video action recognition is computation-heavy.�
  • Traditional models use 3D convolutions or video transformers.�
  • Goal: Improve efficiency without compromising accuracy.�
  • Can we use 2D vision transformers on videos via a clever frame-to-image transformation?

​

�

Yes.

7 of 43

2. Motivation

3/22

  • Video action recognition is computation-heavy.�
  • Traditional models use 3D convolutions or video transformers.�
  • Goal: Improve efficiency without compromising accuracy.�
  • Can we use 2D vision transformers on videos via a clever frame-to-image transformation?

​

  • Solution: Reformulate video task as an image classification problem using Super Images.

​

�

Yes.

8 of 43

3. Creation of Super Image

4/22

3 x 3 Super Image

Video

9 of 43

4. Related Works

5/22

Can an Image Classifier suffice for action recognition?

SIFAR | ICLR | 2022

​

Swin Transformer

Figure : 4.1

10 of 43

5. Our Approach

6/22

Instead of using Swin Transformer, we adopt the Hiera (Hierarchical Vision Transformer), a more recent transformer-based classifier, as the backbone of our model.

​

11 of 43

5. Our Approach

6/22

Instead of using Swin Transformer, we adopt the Hiera (Hierarchical Vision Transformer), a more recent transformer-based classifier, as the backbone of our model.

�

Why Hiera?

​

12 of 43

5. Our Approach

6/22

Instead of using Swin Transformer, we adopt the Hiera (Hierarchical Vision Transformer), a more recent transformer-based classifier, as the backbone of our model.

�

Why Hiera?

​

  • Simpler and more efficient architecture (52M Params)

​

  • Scales well to high-resolution inputs (e.g., 672×672, 896×896)

​

  • Lower computational overhead compared to Swin or MViT

​

13 of 43

5. Our Approach

7/22

Our Working Pipeline

Input Video

↓

Frame Extraction

↓

Super Image Creation

↓

Hiera Transformer (as Image Classifier)

↓

Predicted Action Class

​

​

14 of 43

5. Our Approach

7/22

Our Working Pipeline

SIFAR: Working Pipeline

Input Video

↓

Frame Extraction

↓

Super Image Creation

↓

Hiera Transformer (as Image Classifier)

↓

Predicted Action Class

​

​

Input Video

↓

Frame Extraction

↓

Super Image Creation

↓

Swin Transformer (as Image Classifier)

↓

Predicted Action Class

​

​

15 of 43

6. Learning Framework

8/22

Phase 1 : Two varieties of Super Images (input of the transformer)

Phase 2 : Hiera Transformer (as Super Image Classifier)

​

16 of 43

6. Learning Framework

9/22

Phase 1 : Two varieties of Super Images

Video

4x4 Super Image (896 × 896)

3 x3 Super Image (672 x 672)

17 of 43

6. Learning Framework

10/22

Phase 1 : Two varieties of Super Images

​

16 video frames arranged in a 4 × 4 grid. Each frame is cropped to 224 × 224 pixels, resulting in a final image resolution of 896 × 896.

8 video frames arranged in a 3 × 3 grid with one padded slot. Each frame is resized to 224 × 224 pixels, resulting in a final image resolution of 672 × 672.

18 of 43

6. Learning Framework

10/22

Phase 1 : Two varieties of Super Images

​

16 video frames arranged in a 4 × 4 grid. Each frame is cropped to 224 × 224 pixels, resulting in a final image resolution of 896 × 896.

8 video frames arranged in a 3 × 3 grid with one padded slot. Each frame is resized to 224 × 224 pixels, resulting in a final image resolution of 672 × 672.

224 x 224

224 x 224

19 of 43

6. Learning Framework

11/22

Phase 2: Hiera Transformer

<disk throw, 0.90>

Figure: 6.1 Super Image Classification using Hiera

Super Image as input

Output Label

20 of 43

12/22

7. Experiments and Results

21 of 43

13/22

7.1. Dataset Details

Dataset

Total videos

Source

Resolution(p)

Frame per sec

Avg length of clip (sec)

#classes

SSV2

220,847

Crowdsourced

240 – 720

12–30

4

174

Kinetics-400

306,245

YouTube

–

–

10

400

Table 1: Description of Something-Something V2 (SSV2) and Kinetics-400 datasets

22 of 43

13/22

7.1. Dataset Details

Dataset

Total videos

Source

Resolution(p)

Frame per sec

Avg length of clip (sec)

#classes

SSV2

220,847

Crowdsourced

240 – 720

12–30

4

174

Kinetics-400

306,245

YouTube

–

–

10

400

Table 1: Description of Something-Something V2 (SSV2) and Kinetics-400 datasets

“Bowling” in Kinetics-400

“Throwing something” in SSV2

Diversity of Classes Present in Kinetics 400 and SSV2

23 of 43

14/22

7.2. Training Description

​

Optimizer:

​

AdamW, LR (e.g. 5e-5, 4e-4), Warmup epochs (e.g. 5), Cosine LR decay

​

Training Config:

​

Epochs (e.g. 30, 50), Batch Size (e.g. 48, 128) Weight Decay (e.g. 0.05, 0.1)

​

Augmentation and Regularization:

​

Mixup, Cut Mix, Label Smoothing, Drop Path

​

Hardware:

​

NVIDIA H100 GPU, NVIDIA L40 GPU, NVIDIA Tesla V100 GPU (MIT-Satori)

​

​

24 of 43

15/22

7.3. Evaluation Description

Multi-Crop Testing (3 crops):� Center, left, and right spatial crops → Improves robustness to object location variation.�

Multi-Clip Testing (5 clips):� Five temporal segments per video → Better captures motion and context.�

Why It Matters:� Enhances accuracy and reliability by reducing sensitivity to Spatial and Temporal biases.

​

25 of 43

15/22

7.3. Evaluation Description

Multi-Crop Testing (3 crops):� Center, left, and right spatial crops → Improves robustness to object location variation.�

Multi-Clip Testing (5 clips):� Five temporal segments per video → Better captures motion and context.�

Why It Matters:� Enhances accuracy and reliability by reducing sensitivity to Spatial and Temporal biases.

​

Evaluation Metrics: (in accordance with SIFAR)

  • Top-1 Accuracy: Measures how often the top predicted label is correct.�
  • Top-5 Accuracy: Measures if the correct label is among the top 5 predictions.

​

​

26 of 43

16/22

Results

Table 2 : Our results on Kinetics-400 and SSV2

27 of 43

16/22

Results

Table 2 : Our results on Kinetics-400 and SSV2

28 of 43

16/22

Results

Table 2 : Our results on Kinetics-400 and SSV2

29 of 43

16/22

Results

Table 2 : Our results on Kinetics-400 and SSV2

30 of 43

16/22

Results

Table 2 : Our results on Kinetics-400 and SSV2

31 of 43

17/22

Results

Table 3: Comparison between SIFAR and Hiera on the Something-Something V2 (SSV2) dataset

32 of 43

17/22

Results

Table 3: Comparison between SIFAR and Hiera on the Something-Something V2 (SSV2) dataset

33 of 43

18/22

Results

Table 4: Comparison between SIFAR and Hiera on the Kinetics-400 dataset.

34 of 43

18/22

Results

Table 4: Comparison between SIFAR and Hiera on the Kinetics-400 dataset.

35 of 43

19/22

Results

Table 5: Compact summary of some significant experiments on SSV2

36 of 43

19/22

Results

Table 5: Compact summary of some significant experiments on SSV2

37 of 43

20/22

Results

Table 6: Compact summary of some significant experiments on Kinetics-400

38 of 43

20/22

Results

Table 6: Compact summary of some significant experiments on Kinetics-400

39 of 43

21/22

7.4. Conclusion

Summary of Contributions�

  • Use of Hiera Transformer: Employed a lightweight, hierarchical ViT to process super images efficiently.�
  • Simplified, Scalable Pipeline: No 3D convs or temporal attention—faster, easier to scale, and well-suited for high-res and real-time use.

​

  • Large-Scale Experiments: Performed experiments on Kinetics-400 and SSV2 datasets, exploring variations key hyperparameters to assess model performance and design trade-offs.�

​

40 of 43

22/22

7.4. Conclusion

Limitations & Future Directions

  • Kinetics-400 Performance: Lags behind SOTA—needs enhanced long-term temporal modeling.�
  • Augmentation Techniques: Currently limited to Mixup/CutMix—explore task-specific or adaptive augmentations.�

​

  • Flash Attention Integration: Use Flash Attention to reduce memory and speed up training/inference.

41 of 43

References

[1] Fuli Yu, Alexander Kirillov, Trevor Darrell, and Rui Collobert. Hiera: A Hierarchical Vision Transformer without the Bells-and-Whistles. NeurIPS, 2023.

​

[2] Haoqi Fan, Bo Xiong, Karttikeya Mangalam, Yanghao Li, Zhicheng Yan, Jitendra Malik, and Christoph Feichtenhofer. Multiscale Vision Transformers. ICCV, 2021.

​

[3] Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin Transformer: Hierarchical Vision Transformer using Shifted Windows. ICCV, 2021.

​

[4] Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, and Cordelia Schmid. SViT: Self-supervised Vision Transformers for Action Recognition. CVPR, 2021.

​

[5] Quanfu Fan, Richard Chen, and Rameswar Panda. Can an Image Classifier Suffice for Action Recognition? International Conference on Learning Representations (ICLR), 2022.

�[6] Karen Simonyan Brian Zhang Chloe Hillier Sudheendra Vijayanarasimhan Fabio Viola Tim Green Trevor Back Paul Natsev Mustafa Suleyman Andrew Zisserman Will Kay, Joao Carreira. The Kinetics Human Action Video Dataset. Computer Vision and Pattern Recognition, 2017.

�[7] Andrew Zisserman Joao Carreira. Quo Vadis, Action Recognition? A New Model and the Kinetics Dataset. Computer Vision and Pattern Recognition, 2017.

​

[8] Fabian Caba Heilbron; Victor Escorcia; Bernard Ghanem; Juan Carlos Niebles. ActivityNet: A large-scale video benchmark for human activity understanding. IEEE, 2015.

​

�[9] Mubarak Shah. Khurram Soomro, Amir Roshan Zamir. UCF101 - Action Recognition Data Set. Center for Research in Computer Vision, 2013.

​

[10] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. Attention Is All You Need. Advances in Neural Information Processing Systems (NeurIPS), 2017.

​

[11] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby.�

​

42 of 43

Thank You!

Any Questions

43 of 43

Appendix

​

​

Positional Embedding Interpolation

Enables model to handle variable input image sizes.

​

​

einops Library

Used to rearrange frames into a super image.

​

​

Group Augmentation

Applied on frame groups (8 or 16) for better temporal learning.

​

​

GPU Limitation

Large-scale experiments required high GPU resources.

​

ResNet-50 Results (8 Frames)

Kinetics-400: 69.67% (Top-1)

SSV2: 37.89% (Top-1)�

​