1 of 1

CHAPVIDMR: Chapter-based Video Moment Retrieval using Natural Language Queries

Uday Agarwal*1, Yogesh Kumar*1, Abu Shahid*1, Prajwal Gatti2, Manish Gupta3, Anand Mishra1

1IIT Jodhpur, 2University of Bristol, 3Microsoft

(*equal contribution)

Let’s use YouTube Chapters to answer user queries with CHAPVIDMR !

Introduction

Leveraging semantic content in YouTube video chapters to precisely identify multiple moments in a video relevant to the associated query.

CHAPVIDMR Dataset

Intro

Camera Audio Levels

External Recorder

Wind Filter

Conclusion

Query : What should I use to protect my microphone from wind noise and prevent audio distortion during a shoot?

Safety Channel

Input Video :

Task 1 : Chapter Classification-based Moment Retrieval

Task 2 : Segmentation-based Moment Retrieval

0:00

1:05

2:18

2:43

2:57

3:38

4:27

1. Avengers Tower Battle Overview

2. Details

3. Minifigs

4. Speed Build

5. Comparison

6. Final Thoughts

Query 1 [chapters (2,3)] : How does the Infinity Gauntlet included in the LEGO Avengers Tower set compare to Red Skull's rocket launcher in terms of design and playability?

Query 2 [chapters (5,6)]: Considering the price difference, which Avengers Tower set offers better value for money in terms of features and minifigures?

Query 3 [chapters (1,4)]: What makes building the new LEGO Avengers Tower a fun experience compared to other similar sets?

Query 1 [chapters (2,5)] : Considering the aerodynamic features and weight of the Victor Jet Speed S12, which type of badminton player would it best suit?

Query 2 [chapters 3,7)]: What are some drawbacks of the Victor Jet Speed S12 that might influence a decision to not purchase it?

Query 3 [chapters (1,9)]: How does the Volant Rogue S1 racket compare in terms of versatility and all-round play to the Victor Jet Speed S12 based on its rating metrics?

1. Rating metrics

2. Racket specs

3. Thoughts about the racket

4. Racket ratings

9. Final thoughts

5. Player type recommendations

Methods

Summary

  • We propose, the dataset CHAPVIDMR which extends VMR by utilizing YouTube video chapters and retrieving multiple moments.

  • We introduce a dataset curation pipeline that is cost-effective.

  • We benchmarked CHAPVIDMR on Chapter Classification based VMR and Segmentation based VMR tasks.

GPT Prompt

API

(ii) Representative Frames of each chapter

aZkTQwBxaHY

UFV6wukB_Rg

gWF342u6joo

(i) Chapter Visual

Captions

(iii) Chapter

Subtitles

(i) Chapter Visual

Captions

(iii) Chapter

Subtitles

(iv) Rules

(ii) Representative Frames of each chapter

Chapter clips

Offline

Videos

YouTube Video Ids

Meta Data

...With your DSLR or mirrorless-style camera.. and this automatically boost the preamp..to adjust to the change in volume

… two individuals likely engaged in a tech or photography-related..

  • Chapter name
  • Video descriptions
  • Subtitles
  • Video name

  • Queries must be natural ...
  • Query should be single sentence…
  • Pick non consecutive sentences..

mPLUG-Owl

… two individuals likely engaged in a tech or photography-related..

...With your DSLR or mirrorless-style camera.. and this automatically boost the preamp..to adjust to the change in volume

Data Generation Pipeline

['Wind Filter', 'Safety Channel']

What should I use…distortion during a shoot?

i1wdVtDolU

['How to…', '..enter password']

What steps should…access default password?

MB-SNjdGYiU

[0:30, 0:50], [0:10, 0:20]

[2:42, 2:57], [3:39, 4:27]

Query

Chapter

Time stamp

Video id

['Fidget toy…', '...Meet Fidgets']

What are some examples …causing distractions?

[1:20, 2:0],

[2:13, 3:02]

71PB_Rulk5M

Text query

Video Chapter

Chapter Text Metadata

Chapter Audio

SQV

SQT

SQA

“What should I … during a shoot?”

“… we are going to …”

Video Encoder

Text Encoder

Audio Encoder

Text Encoder

Softmax

Text query

“What should I … during a shoot?”

Video

Segmentation Model

0.01

0.5

0.4

0.07

0.02

top ranked retrieved chapters/segments

0.01

0.4

0.5

0.03

0.04

0.02

ׂ

Code and Dataset

We thank Microsoft for supporting this work through Microsoft Academic Partnership Grant (MAPG).

Dataset Comparison

Paper

Experiments

ICVGIP 2024