CHAPVIDMR: Chapter-based Video Moment Retrieval using Natural Language Queries
Uday Agarwal*1, Yogesh Kumar*1, Abu Shahid*1, Prajwal Gatti2, Manish Gupta3, Anand Mishra1
1IIT Jodhpur, 2University of Bristol, 3Microsoft
(*equal contribution)
Let’s use YouTube Chapters to answer user queries with CHAPVIDMR !
Introduction
Leveraging semantic content in YouTube video chapters to precisely identify multiple moments in a video relevant to the associated query.
CHAPVIDMR Dataset
…
Intro
Camera Audio Levels
External Recorder
Wind Filter
Conclusion
Query : What should I use to protect my microphone from wind noise and prevent audio distortion during a shoot?
Safety Channel
Input Video :
Task 1 : Chapter Classification-based Moment Retrieval
❌
✔
…
Task 2 : Segmentation-based Moment Retrieval
…
0:00
1:05
2:18
2:43
2:57
3:38
4:27
❌
❌
❌
✔
1. Avengers Tower Battle Overview
2. Details
3. Minifigs
4. Speed Build
5. Comparison
6. Final Thoughts
Query 1 [chapters (2,3)] : How does the Infinity Gauntlet included in the LEGO Avengers Tower set compare to Red Skull's rocket launcher in terms of design and playability?
Query 2 [chapters (5,6)]: Considering the price difference, which Avengers Tower set offers better value for money in terms of features and minifigures?
Query 3 [chapters (1,4)]: What makes building the new LEGO Avengers Tower a fun experience compared to other similar sets?
…
Query 1 [chapters (2,5)] : Considering the aerodynamic features and weight of the Victor Jet Speed S12, which type of badminton player would it best suit?
Query 2 [chapters 3,7)]: What are some drawbacks of the Victor Jet Speed S12 that might influence a decision to not purchase it?
Query 3 [chapters (1,9)]: How does the Volant Rogue S1 racket compare in terms of versatility and all-round play to the Victor Jet Speed S12 based on its rating metrics?
1. Rating metrics
2. Racket specs
3. Thoughts about the racket
4. Racket ratings
9. Final thoughts
5. Player type recommendations
…
Methods
…
Summary
GPT Prompt
API
(ii) Representative Frames of each chapter
aZkTQwBxaHY
UFV6wukB_Rg
…
gWF342u6joo
(i) Chapter Visual
Captions
(iii) Chapter
Subtitles
(i) Chapter Visual
Captions
(iii) Chapter
Subtitles
(iv) Rules
(ii) Representative Frames of each chapter
Chapter clips
Offline
Videos
YouTube Video Ids
Meta Data
…
...With your DSLR or mirrorless-style camera.. and this automatically boost the preamp..to adjust to the change in volume
… two individuals likely engaged in a tech or photography-related..
mPLUG-Owl
… two individuals likely engaged in a tech or photography-related..
...With your DSLR or mirrorless-style camera.. and this automatically boost the preamp..to adjust to the change in volume
Data Generation Pipeline
| | | |
| | | |
| | | |
| | | |
| | | |
['Wind Filter', 'Safety Channel']
What should I use…distortion during a shoot?
i1wdVtDolU
['How to…', '..enter password']
What steps should…access default password?
MB-SNjdGYiU
[0:30, 0:50], [0:10, 0:20]
…
…
…
[2:42, 2:57], [3:39, 4:27]
…
Query
Chapter
Time stamp
Video id
['Fidget toy…', '...Meet Fidgets']
What are some examples …causing distractions?
[1:20, 2:0],
[2:13, 3:02]
71PB_Rulk5M
Text query
Video Chapter
Chapter Text Metadata
Chapter Audio
SQV
SQT
SQA
“What should I … during a shoot?”
“… we are going to …”
Video Encoder
Text Encoder
Audio Encoder
Text Encoder
…
Softmax
…
Text query
“What should I … during a shoot?”
Video
Segmentation Model
0.01
0.5
0.4
0.07
0.02
top ranked retrieved chapters/segments
0.01
0.4
0.5
0.03
0.04
0.02
ׂ
…
Code and Dataset
We thank Microsoft for supporting this work through Microsoft Academic Partnership Grant (MAPG).
Dataset Comparison
Paper
Experiments
ICVGIP 2024