S Sakshi ,Utkarsh Tyagi ,Sonal Kumar ,Ashish Seth ,Ramaneswaran Selvakumar
Oriol Nieto ,Ramani Duraiswami , Sreyan Ghosh ,Dinesh Manocha
University of Maryland, College Park, USA Adobe, USA
Equal Contribution Equal Advising
MMAU: A Massive Multi-Task
Audio Understanding and Reasoning Benchmark
*
*
*
*
*
*
*
Motivation
Evaluating Advanced Audio Understanding and Reasoning
Comprehensive Skill Test
Extensive Domain Coverage
Diverse Task Types
Socio-cultural Music Understanding
Speaker Role Mapping
Tongue twisters
Scene Understanding
Understanding
Reasoning
World Knowledge
Speech
Sound
Music
Overview of the MMAU Benchmark. MMAU provides comprehensive coverage across three key domains: speech, sounds, and music, featuring diverse audio samples. It challenges multimodal LLMs with tasks across 27 distinct skills, requiring advanced audio perception, reasoning, and domain-specific knowledge.
Main Contributions
MMAU Vs. Prior Audio Benchmarks
MMAU Core Statistics
Skill Distribution in MMAU
Speech 33%
Music 45%
Sound 22%
Event-Based
Knowledge Retrieval
Phonemic Stress
Pattern Analysis
Musical Texture Interpretation
Instrumentation
Key highlight
Extraction
Conversational Fact
Retrieval
Rhythm and
Tempo Understanding
Eco-Acoustic
Knowledge
Harmony and Chord
Progressions
Sound-Based Event
Recognition
Melodic Structure Interpretation
(Left) Distribution of skills required for information extraction questions and (Right) for reasoning questions in the MMAU benchmark across the multiple domains. Each question in MMAU demands the model to apply one or more of these skills.
Speech 34%
Music 25%
Sound 42%
Counting
Emotion State summarisation
Emotion Flip
Detection
Temporal Event
Reasoning
Acoustic Scene
Reasoning
Event-Based Sound
Reasoning
Ambient Sound Interpretation
Acoustic Source
Inference
Lyrical Reasoning
Musical Genre
Reasoning
Emotional Tone
Interpretation
Multi Speaker Role
Mapping
Dissonant Emotion
Interpretation
Phonological Sequence
Decoding
Temporal Reasoning
Socio-cultural Interpretation
Examples of Question-Answer pairs in MMAU
Examples from the MMAU benchmark illustrating the diverse range of reasoning and information extraction tasks across the domains of sound, speech, and music.
MMAU Benchmark Construction Pipeline
A rigorous seven-step pipeline is used to curate MMAU
Source Selection
Task Curation
Expert Annotation & Filtering
Option Augmentation
Expert Review
MMAU
Performance comparison of various models on MMAU
Main Results
Main Results Contd ..
Deep Dive into Skill-Specific LALMs Performance
Understanding
Temporal Reasoning
Interpretation
Genre and Style Reasoning
Instrumentation
Rhythm and Tempo
Historical and Cultural Reasoning
Easy
Medium
Hard
Emotional Tone
Harmony and Chord Progressions
Phonemic Stress Pattern Analysis
Are LALMs Really Listening?
60
40
30
20
10
0
50
MuLLaMa
SALMONN
GAMA
Qwen2-Instruct
Gemini 1.5 Pro
Accuracy
Audio
Noise
Where are the LALMs Falling Short?
Distribution of error types across 500 instances for Qwen2-Audio-Instruct (Left) and Gemini 2.0 Flash (Right). The dominant error type is Perceptual Errors.
Future Work
Thank You