1 of 17

S Sakshi ,Utkarsh Tyagi ,Sonal Kumar ,Ashish Seth ,Ramaneswaran Selvakumar

Oriol Nieto ,Ramani Duraiswami , Sreyan Ghosh ,Dinesh Manocha

University of Maryland, College Park, USA Adobe, USA

Equal Contribution Equal Advising

MMAU: A Massive Multi-Task

Audio Understanding and Reasoning Benchmark

*

*

*

*

*

*

*

2 of 17

Motivation

    • Large Language Models (LLMs) and Large Multimodal Models (LMMs) are driving AI closer to AGI with expert-level performance across diverse tasks.

    • Existing benchmarks often lack the complexity and real-world relevance required to assess advanced perception and reasoning abilities.

    • Despite the importance of audio perception in human intelligence, current evaluations of Large Audio-Language Models (LALMs) focus on basic tasks (e.g., ASR, classification), failing to capture the complex reasoning necessary for achieving AGI.

3 of 17

Evaluating Advanced Audio Understanding and Reasoning

Comprehensive Skill Test

Extensive Domain Coverage

Diverse Task Types

Socio-cultural Music Understanding

Speaker Role Mapping

Tongue twisters

Scene Understanding

Understanding

Reasoning

World Knowledge

Speech

Sound

Music

Overview of the MMAU Benchmark. MMAU provides comprehensive coverage across three key domains: speech, sounds, and music, featuring diverse audio samples. It challenges multimodal LLMs with tasks across 27 distinct skills, requiring advanced audio perception, reasoning, and domain-specific knowledge.

4 of 17

Main Contributions

    • Introduce MMAU, the first benchmark for evaluating advanced audio perception and reasoning in LALMs, with 10k expertly annotated instances covering speech, sounds, and music.

    • Assessment of 18 models showing that even advanced LALMs achieve only 53% accuracy, revealing gaps in audio understanding.

    • Indepth analysis to reveal insights into model responses, highlighting challenges in audio input processing, skill-wise performance, and the impact of audio captions for text-only models.

5 of 17

MMAU Vs. Prior Audio Benchmarks

6 of 17

MMAU Core Statistics

    • Overall 10K MCQ questions
      • test-mini (public) : 1K
      • test (private): 9K

    • Assesses models across 27 distinct skills across two question types
      • Information Extraction
      • Reasoning

7 of 17

Skill Distribution in MMAU

Speech 33%

Music 45%

Sound 22%

Event-Based

Knowledge Retrieval

Phonemic Stress

Pattern Analysis

Musical Texture Interpretation

Instrumentation

Key highlight

Extraction

Conversational Fact

Retrieval

Rhythm and

Tempo Understanding

Eco-Acoustic

Knowledge

Harmony and Chord

Progressions

Sound-Based Event

Recognition

Melodic Structure Interpretation

(Left) Distribution of skills required for information extraction questions and (Right) for reasoning questions in the MMAU benchmark across the multiple domains. Each question in MMAU demands the model to apply one or more of these skills.

Speech 34%

Music 25%

Sound 42%

Counting

Emotion State summarisation

Emotion Flip

Detection

Temporal Event

Reasoning

Acoustic Scene

Reasoning

Event-Based Sound

Reasoning

Ambient Sound Interpretation

Acoustic Source

Inference

Lyrical Reasoning

Musical Genre

Reasoning

Emotional Tone

Interpretation

Multi Speaker Role

Mapping

Dissonant Emotion

Interpretation

Phonological Sequence

Decoding

Temporal Reasoning

Socio-cultural Interpretation

8 of 17

Examples of Question-Answer pairs in MMAU

Examples from the MMAU benchmark illustrating the diverse range of reasoning and information extraction tasks across the domains of sound, speech, and music.

9 of 17

MMAU Benchmark Construction Pipeline

A rigorous seven-step pipeline is used to curate MMAU

Source Selection

Task Curation

Expert Annotation & Filtering

Option Augmentation

Expert Review

MMAU

10 of 17

Performance comparison of various models on MMAU

11 of 17

Main Results

    • MMAU poses a significant challenge. The best-performing LALM achieves only 53% accuracy, while a cascaded captioning + LLM approach reaches 59%, far below human performance at 82%.
    • Minimal gap between open-source and proprietary models. Qwen2, the top open-access model, performs almost on par with proprietary Gemini-Pro, with only a 0.47% difference, while the fully open-source GAMA lags by 21%.
    • Generalized vs. Specialized Models. Models trained across multiple domains (e.g., Qwen2-Audio, LTU-AS, Gemini) outperform specialized models, indicating that diverse training data enhances audio understanding.

12 of 17

Main Results Contd ..

    • Models perform best on sound and worst on speech. With scores of 18% (speech), 30% (sound), and 23% (music), models excel at sound but struggle with speech reasoning, suggesting ongoing challenges in complex audio perception.
    • Cascaded approaches outperform others. Captioning audio before prompting LLMs yields the highest accuracy, highlighting the benefits of improving both audio and language reasoning separately.

13 of 17

Deep Dive into Skill-Specific LALMs Performance

    • Accuracy for Gemini 2.0 Flash across easy, medium, and hard questions, categorized by skills.

    • LALMs excel in some skills across all difficulty levels (e.g., Phonemic Stress Pattern Analysis) but struggle with others (e.g., Temporal Reasoning) regardless of difficulty.

Understanding

Temporal Reasoning

Interpretation

Genre and Style Reasoning

Instrumentation

Rhythm and Tempo

Historical and Cultural Reasoning

Easy

Medium

Hard

Emotional Tone

Harmony and Chord Progressions

Phonemic Stress Pattern Analysis

14 of 17

Are LALMs Really Listening?

    • To assess LALMs' attention to audio, we replace the original audio in MMAU's test set with Gaussian noise and compare performance

    • MuLLaMa and SALMONN show little change, indicating limited audio reliance, while others drop significantly, suggesting greater dependence on audio inputs.

60

40

30

20

10

0

50

MuLLaMa

SALMONN

GAMA

Qwen2-Instruct

Gemini 1.5 Pro

Accuracy

Audio

Noise

15 of 17

Where are the LALMs Falling Short?

Distribution of error types across 500 instances for Qwen2-Audio-Instruct (Left) and Gemini 2.0 Flash (Right). The dominant error type is Perceptual Errors.

16 of 17

Future Work

    • MMAU-pro is almost here! Larger skill-set, multi-hop reasoning, multiple audios, variable length (and long up to 10 minutes!), speech + sounds, and much more!

    • Improving certain aspects from community feedback:
      1. Reduce language priors in the question
      2. Have more options and specifically a "none of the above option"
      3. Include open-ended QAs beyond just MCQ
      4. Work on improving the evaluation algorithm.

17 of 17

Thank You