1 of 19

Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listening

Padakanti Srijith Khushbu Pahwa Radhika Mamidi

Manish Gupta Bapi S. Raju Subba Reddy Oota

2 of 19

2

Language models (LMs) predict brain activity evoked by complex

language (e.g. listening a story) to an impressive degree

Once

upon

a

time

Brain alignment of a LM ⇒ how similar its representations are to a human brain

Wehbe et al. 2014,

Jain and Huth 2018,

Gauthier and Levy 2019

Toneva and Wehbe 2019,

Caucheteux et al. 2020,

Toneva et al. 2020

Jain et al. 2020,

Schrimpf et al. 2021,

Goldstein et al. 2022

...

3 of 19

3

Language models (LMs) predict brain activity evoked by complex

language (e.g. listening a story) to an impressive degree

Jain and Huth. Incorporating context into language encoding models for fMRI. (NeurIPS 2018)

Toneva and Wehbe. Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain). (NeurIPS 2019)

Brain alignment of a LM ⇒ Which language models truly capture brain-relevant semantics?

brain alignmenti = Pearson corr(true vi, pred vi)

4 of 19

4

Oota et al. 2024 ACL

Text vs. Speech language models: Alignment in late language regions

Text-based

language models

Percentage of decrease:

Brain Alignment

Percentage of decrease:

Brain Alignment

AG

AG

IFG

IFG

Speech-based

language models

IFGOrb

IFGOrb

MFG

MFG

LTC

LTC

AC

AC

EVC

EVC

  • Strong alignment of Text models in late language regions is not due to low-level features, but driven by brain-relevant semantics
  • Speech models alignment in late language regions is entirely due to low-level stimulus features, lacking brain-relevant semantics

5 of 19

5

Multimodal (Text+Speech) models enables learning audio concepts from natural language supervision

CLAP

Pengi

Can speech embeddings from multimodal models capture brain-relevant semantics through cross-modal interactions?

6 of 19

6

Multi-modal vs. Unimodal models: Brain alignment

  • How well do text/speech embeddings from multi-modal models predict speech-evoked brain activity over unimodal models?
  • Is there asymmetric knowledge transfer across modalities in multimodal models, or do multimodal-text and multimodal-speech perform equally well?

7 of 19

7

Which modality of representations in multi-modal models leads to high brain alignment?

Investigate via a residual approach

Toneva et al. 2022 Nature Computational Science, Oota et al. 2023 NeurIPS

Ridge Regression (g)

 

 

Ridge Regression (r)

 

Ridge Regression (g’)

 

 

 

 

residuals

 

8 of 19

8

Datasets & Model

  • Brain: MEG recordings from MEG-MASC dataset [Gwilliams et al. 2022]
    • Listening to the four naturalistic stories
    • N=27 (# subjects)

  • 4 text-based language models
    • BERT-base, XLNet, FLAN-T5, LLaMA-2

  • 3 speech-based language models
    • Wav2Vec2.0, WavLM, Whisper

  • 3 multi-modal Transformer models
    • CLAP, Pengi, SpeechT5

To quantify model predictions, we have an estimate of the explainable variance and use that to measure normalized brain alignment.

9 of 19

Low-level Stimulus Features

Speech Features

fbank

MFCC

Mel-Spectrogram

Phonological

Articulation

Phonation

To examine brain-relevant semantics beyond low-level features, we used the same residual approach, controlling for low-level features from multimodal models.

10 of 19

10

How well do text/speech embeddings from multi-modal models predict speech-evoked brain activity over unimodal models?

11 of 19

Result-1: Multi-modal vs. Unimodal models & brain alignment

  • Multimodal Text Embeddings: show a peak around 200 ms, suggesting that they benefit from speech embeddings.
  • Text embeddings (multimodal & unimodal) outperform speech embeddings beyond 350 ms - aligning with the timeframe associated with semantic word processing.

12 of 19

Result-1: Multi-modal vs. Unimodal models & brain alignment

  • Multimodal Text Embeddings: show a peak around 200 ms, suggesting that they benefit from speech embeddings.
  • Text embeddings (multimodal & unimodal) outperform speech embeddings beyond 350 ms - aligning with the timeframe associated with semantic word processing.

Is there asymmetric knowledge transfer across modalities in multimodal models?

13 of 19

13

How is the alignment between brain recordings and multimodal model representations affected by the elimination of unimodal model features?

14 of 19

Result-2: Modality-specific contribution in knowledge transfer in multimodal models

  • Removing Unimodal Speech embeddings from Multimodal Text:
    • significantly impacts predictivity around ~200ms, suggests that multimodal text embeddings incorporate additional speech-derived information.
  • Removing Unimodal Text embeddings from Multimodal Speech:
    • primarily impacts around 200-350 ms, a timeframe associated with lexical-semantic (word-level) processing.

Multimodal models show asymmetric cross-modal integration: lexical-semantic knowledge transfers from text to speech, while low-level acoustic information transfers from speech to text.

15 of 19

15

How is the alignment between brain recordings and multimodal model representations affected by the elimination of low-level stimulus features?

16 of 19

Result-3: Brain relevant semantics in multimodal models

  • Multimodal Text Embeddings:
    • high brain alignment due to brain-relevant semantics
  • Multimodal Speech Embeddings:
    • brain alignment mostly due to low-level speech features

17 of 19

Qualitative analysis: Topographic maps

18 of 19

Conclusions for neuro-AI research field

  1. Multimodal Speech models are useful for modeling early auditory and lexical word processing. Need to investigate speech models to learn more about word processing

  • Multimodal Text models are useful for modeling both early auditory and high-level semantic information processing

  • More work to do for a complete end-to-end multimodal model capable of bi-directional cross-knowledge transfer between text and speech

19 of 19

Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listening (EMNLP-2025)

Manish Gupta

Bapi S. Raju

Radhika Mamidi

Khushbu Pahwa

Srijith Padakanti

Subba Reddy Oota