Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listening
Padakanti Srijith Khushbu Pahwa Radhika Mamidi
Manish Gupta Bapi S. Raju Subba Reddy Oota
2
Language models (LMs) predict brain activity evoked by complex
language (e.g. listening a story) to an impressive degree
Once
upon
a
time
Brain alignment of a LM ⇒ how similar its representations are to a human brain
Wehbe et al. 2014,
Jain and Huth 2018,
Gauthier and Levy 2019
Toneva and Wehbe 2019,
Caucheteux et al. 2020,
Toneva et al. 2020
Jain et al. 2020,
Schrimpf et al. 2021,
Goldstein et al. 2022
...
3
Language models (LMs) predict brain activity evoked by complex
language (e.g. listening a story) to an impressive degree
Jain and Huth. Incorporating context into language encoding models for fMRI. (NeurIPS 2018)
Toneva and Wehbe. Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain). (NeurIPS 2019)
Brain alignment of a LM ⇒ Which language models truly capture brain-relevant semantics?
brain alignmenti = Pearson corr(true vi, pred vi)
4
Oota et al. 2024 ACL
Text vs. Speech language models: Alignment in late language regions
Text-based
language models
Percentage of decrease:
Brain Alignment
Percentage of decrease:
Brain Alignment
AG
AG
IFG
IFG
Speech-based
language models
IFGOrb
IFGOrb
MFG
MFG
LTC
LTC
AC
AC
EVC
EVC
5
Multimodal (Text+Speech) models enables learning audio concepts from natural language supervision
CLAP
Pengi
Can speech embeddings from multimodal models capture brain-relevant semantics through cross-modal interactions?
6
Multi-modal vs. Unimodal models: Brain alignment
7
Which modality of representations in multi-modal models leads to high brain alignment?
Investigate via a residual approach
Toneva et al. 2022 Nature Computational Science, Oota et al. 2023 NeurIPS
Ridge Regression (g)
Ridge Regression (r)
Ridge Regression (g’)
residuals
8
Datasets & Model
To quantify model predictions, we have an estimate of the explainable variance and use that to measure normalized brain alignment.
Low-level Stimulus Features
Speech Features
fbank
MFCC
Mel-Spectrogram
Phonological
Articulation
Phonation
To examine brain-relevant semantics beyond low-level features, we used the same residual approach, controlling for low-level features from multimodal models.
10
How well do text/speech embeddings from multi-modal models predict speech-evoked brain activity over unimodal models?
Result-1: Multi-modal vs. Unimodal models & brain alignment
Result-1: Multi-modal vs. Unimodal models & brain alignment
Is there asymmetric knowledge transfer across modalities in multimodal models?
13
How is the alignment between brain recordings and multimodal model representations affected by the elimination of unimodal model features?
Result-2: Modality-specific contribution in knowledge transfer in multimodal models
Multimodal models show asymmetric cross-modal integration: lexical-semantic knowledge transfers from text to speech, while low-level acoustic information transfers from speech to text.
15
How is the alignment between brain recordings and multimodal model representations affected by the elimination of low-level stimulus features?
Result-3: Brain relevant semantics in multimodal models
Qualitative analysis: Topographic maps
Conclusions for neuro-AI research field
Aligning Text/Speech Representations from Multimodal Models with MEG Brain Activity During Listening (EMNLP-2025)
Manish Gupta
Bapi S. Raju
Radhika Mamidi
Khushbu Pahwa
Srijith Padakanti
Subba Reddy Oota