Subba Reddy Oota Khushbu Pahwa Mounika Marreddy
Maneesh Singh Manish Gupta Bapi S. Raju
Multi-modal brain encoding models for multi-modal stimuli
2025
2
Language models (LMs) predict brain activity evoked by complex
language (e.g. listening a story) to an impressive degree
Once
upon
a
time
Brain alignment of an LM ⇒ how similar its representations are to a human brain
Wehbe et al. 2014,
Jain and Huth 2018,
Gauthier and Levy 2019
Toneva and Wehbe 2019,
Caucheteux et al. 2020,
Toneva et al. 2020
Jain et al. 2020,
Schrimpf et al. 2021,
Goldstein et al. 2022
...
3
Language models (LMs) predict brain activity evoked by complex
language (e.g. listening a story) to an impressive degree
Jain and Huth. Incorporating context into language encoding models for fMRI. (NeurIPS 2018)
Toneva and Wehbe. Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain). (NeurIPS 2019)
brain alignmenti = Pearson corr(true vi, pred vi)
4
Multi-modal Transformer models can predict visual brain activity impressively well, even with text modality representations
How accurately do multi-modal models predict brain activity evoked by multi-modal stimuli?
Multi-modal vs. Unimodal models: Brain alignment
Visual regions
Auditory regions
Language regions
Video included with Audio
“The wolf of wall street”
movie video clip
Multi-modal naturalistic stimulus
fMRI
Actual brain
activations
Uni-modal Video Model (VM)
Uni-modal Speech Model (SM)
Cross-modality
Model (CM)
Jointly-pretrained Model (JM)
+
Ridge Regression (g)
Ridge Regression (h)
Video Encoder
Audio Encoder
6
Cross-modality
Model (CM)
Ridge Regression (g)
Video Encoder
Audio Encoder
Which modality of representations in multi-modal models leads to high brain alignment?
Investigate via a residual approach
Toneva et al. 2022 Nature Computational Science, Oota et al. 2023 NeurIPS
Video included with Audio
“The wolf of wall street”
movie video clip
Multi-modal naturalistic stimulus
7
Datasets & Models
To quantify model predictions, we have an estimate of the explainable variance and use that to measure normalize brain alignment.
Video included with Audio
“The wolf of wall street”
movie video clip
Multi-modal naturalistic stimulus
8
Multi-modal stimulus: How do multi-modal and unimodal models differ in their ability to predict brain activity in late language regions, higher visual regions and early sensory regions?
Result-1: Multi-modal vs. Unimodal models & brain alignment
Result-1: Multi-modal vs. Unimodal models & brain alignment
Which brain regions process unimodal versus multi-modal information?
Result-2: Which brain regions process unimodal versus multi-modal information?
Result-2: Which brain regions process unimodal versus multi-modal information?
Which modality of representations in multi-modal models lead to high brain alignment?
13
How is the alignment between brain recordings and multi-modal model representations affected by the elimination of modality-specific features?
Result-3: Modality-specific contribution in language and visual regions
Result-3: Modality-specific contribution in language and visual regions
Qualitative Analysis: Effect of removal of modality-specific features
Conclusions
Subba Reddy Oota
Multi-modal brain encoding models for multi-modal stimuli (ICLR-2025)
Manish Gupta
Bapi S. Raju
Mounika Marreddy
Khushbu Pahwa
Maneesh Singh