Subba Reddy Oota Akshett Jindal Ishani Mondal Khushbu Pahwa
Satya Sai Srinath Manish Shrivastava Maneesh Singh Bapi S. Raju Manish Gupta
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the brain)
2025
2
Language models (LMs) are trained to predict missing words
Language model
The
quick
brown
fox
[MASK]
jumps
3
Language models (LMs) predict brain activity evoked by complex
language (e.g. listening a story) to an impressive degree
Once
upon
a
time
Brain alignment of an LM ⇒ how similar its representations are to a human brain
Wehbe et al. 2014,
Jain and Huth 2018,
Gauthier and Levy 2019
Toneva and Wehbe 2019,
Caucheteux et al. 2020,
Toneva et al. 2020
Jain et al. 2020,
Schrimpf et al. 2021,
Goldstein et al. 2022
...
4
Language models (LMs) predict brain activity evoked by complex
language (e.g. listening a story) to an impressive degree
Jain and Huth. Incorporating context into language encoding models for fMRI. (NeurIPS 2018)
Toneva and Wehbe. Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain). (NeurIPS 2019)
Brain alignment of a LM ⇒ Advances in model size, instruction-tuning, and multimodality have improved alignment with neural data.
brain alignmenti = Pearson corr(true vi, pred vi)
5
Multimodal instruction tuning enables models to generalize to new tasks by following unseen instructions
Do multimodal instruction-tuned models prompted with natural language improve brain alignment and capture instruction-specific representations?
?
How does the brain integrate information during the processing of visual images?
?
How do multimodal instruction-tuned LLMs process visual images when guided by natural language task instructions?
6
Multi-modal Instruction-tuned LLMs (MLLMs): brain alignment
NSD dataset naturalistic Image stimulus
Task-specific instructions
Image Captioning:
What is the caption of the image?
Image Understanding:
Describe the most dominant color in the image.
Visual Relationship:
What objects are being used by the largest animal in this image?
Multimodal instruction-tuned model (MLLM)
Instruction Embedding
estimate alignment
Early Visual
7
Datasets & Models
To quantify model predictions, we have an estimate of the explainable variance and use that to measure normalize brain alignment.
NSD dataset naturalistic Image stimulus
8
Task-specific natural instructions
These tasks which are generally applicable to any image regardless of the contents in the image
9
How do MLLMs, unimodal and multi-modal models differ in their ability to predict brain activity in higher visual and early visual regions?
Result-1: MLLMs vs. Unimodal vs. Multi-modal models and brain alignment
10
11
Which task-specific instructions are highly correlated to visual function localizers?
Result-1: MLLMs vs. Unimodal vs. Multi-modal models and brain alignment
12
Result-2: Which task-specific instructions are highly correlated to visual function localizers?
S1: InstructBLIP
S1: mPLUG-Owl
13
Result-2: Which task-specific instructions are highly correlated to visual function localizers?
S1: InstructBLIP
S1: mPLUG-Owl
Do task-specific instructions from MLLMs account for visual concepts understanding?
14
Result-3: MLLMs capture count- and recognition-related visual concepts effectively across instructions
15
What is the unique and shared variance of each task-specific instruction to brain responses?
16
Result-4: Partitioning explained variance between task-specific instructions
17
Conclusions
18
Subba Reddy Oota
Correlating instruction-tuning (in multimodal models) with vision-language
processing (in the brain) (ICLR-2025)
Manish Gupta
Bapi S. Raju
Satya Sai Srinath
Khushbu Pahwa
Maneesh Singh
Ishani Mondal
Manish Shrivastava
Akshett Jindal