1 of 13

Subba Reddy Oota Manish Gupta Mariya Toneva

Joint processing of linguistic properties in brains and language models

2 of 13

2

Language models (LMs) predict brain activity evoked by complex

language (e.g. listening a story) to an impressive degree

Once

upon

a

time

Jain and Huth. Incorporating context into language encoding models for fMRI. (NeurIPS 2018)

Toneva and Wehbe. Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain). (NeurIPS 2019)

3 of 13

3

Language models (LMs) predict brain activity evoked by complex

language (e.g. listening a story) to an impressive degree

Once

upon

a

time

Jain and Huth. Incorporating context into language encoding models for fMRI. (NeurIPS 2018)

Toneva and Wehbe. Interpreting and improving natural-language processing (in machines) with natural language-processing (in the brain). (NeurIPS 2019)

Brain alignment of a LM ⇒ Why do language models have better brain alignment? What are the reasons?

4 of 13

4

Language models (LMs) are trained to predict missing words

Language model

The

quick

brown

fox

[MASK]

jumps

5 of 13

5

Language models (LMs) are trained to predict missing words

Language model

The

quick

brown

fox

[MASK]

jumps

BERT composes a hierarchy of linguistic signals ranging from surface to semantic features.

Surface

Syntactic

Semantic

6 of 13

6

What are the reasons behind the success of LMs?

BERT composes a hierarchy of linguistic signals ranging from surface to semantic features.

7 of 13

7

The strongest alignment with high-level language brain regions has consistently been observed in middle layers

BERT

XLM

Toneva et al. 2019

Caucheteux et al. 2022

Across several types of large NLP systems, best alignment with fMRI in middle layers

8 of 13

8

What are the reasons for this observed brain alignment?

Investigate via a perturbation approach

fMRI

Linguistic property

Language model

Significant �difference Ling. prop. affects alignment

Residual =

Original encoding performance

Residual encoding performance

Naturalistic stimulus

This is Los Angeles. And it's the

9 of 13

9

Removal of each linguistic property leads to a significant decrease in brain alignment across layers.

Does the removal of a linguistic property affects the alignment between LM and brain across all layers?

10 of 13

10

Removal of each linguistic property leads to a significant decrease in brain alignment across layers.

11 of 13

11

Which linguistic properties have the most influence on the trend of brain alignment across BERT layers?

Syntactic

Semantic

ROI-Level Analysis

Syntactic properties have the largest effect on the trend of brain alignment across model layers

Corrtask (accuracytask – accuracytask-residual , brain alignment of BERT – brain alignmenttask-residual)

12 of 13

12

Qualitative Analysis: Effect of each linguistic property

effect of surface property

effect of syntactic property

effect of semantic property

Top Constituents has the largest effect on the trend in brain alignment across BERT layers for all language regions

Several linguistic properties may play a significant role in local trends:

Object Number for ATL and IFGOrb regions, Tense for PCC regions, Word Length and Subject Number for PFm sub-region

13 of 13

Bridging AI and Neuroscience (BrAIN) group

Subba Reddy Oota

Mariya Toneva

Joint Processing of linguistic properties in brains and language models (NeurIPS 2023)

Manish Gupta