1 of 205

Language and the Brain: Deep Learning for Brain Encoding and Decoding

Subba Reddy Oota1, Manish Gupta2,3, Raju S. Bapi2

1Inria Bordeaux, France; 2IIIT Hyderabad, India; 3Microsoft, India

subba-reddy.oota@inria.fr, gmanish@microsoft.com, raju.bapi@iiit.ac.in

2 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

2

3 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

3

4 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
    • Brain Encoding/Decoding: Techniques and Research Goals
    • Introduction to popular text datasets
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

4

5 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
    • Brain Encoding/Decoding: Techniques and Research Goals
    • Introduction to popular text datasets
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

5

6 of 205

Neuroscience

  • Field of science that studies the structure and function of the nervous system of different species.
  • Involves answering interesting questions
    • How learning occurs during adolescence, and how it differs from the way adults learn and form memories.
    • Which specific cells in the brain (and what connections they form with other cells), have a role in how memories are formed.
    • How animals cancel out irrelevant information arriving from the senses and focus only on information that matters.
    • How do humans make decisions.
    • How humans develop speech and learn languages.
  • Neuroscientists study diverse topics that help us understand how the brain and nervous system work.

IJCNN 2023: DL for Brain Encoding and Decoding

6

7 of 205

Brain encoding and decoding in cognitive neuroscience

  •  

IJCNN 2023: DL for Brain Encoding and Decoding

7

8 of 205

Brain encoding and decoding

  •  

IJCNN 2023: DL for Brain Encoding and Decoding

8

9 of 205

Techniques for studying the brain function

  • fMRI: high spatial but low time resolution.
    • Good to study a specific location in the brain
    • Unsuitable for sentence-level analysis. fMRI takes about two seconds to complete a scan. This is far lower than the speed at which humans can process language.
    • Cannot capture syntactic information (Gauthier and Levy, 2019)
  • EEG: high time but low spatial resolution.
    • Can preserve rich syntactic information (Hale et al., 2018)
    • But cannot use for source analysis.
  • fNIRS: compromise option
    • Time resolution better than fMRI
    • Spatial resolution better than EEG
    • Balance of spatial and temporal resolution may not be enough to compensate for the loss in both.

IJCNN 2023: DL for Brain Encoding and Decoding

9

Single Micro-Electrode (ME), Micro-Electrode array (MEA), Electro-Cortico Graphy (ECoG), Positron emission tomography (PET), functional MRI (fMRI), Magneto-encephalography (MEG), Electro-encephalography (EEG), Near-Infrared Spectroscopy (NIRS)

10 of 205

fMRI

  • No injections, surgery, the ingestion of substances, or exposure to ionizing radiation.
  • The primary form of fMRI uses the blood-oxygen-level dependent (BOLD) contrast, discovered by Seiji Ogawa in 1990.
    • Measures brain activity by detecting changes associated with blood flow.
    • When an area of the brain is in use, blood flow to that region also increases.
  • Hemodynamic response (HRF)
    • It takes a while for the vascular system to respond to the brain's need for glucose.
    • Blood flow lags the neuronal events triggering it by about 5 seconds.

IJCNN 2023: DL for Brain Encoding and Decoding

10

An fMRI image with yellow areas showing increased activity compared with a control condition

11 of 205

Computational Cognitive Science Research goals

  • Predictive Accuracy
    • Compare feature sets: Which feature set provides the most faithful reflection of the neural representational space?
    • Test feature decodability: “Does neural data Y contain information about features X?”
    • Build accurate models of brain data: Aim is to enable simulations of neuroscience experiments.
  • Interpretability
    • Examine individual features: Which features contribute the most to neural activity?
    • Test correspondences between representational spaces
      • “CNNs vs ventral visual stream” or “Two text representations”
    • Interpret feature sets
      • Do features X, generated by a known process, accurately describe the space of neural responses Y?
      • Do voxels respond to a single feature or exhibit mixed selectivity?
    • How does the mapping relate to other models or theories of brain function?

IJCNN 2023: DL for Brain Encoding and Decoding

11

12 of 205

Computational Cognitive Science Research goals

  • Biological plausibility
    • Simulate linear readout
      • If the features can be extracted with a linear mapping model, it means that they require few additional computations in order to be used downstream.
    • Incorporate measurement-related considerations
      • Rather than assuming a fixed HRF across voxels and/or conditions, what are better ways?

IJCNN 2023: DL for Brain Encoding and Decoding

12

13 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
    • Brain Encoding/Decoding: Techniques and Research Goals
    • Introduction to popular text datasets
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

13

14 of 205

Forms of stimulus presentation and data collection

  • Type: fMRI, EEG, MEG, …
  • TR: Sampling time.
  • Fixation points: location, color, shape.
  • Form of stimuli presentation: text, video, audio, images.
  • Task: question answering, property generation, understanding, …
  • Time given to participants: 1 minute to list properties, …
  • Type of participants: males/females, sighted/blind, …
  • Number of times the response to stimuli was recorded.
  • Language

IJCNN 2023: DL for Brain Encoding and Decoding

14

15 of 205

Text Stimulus Datasets

IJCNN 2023: DL for Brain Encoding and Decoding

15

Dataset

Type

Language

Stimulus

#Subjects

Paradigm

Size

Task

Wehbe et al., 2014

fMRI

English

Chapter 9 of Harry Potter and the Sorcerer's Stone

9

Reading stories

5000 word chapter was presented in 45 minutes.

Story understanding

Handjaras et al., 2016

fMRI

Italian

Verbal, pictorial or auditory presentation of 40 concrete nouns

20

Reading, viewing or listening

40 nouns * 4 times.

Property Generation

Anderson et al., 2017

fMRI

Italian

70 concrete and abstract nouns from law/music.

7

Reading

70 nouns * 5 times.

Imagine a situation that they personally associate with the noun

Zurich Cognitive Language Processing Corpus (ZuCo): Hollenstein et al., 2018

EEG and eye-tracking

English

Sentences from movie reviews or Wikipedia

12

Reading natural sentences

21,629 words in 1107 sentences and 154,173 fixations

Rate movie quality, answer control questions, check for existence of a relation

Anderson et al., 2019

fMRI

English

240 active voice sentences describing everyday situations

14

Reading

240 sentences seen 12 times (by 10 subjects) and 6 times (by 4 subjects)

Passive reading

BCCWJ-EEG: Oseki and Asahara, 2020

EEG

Japanese

20 newspaper articles

40

Reading

1 time reading for ~30-40 minutes

Passive reading

16 of 205

Data for concrete nouns from sighted/blind subjects

  • Participants were asked to verbally enumerate in one minute the properties (features) that describe the entities the words refer to.
  • 4 groups of participants
    • 5 sighted individuals were presented with a pictorial form of the nouns
    • 5 sighted individuals with a verbal visual (i.e., written Italian words) form
    • 5 sighted individuals with a verbal auditory (i.e., spoken Italian words) form
    • 5 congenitally blind with a verbal auditory form.

IJCNN 2023: DL for Brain Encoding and Decoding

16

17 of 205

70 - Italian word stimuli fMRI data

  • Taxonomic categories in law and music domain
    • Ur-abstract: that are classified as abstract in WordNet
    • Attribute: A construct whereby objects or individuals can be distinguished
    • Communication: Something that is communicated by, to or between groups
    • Event/action: Something that happens at a given place and time
    • Person/Social role: Individual, someone, somebody, mortal
    • Location: Points or extents in space
    • Object/Tool: A class of unambiguously concrete nouns

IJCNN 2023: DL for Brain Encoding and Decoding

17

18 of 205

Zurich Cognitive Language Processing Corpus (ZuCo)

  • Personal reading speed.
    • Sentences were presented to the subjects in a naturalistic reading scenario
    • Complete sentence is presented on the screen
    • Subjects read each sentence at their own speed, i.e., the reader determines for how long each word is fixated and which word to fixate next.

IJCNN 2023: DL for Brain Encoding and Decoding

18

19 of 205

Text Stimulus Datasets References

  • Wehbe, Leila, Brian Murphy, Partha Talukdar, Alona Fyshe, Aaditya Ramdas, and Tom Mitchell. "Simultaneously uncovering the patterns of brain regions involved in different story reading subprocesses." PloS one 9, no. 11 (2014): e112575.
  • Hollenstein, Nora, Jonathan Rotsztejn, Marius Troendle, Andreas Pedroni, Ce Zhang, and Nicolas Langer. "ZuCo, a simultaneous EEG and eye-tracking resource for natural sentence reading." Scientific data 5, no. 1 (2018): 1-13.
  • Handjaras, Giacomo, Emiliano Ricciardi, Andrea Leo, Alessandro Lenci, Luca Cecchetti, Mirco Cosottini, Giovanna Marotta, and Pietro Pietrini. "How concepts are encoded in the human brain: a modality independent, category-based cortical organization of semantic knowledge." Neuroimage 135 (2016): 232-242.
  • Anderson, Andrew J., Douwe Kiela, Stephen Clark, and Massimo Poesio. "Visually grounded and textual semantic models differentially decode brain activity associated with concrete and abstract nouns." Transactions of the Association for Computational Linguistics 5 (2017): 17-30.
  • Anderson, Andrew James, Jeffrey R. Binder, Leonardo Fernandino, Colin J. Humphries, Lisa L. Conant, Rajeev DS Raizada, Feng Lin, and Edmund C. Lalor. "An integrated neural decoder of linguistic and experiential meaning." Journal of Neuroscience 39, no. 45 (2019): 8969-8987.
  • Oseki, Yohei, and Masayuki Asahara. "Design of BCCWJ-EEG: Balanced corpus with human electroencephalography." In Proceedings of the 12th Language Resources and Evaluation Conference, pp. 189-194. 2020.

IJCNN 2023: DL for Brain Encoding and Decoding

19

20 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

20

21 of 205

Text Stimulus Representations

  • Basic NLP Representations
    • Corpus co-occurrence counts
    • Topic models
    • Linguistic: POS, dependencies, roles.
  • Discourse
    • Characters, motion, speech, emotions, non-motion verbs
  • Deep Learning based Representations
    • Embeddings
    • Longer context using LSTMs
    • Transformers
  • Experiential attributes
    • Rated on 0-6 scale
    • Binary

IJCNN 2023: DL for Brain Encoding and Decoding

21

22 of 205

Basic NLP Representations for Word Stimuli

  • Corpus co-occurrence counts
    • 25 verbs (Mitchell et al., 2008; Pereira et al., 2013)
      • Verbs: see, hear, listen, taste, smell, eat, touch, nib, lift, manipulate, run, push, fill, move, ride, say, fear, open, approach, near, enter, drive, wear, break, and clean.
      • These verbs generally correspond to basic sensory and motor activities, actions per formed on objects, and actions involving changes to spatial relationships.
      • For each (verb, stimulus word w), feature value = normalized co-occurrence count of w with any of three forms of the verb (e.g., taste, tastes, or tasted) over the text corpus.
    • 985 common English words (such as above, worry, and mother) in (Huth et al., 2016).
  • Topic models (Pereira et al., 2013)
    • Get relevant Wiki pages (e.g., “airplane” is “Fixed-Wing Aircraft”) and other linked pages (e.g. “Aircraft cabin”)
    • LDA topic modelling on 3500 pages with #topics from 10 to 100, in increments of 5, setting the α parameter to 25/#topics.
    • LSA topic modelling (Wang et al., 2017)

IJCNN 2023: DL for Brain Encoding and Decoding

22

23 of 205

Basic NLP Representations for Word Stimuli

  • Word length
  • Is the word related to one of the 28 unique parts of speech and 17 unique dependency relationships?
  • Position of word in the sentence
  • Roles
    • Main verb
    • Agent or experiencer
    • Patient or recipient
    • Predicate of a sentence (The window was dusty)
    • Modifier (The angry activist broke the chair)
    • Complement in adjunct and propositional phrase, including direction, location, and time (The restaurant was loud at night).

IJCNN 2023: DL for Brain Encoding and Decoding

23

24 of 205

Discourse features (for Harry Potter dataset)

  • Characters: Resolve all pronouns to the character to whom they refer, and make binary features to signal which of the 10 characters are mentioned.
  • Motions: Identify a set of motions that occurred frequently in the chapter (e.g. fly, manipulate, collide physically, etc.).
  • Speech: Indicate the parts of the story that correspond to direct speech between the characters. Used the presence of dialog as a feature. 
  • Emotions: Identified a set of emotions that were felt by the characters in the chapter (e.g. annoyance, nervousness, pride, etc.).
  • Verbs: Identified a set of actions that occurred frequently in the chapter that were distinct from motion (e.g. hear, know, see, etc.).

IJCNN 2023: DL for Brain Encoding and Decoding

24

25 of 205

DL Representations: Using embeddings for word stimuli

  • GloVe 300D vectors (Pereira et al., 2016; Wang et al., 2017; Pereira et al., 2018; Anderson et al., 2019)
  • 1000D Non-negative sparse embeddings (Wehbe et al., 2014).
  • 300D embeddings by training a skip-gram model using negative sampling (SGNS) on Italian and English Wikipedia dumps using Gensim. (Anderson et al., 2017a)
  • FastText (Berezutskaya et al., 2020)
  • Comparison across multiple embedding methods
    • GloVe, word2vec, WordNet2Vec, FastText, ELMo (Hollenstein et al., 2019)
    • word2Vec, fastText, GloVe, Dependency-based word2vec, RWSGwn, ConceptNet, ELMo, averaged and concatenated combinations (Wang et al., 2020)

IJCNN 2023: DL for Brain Encoding and Decoding

25

26 of 205

DL Representations: Using longer context for word stimuli

  • Multi-task LSTMs
    • Predict next word and POS of next word.
  • ELMo embeddings: LSTM based pretrained language model

IJCNN 2023: DL for Brain Encoding and Decoding

26

27 of 205

DL Representations: Using sentence embeddings

  • Unstructured Models: Ignore sentence structure
    • Simple Pooling Methods
      • Average/max/concat(max, avg) pooling over word embeddings.
    • Advanced Pooling Methods
      • FastSent (Hill, Cho, and Korhonen 2016) sums word embeddings in a sentence as its representation to predict the surrounding sentences.
      • SIF (Arora, Liang, and Ma 2016) adapts the naïve averaging of word embeddings to weighted averaging.
  • Structured Models
    • Unsupervised Methods: Skip-thought, QuickThought.
    • Supervised Methods: InferSent, GenSen (Subramanian et al. 2018), Universal Sentence Encoder

IJCNN 2023: DL for Brain Encoding and Decoding

27

28 of 205

DL Representations: Transformer-based methods for text stimuli (Layer #, context length, architecture)

IJCNN 2023: DL for Brain Encoding and Decoding

28

Transformer-XL is the only model that continues to increase performance as the context length is increased. In all networks, the middle layers perform the best for contexts longer than 15 words. The deepest layers across all networks show a sharp increase in performance at short-range context (fewer than 10 words), followed by a decrease in performance. [Toneva and Wehbe, 2019]

29 of 205

DL Representations: Transformer-based methods for text stimuli (NLP task finetuning and scrambled LM)

  • Scrambled LM
    • Randomly shuffle words from the corpus samples, to remove all first order cues to syntactic structure.
    • LM-scrambled: words are shuffled within sentences
    • LM-scrambled-para: words are shuffled within their containing paragraphs in the corpus.
  • LM_pos: predict only the part of speech of a masked word, rather than the word itself.
  • Scrambled LMs work best!

IJCNN 2023: DL for Brain Encoding and Decoding

29

30 of 205

DL Representations: Transformer-based methods for text stimuli (NLP task finetuning)

IJCNN 2023: DL for Brain Encoding and Decoding

30

Tasks

Paraphrase, Summarization, Question Answering, Sentiment Analysis, NER, Word Sense Disambiguation, Natural Language Inference, Semantic Role Labeling, Coreference Resolution, Shallow Syntax Parsing

Pereira dataset: CR, NER, and SS perform the best.

Dendrogram constructed using similarity on representations from task-specific Transformer encoder models with stimuli from the dataset passed as input.

31 of 205

DL Representations: Transformer-based methods for text stimuli (Multi-task setup)

  • Settings
    • Finetune BERT vs not
    • Finetune BERT using one representative subject and train dense layer for each subject, vs finetune BERT for each subject.
    • Finetune BERT on MEG for all subjects, then finetune BERT on fMRI.
    • Multi-task finetune BERT for fMRI+MEG prediction task
  • Results
    • Fine-tuned models predict fMRI data better than vanilla BERT
    • Relationships between text and brain activity generalize across experiment participants.
    • Using MEG data can improve fMRI predictions.
    • A single model can be used to predict fMRI activity across multiple experiment participants.

IJCNN 2023: DL for Brain Encoding and Decoding

31

32 of 205

DL Representations: Comparing Transformers and extracting syntax vs semantics

  • Representations:
    • Lexical: representation that is context-invariant. E.g., word embeddings.
    • Compositional: “contextualized” representation generated by a system combining multiple words. E.g., parse trees
    • Syntax: representation associated with the structure of sentences independently of their meaning
    • Semantics: representation of a language system that are not syntactic.

  •  

IJCNN 2023: DL for Brain Encoding and Decoding

32

33 of 205

Experiential attributes model for text stimuli

  • Represents words in terms of human (Amazon Mechanical Turk) ratings of their degree of association with different attributes of experience
    • “On a scale of 0 to 6, to what degree do you think of a banana as having a characteristic or defining color?”
    • Anderson et al., 2019: 65 attributes spanning sensory, motor, affective, spatial, temporal, causal, social, and abstract cognitive experiences.
  • Value-add on top of text models: a lot of experiential information goes unstated in natural verbal communication.
    • E.g., it is rarely useful to communicate the color of bananas because it is obvious to all those with experience of bananas.
    • E.g., it would be unusual to specify that dropping things involves movement.
  • Nishida et al., 2020 use a subset of 20 attributes.

IJCNN 2023: DL for Brain Encoding and Decoding

33

34 of 205

Binary attribute representations

  • Each stimulus is represented using a binary vector capturing membership to one of the eight semantic categories.

  • 42 neurally plausible semantic features (NPSFs)
    • Perceptual and affective characteristics of an entity (10 NPSFs coded such features, such as man-made, size, color, temperature, positive affective valence, high affective arousal), animate beings (person, human-group, animal), and time and space properties (e.g. unenclosed setting, change of location)

IJCNN 2023: DL for Brain Encoding and Decoding

34

35 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

35

36 of 205

Thanks!

IJCNN 2023: DL for Brain Encoding and Decoding

36

37 of 205

Language and the Brain: Deep Learning for Brain Encoding and Decoding

Subba Reddy Oota1, Manish Gupta2,3, Raju S. Bapi2

1Inria Bordeaux, France; 2IIIT Hyderabad, India; 3Microsoft, India

subba-reddy.oota@inria.fr, gmanish@microsoft.com, raju.bapi@iiit.ac.in

38 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

38

39 of 205

39

LMs are trained to predict missing words

Language model

The

quick

brown

fox

[MASK]

jumps

40 of 205

Background

  • Language Model?
    • Task of predicting what word comes next

41 of 205

You use Language Models every day!

42 of 205

42

[Trichelair et al.]

I put the heavy table on the book and it broke.

But… LMs may be learning shallow heuristics to solve tasks

What broke?

“the book”

“the table”

43 of 205

43

Lots of progress made, but still a long way to go

How do we build systems with a deeper understanding of language?

Still a long way to go

Lots of progress

44 of 205

Transformer

Harry never thought he would

Harry never thought he ???

45 of 205

LLMs: Pretraining for three types of architectures

46 of 205

Task-specific language models

47 of 205

Language models are everywhere

Sentiment

Question Answering

Summarization

Coreference Resolution

48 of 205

Instruction Models:

  • Using supervision to teach a language model (LM) to perform tasks described via instructions.
  • The LM will learn to follow instructions and do so even for unseen tasks.
  • Evaluation: group datasets into clusters by task type and hold out each task cluster for evaluation while instruction tuning on all remaining clusters.

NLU tasks in blue; NLG tasks in teal

49 of 205

Multiple Instruction Templates for Each NLP Task

  • Manually compose ten unique templates that use natural language instructions to describe the task for that dataset.
    • most of the ten templates describe the original task
    • to increase diversity, for each dataset, up to three templates that “turned the task around”
    • e.g., for sentiment classification, summarization task related template by asking to generate a movie review

50 of 205

In-Context Learning and Chain-of-Thought

51 of 205

Data-driven encoding models evaluate the relationships between brains and deep learning models

fMRI

A priori locations in DL system and brain

Deep learning system

how are they related?

Multimodal naturalistic stimulus

Data-driven encoding model

52 of 205

Deep learning models enable data-driven encoding models for naturalistic stimuli

more stimulus properties that affect brain activity

more naturalistic stimuli

<0,1,...0>

simple stim. representations explain less variance in brain activity

53 of 205

Deep learning models enable data-driven encoding models for naturalistic stimuli

54 of 205

Encoding (Well-posed) vs Decoding (Ill-posed) in Neuroscience

  • Encoding: How is the stimulus represented in the brain?
  • Decoding: Can we reconstruct the stimulus, given the brain response?
  • What information is to be decoded from the brain?
  • Can we read what the subject is thinking while watching the stimulus?

Decoding is ill-posed?

55 of 205

Brain Encoding?

Present

Stimulus

Stimulus

Ridge Regression

Input

Input

Output

X

Y

W

Pearson Correlation (R) = Corr(Y, W(X))

56 of 205

Encoding: training independent models

  • Independent model per participant
  • Independent model per voxel / sensor-timepoint

P1

P2

PN

P1, v1

P1, v2

P1, vm

57 of 205

Mechanistic understanding of information processing in the brain: 4 big questions

57

How

Where

When

What

58 of 205

With MEG we can analyze sub-word time course

  • MEG recording data at very fast temporal resolution
  • So, we can look at sub-word process
  • fMRI recording data at very high-spatial resolution

Where

When

59 of 205

How does brain represents complex meaning? (Where, When and What)

60 of 205

Word Context

1-word context

2-word context

3-word context

4-word context

5-word context

Past context

Future context

61 of 205

Normalized Predictivity

62 of 205

Recent work utilizing progress in LLMs for encoding

  • Using representations of stimuli from deep learning systems
  • Language:
    • Wehbe et al. 2014;
    • Jain and Huth, 2018;
    • Toneva and Wehbe, 2019;
    • Caucheteux and King, 2020/2022;
    • Schrimpf et al. 2020/2021;
    • Goldstein et al. 2021/2022;
    • Oota et al. 2022/2023;

63 of 205

Language: work utilizing DL progress

63

  • Stimuli: one chapter of Harry Potter
  • Stimulus representation: derived from pretrained NLP systems
  • Brain recording & modality: fMRI, reading

across several types of large NLP systems, best alignment with fMRI in middle layers

64 of 205

Language: work utilizing DL progress

  • Stimuli: sentences, passages, short story
  • Stimulus representation: derived from pretrained NLP systems (BERT, GPT-2, T5 , and XLM)
  • Brain recording & modality: fMRI & ECoG, reading & listening

some NLP systems can predict fMRI and ECoG up to 100% of estimated noise ceiling

65 of 205

Language: work utilizing DL progress

65

  • Stimuli: sentences
  • Stimulus representation: derived from pretrained NLP systems (BERT and GPT-2)
  • Brain recording & modality: MEG & fMRI, reading

best alignment with fMRI & MEG in middle layers

better performance at predicting next word -> better prediction of fMRI & MEG

66 of 205

Language: work utilizing DL progress

66

  • Stimuli: sentences
  • Stimulus representation: derived from pretrained NLP systems (BERT and GPT-2)
  • Brain recording & modality: MEG & fMRI, reading

best alignment with fMRI & MEG in middle layers

better performance at predicting next word -> better prediction of fMRI & MEG

67 of 205

Language: work utilizing DL progress

67

  • Stimuli: sentences
  • Stimulus representation: derived from pretrained NLP systems (GPT-2 XL)
  • Brain recording & modality: fMRI, reading

model-selected ‘out-of- distribution’ sentences indeed drive and suppress activity of human language areas in new individuals

68 of 205

Challenges in using DL for cognitive science

  • Not designed to specifically model brain processing

68

NLP systems: Designed to predict upcoming words

Harry never thought ???

Harry never thought he ???

Harry never thought he would ???

...

69 of 205

Challenges in using DL for cognitive science

  • Not designed to specifically model brain processing
    • Training DL models using brain recordings
    • Task-based modeling

69

70 of 205

Challenges in using DL for cognitive science

  • Not designed to specifically model brain processing
    • Training DL models using brain recordings
    • Task-based modeling
  • Can be difficult to interpret due to multiple sources of information

70

part-of-speech

semantic role

dependence on other words

...

+

+

+

?

71 of 205

Challenges in using DL for cognitive science

  • Not designed to specifically model brain processing
    • Training DL models using brain recordings
    • Task-based modeling
  • Can be difficult to interpret due to multiple sources of information
    • Disentangling contributions of different info sources to brain predictions

71

72 of 205

Challenges in using DL for cognitive science

  • Not designed to specifically model brain processing
    • Training DL models using brain recordings
    • Task-based modeling
  • Can be difficult to interpret due to multiple sources of information
    • Disentangling contributions of different info sources to brain predictions

72

73 of 205

Training DL models using brain recordings

73

Brain-optimized NLP model predicts unseen fMRI recordings better, especially in canonical language regions

A priori locations in NLP system and brain

NLP system

Chapter of a book

𝑥 alignment

error propagation

fMRI

  • Stimuli: one chapter of Harry Potter
  • Stimulus representation: brain-optimized NLP model
  • Brain recording & modality: fMRI & MEG, reading

74 of 205

Inducing Brain Relevant Bias

75 of 205

Challenges in using DL for cognitive science

  • Not designed to specifically model brain processing
    • Training DL models using brain recordings
    • Task-based modeling
  • Can be difficult to interpret due to multiple sources of information
    • Disentangling contributions of different info sources to brain predictions

75

76 of 205

Tasks affect processing

76

bear

X

veg?

bear

X

tool?

800ms

306 sensors

800ms

306 sensors

Systematic difference due to different question tasks

Attention emphasizes task-relevant information

Mechanism?

Can we model as a function of the task AND stimulus?

77 of 205

Tasks affect processing

77

question task effect word effect

significant prediction performance

The end of semantic processing of a word is task-dependent

  • Stimuli: concrete nouns + line drawings
  • Task: answer Yes/No questions about noun
  • Stimulus representation: human judgments
  • Brain recording & modality: MEG, reading

78 of 205

Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity

Subba Reddy Oota1,2, Jashn Arora2, Veeral Agarwal2, Mounika Marreddy2, Manish Gupta2,3, Bapi Raju Surampudi2

July 13, 2022

NAACL-HLT 2022

1Inria Bordeaux France, 2IIIT-Hyderabad, 3Microsoft India

79 of 205

Can task-specific language models better predict fMRI brain activity?

Tasks

  • Paraphrase
  • Summrisation
  • Question Answering
  • Sentiment Analysis
  • NER
  • Word Sense Disambiguation
  • Natural Language Inference
  • Semantic Role Labeling
  • Coreference Resolution
  • Shallow Syntax
  • Pretrained BERT

Devlin et al. 2019, Bowon et al. 2020

Syntactic

80 of 205

Can task-specific language models have similar predictive performance in reading and listening?

81 of 205

Tasks affect processing

81

  • Stimuli: passages and narratives
  • Stimulus representation: task-optimized NLP models for a range of tasks
  • Brain recording & modality: fMRI, reading & listening of different stimuli

Reading fMRI best explained by coref. resolution, NER, shallow syntax parsing

Listening fMRI best explained by paraphrasing, summarization, NLI

82 of 205

Challenges in using DL for cognitive science

  • Not designed to specifically model brain processing
    • Training DL models using brain recordings
    • Task-based modeling
  • Can be difficult to interpret due to multiple sources of information
    • Disentangling contributions of different info sources to brain predictions

82

83 of 205

Disentangling contributions of different info sources to brain predictions

83

“Mary finished the apple”

supra-word meaning may contain concept of:

  • eating
  • apple core

supra-word

meaning

Isolating supra-word meaning is a type of intervention

84 of 205

Disentangling contributions of different info sources to brain predictions

84

full context

supra-word

Bilateral PTL and ATL process supra-word meaning

Word-level information important for prediction of most language regions

  • Stimuli: one chapter of Harry Potter
  • Stimulus representation: disentangled embeddings from pretrained NLP models
  • Brain recording & modality: fMRI & MEG, reading

85 of 205

Disentangling contributions of different info sources to brain predictions

85

Syntactic structure-based features explain additional variance in language regions over complexity metrics

Regions predicted by syntactic and semantic are difficult to distinguish

  • Stimuli: one chapter of Harry Potter
  • Stimulus representation: syntactic tree representations & pretrained NLP model
  • Brain recording & modality: fMRI, reading

86 of 205

Disentangling contributions of different info sources to brain predictions

86

Constituency tree structure is better in temporal cortex and MFG, while Dependency structure is better in AG and PCC,

Regions predicted by syntactic and semantic are difficult to distinguish

  • Stimuli: Narratives
  • Stimulus representation: syntactic tree representations & pretrained NLP model
  • Brain recording & modality: fMRI, listening

87 of 205

87

Subba reddy Oota, Manish Gupta, Mariya Toneva

Joint processing of linguistic properties in brains and language models

88 of 205

Hierarchy of Linguistic Info - Setting

  • Conneau et al., ACL’18 - Build diagnostic classifier to predict if a linguistic property is encoded in the given sentence representation.
  • Features:
    • Surface – Sentence Length, Word Content
    • Syntactic – Bigram shift, Tree depth, Top constituent
    • Semantic – Tense, Subject Number, Object Number, Coordination Inversion and Semantic Odd Man Out.

88

BERT layer

Simple classifier

predict sentence length

If the prediction accuracy is good, then the model might be capturing the sentence length feature

89 of 205

Hierarchy of Linguistic Info - Result

89

90 of 205

Takeaway

  • BERT composes a hierarchy of linguistic signals ranging from surface to semantic features.

90

BERT does capture many

structural properties of the English language.

91 of 205

Disentangling contributions of different info sources to brain predictions

91

Top constituents and Tree Depth contribute the most to the alignment trend across layers

  • Stimuli: Narrative Stories
  • Stimulus representation: pretrained NLP model and removal of linguistic properties
  • Brain recording & modality: fMRI, Listening
  • Questions: What linguistic properties underlie brain alignment, across all layers but also specifically in middle layers?

92 of 205

Disentangling contributions of different info sources to brain predictions

92

Past word context iscrucial in obtaining significant results.

  • Stimuli: four naturalistic stories
  • Stimulus representation: basic syntactic tree representations & pretrained NLP model
  • Brain recording & modality: MEG, Listening

93 of 205

Khai Loong Aw Mariya Toneva

Max Planck Institute for Software Systems (MPI-SWS)

Training language models to summarize narratives�improves brain alignment

How to build better Language models?

Biological

Artificial

94 of 205

94

Compare against actual brain recordings

(brain alignment)

Use model’s internal layer�activations to predict brain activity on held-out data

Model trained with�language modeling

Model trained to�summarize narratives

input

input

activations

activations

book�chapter

95 of 205

95

Model trained with�language modeling

Model trained to�summarize narratives

input

input

activations

activations

Compare against actual brain recordings

(brain alignment)

Use model’s internal layer�activations to predict brain activity on held-out data

96 of 205

96

Result: Summarize narratives → Greater brain alignment 🧠

Training language models to summarize narratives improves brain alignment

this is the title of our paper!

brain alignment (Pearson correlation)

97 of 205

97

Result: Brain alignment improves for all discourse features

Booksum models’ representations of Characters, Emotions and Motions are more aligned to the brain than the base models’ representations.

brain alignment (Pearson correlation)

98 of 205

98

Subba reddy Oota, Fatma Deniz, Mariya Toneva

What aspects of NLP models and brain datasets affect brain-NLP alignment?

99 of 205

99

Text models predict fMRI recordings significantly better than speech models

  • Stimuli: Narrative Stories
  • Stimulus representation: pretrained NLP model and speech models
  • Brain recording & modality: fMRI, Reading, Listening
  • Questions: Is the choice of stimulus modality (reading vs. listening) important for the study of brain alignment?
  • Are all naturalistic fMRI datasets equally good for brain encoding?
  • How does the type of model (text vs. speech and encoder vs. decoder) affect the resulting alignment?

100 of 205

Recent work utilizing progress in LLMs for encoding

  • Using representations of stimuli from deep learning systems
  • Language:
    • Wehbe et al. 2014; Jain and Huth, 2018; Toneva and Wehbe, 2019; Caucheteux and King, 2020/2022; Schrimpf et al. 2020/2021; Goldstein et al. 2021/2022; Oota et al. 2022;
  • Vision:
    • Yamins et al. 2014; Cichy et al. 2016; Konkle and Alvarez, 2020/2022; Zhuang et al. 2022
  • Audio:
    • Kell et al. 2018; Vaidya, Jain, and Huth 2022; Millet et al. 2022, Tuckte et al. 2022, Oota et al. 2023
  • Multi-Modal:
    • Oota et al. 2022;

101 of 205

Audio: work utilizing DL progress

  • Stimuli: Moth Radio Hour
  • Stimulus representation: derived from pretrained self-supervised speech models (HuBERT, Wav2Vec2.0, APC)
  • Brain recording & modality: fMRI, listening

Middle layers of self-supervised speech models predict auditory cortex the best

102 of 205

Audio: work utilizing DL progress

  • Stimuli: audio books
  • Stimulus representation: derived from pretrained self-supervised speech model (Wav2Vec2.0)
  • Brain recording & modality: fMRI, listening in 3 languages (Eng, Fr, Mandarin)

Self-supervised speech models exhibit specialization for native sounds in the STS and MTG;

IFG and AG show more general specialization for speech rather than native-language

103 of 205

Neural Architecture of Speech

Subba Reddy Oota1, Khushbu Pahwa2, Mounika Marreddy3, Manish Gupta3,4, Bapi Raju Surampudi3

June 6, 2023

ICASSP 2023

1Inria Bordeaux France, 2University of California LA, 3IIIT-Hyderabad, 4Microsoft India

104 of 205

Speech representation learning methods

105 of 205

Speech Models

Generative approaches

Predictive approaches

Contrastive approaches

Traditional approaches

106 of 205

Encoding Performance of Speech Models

107 of 205

Model Encoding Performance (Data2Vec)

108 of 205

Layer Selectivity

109 of 205

How do we assess models’ performance?

110 of 205

Neural Architecture of Speech: MEG Encoding

110

Previous studies using statistical correlation have observed using controlled settings (like piano tones) that the response to auditory stimulus peaks at around 200ms

  • Stimuli: Narrative Stories
  • Stimulus representation: SSL speech models and speech descriptors
  • Brain recording & modality: MEG, Listening

111 of 205

Neural Architecture of Speech: MEG Encoding

111

Data2Vec: Early layers contribute the most to auditory response while later layers contribute both auditory and language information

112 of 205

Recent work utilizing progress in LLMs for encoding

  • Using representations of stimuli from deep learning systems
  • Language:
    • Wehbe et al. 2014; Jain and Huth, 2018; Toneva and Wehbe, 2019; Caucheteux and King, 2020/2022; Schrimpf et al. 2020/2021; Goldstein et al. 2021/2022; Oota et al. 2022;
  • Vision:
    • Yamins et al. 2014; Cichy et al. 2016; Konkle and Alvarez, 2020/2022; Zhuang et al. 2022
  • Audio:
    • Kell et al. 2018; Vaidya, Jain, and Huth 2022; Millet et al. 2022, Tuckte et al. 2022, Oota et al. 2023
  • Multi-Modal:
    • Oota et al. 2022;

113 of 205

Image reconstruction with latent diffusion models from human brain activity

  • Stimuli: natural images
  • Stimulus representation: derived from CLIP and AlexNet models
  • Brain recording & modality: fMRI, Viewing images, NSD dataset

114 of 205

Image reconstruction with latent diffusion models from human brain activity

115 of 205

Seeing Beyond the Brain: Conditional Diffusion Model with Sparse Masked Modeling for Vision Decoding

  • Stimuli: natural images
  • Stimulus representation: derived from ViT model
  • Brain recording & modality: fMRI, Viewing images, NSD dataset

116 of 205

Recent work utilizing progress in LLMs for encoding

  • Using representations of stimuli from deep learning systems
  • Language:
    • Wehbe et al. 2014; Jain and Huth, 2018; Toneva and Wehbe, 2019; Caucheteux and King, 2020/2022; Schrimpf et al. 2020/2021; Goldstein et al. 2021/2022; Oota et al. 2022;
  • Vision:
    • Yamins et al. 2014; Cichy et al. 2016; Konkle and Alvarez, 2020/2022; Zhuang et al. 2022
  • Audio:
    • Kell et al. 2018; Vaidya, Jain, and Huth 2022; Millet et al. 2022, Tuckte et al. 2022, Oota et al. 2023
  • Multi-Modal:
    • Oota et al. 2022;

117 of 205

Visio-Linguistic Brain Encoding

Subba Reddy Oota1,2, Jashn Arora2, Vijay Rowtula2, Manish Gupta2,3, Bapi Raju Surampudi2

August 13, 2022

COLING 2022

1Inria Bordeaux France, 2IIIT-Hyderabad, 3Microsoft India

118 of 205

Can image-based and multi-model Transformers accurately perform fMRI encoding?

Dosovitskiy et al. 2021, Tan et al. 2019, Harold Li et al. 2019

119 of 205

Models used: Multi-Modal Transformers

CLIP

LXMERT

VisualBERT

Radford et al. 2021, Tan et al. 2019, Harold Li et al. 2019

120 of 205

Dataset Details

120

Periera

Periera et al. 2018, Nadine et al. 2019

BOLD5000

Concept+Picture (Bird)

121 of 205

Encoding performance (BOLD5000)

122 of 205

Incorporating natural language into vision models

123 of 205

DNNs & The Brain: Multi-modal, Multi-task

  • Brain response to a stimulus is multi-modal, multi-task related
    • Cross-view and multi-view decoding (Oota et al 2022)
    • Visio-linguistic encoding (fusion of vision and language information) (Oota et al 2022)
    • Multimodal foundation model (Fei et al 2022)

Fei, Lu, Gao et al (2022). Towards artificial general intelligence via a multimodal foundation model. Nature Communications 13:3094

doi.org/10.1038/s41467-022-30761-2

124 of 205

DNNs & Brain: Multi-modal, Multi-task

  • Brain response to a stimulus is multi-modal, multi-task related
    • Cross-view and multi-view decoding (Oota et al 2022)
    • Visio-linguistic encoding (fusion of vision and language information) (Oota et al 2022)
    • Multimodal foundation model (Fei et al 2022)

Fei, Lu, Gao et al (2022). Towards artificial general intelligence via a multimodal foundation model. Nature Communications 13:3094 doi.org/10.1038/s41467-022-30761-2

125 of 205

DNNs & Brain Damage

  • DL models of encoding and decoding have not yet been put through the brain-damage experiments. Ex. Semantic Dementia

Snowden, Harris, Thompson, Kobylecki, Jones, Richardson, Neary (2018). Semantic dementia and the left and right temporal lobes, Cortex, 107(188-203).

https://doi.org/10.1016/j.cortex.2017.08.024.

Rt Ant Temporal Lobe Damage (Patient 8)

Animal habitat task.

The patient is asked:

Where would you find this?

Do DL Models exhibit such degradation with damage to units?

126 of 205

Future Works

127 of 205

Instruction Models and Brain Alignment

127

  • Does all instructions are useful for human brain alignment?
  • Can we automate the prompts with brain data?

FLAN: Fine-Tuned Language Models are Zero-Shot Learners

BLOOM

PaLM: Scaling Language Modeling with Pathways

FLAN-T5

InstructGPT

Gopher

LaMDA

GPT-3

GLaM

Tk-Instruct

Llama

128 of 205

Human guided instructions into vision models and Brain Alignment

128

  • Which human guided instructions yield brain language and visual hierarchy ?

InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

129 of 205

Human Alignment of Neural Network Representations

Lukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen

https://arxiv.org/abs/2211.01201

Research Question:

  • What factors affect the alignment between the representations learned by neural networks and human mental representations inferred from behavioral responses? Do models that are better at classifying images naturally learn more human-like conceptual representations?

  • Human behavior: odd-one-out judgements
  • Metrics: zero-shot odd-one-out accuracy, probing, RSA (repr. Sim. analysis)

Main Contributions/Findings:

  • Model scale and architecture have mostly no effect on the alignment with human behavioral responses
  • Training dataset and objective function both have a much larger impact
  • Some human concepts (food/animals) are well-represented whereas other objects (royal/sports-related) are not

130 of 205

Training Vision models using brain recordings

130

Does brain-optimized vision models align better with human judgements?

Inducing Brain relevant bias into image Transformers

131 of 205

Multi-modal Brain CLIP

132 of 205

A big thank you!

Tutorial, Code and Material:

Deep Learning for Brain Encoding and Decoding, Cogsci-2022

https://tinyurl.com/DL4Brain

Upcoming Tutorials:

  • Deep Neural Networks and Brain Alignment: Brain Encoding and Decoding, IJCAI-2023 (A* conference)

133 of 205

Language and the Brain: Deep Learning for Brain Encoding and Decoding

Subba Reddy Oota1, Manish Gupta2,3, Raju S. Bapi2

1Inria Bordeaux, France; 2IIIT Hyderabad, India; 3Microsoft, India

subba-reddy.oota@inria.fr, gmanish@microsoft.com, raju.bapi@iiit.ac.in

134 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

134

135 of 205

Outline

  • Introduction to Brain Decoding [5 mins]
  • Decoding models [5 mins]
    • Linear Models
    • Non-Linear Models (including DNNs)
  • Language [20 mins]
    • Periera et al. 2018, Gauthier et al. 2019, Huth et al. 2023, Oota et al. 2022

IJCNN 2023: DL for Brain Encoding and Decoding

135

136 of 205

Encoding vs. Decoding

IJCNN 2023: DL for Brain Encoding and Decoding

136

Haiguang Wen et al, 2017

Encoding

Decoding

Stimulus

Representation

Stimulus

Representation

fMRI

fMRI

137 of 205

What is Brain Decoding?

  • Can we reconstruct the stimulus, given the brain response?
  • Can you read the mind with fMRI?
  • Or at least tell what the person saw?

IJCNN 2023: DL for Brain Encoding and Decoding

137

Visual Task

Language Task

Smith et al., 2011, Wang et al. 2019

138 of 205

Linguistic Decoding

IJCNN 2023: DL for Brain Encoding and Decoding

138

input

output

Zou et al., 2022

139 of 205

Encoder-Decoder Models in AI

IJCNN 2023: DL for Brain Encoding and Decoding

139

Encoder

Decoder

Encoder

Decoder

Youssef et al. 2018

140 of 205

Outline

  • Introduction to Brain Decoding [5 mins]
  • Decoding models [5 mins]
    • Linear Models
    • Non-Linear Models (including DNNs)
  • Language [20 mins]
    • Periera et al. 2018, Gauthier et al. 2019, Huth et al. 2023, Oota et al. 2022

IJCNN 2023: DL for Brain Encoding and Decoding

140

141 of 205

Linear Decoder Models

IJCNN 2023: DL for Brain Encoding and Decoding

141

Ridge / Logistic Regression

Stimulus Representation

Stimulus Classification

Horikawa et al. 2018

142 of 205

Non-Linear Decoder

IJCNN 2023: DL for Brain Encoding and Decoding

142

Vu et al. 2018

Deep CNNs

143 of 205

Word-Level Brain Decoding

  • Build decoders to associate brain activities with word stimulus via distributed representations (GloVe).

Toward a universal decoder of linguistic meaning from brain activation https://www.nature.com/articles/s41467-018-03068-4

IJCNN 2023: DL for Brain Encoding and Decoding

144 of 205

Evaluating Decoding Models: Rank Accuracy

IJCNN 2023: DL for Brain Encoding and Decoding

144

Y1

 

Y2

Yn

Periera et al. 2018

ith Concept Word

Correaltion

 

rank = rsort(corr_scores).index(correlation)

All the correlation scores in descending order

145 of 205

Outline

  • Introduction to Brain Decoding [5 mins]
  • Decoding models [5 mins]
    • Linear Models
    • Non-Linear Models (including DNNs)
  • Language [20 mins]
    • Periera et al. 2018, Gauthier et al. 2019, Huth et al. 2023, Oota et al. 2022

IJCNN 2023: DL for Brain Encoding and Decoding

145

146 of 205

Linguistic Brain Decoding

  • Toward Word-level Universal Brain Decoder
  • Does injecting linguistic structure into language models lead to better alignment with brain recordings?
  • Multi-view and Cross-view Decoding

IJCNN 2023: DL for Brain Encoding and Decoding

146

Periera et al. 2018, Gauthier et al. 2019, Huth et al. 2023, Oota et al. 2022

147 of 205

Classical Decoders

  • Classical decoding solutions extracting linguistic meaning from imaging data have been largely limited to
    • concrete nouns,
    • using similar stimuli for training and testing,
    • small number of semantic categories.

IJCNN 2023: DL for Brain Encoding and Decoding

147

Mitchell et al. 2008

148 of 205

Toward a universal decoder

  • Presented a new approach for building a brain decoding system:
    • words and sentences are represented as vectors in a semantic space constructed from massive text corpora.
    • wide variety of both concrete and abstract topics from two separate datasets.
    • subject reads naturalistic linguistic stimuli on potentially any topic, including abstract ideas (ex., pleasure, justice, love, etc).

IJCNN 2023: DL for Brain Encoding and Decoding

148

Pereira et al. 2018

GloVE

149 of 205

Dataset Details (Experiment-1)

IJCNN 2023: DL for Brain Encoding and Decoding

149

Concept + Sentence View

Concept Word

Concept + Picture View

Concept + Wordcloud View

Periera et al. 2018

150 of 205

Dataset Details (Experiment-1)

  • 180 Concepts
    • 128 nouns
    • 22 verbs
    • 29 adjectives
    • 1 function word
  • 16 subjects
  • AAL atlas (180 regions)
  • Gordon atlas (333 regions)

IJCNN 2023: DL for Brain Encoding and Decoding

150

Periera et al. 2018

151 of 205

Dataset Details (Experiments 2 and 3)

IJCNN 2023: DL for Brain Encoding and Decoding

151

Topic

Concept

Topic

Periera et al. 2018

152 of 205

Informative Voxel Selection

Cogsci-2022: DL for Brain Encoding and Decoding

152

Voxel + 26 neighbors in 3D

Input

Ridge Regression

Output

Stimulus:

Apartment

Present

GloVE

Present

Stimulus:

Apartment

Pearson Correlation (R) = Corr(Y, W(X))

Correlation across feature dimensions

V1 – R1

V2 – R2

….

Vn – R3

Select 5000 voxels based on top-5000 correlation scores

3D Image

X

Y

W

153 of 205

Experimental Setup

  • 18-Fold cross-validation
  • Informative voxel selection in each fold for three views
  • Ridge regression

IJCNN 2023: DL for Brain Encoding and Decoding

153

154 of 205

Brain Decoder Schematic? (concept+picture)

IJCNN 2023: DL for Brain Encoding and Decoding

154

Present

Stimulus:

Apartment

Stimulus:

Apartment

Present

Ridge Regression

GloVE

Periera et al. 2018, Pennington et al. 2014

155 of 205

Brain Decoder Schematic? (concept+picture)

IJCNN 2023: DL for Brain Encoding and Decoding

155

Wang et al. 2019

156 of 205

Brain Decoder Schematic? (concept+sentence)

IJCNN 2023: DL for Brain Encoding and Decoding

156

Present

Stimulus:

Apartment

Stimulus:

Apartment

Present

Ridge Regression

GloVE

Periera et al. 2018, Pennington et al. 2014

157 of 205

Brain Decoder (Exp 2& 3: Different Topics)

IJCNN 2023: DL for Brain Encoding and Decoding

157

Present

Stimulus

Present

Ridge Regression

GloVE

Periera et al. 2018, Pennington et al. 2014

Stimulus

Testing

  • Different topics (e.g., a sentence about a piano in training vs. a butterfly in testing)

A butterfly is a flying insect with four large wings.

The piano is a popular musical instrument …

The piano is a popular musical instrument …

158 of 205

Brain Decoder ( Different passages from same topic)

IJCNN 2023: DL for Brain Encoding and Decoding

158

Present

Stimulus

Present

Ridge Regression

GloVE

Periera et al. 2018, Pennington et al. 2014

Stimulus

Testing

  • Different passages from the same topic (e.g., a sentence about a dragonfly in training vs. a butterfly in testing)

A butterfly is a flying insect …

Mosquitos are thin, small flying …

Mosquitos are thin, small flying …

Insect

159 of 205

Brain Decoder (Different sentences within the same passage)

IJCNN 2023: DL for Brain Encoding and Decoding

159

Present

Stimulus

Present

Ridge Regression

GloVE

Periera et al. 2018, Pennington et al. 2014

Stimulus

Testing

  • Different sentences within the same passage (e.g., two sentences about a piano: one sentence in training and second one in testing)

The piano is a popular musical instrument …

The piano is a popular musical instrument …

The piano has an enormous …

160 of 205

Pairwise and Rankwise Results

IJCNN 2023: DL for Brain Encoding and Decoding

160

Periera et al. 2018

161 of 205

Distribution of Informative Voxels

IJCNN 2023: DL for Brain Encoding and Decoding

161

Periera et al. 2018

162 of 205

Average decoding performance

IJCNN 2023: DL for Brain Encoding and Decoding

162

50K informativevoxels

Periera et al. 2018

Voxels contributing to decoding are widely distributed

163 of 205

Insights

  • Presented a viable approach for building a universal decoder, capable of extracting a representation of mental content from linguistic materials.
  • The semantic resolution of brain-based decoding of mental content will continue to improve rapidly,
    • given the progress in the development of distributed semantic representations

IJCNN 2023: DL for Brain Encoding and Decoding

163

Periera et al. 2018

164 of 205

Linguistic Brain Decoding

  • Toward Word-level Universal Brain Decoder
  • Linking artificial and human neural representations of language
  • Multi-view and Cross-view Decoding

IJCNN 2023: DL for Brain Encoding and Decoding

164

Periera et al. 2018, Gauthier et al. 2019, Huth et al. 2023, Oota et al. 2022

165 of 205

Linking artificial and human neural representations of language

IJCNN 2023: DL for Brain Encoding and Decoding

165

Ridge Regression

Gauthier et al. 2019

  • Evaluate the link between human brain activity and neural network models as the models are optimized for different tasks.
  • To investigate why these mappings are successful?
  • Uncovering the parallel representational contents shared between human brains and neural networks

166 of 205

Cogsci-2022: DL for Brain Encoding and Decoding

166

Devlin et al. 2019

Pretrained vs. Task-specific language models

167 of 205

IJCNN 2023: DL for Brain Encoding and Decoding

167

Natural Language Understaning Tasks

  • Paraphrase
  • Question Answering
  • Sentiment Analysis
  • Natural Language Inference

Devlin et al. 2019, Bowon et al. 2020

Pretrained vs. Task-specific language models

Squad-2.0: Question Answering

168 of 205

Custom Tasks

  • Scrambled language modeling:
    • LM-scrambled: deals with sentence inputs where words are shuffled within sentences
    • LM-scrambled-para, uses inputs where words are shuffled within their containing paragraphs in the corpus.

IJCNN 2023: DL for Brain Encoding and Decoding

168

Fingers are used for grasping, writing, grooming and other activities.

grasping are used for Fingers, grooming, writing and other activities.

This is Los Angeles. And it's the height of summer. In a small bungalow off of La Cienega, Clara serves homemade chili and chips in red plastic bowls -- wine in blue plastic.

This is Los Angeles. And the height it's of summer. In a bungalow off small of La Cienega, Clara serves homemade chili and chips in red plastic bowls -- wine in blue plastic.

Gauthier et al. 2019

169 of 205

Custom Tasks

  • Part-of-speech language modeling:
    • LM-pos, to select against fine-grained semantic representation of inputs.
    • Masked word is POS tag

IJCNN 2023: DL for Brain Encoding and Decoding

169

Dotted lines: Incorrect dependecies

Solid lines: correct dependecies

Gauthier et al. 2019

170 of 205

Reading data target: human brain recordings

IJCNN 2023: DL for Brain Encoding and Decoding

170

Periera et al. 2018 fMRI

  • Periera dataset
    • reading sentences
    • 5 subjects
    • 627 sentences (experiment 2 + 3)

Example: ''A clarinet is a woodwind musical instrument''

171 of 205

Brain decoding performance

IJCNN 2023: DL for Brain Encoding and Decoding

171

Scrambled language models have shown better performance!!

Gauthier et al. 2019

172 of 205

Brain decoding performance trajectories over fine-tuning time

IJCNN 2023: DL for Brain Encoding and Decoding

172

Gauthier et al. 2019

173 of 205

Representational similarity of the sentence encodings

IJCNN 2023: DL for Brain Encoding and Decoding

173

Correlation between pair of model representaions

Gauthier et al. 2019

174 of 205

Summary

  • Set of scrambled language modeling tasks which best match the structure of brain activations among the models tested.
    • models optimized for LM- scrambled and LM-scrambled-para — the models which improve in brain decoding performance

IJCNN 2023: DL for Brain Encoding and Decoding

174

Gauthier et al. 2019

175 of 205

Linguistic Brain Decoding

  • Toward Word-level Universal Brain Decoder
  • Linking artificial and human neural representations of language (contd)
  • Multi-view and Cross-view Decoding

IJCNN 2023: DL for Brain Encoding and Decoding

175

Periera et al. 2018, Gauthier et al. 2019, Huth et al 2023, Oota et al. 2022

176 of 205

IJCNN 2023: DL for Brain Encoding and Decoding

176

Continuous Language Decoder

Tang, LaBel, Jain & Huth (2023)

177 of 205

IJCNN 2023: DL for Brain Encoding and Decoding

177

Continuous Language Decoder

Tang, LaBel, Jain & Huth (2023)

178 of 205

IJCNN 2023: DL for Brain Encoding and Decoding

178

Continuous Language Decoder

Tang, LaBel, Jain & Huth (2023)

179 of 205

Summary

  • Continuous language representations of semantic meaning can be decoded (reconstructed) from non-invasive brain recordings (fMRI),
  • Given novel brain recordings, decoder generates intelligible word sequences that recover the meaning of perceived speech, imagined speech, and even silent videos, demonstrating that a single language decoder can be applied to a range of semantic tasks.
  • Exciting possibility enabling future multipurpose brain-computer interfaces!

IJCNN 2023: DL for Brain Encoding and Decoding

179

Tang, LaBel, Jain & Huth (2023)

180 of 205

Linguistic Brain Decoding

  • Toward Word-level Universal Brain Decoder
  • Linking artificial and human neural representations of language
  • Multi-view and Cross-view Decoding

IJCNN 2023: DL for Brain Encoding and Decoding

180

Periera et al. 2018, Gauthier et al. 2019, Huth et al. 2023, Oota et al. 2022

181 of 205

Multi-view and Cross-ViewBrain Decoding

  • Human brains have the unique capability of language acquisition:
    • the process of learning the language
    • understand the meaning of concepts from multiple modalities such as images, text, speech, and videos.
  • Prior works focus on single-view brain decoding using traditional feature engineering.
  • However, how the brain captures the meaning of linguistic stimuli across multiple views is still a critical open question in neuroscience.
  • Consider three different views of the concept bird:
    • (1) sentence using the target word,
    • (2) picture presented with the target word label, and
    • (3) word cloud containing the target word along with other semantically related words.
  • Earlier works have explored which of these three different views provides richer information to understand the concept.

IJCNN 2023: DL for Brain Encoding and Decoding

181

Oota et al. 2022

182 of 205

Multi-view decoding

IJCNN 2023: DL for Brain Encoding and Decoding

182

Wordcloud View

Train

Sentence View

Picture View

Wordcloud View

Oota et al. 2022

Picture View

Train

Sentence View

Train

183 of 205

Multi-view decoding results

IJCNN 2023: DL for Brain Encoding and Decoding

183

Picture View

Train

BERT Representaions

Shuffled the Target Concepts

Test

Sentence View

Train

WordCloud View

Train

Pictures Best Accuracy

Sentences Best Accuracy

Oota et al. 2022

184 of 205

Distribution of Informative Voxels

IJCNN 2023 : DL for Brain Encoding and Decoding

184

Oota et al. 2022

185 of 205

Distribution of Information Voxels: Language Network

IJCNN 2023 : DL for Brain Encoding and Decoding

185

Oota et al. 2022

186 of 205

Distribution of Informative Voxels: Visual Network

IJCNN 2023: DL for Brain Encoding and Decoding

186

Oota et al. 2022

187 of 205

Cross-view Decoding

IJCNN 2023: DL for Brain Encoding and Decoding

187

Picture View

Train

Caption

Test

Picture View

Train

Visual words

Test

Wordcloud View

Train

Sentence

Test

Sentence View

Train

Keywords

Test

Oota et al. 2022

188 of 205

Cross-view Decoding results

IJCNN 2023: DL for Brain Encoding and Decoding

188

BERT Representaions

Shuffled the Target Concepts

Oota et al. 2022

189 of 205

Distribution of Informative Voxels

IJCNN 2023: DL for Brain Encoding and Decoding

189

Oota et al. 2022

190 of 205

Distribution of Informative Voxels: Language Network

IJCNN 2023: DL for Brain Encoding and Decoding

190

Oota et al. 2022

191 of 205

Distribution of Information Voxels: Visual Network

IJCNN 2023: DL for Brain Encoding and Decoding

191

Oota et al. 2022

192 of 205

Summary

  • Cross-view and Multi-view decoding tasks establish that the information contained in the brain response is rich and capable of driving multiple downstream tasks.

IJCNN 2023: DL for Brain Encoding and Decoding

192

Oota et al. 2022

193 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

193

194 of 205

References

IJCNN 2023: DL for Brain Encoding and Decoding

194

195 of 205

References

  • Nishimoto, Shinji, et al. "Reconstructing visual experiences from brain activity evoked by natural movies." Current biology 21.19 (2011): 1641-1646.
  • Anumanchipalli, Gopala K., Josh Chartier, and Edward F. Chang. "Speech synthesis from neural decoding of spoken sentences." Nature 568.7753 (2019): 493-498.
  • Schrimpf, Martin, et al. "The neural architecture of language: Integrative modeling converges on predictive processing." Proceedings of the National Academy of Sciences 118.45 (2021): e2105646118.
  • Wehbe, Leila, et al. "Simultaneously uncovering the patterns of brain regions involved in different story reading subprocesses." PloS one 9.11 (2014): e112575.

IJCNN 2023: DL for Brain Encoding and Decoding

195

196 of 205

Language and the Brain: Deep Learning for Brain Encoding and Decoding

Subba Reddy Oota1, Manish Gupta2,3, Raju S. Bapi2

1Inria Bordeaux, France; 2IIIT Hyderabad, India; 3Microsoft, India

subba-reddy.oota@inria.fr, gmanish@microsoft.com, raju.bapi@iiit.ac.in

197 of 205

Agenda

  • Introduction to Brain encoding and decoding [10 min]
  • Text Stimulus Representations [30 min]
  • Deep Learning for Brain Encoding [40 min]
  • Deep Learning for Brain Decoding [30 min]
  • Summary and Future Trends [10 min]

IJCNN 2023: DL for Brain Encoding and Decoding

197

198 of 205

Outline

  1. Summary [5min]
  2. Future trends [5min]

198

IJCNN 2023: DL for Brain Encoding and Decoding

199 of 205

Summary

  • Exciting times: publicly accessible neuroimaging data of various tasks starting to be avaliable now!
    • Opportunities:
      • Data ahead of theory, so it’s an open field for theoretical and methodological innovation!
      • DL is helpful in uncovering patterns in brain responses and may lead to theories of information organization in the brain.
    • Challenges:
      • Hypothesis-driven data collection might be more helpful
      • Individual variability is the norm in neuroimaging data!
      • Neuroimaging data is more complex, noisy as compared to classical datasets used by DL researchers

199

IJCNN 2023: DL for Brain Encoding and Decoding

200 of 205

Summary

  • This Tutorial:
  • Stimulus representation schemes
    • Vision: CNN-based
    • Language: Transformer-based
  • Datasets available (Reading/Listening/Viewing tasks in EEG, MEG, fMRI)
  • Encoding
    • Linear and Non-linear models
  • Decoding
    • Linear and Non-linear models
  • Advance methods
    • Tuning/Training DL models using brain recordings
    • Task-based modeling

200

IJCNN 2023: DL for Brain Encoding and Decoding

201 of 205

Outline

  1. Summary [5min]
  2. Future trends: DNNs & The Brain [10min]

201

IJCNN 2023: DL for Brain Encoding and Decoding

202 of 205

DNNs & The Brain: Multi-modal, Multi-task

  • Brain response to a stimulus is multi-modal, multi-task related
    • Cross-view and multi-view decoding (Oota et al 2022a)
    • Visio-linguistic encoding (fusion of vision and language information) (Oota et al 2022b)
    • Task-based representations give better brain alignment (Neural Taskonomy: Oota et al 2022c)
    • Multimodal foundation model (Fei et al 2022)

Fei, Lu, Gao et al (2022). Towards artificial general intelligence via a multimodal foundation model. Nature Communications 13:3094

doi.org/10.1038/s41467-022-30761-2

203 of 205

DNNs & Brain: Multi-modal, Multi-task

  • Brain response to a stimulus is multi-modal, multi-task related
    • Cross-view and multi-view decoding (Oota et al 2022)
    • Visio-linguistic encoding (fusion of vision and language information) (Oota et al 2022)
    • Multimodal foundation model (Fei et al 2022)

Fei, Lu, Gao et al (2022). Towards artificial general intelligence via a multimodal foundation model. Nature Communications 13:3094 doi.org/10.1038/s41467-022-30761-2

204 of 205

DNNs & Brain Damage

  • DL models of encoding and decoding have not yet been put through the brain-damage experiments. Ex. Semantic Dementia

Snowden, Harris, Thompson, Kobylecki, Jones, Richardson, Neary (2018). Semantic dementia and the left and right temporal lobes, Cortex, 107(188-203).

https://doi.org/10.1016/j.cortex.2017.08.024.

Rt Ant Temporal Lobe Damage (Patient 8)

Animal habitat task.

The patient is asked:

Where would you find this?

Do DL Models exhibit such degradation with damage to units?

205 of 205

A big thank you!

Tutorial, Code and Material:

Material from IJCNN 2023 Tutorial would be uploaded soon!

Upcoming Tutorials:

  • Deep Neural Networks and Brain Alignment: Brain Encoding and Decoding IJCAI-2023 (Macau, Aug 2023)

(Past): Deep Learning for Brain Encoding and Decoding, Cogsci-2022

https://tinyurl.com/DL4Brain