1 of 24

Speech Taskonomy: Which Speech Tasks are the most Predictive of fMRI Brain Activity?

1

Subba Reddy Oota1

Veeral Agarwal2

Mounika Marreddy2

Manish Gupta2,3

Bapi Raju Surampudi2

1Inria Bordeaux, France 2IIIT-Hyderabad, India 3Microsoft, India

2 of 24

What is fMRI?

2

A listening task in the scanner

Narrative Story

fMRI Brain Activity

3 of 24

Brain Encoding vs Decoding

3

Encoding

Stimulus

Representation

fMRI

Stimulus

Representation

Decoding

fMRI

4 of 24

4

Method to study alignment of SM & brain representations

Brain alignment of a SM ⇒ how similar its representations are to a human brain’s

Speech Model (SM)

5 of 24

Brain Encoding?

Present

Stimulus

Stimulus

Ridge Regression

Input

Input

Output

X

Y

W

Pearson Correlation (R) = Corr(Y, W(X))

6 of 24

Encoding: training independent models

  • Independent model per participant
  • Independent model per voxel / sensor-timepoint

P1

P2

PN

P1, v1

P1, v2

P1, vm

7 of 24

Recent work utilizing progress in self-supervised speech models for encoding

  • Stimuli: Moth Radio Hour
  • Stimulus representation: derived from pretrained self-supervised speech models (HuBERT, Wav2Vec2.0, APC)
  • Brain recording & modality: fMRI, listening

Middle layers of self-supervised speech models predict auditory cortex the best

8 of 24

Audio work utilizing DL progress

  • Stimuli: audio books
  • Stimulus representation: derived from pretrained self-supervised speech model (Wav2Vec2.0)
  • Brain recording & modality: fMRI, listening in 3 languages (Eng, Fr, Mandarin)
  • Self-supervised speech models exhibit specialization for native sounds in the STS and MTG.
  • IFG and AG show more general specialization for speech rather than native-language

9 of 24

Audio work utilizing DL progress

  • Stimuli: Moth-Radio-Hour
  • Stimulus representation: derived from 5 basic + 25 pretrained self-supervised speech models
  • Brain recording & modality: fMRI

Contrastive and predictive models encode the information better than the generative and the traditional low-level acoustic baselines, and VGGish models.

Data2Vec

10 of 24

Challenges in using DL for cognitive science

  • Not designed to specifically model brain processing
    • Training DL models using brain recordings
    • Task-based modeling
  • Can be difficult to interpret due to multiple sources of information
    • Disentangling contributions of different info sources to brain predictions

10

11 of 24

Tasks affect processing

11

  • Stimuli: images of natural scenes
  • Stimulus representation: task-optimized CNNs for a range of tasks
  • Brain recording & modality: fMRI, vision

Semantic Low-dim. Geometric 2D 3D

Vision tasks with higher transferability make similar predictions for brain responses from different regions

12 of 24

Tasks affect processing

12

  • Stimuli: passages and narratives
  • Stimulus representation: task-optimized NLP models for a range of tasks
  • Brain recording & modality: fMRI, reading & listening of different stimuli
  • Reading fMRI best explained by coref. resolution, NER, shallow syntax parsing
  • Listening fMRI best explained by paraphrasing, summarization, NLI

13 of 24

Can task-specific speech models better predict fMRI brain activity?

Model fine-tuned to�downstream speech tasks

SUPERB (Speech Processing Universal PERformance Benchmark)

https://superbbenchmark.org

PR

SID

KS

SV

ASR

SD

ER

IC

14 of 24

14

Pretrained�speech model

Model fine-tuned to�downstream speech tasks

input

input

activations

activations

Wav2Vec 2.0-base

15 of 24

15

Pretrained�speech model

Model fine-tuned to�downstream speech tasks

input

input

Wav2Vec 2.0-base

Use model’s internal layer�activations to predict brain activity on held-out data

activations

activations

Compare against actual brain recordings

(brain alignment)

16 of 24

16

Pretrained�speech model

Model fine-tuned to�downstream speech tasks

input

input

Wav2Vec2.0-base

Use model’s internal layer�activations to predict brain activity on held-out data

activations

activations

Compare against actual brain recordings

(brain alignment)

Hypothesis: If the pretrained model fine-tuned to downstream speech task has greater brain alignment than pretrained model, the downstream task is capturing more brain-relevant information

17 of 24

Listening data target: human brain recordings

17

  • We use Pieman story listening:
    • 82 subjects,
    • 282 TRs (repetition time)
    • here it is 1.5 sec.

Example: ''I began my illustrious carrier in journalism…''

18 of 24

18

Brain alignment – 4-fold Cross-Validation + Ridge regression

  1. 4-fold Cross-Validation
  1. Linear regression regularized�with ridge penalty

Input: R282 x d

d = embedding size, e.g. 4608

Prediction: Rn x v

n = number of fMRI intervals

v = number of voxels in participant’s brain

19 of 24

19

Brain alignment – Methodology

  • 300 TRs x 768 ⇒ SM representations

    • Wav2Vec 2.0-base
    • SUPERB Benchmark downstream tasks : eight tasks

  • 282 TRs x 768

    • Remove 18 TRs ⇒ 10 TRs in the beginning and 8 TRs in the ending (silent music)

  • 282 fMRI time intervals x 4608

    • Concatenate SM representations for previous 6 TRs ⇒ fMRI response from brain activity peaks about 8-10 seconds after stimulus onset

  • 282 fMRI intervals x 4608

    • Ridge regression (RR) ⇒ for each voxel, 282 data rows of 4608 parameters to predict 1 output
    • 4-fold Cross-Validation to improve reliability

  • 282 fMRI intervals x 52400 voxels ⇒ fMRI predictions (same dimensions as actual brain activity)

20 of 24

Results: Whole-Brain

20

ASR best encodes speech stimuli for brain response prediction.

ASR task has the best brain alignment in the middle layers.

  • Certain speech tasks (ASR, ER, SID and IC) that are important for improved brain alignment over pretrained Wav2Vec2.0.
  • SD and SV are not important in listening to stories.

21 of 24

Region level alignments

    • All speech tasks are better aligned with EAC compared to AAC and IFG regions.
    • Finetuning on ER, SID and IC leads to the best alignment for the early auditory cortex
    • Finetuning on ASR provides the best encoding for the auditory associative cortex and language regions.

Speech Taskonomy: Which Speech Tasks are the most Predictive of fMRI Brain Activity? Subba Reddy Oota, Veeral Agarwal, Mounika Marreddy, Manish Gupta, Raju Bapi. InterSpeech 2023

21

22 of 24

Sub-region level alignments

    • EAC: A1 has a higher Pearson correlation than other sub-ROIs.
    • Language ROIs 44 and 45, together with STSda and STSdp in AAC, are part of the well-known language network associated with narrative comprehension
    • ASR finetuned model performs best in these regions.

22

23 of 24

Limitations

  • We leveraged models finetuned using datasets of different sizes across tasks.
  • While a fair comparison of dataset sizes across tasks is impossible,
    • we understand that this could have resulted in some bias in our results.

24 of 24

Thank You

Questions?