Speech Taskonomy: Which Speech Tasks are the most Predictive of fMRI Brain Activity?
1
Subba Reddy Oota1
Veeral Agarwal2
Mounika Marreddy2
Manish Gupta2,3
Bapi Raju Surampudi2
1Inria Bordeaux, France 2IIIT-Hyderabad, India 3Microsoft, India
What is fMRI?
2
A listening task in the scanner
Narrative Story
fMRI Brain Activity
Brain Encoding vs Decoding
3
Encoding
Stimulus
Representation
fMRI
Stimulus
Representation
Decoding
fMRI
4
Method to study alignment of SM & brain representations
Brain alignment of a SM ⇒ how similar its representations are to a human brain’s
Speech Model (SM)
Brain Encoding?
Present
Stimulus
Stimulus
Ridge Regression
Input
Input
Output
X
Y
W
Pearson Correlation (R) = Corr(Y, W(X))
Encoding: training independent models
P1
…
P2
PN
P1, v1
P1, v2
…
P1, vm
Recent work utilizing progress in self-supervised speech models for encoding
Middle layers of self-supervised speech models predict auditory cortex the best
Audio work utilizing DL progress
Audio work utilizing DL progress
Contrastive and predictive models encode the information better than the generative and the traditional low-level acoustic baselines, and VGGish models.
Data2Vec
Challenges in using DL for cognitive science
10
Tasks affect processing
11
Semantic Low-dim. Geometric 2D 3D
Vision tasks with higher transferability make similar predictions for brain responses from different regions
Tasks affect processing
Oota, Subba Reddy, Jashn Arora, Veeral Agarwal, Mounika Marreddy, Manish Gupta, and Bapi Raju Surampudi. "Neural Language Taskonomy: Which NLP Tasks are the most Predictive of fMRI Brain Activity?." arXiv preprint arXiv:2205.01404 (2022).
12
Can task-specific speech models better predict fMRI brain activity?
Model fine-tuned to�downstream speech tasks
SUPERB (Speech Processing Universal PERformance Benchmark)
https://superbbenchmark.org
PR
SID
KS
SV
ASR
SD
ER
IC
14
Pretrained�speech model
Model fine-tuned to�downstream speech tasks
input
input
activations
activations
Wav2Vec 2.0-base
15
Pretrained�speech model
Model fine-tuned to�downstream speech tasks
input
input
Wav2Vec 2.0-base
Use model’s internal layer�activations to predict brain activity on held-out data
activations
activations
Compare against actual brain recordings
(brain alignment)
16
Pretrained�speech model
Model fine-tuned to�downstream speech tasks
input
input
Wav2Vec2.0-base
Use model’s internal layer�activations to predict brain activity on held-out data
activations
activations
Compare against actual brain recordings
(brain alignment)
Hypothesis: If the pretrained model fine-tuned to downstream speech task has greater brain alignment than pretrained model, the downstream task is capturing more brain-relevant information
Listening data target: human brain recordings
17
Example: ''I began my illustrious carrier in journalism…''
18
Brain alignment – 4-fold Cross-Validation + Ridge regression
Input: R282 x d
d = embedding size, e.g. 4608
Prediction: Rn x v
n = number of fMRI intervals
v = number of voxels in participant’s brain
19
Brain alignment – Methodology
Results: Whole-Brain
20
ASR best encodes speech stimuli for brain response prediction.
ASR task has the best brain alignment in the middle layers.
Region level alignments
Speech Taskonomy: Which Speech Tasks are the most Predictive of fMRI Brain Activity? Subba Reddy Oota, Veeral Agarwal, Mounika Marreddy, Manish Gupta, Raju Bapi. InterSpeech 2023
21
Sub-region level alignments
22
Limitations
Thank You
Questions?