Enhancing Stroke-Induced Aphasia Assessment Using Interpretable Linguistic Biomarkers and Large Language Models
�2025 Voice AI Symposium��Yan Cong�cong4@purdue.edu �Purdue University
1
Outline
2
Outline
3
Background
“A biological marker or biomarker is a characteristic that is objectively measured and evaluated as an indicator of normal biologic processes, pathologic processes, or biological responses to a therapeutic intervention.”
(Source: concept clearance by the NIMH, Brady 2014;
Corcoran, Mittal, …, Cecchi, Wolff, 2020
Language as a biomarker for psychosis: a natural language processing approach)
4
Background
"In the context of speech (text + sound), a linguistic biomarker would be a feature or set of features that is associated with a clinical outcome and can be used either to detect a pathological state, or to monitor and classify its severity and stages of impairment. "
(Source: concept clearance by the NIMH, Brady 2014;
Corcoran, Mittal, …, Cecchi, Wolff, 2020
Language as a biomarker for psychosis: a natural language processing approach)
5
Background
(Inspirations: Sunny X. Tang; Phillip Wolff; Sunghye Cho; among others)
6
Background
(Inspirations: Sunny X. Tang; Phillip Wolff; Sunghye Cho; among others)
7
LLM-surprisal and its relation to the clinical manifestation of aphasia
Research gap
Aphasia: an acquired neurogenic language disorder, most often caused by stroke
Language measurement: critical and non-invasive in predicting and treating language-related disorders and impairments
8
Research gap
9
Today’s plan
10
Jiyeon Lee
Arianna N. LaCroix
LLM-surprisal as a promising and interpretable linguistic marker of aphasia
Current study
Utilize pre-trained large language models (LLMs) derived surprisals to detect aphasia in Chinese speakers, and examine how surprisals relate to the clinical manifestation of aphasia.
11
Methods
12
A working definition of Surprisal
13
LLM-surprisal
LLM-surprisal
LLM-surprisal in clinical speech-transcribed text
Using GPT2:
16
Rezaii et al., 2023
Computational psycholinguistic study of Surprisal
17
Methods
18
Data
All datasets were drawn from the AphasiaBank (MacWhinney et al., 2011 https://talkbank.org/DB/).
Participants: monolingual English or Mandarin Chinese speakers, with a Western Aphasia Battery Revised Aphasia Quotient (WAB-R-AQ, Kertesz, 2007) of 92 or lower in the aphasia group.
19
Data
Chinese dataset
English dataset
20
Model
Three tasks in both English and Chinese datasets:
Logistic regression classifiers classify aphasia and control (task 1) and Broca’s and anomic aphasia (task 2).
Elastic net regressions predict WAB-R-AQ scores (task 3).
21
Model
Each LLM read in an utterance in text and output a surprisal score.
Mean surprisal: token-wise surprisals averaged over the utterance.
Hypothesis: higher surprisals, as an indicator of larger amount of grammatical unacceptability, are associated with higher severity of aphasia.
22
Model
Three pre-trained LLMs:
23
Features
Predictor variable:
A preliminary experiment focusing on one utility (i.e., LLMs surprisal) in a cross-linguistic clinical setting
24
Results
25
Results
More effective in subtyping than detecting the presence of aphasia in Chinese speakers.
Inverse pattern for detecting aphasia in English speakers.
- Crosslinguistic difference
- Character-level tokenization
26
Results
27
Results
English tasks: the two decoder LLMs showed negative effects, Llama2 showed the strongest coefficients.
Chinese tasks: utterance length matters, all LLMs showed negative coefficients, Llama2 gave the largest coefficients.
- scaling improves performance in both English and Chinese tasks.
- clinical application: a critical need to pre-train LLMs in the target language
28
Qualitative error analysis
29
Qualitative error analysis
30
Qualitative error analysis
31
Conclusion & Discussion
32
Conclusion & Discussion
Enhancing Stroke-Induced Aphasia Assessment Using Interpretable linguistic biomarkers and Large Language Models
33
Work-in-progress
Enhancing Stroke-Induced Aphasia Assessment Using Interpretable linguistic biomarkers and Large Language Models
34
Acknowledgements
We acknowledge the AphasiaBank (https://talkbank.org/DB/#), Brian MacWhinney and Davida Fromm, for the valuable resource.
We acknowledge Emily Tumacder’s and Cameron Pilla’s help with compiling the datasets and optimizing the machine learning pipeline. We thank Emmanuele Chersoni, Sunny X. Tang, Phillip Wolff, and Sunghye Cho for their inspirations.
All errors remain mine.
35
Acknowledgements
36
Sunny X. Tang
Jiyeon Lee
Arianna N. LaCroix
Emmanuele Chersoni
Phillip Wolff