1 of 36

Enhancing Stroke-Induced Aphasia Assessment Using Interpretable Linguistic Biomarkers and Large Language Models

�2025 Voice AI Symposium��Yan Cong�cong4@purdue.edu �Purdue University

1

2 of 36

Outline

  • Brief background & research gap
  • Current study & Hypothesis
  • Methodology
  • Results & Discussion

2

3 of 36

Outline

  • Brief background & research gap
  • Current study & Hypothesis
  • Methodology
  • Results & Discussion

3

4 of 36

Background

A biological marker or biomarker is a characteristic that is objectively measured and evaluated as an indicator of normal biologic processes, pathologic processes, or biological responses to a therapeutic intervention.”

(Source: concept clearance by the NIMH, Brady 2014;

Corcoran, Mittal, …, Cecchi, Wolff, 2020

Language as a biomarker for psychosis: a natural language processing approach)

4

5 of 36

Background

"In the context of speech (text + sound), a linguistic biomarker would be a feature or set of features that is associated with a clinical outcome and can be used either to detect a pathological state, or to monitor and classify its severity and stages of impairment. "

(Source: concept clearance by the NIMH, Brady 2014;

Corcoran, Mittal, …, Cecchi, Wolff, 2020

Language as a biomarker for psychosis: a natural language processing approach)

5

6 of 36

Background

  • Identifying features with relevance for computational modeling
    • Objective quantification of language-related changes in language-related disorder and impairment
  • How should we define the clinical measure of interest?
    • Identify features behind clinical observations of language disorders
      • Have distinct clinical significance
      • Suitable for computational modeling using distinct strategies

(Inspirations: Sunny X. Tang; Phillip Wolff; Sunghye Cho; among others)

6

7 of 36

Background

  • Identifying features with relevance for computational modeling
    • Objective quantification of language-related changes in language-related disorder and impairment
  • How should we define the clinical measure of interest?
    • Identify features behind clinical observations of language disorders
      • Have distinct clinical significance
      • Suitable for computational modeling using distinct strategies

(Inspirations: Sunny X. Tang; Phillip Wolff; Sunghye Cho; among others)

7

LLM-surprisal and its relation to the clinical manifestation of aphasia

8 of 36

Research gap

Aphasia: an acquired neurogenic language disorder, most often caused by stroke

Language measurement: critical and non-invasive in predicting and treating language-related disorders and impairments

  • Large language models (LLMs) update the state-of-the-art of language measurement
  • Operationalize natural language processing (NLP) and LLMs in a clinical context (ecological validity)
  • Accelerate more economical, accurate, and effective clinical trials and decision-making

8

9 of 36

Research gap

  • Aphasia studies with NLP approaches regarding monolingual English speakers (Salem et al., 2023; Purohit et al., 2023; Sanguedolce et al., 2023; Ortiz-Perez et al., 2023).
  • Other population (Smaïli et al., 2022; Chatzoudis et al., 2022; Balagopalan et al., 2020).
  • Aphasia studies in Chinese speakers adopt NLP methods (Balagopalan et al., 2020; Shivkumar et al., 2020; Mahmoud et al., 2020; Qin et al., 2022).

9

10 of 36

Today’s plan

10

Jiyeon Lee

Arianna N. LaCroix

LLM-surprisal as a promising and interpretable linguistic marker of aphasia

11 of 36

Current study

Utilize pre-trained large language models (LLMs) derived surprisals to detect aphasia in Chinese speakers, and examine how surprisals relate to the clinical manifestation of aphasia.

11

12 of 36

Methods

  1. Implement LLMs surprisals for aphasia detection in Chinese speakers
  2. Compare LLMs surprisals in Chinese datasets with those in English

  • More about surprisal
  • Data
  • Model
  • Features

12

13 of 36

A working definition of Surprisal

  • The level of unexpectedness associated with a sequence of words in a given context (Hale, 2001; Levy, 2008; Tunstall et al., 2022; Futrell et al., 2018; Van Schijndel and Linzen, 2018; Wilcox et al., 2018; Michaelov and Bergen, 2020, 2022a; Michaelov et al., 2023;Rezaii et al., 2023,2024, among others)
    • S(wi) = -logP(wi|w1...wi-1)
  • Operationalize the measurement of Surprisal: minicons API (Misra 2022)
    • Sequence Surprisal (i.e., sentence), normalized by number of tokens

13

14 of 36

LLM-surprisal

  1. The keys to the cabinet are on the table.
  2. The keys to the cabinet is on the table.
  3. Olivia bought a German shepherd. The dog was docile and friendlyHowever, it bit her hand.
  4. Olivia bought a German shepherd. The dog was unpredictable and violentHowever, it bit her hand.

15 of 36

LLM-surprisal

  1. The keys to the cabinet are on the table. [GPT2 Surprisal 38.88].
  2. The keys to the cabinet is on the table. [GPT2 Surprisal 42.76].
  3. Olivia bought a German shepherd. The dog was docile and friendlyHowever, it bit her hand. [GPTNeo Surprisal 6.38].
  4. Olivia bought a German shepherd. The dog was unpredictable and violentHowever, it bit her hand. [GPTNeo Surprisal 9.77].

16 of 36

LLM-surprisal in clinical speech-transcribed text

  1. I forget the name of it.
  2. There is a boombox playing music.
  3. That is my particular unique problem.

Using GPT2:

  1. 4.9
  2. 6.1
  3. 7.5

16

Rezaii et al., 2023

17 of 36

Computational psycholinguistic study of Surprisal

  • Reproduce human semantic priming effects using BERT word predictions (Misra et al., 2020)
  • LLMs surprisals and the collateral facilitation effect (Michaelov and Bergen, 2022a)
  • Reproduce the effect of context in reducing the N400 amplitude for inconsistent words (Michaelov et al., 2023)
  • Morphosyntactic attraction effects (Ryu and Lewis, 2021)
  • Using LLMs surprisals to account for psycholinguistic phenomena (Futrell et al., 2018; Van Schijndel and Linzen, 2018; Wilcox et al., 2018; a.o.)

17

18 of 36

Methods

  1. Implement LLMs surprisals for aphasia detection in Chinese speakers
  2. Compare LLMs surprisals in Chinese datasets with those in English

  • More about surprisal
  • Data
  • Model
  • Features

18

19 of 36

Data

All datasets were drawn from the AphasiaBank (MacWhinney et al., 2011 https://talkbank.org/DB/).

Participants: monolingual English or Mandarin Chinese speakers, with a Western Aphasia Battery Revised Aphasia Quotient (WAB-R-AQ, Kertesz, 2007) of 92 or lower in the aphasia group.

19

20 of 36

Data

Chinese dataset

  • matched sample (on age, education, sex)
  • 1756 observations for each group (healthy control and aphasia)
  • tasks picture description and story retelling
  • subtyping aphasia: randomly sampled balanced sets for Broca’s and anomic aphasia (N=86).

English dataset

  • same methods
  • N=1586 in aphasia detection and severity measurement
  • N=86 in aphasia subtyping.

20

21 of 36

Model

Three tasks in both English and Chinese datasets:

  1. Detecting the presence of aphasia
  2. Detecting aphasia subtypes
  3. Detecting aphasia severity

Logistic regression classifiers classify aphasia and control (task 1) and Broca’s and anomic aphasia (task 2).

Elastic net regressions predict WAB-R-AQ scores (task 3).

21

22 of 36

Model

Each LLM read in an utterance in text and output a surprisal score.

Mean surprisal: token-wise surprisals averaged over the utterance.

Hypothesis: higher surprisals, as an indicator of larger amount of grammatical unacceptability, are associated with higher severity of aphasia.

22

23 of 36

Model

Three pre-trained LLMs:

  1. GPT2 (Radford et al., 2019; Zhao et al., 2019, 2023b)
  2. Llama2-7B (Touvron et al., 2023)
  3. BERT (bert-base-chinese for Chinese and bert-base-uncased for English) (Devlin et al., 2019, 2018)

23

24 of 36

Features

Predictor variable:

  1. utterance length (MacWhinney et al., 2011; Fromm and MacWhinney, 2023; Fromm et al., 2022, 2020)
  2. utterance level mean surprisal computed by pre-trained LLMs (Rezaii et al., 2023a).

A preliminary experiment focusing on one utility (i.e., LLMs surprisal) in a cross-linguistic clinical setting

24

25 of 36

Results

25

26 of 36

Results

More effective in subtyping than detecting the presence of aphasia in Chinese speakers.

Inverse pattern for detecting aphasia in English speakers.

- Crosslinguistic difference

- Character-level tokenization

26

27 of 36

Results

27

28 of 36

Results

English tasks: the two decoder LLMs showed negative effects, Llama2 showed the strongest coefficients.

Chinese tasks: utterance length matters, all LLMs showed negative coefficients, Llama2 gave the largest coefficients.

- scaling improves performance in both English and Chinese tasks.

- clinical application: a critical need to pre-train LLMs in the target language

28

29 of 36

Qualitative error analysis

29

30 of 36

Qualitative error analysis

30

31 of 36

Qualitative error analysis

  • Extremely short utterances turn out to give rise to large surprisal scores in both Chinese and English datasets.
  • For BERT, the effect of utterance length is not salient.
  • Specific words that may lead to outstanding LLMs surprisals: interjection, filler words, low frequency words, and sentence final particles (in Chinese).
  • Level of pre-processing matters.

31

32 of 36

Conclusion & Discussion

  • Leverage pre-trained LLMs to detect the presence, subtypes, and severity of aphasia in English and Mandarin Chinese speakers.
  • Without fine-tuning, taking pre-trained LLMs off-the-shelf can inform us how surprisals distribute in aphasic individuals whose first language is not English.
  • English LLMs exhibit decent accuracy in detecting the presence of aphasia; the Chinese counterparts demonstrate satisfactory performance in subtyping aphasia.
  • Pre-trained LLMs have clinical potential (e.g., automatic aphasia diagnosis), especially in the context of multilingual populations.

32

33 of 36

Conclusion & Discussion

Enhancing Stroke-Induced Aphasia Assessment Using Interpretable linguistic biomarkers and Large Language Models

  • LLM-surprisal as a promising and interpretable linguistic marker of aphasia
    • Language measurement: critical and non-invasive in predicting and treating language-related disorders and impairments
    • LLMs update the state-of-the-art of language measurement
    • Operationalize NLP and LLMs in a clinical context (ecological validity)
    • Accelerate more economical, accurate, and effective clinical trials and decision-making

33

34 of 36

Work-in-progress

Enhancing Stroke-Induced Aphasia Assessment Using Interpretable linguistic biomarkers and Large Language Models

  • Synthetic data augmentation
  • Fine-tuning on domain-specific aphasia data
  • LLMs implementation in language recovery and treatment

34

35 of 36

Acknowledgements

We acknowledge the AphasiaBank (https://talkbank.org/DB/#), Brian MacWhinney and Davida Fromm, for the valuable resource.  

We acknowledge Emily Tumacder’s and Cameron Pilla’s help with compiling the datasets and optimizing the machine learning pipeline. We thank Emmanuele Chersoni, Sunny X. Tang, Phillip Wolff, and Sunghye Cho for their inspirations.

All errors remain mine.

35

36 of 36

Acknowledgements

36

Sunny X. Tang

Jiyeon Lee

Arianna N. LaCroix

Emmanuele Chersoni

Phillip Wolff