Responsible Development and Translation of Clinical Speech Analytics
Instruction for preparing the tutorial slides
With a subtitle if you need it
Here is a text slide
With a subtitle if you need it
When using the template, try to not move the headlines up and down. It will add a level or professional design if the top margin of your headline is not moving around between slide transitions.
Presenters
Agenda
Morning Session
9:00 – 9:10 a.m.� Introduction�� 9:10 – 9:50 a.m.� Design of speech elicitation tasks
●Saliency of speech across different conditions
●Categorization of elicitation tasks
●Matching elicitation tasks with clinical conditions
� 9:50 – 10:30 a.m.� Speech data acquisition
●Influence of hardware devices
●Sources of noise and how it impacts speech features
●Guidelines to control parameters in data acquisition
●Hardware validation framework
Agenda
Morning Session
11:00 – 12:00 p.m.�Methodological shift from speech features to speech measures
●Methodological shift from features to measures
●Existing speech measures
●Validation framework of speech measures
●Case study: FDA breakthrough designation
Agenda
Afternoon Session
2:00 – 3:15 p.m.� Learning generalizable and interpretable� speech representations and clinical ML models
● Development speech measures with ML
● Analytical validation of speech measures
● Development clinical ML models using speech measures
● Clinical validation of developed models
● Recommendations from a dataset perspective to support development of speech measures and models
Agenda
Afternoon Session
3:30 – 4:30 p.m.� Ethical, privacy and security considerations
●Data bias, data privacy, model security, improper ML model use
●Directions toward responsible development of clinical speech analytics
●Participant acceptability, motivation, convenience, accessibility, and usability.
� 4:30 – 5:00 p.m.� Discussion
Introduction
Speech
Disease
Neurological disorders, mental health disorders, speech-motor disorders, voice disorders
Speech acoustics and natural language
Activities
Impaired abilities to communicate with others (e.g. reduced intelligibility, impaired social skills)
Participation
Unable to work, spend time with friends, etc.
The promise of speech-based biomarkers
The gap between promise and reality
but there are few (if any?) clinical speech models that have been deployed
Why is this?
What can we do about it?
The current approach to speech-based biomarkers
Extract standard speech features using existing tools
Error
Rate
Predict a clinical variable of interest
OpenSMILE
Feature Extraction
Extract features using existing tools
wav2vec
NLP features
Mel spectra
Healthy vs. MCI
Speech Database
If the resulting accuracy is “good”
PUBLISH the final model
Else
MOVE ON TO NEXT PROJECT
13
1Stegmann, Hahn, Liss, Shefner, Rutkove, Kawabata, Bhandari, Shelton, Duncan, Berisha. The Repeatability of Commonly Used Speech and Language Features for Clinical Applications. Digital Biomarkers. Jan, 2021.
2Alhanai et al. "Spoken language biomarkers for detecting cognitive impairment." 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2017.
3Berisha et al “Digital medicine and the curse of dimensionality.” Nature Digital Medicine, October 2021.
4Berisha et al, “Are reported accuracies in the clinical speech machine learning literature overoptimistic”, to appear at Interspeech 22.
Three problems with commonly-used features
Berisha et al. "Digital medicine and the curse of dimensionality." Nature npj Digital Medicine 4.1 (2021): 1-8.
Berisha et al, “Are reported accuracies in the clinical speech machine learning literature overoptimistic”, Interspeech 22.
This approach leads to overoptimism in the literature
4% decrease in reported accuracy per unit increase in log(SampleSize)
Sample Size
Accuracy
Building clinical speech analytics isn’t just an ML problem
How to design speech elicitation tasks?
How to collect the data in a robust way?
How to design interpretable representations for clinical speech models and how to validate them?
How do we train interpretable clinical speech ML models?
What are important ethical, safety, and security considerations?
Design of
Speech Elicitation Tasks
Saliency of Speech across Different Conditions
The production of spoken language
Saliency of Speech across Different Conditions
How speech sounds and what is said?
Disease or Condition | Speech Process Affected | Content: “What is said…” Natural Language Processing | Form: “How it’s said…” Acoustic Analysis |
Mental Health (schizophrenia, bipolar depression) | CONCEPTUALIZATION | Reduced or increased speech output, incoherent speech, atypical sentence structure | Atypical speaking rate (too fast or slow), atypical prosody |
Cognitive/Language (ADRD, Aphasia) | FORMULATION | Low lexical complexity, simple or atypical syntax, | Frequent and long pauses, slow speaking rate, imprecise articulation |
Motor (Parkinson’s, ALS) | ARTICULATION | Reduced speech output | Slurred, slow or mumbled speech, atypical prosody, dysphonia, hypernasality |
Speech Elicitation Tasks
Critical Considerations for Design
·
1.Goal/Purpose of the AI Model
2.Stage of Speech Production Impacted: Conception, Formulation, Articulation
3.Patient/Participant Considerations
4.Language and/or Cultural Considerations
5.Speech Processing and Metrics
Critical Considerations for Designing Speech Elicitation Tasks
2. Conception, Formulation, or Articulation?
Critical Considerations for Designing Speech Elicitation Tasks
3. Patient/Participant Considerations?
Critical Considerations for Designing Speech Elicitation Tasks
4. Language and/or Cultural Considerations?
Critical Considerations for Designing Speech Elicitation Tasks
4. Speech Processing and Metrics?
Critical Considerations for Designing Speech Elicitation Tasks
Speech Elicitation Tasks
Speech task landscape
Matching Elicitation Tasks with ML Use-Cases
Fit-for-Purpose Design
Example 1: �Prospective collection of speech samples to classify between mild cognitive impairment and no cognitive impairment in a genetically at-risk population of 50–80-year-old men and women in a health care system in New England major cities.
GOAL | Cross-sectional classification of +/- MCI |
DEFICITS/SEVERITY | Formulation: cognition, memory, executive function, processing speed, word-finding/ Unimpaired to mildly impaired |
PARTICIPANT LIMITATIONS | None to minimal |
LANGUAGE/CULTURE | All American English native speakers |
SPOKEN RESPONSE PROCESSING/METRICS | ASR generated transcripts for NLP analysis |
Matching Elicitation Tasks with ML Use-Cases
Fit-for-Purpose Design
Example 2: �Prospective collection of speech samples to identify impact of intervention on disease progression in chronic obstructive pulmonary disease in women between the ages of 60-75 years old.
GOAL | Within-subject, longitudinal change detection |
DEFICITS/SEVERITY | Articulation: Respiratory insufficiency for speech/ mild to severe |
PARTICIPANT LIMITATIONS | Time-to-fatigue with increased severity |
LANGUAGE/CULTURE | Follow instructions in native language |
SPOKEN RESPONSE PROCESSING/METRICS | Digitized acoustic signal for extraction of primarily temporal measures (speaking rate, pause rate, phonation duration) |
Matching Elicitation Tasks with ML Use-Cases
Fit-for-Purpose Design
Example 3:
Existing database of speech samples collected via telephone from patients with Parkinson’s disease. The goal of the study for which these data were collected was for voice detection of dysarthria progression over time.
GOAL | Within-subject, longitudinal changes in dysarthria severity |
DEFICITS/SEVERITY | Articulation: articulatory precision from mild to severe deficits |
PARTICIPANT LIMITATIONS | Follow instructions to produce “pa ta ka” as quickly and crisply as possible |
LANGUAGE/CULTURE | Follow instructions in native language |
SPOKEN RESPONSE PROCESSING/METRICS | Digitized acoustic signal for extraction of syllable count and rate, and goodness of articulation |
Would this dataset be valuable for detecting progression from PD to PD+Lewy body dementia?
Remarks and Questions
Speech Data Acquisition
Recording Data is easy, isn’t it?
If you acquire a professional recording booth (with technician)
Image Credit: https://whisperroom.com
Recording Pipeline?
A simple one
Recording Quality
Big problem in later analyses
Important Microphone Specs
Technical Aspects
Microphone sensitivity
microphone's ability to convert sound pressure to an electric voltage
Image Credit: https://www.dpamicrophones.com
Video
Microphone sensitivity
microphone's ability to convert sound pressure to an electric voltage
Image Credit: https://www.dpamicrophones.com
Video
Lower Red Curve:
Upper Blue Curve:
Microphone sensitivity
microphone's ability to convert sound pressure to an electric voltage
Image Credit: https://www.dpamicrophones.com
Video
How much voltage you get for applying one pascal sound pressure
Assumed to produce �1V / 1Pa
Example: Sennheiser MKE600
The lower the negative value the better the sensitivity
Microphone sensitivity
Select the proper microphone
Image Credit: https://www.dpamicrophones.com
Video
Signal-to-Noise Ratio
Determines how “clean” the output signal is
Video
Signal-to-Noise Ratio
Real Examples
Sennheiser MKE 600 �Shutgun Microphone
Comica �Traxshot Microphone
aka Field Microphone
Frequency Response
How are different frequencies perceived?
Reproduces sound with little or no coloration/variation from original sound
Flat response microphone (Shure SM81)
Decreased sensitivity for low frequencies reduce pick up of room noise or vibration
As well as counteracts that build up of bass that can occur when the recording distance is low
Increased sensitivity in upper mid-range add clarity to vocals
Shaped response microphone (Shure KSM42)
Frequency Response
Low-Frequency Roll-off Control
adjusted
Low frequency roll-off inactive
Low frequency roll-off active
Frequency Response
How are different frequencies perceived from different directions?
Image Credit: https://www.dpamicrophones.com
Directivity
“Control” which directions are recorded
Video
Directivity
“Control” which directions are recorded
Image Credit: https://www.dpamicrophones.com
Types of directional microphones
Pop Filter
Reduce the distortion from plosives
Image Credit: https://www.audiomentor.com
Shotgun vs. Condenser vs. Lavalier
Specifications and Specialities
Image Credit: https://www.dpamicrophones.com
Sennheiser �MKE 600 �
Shure �KSM42
Rode Lavalier GO
Impact of technical measurement factors
Why microphone specs matter
Paper
Impact of technical measurement factors
Why microphone specs matter
Participant characteristics (n = 42)
Standardized differences in timing features
Sex | female | 23 |
male | 19 | |
Age (years) | median | 28 |
IQR | 23-32 | |
English L1 | Yes | 29 |
No | 13 | |
Height (m) | median | 1.70 |
IQR | 1.63-1.80 |
Impact of technical measurement factors
Why microphone specs matter
Challenge: Impact of technical measurement factors
Challenge: Impact of technical measurement factors
Speech analytical pipeline is susceptible to many types of variability
In-the-wild longitudinal monitoring introduces multiple technical, acoustic and human factors into the recording process
Can the associated variability in the recorded speech signal be erroneously interpreted as related to a change in health state?
Recorded the speech of 42 healthy volunteers recorded consecutively in rooms with low and high reverberation
Simultaneous recordings on one budget and two higher-end smartphones and a condenser microphone
Challenge: Impact of technical measurement factors
Participant characteristics (n = 42)
Standardized differences in timing features
Errors bars are 95% confidence intervals
Negative differences represent lower feature values in the presence of higher reverberation
Sex | female | 23 |
male | 19 | |
Age (years) | median | 28 |
IQR | 23-32 | |
English L1 | Yes | 29 |
No | 13 | |
Height (m) | median | 1.70 |
IQR | 1.63-1.80 |
Challenge: Impact of technical measurement factors
Speech analytical pipeline is susceptible to many types of variability
Standardized differences in acoustic features
Positive differences represent higher values in the presence of higher reverberation.
Error bars represent 95% confidence intervals.
Voice quality features are seemingly the most affected by reverberation
Amplify concerns about the robustness and validity of these features for health assessments
Improved Recording Pipleline?
High-quality shot-gut mics used
Improved Recording Pipleline?
High-quality headsets used
Recording can still include errors
Why adjustment of recording level (GAIN) is important
Watch the sound level meter!
Clipping!
The Effects of noise on Acoustic Parameters
Why noise should be avoided
The Effects of noise on Acoustic Parameters
Why noise should be avoided
The Effects of noise on Acoustic Parameters
Why noise should be avoided
Recording can still include errors
Why sampling rate matters
Sampling is the reduction of a continuous-time signal to a discrete-time signal
Image Credit: https://www.youtube.com/watch?v=Z0EMObqS90U
Range of human hearing:
20 Hz- 20,000 Hz
Recording can still include errors
Why sampling rate matters
Speech Sampling, oriented towards its application:
Human Speech intelligibility is between 300 and 3,400Hz
Video
Recording can still include errors
Why recording format and bitrate matters
Image Credit: https://filesconverter.com/best-audio-format
Recording can still include errors
Why recording format and bitrate matters
I. Siegert et al., (2016). Measuring the impact of audio compression on the spectral quality of speech data. ESSV, 2016
Compression and acoustic parameters
Why low bitrates influence acoustic measurement
Pitch dynamism quotient�
(PDQ)⇒
Paper
Compression and emotion recognition
Why recording format and bitrate matters
Speech Emotion Recognition Experiments�
Paper
Recording Data:
Towards mobile health application
Speech analytical pipeline is susceptible to many types of variability
Remarks and Questions
Link to survey on speech data acquisition
From
Speech Features to
Speech Measures
Reproducibility Challenge
As a community we need to build-up reproducible and reliable clinical evidence
Challenge: Speech analysis has to be reliable at a level that is acceptable for clinical decision making
Issues
Reproducibility Challenge
Harmonisation across studies
Challenges
Reproducibility Challenge
As a community we need to build-up reproducible and reliable clinical evidence
Reproducibility and harmonisation
Starting point: Add additional features and improve reliability through hypothesis driven research
Reproducibility Challenge
Minimal speech-health feature set
Inclusion criteria:
Minimal Speech Health Data Set
Timing/Fluency
Speed: speed with which speech is performed
Breakdown: the pauses and silences that disrupt the flow of speech
Repair fluency: hesitations, repetitions, and reformulations
Minimal Speech Health Data Set
Timing/Fluency
Phonation ratio: phonation time divided by duration
Speech rate: Number of linguistic units divided by duration
Articulation Rate: Number of linguistic units divided by phonation time
Pause Rate: Number of Pause divided by duration
Mean Pause Duration: Mean duration of all pauses longer than 300ms
Minimal Speech Health Data Set
Timing/Fluency
Extraction:
Our extraction utilises the code originally developed for L2 fluency, with slight modifications regarding efficiency and error catching.
This code uses an intensity threshold to automatically identify pause boundaries,
Syllables identified using a combination of intensity thresholds and voicing information.
Notes:
Other methods, such as forced alignment, can be used to extract these features, this would mean they are reliant on additional third-party software.
Repair fluency properties difficult to extract without ASR
Minimal Speech Health Data Set
Speech production subsystems and representative acoustic features
Minimal Speech Health Data Set
Respiration
Process of moving air in and out of the lungs
Stability: Speech production requires stable air pressure throughout our vocal tract
Power: When shouting or speaker longer sentences, we need a bigger inhalation
Minimal Speech Health Data Set
Respiration
Features:
Extraction:
Utilises Praat Sound: To Intensity… function:
Intensity calculated as the root-mean-square amplitude of the provided audio signal:
Note
Other respiratory measures (E.g., the detection of respiratory events within a speech signal) require third-party software in their extraction
Minimal Speech Health Data Set
Phonation
Production of sound at the level of the vocal folds
Voiced Speech: produced when vocal folds are vibrating
Unvoiced Speech: vocal folds are lax and open
Related perceptual properties include: pitch and vocal quality (E.g., hoarseness, breathiness, roughness)
Minimal Speech Health Data Set
Phonation
Features:
Minimal Speech Health Data Set
Phonation
Fundamental frequency (F0) vs. Pitch?
Features:
Extraction
Minimal Speech Health Data Set
Phonation
Harmonics-to-noise ratio (HNR)
Extraction
Minimal Speech Health Data Set
Phonation
Long-time average spectrum (LTAS)
Spectral Slope
Spectral Tilt
Minimal Speech Health Data Set
Phonation
Cepstral Peak Prominence
Extraction:
Minimal Speech Health Data Set
Phonation
Jitter and Shimmer
Note that we have not included two of the better-known and widely used phonation measures in jitter and shimmer.
These features are open to errors due to differing sound pressure levels and phonetic content between and within individuals
Should be considered unreliable for analysis of voice pathology
Also subject to measurement errors due to recording conditions
Minimal Speech Health Data Set
Articulatory
Producing unique speech sounds by changing the shape of the vocal tract
Different muscle formations changes the point of constriction in the vocal tract
Minimal Speech Health Data Set
Articulatory
Formants
Formants are the resonances of the vocal tract
The defines consonant and vowel perception
Extraction
Praat: Sound: To Formant (burg)... function which utilises linear predictive coding
Minimal Speech Health Data Set
Articulatory
Spectral features characterise the speech spectrum.
Typical spectral features are high dimensional representations that capture all the information contained in speech, which could include confounds
Minimal Speech Health Data Set
Articulatory
Use the four spectral moments to characterise the spectrum.
Extraction:
Extracted from a Praat Spectrogram object using the relevant function above
Minimal Speech Health Data Set
Code being tested before initial release
Code is implemented in Parselmouth so it can run in Python
Please get in touch if you would like early access
Not linked to theoretical construct
Reliability and generalizability not individually characterized
Different implementations by different groups
Designed to measure specific construct
Psychometric properties characterized; normative data
Standardized for comparability
Selected to improve model accuracy
Clinical meaningfulness?
Liss, Julie, and Visar Berisha. "Operationalizing Clinical Speech Analytics: Moving From Features to Measures for Real-World Clinical Impact." Journal of Speech, Language, and Hearing Research (2024): 1-7.
Allen, Mary J., and Wendy M. Yen. Introduction to measurement theory. Waveland Press, 2001.
Features
Measures
Moving from speech features to speech measures
5.
Stegmann, G.M., Hahn, S., Liss, J., Shefner, J., Rutkove, S., Shelton, K., Duncan, C.J. and Berisha, V., 2020. Early detection and tracking of bulbar changes in ALS via frequent and remote speech analysis. Nature npj Digital Medicine, 3(1), pp.1-
ALS@Home and ALS - FTD studies:
Amyotrophic Lateral Sclerosis
Witt, Silke M., and Steve J. Young. "Phone-level pronunciation scoring and assessment for interactive language learning." Speech communication 30.2-3 (2000): 95-108.
Mathad, Vikram C., et al. "The Impact of Forced-Alignment Errors on Automatic Pronunciation Evaluation." Interspeech. 2021.
Measuring the right constructs
Witt, Silke M., and Steve J. Young. "Phone-level pronunciation scoring and assessment for interactive language learning." Speech communication 30.2-3 (2000): 95-108.
Mathad, Vikram C., et al. "The Impact of Forced-Alignment Errors on Automatic Pronunciation Evaluation." Interspeech. 2021.
How it works?
Acoustic model for likelihood estimation
Acoustic model for alignment
Speech + Transcript
Likelihood ratio computation
Operationalizing the construct
Mathad, Vikram C., et al. "The Impact of Forced-Alignment Errors on Automatic Pronunciation Evaluation." Interspeech. 2021.
Stegmann, G., Charles, S., Liss, J., Shefner, J., Rutkove, S., & Berisha, V. (2023). A speech-based prognostic model for dysarthria progression in ALS. Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration, 1-6.
N=110 (40F); tracked longitudinally for up to 1 year
Analysis | Result |
Correlation with clinical ratings of articulatory precision | r = .90 |
Test-retest reliability | ICC = .97 |
Correlation with ALSFRS-R speech | r = .82 |
Separation of severe and nonsevere speech | ROC AUC = .97 |
Correlation between articulatory precision and ALSFRS-R speech subscale longitudinal slopes | r = .37 |
Correlation between articulatory precision and ALSFRS-R bulbar subscale longitudinal slopes | r = .41 |
Longitudinal change | Slope = -.004 |
How good does the aligner have to be?
TLDR: MFA forced aligner is fine
Validating the algorithm
Establishing clinical meaningfulness: why is the objective measure meaningful to patients?
ALS impacts the speech production mechanism. This has a negative impact on a patient’s communicative participation and quality of life (Börjesson, 2021; Borrie et al, 2022). This is because others have a difficult time understanding them and there is increased listener effort to understand them (Stipancic et al., 2018). The standard way to measure the impact on speech is the ALSFRS-R speech subscale, a 5-point scale that is part of the clinical standard. We devised a new measure for object assessment of precision of articulation. This tool has high correlation with perceptual measures of articulatory precision, establishing construct validity. It has high test-retest reliability, making it appropriate for longitudinal monitoring (Stegmann et al., 2020; Stegmann et al., 2023). It is more sensitive than the ALSFRS-R speech subscale to longitudinal change in a patient’s speech (Stegmann et al., 2024).
Stegmann, G., Krantsevich, C., Liss, J., Charles, S., Bartlett, M., Shefner, J., Rutkove, S., Kawabata, K., Talkar, T. and Berisha, V., 2024. Automated speech analytics in ALS: higher sensitivity of digital articulatory precision over the ALSFRS-R. Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration, pp.1-9.
Establishing clinical utility: improving clinical research
Reducing sample size requirements in clinical trials
Designing new trials specifically focused on speech outcomes
Stegmann, G., Charles, S., Liss, J., Shefner, J., Rutkove, S., & Berisha, V. (2023). A speech-based prognostic model for dysarthria progression in ALS. Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration, 1-6.
Establishing clinical utility: A prognostic model for ALS
FDA and Clinical Trial Outcome
March 2023: FDA Breakthrough Device Designation
“Medical devices receive breakthrough designation from the FDA when they provide for more effective treatment or diagnosis of life-threatening or irreversibly debilitating diseases or conditions. The goal of the program is to provide patients and healthcare providers with timely access to groundbreaking technologies by expediting the approval process.”
February 2023: “The impact of pridopidine on speech measures was notable, likely due to its S1R mechanism of action. Speech is a highly clinically relevant endpoint in ALS studies, and more than 80 percent of ALS patients become speech impaired, which significantly impacts their quality of life.”
From features to measures
With a subtitle if you need it
When using the template, try to not move the headlines up and down. It will add a level or professional design if the top margin of your headline is not moving around between slide transitions.
Remarks and Questions
Afternoon Session
Recap of Morning Session
Recap of Morning Session
2. Speech data acquisition�� → Considerations for microphone specs � - Frequency response, sensitivity, SNR, directionality, microphone type, etc.�� → Differences in measurement across microphone brands�� → Noise / post-processing technique that impact speech features compression, etc.)�� → Variability in data acquisition pipeline�
Recap of Morning Session
3. Speech features and measures�� → Limitations in existing speech features - reproducibility, clinical evidence, � implementation, etc.�� → Minimal Speech Health Data Set - Timing/frequency, speech production � subsystems, respiration, phonation, articulatory
� → Moving from speech features to speech measures - finding right construct,� operationalize the construct, validate the algorithm, establish clinical meaning� � → FDA breakthrough - An ALS example��
Learning Generalizable & Interpretable
�Speech Representations &�Clinical Models
From Speech Data to Clinical Label: Commonly-seen Pipeline
A Revised Pipeline from Speech Data to Clinical Label
Design of Speech Measures
Traditional ways to obtain speech measures��→ Perceptual rating (use a scale to rate coordination, speed, voice quality, etc.)��→ Manual transcription (speech intelligibility, speech rate, etc.)��→ Signal processing methods (F0, jitter, shimmer, cepstral peak prominence, etc.)
Design of Speech Measures
We have speech data and some type of labels…��How to design an ML algorithm to derive speech measure?��- Utilize domain knowledge of speech production, semantic similarity in language.��- Predict clinically-validated measures
Traditional ways to obtain speech measures:�1) Perceptual rating�2) Manual annotation�3) Signal processing
Design New Speech Measures
Using domain knowledge of speech production
Derive measures using probability outputs from acoustic models
Design New Speech Measures
Using domain knowledge of speech production
Design New Speech Measures
Example 1: Objective Hypernasality Measure (OHM) - Algorithm
Task: �Assess hypernasality in speakers with cleft palate (Mathad et al. 2021)
Training data: Healthy speech
Design New Speech Measures
Example 1: Objective Hypernasality Measure (OHM) - Algorithm
Visualization of OHM:
-> Sentence: “Buy baby a bib”�(No nasal consonant & vowels)�
-> Nasal consonant & vowel detected in hypernasal speech (shown in red).�
Can we use the OHM now?�NO, we need validation.
A Measure’s Reliability and Validity
Learnable Speech Measures
Example 1: Objective Hypernasality Measure (OHM) - Validation
Design New Speech Measures
Example 2: Objective Articulation Measure (OAM) - Algorithm
Task: �Evaluation articulation in speakers with dysarthria, cleft lips/palate (Mathad et al. 2023)
Design New Speech Measures
Example 2: Objective Articulation Measure (OAM) - Algorithm
Visualization (CNN Salience Map): ��Trained model is focused on consonants and vowel onsets when classifying the consonants��→ Confirm the model is properly utilizing the inputs.
Design New Speech Measures
Example 2: Objective Articulation Measure (OAM) - Validation
Validation:� �- Compare OAM with perceptual rating of clinical speech
- Compare OAM with existing measures that also measure articulation (GOP)
Design New Speech Measures
NLP: Cosine similarity as a measure
�Speech can be transcribed into text ��→ Natural language processing (NLP) to derive speech measures
�
Design New Speech Measures
Example 3: Using cosine similarity to assess thought disorder - Algorithm
Task: Assess thought disorder (Bilgrami et al. 2022)
Approach: �1) Use pre-trained BERT model to extract sentence-level embeddings �2) Find the minimum semantic coherence based on cosine similarity between sentences.
Design New Speech Measures
Example 3: Using cosine similarity to assess thought disorder - Validation
Validation:��- Construct validity: Compare the minimum semantic coherence with 8 clinical measures �
Learnable Speech Measures
NLP: Cosine similarity as a measure
Cosine similarity between text embeddings → Correlate to clinical measures
�1. Word2Vec → Semantic similarity between flow of ideas → Thought disorder
2. Sentence embedding between adjacent utterance pairs → Coherence measure → Schizophrenia
Predict Validated Measures
Learnable speech measures
Predict Validated Measures
Example 1: Information bottleneck - Algorithm & Validation
W1: Measure prediction;
W2: Predict diagnostic labels.
y: Severity of dysarthric speech
Task: Assess severity of dysarthric speech
Predict Validated Measures
Example 1: Information bottleneck - Algorithm & Validation
Validation: Compare predicted measures with ground truth perceptual labels
Predict Validated Measures
Example 2: Information bottleneck - Algorithm
Binary Output → Presence of dysarthria
Obtaining these measures �DO NOT require manual annotation
Predict Validated Measures
Example 2: Information bottleneck - Validation
Validation:�- Compare selected measures with ground truth values
- Evaluate if the selected measures are truly useful for dysarthria detection
Large absolute SHAP value �→ More contribution to decision making
Clinical ML for Predicting Clinical Variables
An ideal clinical ML model: ��1) Interpretable inputs �
2) Interpretable & generalizable model �
3) Produce reliable and clinically meaningful outputs
How to design a clinical ML model based on speech measures?
Clinical ML for Predicting Clinical Variables
An ideal pipeline of clinical ML model: ��1) Interpretable inputs → Adopt interpretable speech measures �
2) Interpretable & generalizable model → Simple ML model, low-dimensional inputs�
3) Produce reliable and clinically meaningful outputs → Perform clinical validation
Clinical ML for Predicting Clinical Variables
Example 1: Assess verbal memory in mental illnesss (Holmlund et al., 2020)
Clinical ML for Predicting Clinical Variables
Example 1: Assess verbal memory in mental illnesss (Holmlund et al., 2020) �- Clinical validation
Clinical ML for Predicting Clinical Variables
Example 2: Longitudinal Tracking of Forced Vital Capacity Using Speech Acoustics (Stegmann et al. , 2021)
Clinical ML for Predicting Clinical Variables
Example 2: Longitudinal Tracking of Forced Vital Capacity Using Speech Acoustics (Stegmann et al. , 2021) - Clinical Validation
Clinical ML for Predicting Clinical Variables
Example 2: Longitudinal Tracking of Forced Vital Capacity Using Speech Acoustics (Stegmann et al. , 2021) - Clinical Validation
Clinical ML for Predicting Clinical Variables
Example 2: Longitudinal Tracking of Forced Vital Capacity Using Speech Acoustics (Stegmann et al. , 2021) - Clinical Validation
Learning Generalizable & Interpretable Speech Representations & Clinical models
Suggestions
Thoroughly understand data and labels. �- Ground truth measure → Analytical validity �- Repeated tests → Test-retest reliability�- Same subject, measured on different words/sentence → Internal reliability�- Other validated measures that measure the similar construct → Construct validity�- Measures made by multiple raters → Inter-rater reliability�etc. �
Learning Generalizable & Interpretable Speech Representations
Suggestions
Learning Generalizable & Interpretable Speech Representations
Suggestions
- Lots of measures can be designed without disordered data
- Allow more clinical constructs to be utilized in clinical speech analytics → A more comprehensive diagnosing process
- More feature options in clinical model development → Improve analytic performance
- etc.
Design of Clinical ML for Diagnostic Decision
Example: (Sauder et al. 2017)
1. Interpretable inputs: ��→ Smoothed cepstral peak prominence (CPPS): describe breathiness of voice
�2. Model: Logistic regression �
Disorder status ~ coeff1 + coeff2 * Gender + coefff3 * CPP
3. Clinical output: Presence of voice disorder
�4. Interpretation: Regression coefficient �
Expectation from Regulatory Bodies
A perspective from FDA
Endpoint: what happens to people in the clinical trial.��Clinical endpoint: directly measure what matters most to people, �e.g. occurrence of disease, improvement in symptoms, function better.�→ Occurence of disease�→ Severity score, depression score, mental state score �→ Change of score reflects improvement/worsening of disease�
Here is a text slide
With a subtitle if you need it
When using the template, try to not move the headlines up and down. It will add a level or professional design if the top margin of your headline is not moving around between slide transitions.
Remarks and Questions
Ethical, Privacy and Security Considerations
Overview
Responsible Clinical
Speech Analytics
Participant Factors
Privacy�Attacks�Misuse
Bias�Discrimination Fairness
Access Equity
Sustainability
Robust Experimental Designs
Robust
Experimental Designs
Generalizable AI is Ethical AI
Ghasemzadeh et al. (2024)
The experimental design strategy used in machine learning studies influences:
The experimental design strategy is particularly important in clinical datasets.
Data Splitting Experimental Designs
Ghasemzadeh et al. (2024)
C) Train-validation-test D) nested k-fold cross validation (k = 3)
Impact of Experimental Design on Clinical Data
Ghasemzadeh et al. (2024)
Voice Dataset
153 females with phonotraumatic vocal hyperfunction and 136 female controls
6 voice measures: F0, cepstral peak prominence, sound pressure level, H1-H2, voicing duration, and phonatory rest duration.
8 distributional characteristics from data on each day: mean, median, standard deviation, interquartile range, 5th percentile, 95th percentile, skewness, and kurtosis
Each participant provided one 48-dimension datapoint to the study.
Impact of Experimental Design on Clinical Data
Ghasemzadeh et al. (2024)
Impact of Experimental Design on Clinical Data
Ghasemzadeh et al. (2024)
Statistical Comparison with Simulated Data
Ghasemzadeh et al. (2024)
Statistical Power of Nested k-fold CV
Ghasemzadeh et al. (2024)
Statistical Power of Nested k-fold CV
Ghasemzadeh et al. (2024)
Statistical Confidence
Ghasemzadeh et al. (2024)
Nested k-fold cross validation
Ghasemzadeh et al. (2024)
Nested k-fold cross validation
Bias in Clinical Speech Analytics
The Presence of Bias in Clinical Datasets
Berisha and Liss (2024)
Published clinical AI models are demographically biased, with 71% of training data coming from the states of California, Massachusetts, and New York (Kaushal et al (2020; JAMA).
This is problematic, and emphasizes the importance of global equity, because health-related changes in speech can be language specific (Garcia et al., 2023; Brain).
Within-language dialect differences are also known to impact error rates in speech machine learning applications (Wassink et al., 2022).
Understanding of sociolinguistic variation is biased towards western, Indo European languages (Adli & Guy, 2022).
The Risks of Bias in Clinical Speech Analytics
Chen et al. (2023)
Bias in training datasets can lead to skewed AI model outputs, affecting clinical decision-making.
Smaller clinical datasets often fail to represent minority groups adequately, amplifying potential biases in model performance.
Disparities in model performance can result in unequal healthcare outcomes, particularly in underrepresented populations.
Enhancing data diversity and inclusion in the data acquisition phase is critical to developing fair and effective clinical AI systems.
Biases in Data: Research Constructs
Mehrabi et al. (2021)
Measurement bias: face validity of features
Omitted variable bias: an important variable is left out
Representation bias: non-representative sampling omits entire subgroups or
outgroups
Sampling bias: non-uniform sampling of subgroups
Aggregation bias: group level averaging that obscures differences in subgroups
User Biases: Social Constructs
Mehrabi et al. (2021)
Historical bias: an already existing bias perpetuates through to the data
Population bias: dataset shift from development to deployment due to sampling
Temporal bias: dataset shift from development to deployment due to social evolution
Self-selection bias: tool works better for people motivated/have the resources to appear in research
Content production bias: linguistic diversity and diversity of speech register between development speech and deployment speech impacts the data
Algorithmic Biases
Mehrabi et al. (2021)
Algorithmic bias: bias is a result of functions, regularizations, or statistically biased estimators.
Evaluation bias: inappropriate or disproportionate benchmarks for evaluation
Case Example: Evaluation Bias
Benway et al. (2024)
Case Example: Evaluation Bias
Benway et al. (2022)
Case Example: Uncovering Population Bias
Benway and Preston (2024)
Prospective clinical validation can help detect biases and distributional shifts.
Clinical Therapeutics Require Clinical Trials
Healthcare and allied health practitioners cannot ethically use AI therapies without efficacy evidence.
Is there evidence that a given AI therapy is valid and reliable for clinical use?
Is there evidence that a given AI therapy is poised to work without bias in real clinical populations with speakers it has never encountered before?
Clinical Therapeutics Require Clinical Trials
Benway and Preston, (2024)
Discrimination in Clinical Speech Analytics
Risks of Discrimination
Obermeyer et al. (2019)
Clinical speech analytic systems may inadvertently perpetuate or amplify existing biases, leading to discriminatory practices in healthcare.
Discriminatory biases can affect diagnosis accuracy, treatment recommendations, and patient trust in healthcare systems and clinical AI.
Direct Discrimination
Obermeyer et al. (2019)
A non-favorable outcome for an individual is because of a protected attribute
Legally-protected attributes will vary by country/jurisdiction but generally might include:
Ethnicity/Race
Religion
Age
Sex/Gender/Sexual Orientation
Disability
Indirect Discrimination
Systemic Discrimination
Obermeyer et al. (2019)
Individual cases are handled according to non-protected attributes that serve as proxy factors for protected attributes
Discrimination against an individual or subgroup is perpetuated by policies and customs.
Obermeyer et al. (2019)
Fairness in Clinical Speech Analytics
Risks of Unfair Clinical Speech Analytics
Binns (2020)
Unfairness in clinical speech analytics refers to systems that offer unequal benefits or burdens to different groups without justified reasons.
Systems perceived as unfair can violate anti-discrimination laws and lead to legal consequences for healthcare providers and clinical tech developers.
Fairness Awareness or Unawareness
Mehrabi et al. (2021)
Fairness through awareness:
Individuals who are similar on a particular task should receive similar predictions across different subgroupings of protected attributes
Fairness through unawareness:
An algorithm is fair as long as protected information is not represented in the decision-making process
However, unawareness raises the opportunity for indirect discrimination that cannot be easily monitored or detected.
Quantifying Fairness
Mehrabi et al. (2021)
Equalized Odds
True positives and false positives rates should be equal across different subgroupings of protected attributes
Equal Opportunity
True positive rates should be equal across different subgroupings of protected attributes
Demographic/Statistical Parity
The likelihood of a positive outcome should be the same across different subgroupings of protected attributes
Quantifying Fairness
Mehrabi et al. (2021)
Treatment Equality
False positives:false negative ratio should be equal across different subgroupings of protected attributes
Test Fairness
For any predicted probability score S, there is an equal probability of belonging to the protected class across different subgroupings of protected attributes
Counterfactual Fairness
A decision is fair if the decision is the same regardless of which protected attribute subgroup the individual belonged to
Case Example: Fairness
Benway et al. (2024a)
A linear mixed model fit on the combined test dataset indicated that neither the fixed effects of age nor sex nor the age-sex interaction significantly influenced classifier performance.
Privacy, Attacks and Misuse
Privacy
Berisha & Liss (2024)
If speech can be used as a biomarker, and the recorded speech of individuals is public, then individuals are effectively leaving a publicly accessible, longitudinal trail of data that can be used to make claims about one’s personal health or emotional status.
Such claims could impact health privacy of private individuals or, for example, support the spread of disinformation in the case of political figures.
Privacy is therefore paramount.
Data Breaches
Liu et al. (2015)
Clinical speech analytic databases themselves could be compromised, potentially exposing individually identifiable protected health information.
Individual jurisdictions will have regulations related to data privacy and data breaches. For example:
Attacks on Algorithms
Finlayson et al. (2019)
Adversarial attacks pose a significant risk to clinical machine learning models by exploiting vulnerabilities that can alter outputs.
Potential attacks include data poisoning, model inversion, and adversarial examples, each posing unique threats to model reliability and data security.
Compromised models could lead to misdiagnoses, inappropriate treatment suggestions, patient harm, and erosion of trust in clinical speech analytics.
Ensuring robust cybersecurity measures for clinical machine learning models is vital due to the potential for attacks to compromise patient care.
Malicious Misuse
Price et al. (2019)
Analytic systems can be manipulated to generate false claims or diagnoses, especially when used without robust oversight or ethical guidelines.
Using personal data without consent to develop clinical analytics can violate privacy rights and lead to ethical and legal repercussions.
Misrepresentation of data can harm careers, influence treatment plans, and erode trust in healthcare systems.
Misuse through Premature Deployment
Berisha and Liss (2024); Yawer et al. (2023)
Insufficient validation of clinical speech analytics runs the risk of overoptimistic lab testing and clinically inaccurate output in production.
Example: the Cigna StressWaves Test
Models that are deployed prematurely may cause harm to patients through inappropriate treatment, wasted resources, increased anxiety, or false reassurance.
Equitable Access
Environmental Impact of Computation
Bender, Gebru, et al. (2021)
The training of large AI models has an environmental impact due to the energy resources required to support computation.
This environmental impact is most likely to impact climate-vulnerable communities.
Because climate-vulnerable communities are also associated with social factors like lower income and minority languages, the communities bearing the brunt of AI environmental impact are the least likely to benefit from AI advancements.
AI Dataset Microtask Economies
Casilli (2024), Rowe (2023), Munn (2024)
AI-ready datasets must be curated, labeled, and validated.
Who is labeling these data and how fairly are workers being paid?
There is scientific and journalistic evidence that low-wage workers are involved in these tasks.
These low-wage workers also come from the least likely communities to benefit from the technology they are helping to train.
Patient and Public Involvement
Community Based Participatory Research
Active involvement of patients and the public in clinical speech analytics development ensures the technologies address real-world needs and ethical considerations.
Community involvement in clinical speech analytics development can improve the relevance and acceptance of technologies, leading to better adoption of the technology and better trust in the technology.
AI 4 Social Good Framework
Bondi et al. (2021)
Are the communities being designed “for” rather than having a seat at the table?
Does the development of clinical speech analytics mirror the needs and preferences of specific communities with clinical speech needs?
Is this engagement reciprocal and part of the process from the beginning?
How was the community involved in the design, development, and deployment of clinical AI?
Guiding Principles for Participatory AI4SG
Bondi et al. (2021)
How are impacted communities identified and how are they represented in this process? Who represents historically marginalized groups?
How are plans laid out for maintaining and sustaining this work in the long-term, and how would the partnership be ended?
What kind of compensation are interested parties receiving for their time and input? Does that compensation respect them as partners in the process?
How are viewpoints from many different groups understood and incorporated?
What specific concerns are raised during deliberation with interested parties, and how are these addressed?
Determining Capability Sets
Bondi et al. (2021)
What capabilities to interested parties want to achieve?
What are the priorities for the most vulnerable in the community?
Will any of the priorities pay dividends in other important areas (occupational, social, etc)?
Do the values of the project match the values of the community?
Evaluating AI4SG
Bondi et al. (2021)
How does the clinical speech analytics system affect the capabilities prioritized by the community?
Are any capabilities negatively affected by the clinical speech analytic system?
What should the role of the AI researcher be?
Case Study (KCL)
Automating coding of Five Minute Speech Samples
E-Risk Study
Case Study (KCL)
Automating coding of Five Minute Speech Samples
Mothers’ speech samples
When the children were 10 years old, mothers were asked to speak about them to the researcher
“For the next 5 minutes, I would like you to describe [child] to me; what is [child] like?”
This speech was recorded
Case Study (KCL)
Automating coding of Five Minute Speech Samples
Coding mothers’ emotions towards child
Case Study (KCL)
Automating coding of Five Minute Speech Samples
Predicting later mental health issues
Case Study (KCL)
Automating coding of Five Minute Speech Samples
Conducted PPI sessions with parents and young carers
Sessions included the brief outline of the study
Had employed a near-future science fiction who wrote to dystopian stories based our research
These were also read to participants in our workshops
Case Study (KCL)
Automating coding of Five Minute Speech Samples
Parents Workshop
Data & privacy
Will the AI of the future be able to de-anonymise us? What happens to anonymity then?
Consent as a continuous process
Research directions may change over time
The assessment
Parent-blaming attitudes
How can you capture parenting across the lifespan?
Lack of expertise of, and consistency between, clinicians
Discrimination against neurodivergent parents
Key contextual issue: waitlists and lack of NHS support available, health inequities
Predicting mental health outcomes
‘Gaming the system’
Self-fulfilling prophecies
The AI output: reductionism of interpreting health statistics
Limitations of predicting health outcomes and of diagnosis
Issues with the method more than use of AI
Case Study (KCL)
Automating coding of Five Minute Speech Samples
Parents Workshop
Data & privacy
Will the AI of the future be able to de-anonymise us? What happens to anonymity then?
Consent as a continuous process
Research directions may change over time
The assessment
Parent-blaming attitudes
How can you capture parenting across the lifespan?
Lack of expertise of, and consistency between, clinicians
Discrimination against neurodivergent parents
Key contextual issue: waitlists and lack of NHS support available, health inequities
Predicting mental health outcomes
‘Gaming the system’
Self-fulfilling prophecies
The AI output: reductionism of interpreting health statistics
Limitations of predicting health outcomes and of diagnosis
Issues with the method more than use of AI
Case Study (KCL)
Automating coding of Five Minute Speech Samples
Young Carers Workshop
What if parents don’t tell the truth?
‘Gaming the system’
Lying
May not know their child
Not all cultures discuss mental health
Preference for doctors over AI
Can respond to emotion
More human
Doctors have specialised training
Can take context & tone into account
Risks of using the tool
Self-fulfilling prophecy
Parents treat children differently
Case Study (KCL)
Automating coding of Five Minute Speech Samples
Common Themes
Inability to account for other life events (positive or negative)
Dangers of ‘gaming the system’ and/or self-fulfilling prophecies
Concerns about money & access
Potential positive applications of AI:
Time-Ordered Actions Toward Responsible Clinical Speech Analytics
Establish Participatory Partnerships
Early in the process
Enlist community members ahead of the development of clinical speech analytics to inform the technological design and technological fidelity for:
This is most ernest when done during the proposal writing process.
Design Speech Tasks
Factors for consideration
Minimal Speech Health Data Set
Group | Feature | Description | unit |
Timing/Fluency | Phonation Ratio | Phonation time divided by duration | - |
Speaking Rate | Number of syllables divided by duration | syl/sec | |
Articulation Rate | Number of linguistic units divided by phonation time | syl/sec | |
Pause Rate | Number of Pause (> 300ms) divided by duration | pause/sec | |
Mean Pause Duration | Mean duration of all pauses (> 300ms) | sec | |
Respiratory | Intensity | Root-mean-square amplitude of the provided audio signal | dB |
Intensity | Ratio of max to min intensity | - | |
Phontation | Pitch | Rate of vibration of vocal folds | Hz |
Pitch Sigma | Standard deviation of pitch in semitones | Semitone | |
Harmonic to Noise Ratio | Degree of acoustic periodicity | dB | |
Cepstral Peak Prominence | Stability of vocal fold vibration | dB | |
Spectral Slope | Ratio of energy in a spectra between 10-1000Hz over 1000-4000Hz | - | |
Spectral Tilt | Linear slope of energy distribution between 100-5000Hz | - | |
Articulatory | Formant Frequency (1st and 2nd) | Centre frequency of vocal tract resonant peaks | Hz |
| Formant Bandwidths (1st and 2nd) | Bandwidth of vocal tract resonant peak | Hz |
| Spectral Gravity | Spectral centroid | Hz |
| Spectral Standard Deviation | Spread of frequencies around the centroid | Hz |
| Spectral Skewness | Symmetry of frequencies around the centroid | Hz |
| Spectral Kurtosis | Flatness of the spectrum the centroid | Hz |
Select Speech Measures
Factors for consideration
Obtain Informed Consent/Assent
Beauchamp & Childress (2019)
Informed consent is a cornerstone of ethical research and clinical practice, ensuring that participants are fully aware of the nature and risks of the data collection and analysis process.
Obtaining consent is not only an ethical obligation but also a legal requirement under frameworks like GDPR in Europe and HIPAA in the U.S., protecting patient rights and data privacy.
Effective consent processes involve clear communication about how speech data will be used, stored, and shared, ensuring participants' understanding and agreement.
Acquire Speech Data
Recording Device and Environment Bias
whisperroom.com Ingo Siegert, Uni Magdeburg Univ. Maryland Medical Center
Transparently Report Dataset and Biases
Datasheets for Datasets Model Cards for Model Reporting Replicability Cards
Gebru et al. (2018)
Mitchell et al. (2018)
Kapoor & Narayanan (2023)
Develop Generalizable Models
Factors for consideration
Demonstrate Interpretability of Model/Measures
Use Robust Experimental Designs
Ghasemzadeh et al. (2024)
Increase the chances that speech measures and hyperparameters are optimized.
Reduce the chances that reported model performance will change with a different dataset split.
Link to paper ->
Justify Model Improvements
Benway et al. (2024a)
Are model improvements different from chance?
Identifying this helps streamline the development of parsimonious models that do not waste computational or environmental resources.
Search for evidence of biases
What biases are present?
Techniques like hypothesis testing and linear mixed modeling can help identify biases in algorithms before clinical deployment.
Recall previous Case Example from �Benway et al. (2024a)
Clinical Therapeutics Require Clinical Trials
Prospective clinical trials for therapeutic technologies will help identify sources of bias.
Multiple clinical trials will happen in sequence, each focusing on a unique factor of deployment.
Post-deployment Monitoring
Reanalysis of Benway et al. (2024b)
Post-deployment monitoring and continuous model updating are essential to maintain fairness over time, addressing new data or changing population dynamics.
Lab Data Prospectively Collected Data
Balanced Accuracy
Count
Measuring AI Impact
Stahl et al. (2023)
A repository of AI Impact Assessments associated with the systematic review of Stahl and colleagues is available at this QR code →
https://www.zotero.org/groups/4042832/ai_impact_assessments
Remarks and Questions