1 of 219

Responsible Development and Translation of Clinical Speech Analytics

2 of 219

Instruction for preparing the tutorial slides

With a subtitle if you need it

  • Use Slide No.3 as the template. �
  • The current agenda is based on last year’s one, so it’s not finalized.�(You may revise the sub-topic title that better describes your presentations.)�
  • We will unify the slides (format, alignment, etc.) as we finish our parts.�
  • If you have any technical questions about this Google slide, �please reach Herman at siioing@asu.edu.

3 of 219

Here is a text slide

With a subtitle if you need it

When using the template, try to not move the headlines up and down. It will add a level or professional design if the top margin of your headline is not moving around between slide transitions.

4 of 219

Presenters

5 of 219

Agenda

Morning Session

9:00 – 9:10 a.m.� Introduction�� 9:10 – 9:50 a.m.� Design of speech elicitation tasks

●Saliency of speech across different conditions

●Categorization of elicitation tasks

●Matching elicitation tasks with clinical conditions

9:50 – 10:30 a.m.� Speech data acquisition

●Influence of hardware devices

●Sources of noise and how it impacts speech features

●Guidelines to control parameters in data acquisition

●Hardware validation framework

6 of 219

Agenda

Morning Session

11:00 – 12:00 p.m.�Methodological shift from speech features to speech measures

●Methodological shift from features to measures

●Existing speech measures

●Validation framework of speech measures

●Case study: FDA breakthrough designation

7 of 219

Agenda

Afternoon Session

2:00 – 3:15 p.m.� Learning generalizable and interpretable� speech representations and clinical ML models

● Development speech measures with ML

● Analytical validation of speech measures

● Development clinical ML models using speech measures

● Clinical validation of developed models

● Recommendations from a dataset perspective to support development of speech measures and models

8 of 219

Agenda

Afternoon Session

3:30 – 4:30 p.m.� Ethical, privacy and security considerations

●Data bias, data privacy, model security, improper ML model use

●Directions toward responsible development of clinical speech analytics

●Participant acceptability, motivation, convenience, accessibility, and usability.

4:30 – 5:00 p.m.� Discussion

9 of 219

Introduction

10 of 219

Speech

Disease

Neurological disorders, mental health disorders, speech-motor disorders, voice disorders

Speech acoustics and natural language

Activities

Impaired abilities to communicate with others (e.g. reduced intelligibility, impaired social skills)

Participation

Unable to work, spend time with friends, etc.

The promise of speech-based biomarkers

11 of 219

The gap between promise and reality

  • Machine learning (mostly supervised learning) has been proposed as a way to distill the high-dimensional speech signal into clinically meaningful outputs
  • There are many academic papers reporting high accuracy

but there are few (if any?) clinical speech models that have been deployed

Why is this?

What can we do about it?

12 of 219

The current approach to speech-based biomarkers

Extract standard speech features using existing tools

Error

Rate

Predict a clinical variable of interest

OpenSMILE

Feature Extraction

Extract features using existing tools

wav2vec

NLP features

Mel spectra

Healthy vs. MCI

Speech Database

If the resulting accuracy is “good”

PUBLISH the final model

Else

MOVE ON TO NEXT PROJECT

13 of 219

13

1Stegmann, Hahn, Liss, Shefner, Rutkove, Kawabata, Bhandari, Shelton, Duncan, Berisha. The Repeatability of Commonly Used Speech and Language Features for Clinical Applications. Digital Biomarkers. Jan, 2021.

2Alhanai et al. "Spoken language biomarkers for detecting cognitive impairment." 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU). IEEE, 2017.

3Berisha et al “Digital medicine and the curse of dimensionality.” Nature Digital Medicine, October 2021.

4Berisha et al, “Are reported accuracies in the clinical speech machine learning literature overoptimistic”, to appear at Interspeech 22.

  • Problem 1: Existing speech features aren’t repeatable1
  • Problem 2: The features do not necessarily capture clinically-relevant parameters2
  • Problem 3: The features are not clinically interpretable2

  • All this leads to models that don’t generalize3, 4

Three problems with commonly-used features

14 of 219

Berisha et al. "Digital medicine and the curse of dimensionality." Nature npj Digital Medicine 4.1 (2021): 1-8.

Berisha et al, “Are reported accuracies in the clinical speech machine learning literature overoptimistic”, Interspeech 22.

This approach leads to overoptimism in the literature

4% decrease in reported accuracy per unit increase in log(SampleSize)

Sample Size

Accuracy

15 of 219

Building clinical speech analytics isn’t just an ML problem

How to design speech elicitation tasks?

How to collect the data in a robust way?

How to design interpretable representations for clinical speech models and how to validate them?

How do we train interpretable clinical speech ML models?

What are important ethical, safety, and security considerations?

16 of 219

Design of

Speech Elicitation Tasks

  • Saliency of speech across different conditions
  • Categorization of elicitation tasks
  • Matching elicitation tasks with clinical conditions

17 of 219

Saliency of Speech across Different Conditions

The production of spoken language

18 of 219

Saliency of Speech across Different Conditions

How speech sounds and what is said?

Disease or Condition

Speech Process Affected

Content:

“What is said…”

Natural Language Processing

Form:

“How it’s said…”

Acoustic Analysis

Mental Health (schizophrenia, bipolar depression)

CONCEPTUALIZATION

Reduced or increased speech output, incoherent speech, atypical sentence structure

Atypical speaking rate (too fast or slow), atypical prosody

Cognitive/Language (ADRD, Aphasia)

FORMULATION

Low lexical complexity, simple or atypical syntax,

Frequent and long pauses, slow speaking rate, imprecise articulation

Motor (Parkinson’s, ALS)

ARTICULATION

Reduced speech output

Slurred, slow or mumbled speech, atypical prosody, dysphonia, hypernasality

19 of 219

Speech Elicitation Tasks

Critical Considerations for Design

·

1.Goal/Purpose of the AI Model

2.Stage of Speech Production Impacted: Conception, Formulation, Articulation

3.Patient/Participant Considerations

4.Language and/or Cultural Considerations

5.Speech Processing and Metrics

20 of 219

  1. Goal/Purpose of the Model

Critical Considerations for Designing Speech Elicitation Tasks

21 of 219

2. Conception, Formulation, or Articulation?

Critical Considerations for Designing Speech Elicitation Tasks

22 of 219

3. Patient/Participant Considerations?

Critical Considerations for Designing Speech Elicitation Tasks

23 of 219

4. Language and/or Cultural Considerations?

Critical Considerations for Designing Speech Elicitation Tasks

24 of 219

4. Speech Processing and Metrics?

Critical Considerations for Designing Speech Elicitation Tasks

25 of 219

Speech Elicitation Tasks

Speech task landscape

26 of 219

Matching Elicitation Tasks with ML Use-Cases

Fit-for-Purpose Design

Example 1: �Prospective collection of speech samples to classify between mild cognitive impairment and no cognitive impairment in a genetically at-risk population of 50–80-year-old men and women in a health care system in New England major cities.

GOAL

Cross-sectional classification of +/- MCI

DEFICITS/SEVERITY

Formulation: cognition, memory, executive function, processing speed, word-finding/ Unimpaired to mildly impaired

PARTICIPANT LIMITATIONS

None to minimal

LANGUAGE/CULTURE

All American English native speakers

SPOKEN RESPONSE PROCESSING/METRICS

ASR generated transcripts for NLP analysis

27 of 219

Matching Elicitation Tasks with ML Use-Cases

Fit-for-Purpose Design

Example 2: �Prospective collection of speech samples to identify impact of intervention on disease progression in chronic obstructive pulmonary disease in women between the ages of 60-75 years old.

GOAL

Within-subject, longitudinal change detection

DEFICITS/SEVERITY

Articulation: Respiratory insufficiency for speech/ mild to severe

PARTICIPANT LIMITATIONS

Time-to-fatigue with increased severity

LANGUAGE/CULTURE

Follow instructions in native language

SPOKEN RESPONSE PROCESSING/METRICS

Digitized acoustic signal for extraction of primarily temporal measures (speaking rate, pause rate, phonation duration)

28 of 219

Matching Elicitation Tasks with ML Use-Cases

Fit-for-Purpose Design

Example 3:

Existing database of speech samples collected via telephone from patients with Parkinson’s disease. The goal of the study for which these data were collected was for voice detection of dysarthria progression over time.

GOAL

Within-subject, longitudinal changes in dysarthria severity

DEFICITS/SEVERITY

Articulation: articulatory precision from mild to severe deficits

PARTICIPANT LIMITATIONS

Follow instructions to produce “pa ta ka” as quickly and crisply as possible

LANGUAGE/CULTURE

Follow instructions in native language

SPOKEN RESPONSE PROCESSING/METRICS

Digitized acoustic signal for extraction of syllable count and rate, and goodness of articulation

Would this dataset be valuable for detecting progression from PD to PD+Lewy body dementia?

29 of 219

Remarks and Questions

30 of 219

Speech Data Acquisition

  • Influence of hardware devices
  • Source of noise and how it impacts speech features
  • Guidelines to control parameters in data acquisition
  • Hardware validation framework

31 of 219

Recording Data is easy, isn’t it?

If you acquire a professional recording booth (with technician)

Image Credit: https://whisperroom.com

32 of 219

Recording Pipeline?

A simple one

  • Simple in terms of preparation
  • But still many problems:
    • Separate tracks per speaker?
    • Recording quality?
    • Data handling?
    • Access control?

33 of 219

Recording Quality

Big problem in later analyses

  • How to select the proper microphone
  • Use the most expensive one?
  • Read (and understand) specifications
  • Plan the recording task ahead

34 of 219

Important Microphone Specs

Technical Aspects

  • Frequency Response:
    • Wide and flat frequency response to accurately �capture the full range of the human voice
  • Sensitivity
    • High sensitivity to capture quiet sounds and nuances �in the voice
  • Signal-to-Noise Ratio (SNR):
    • High SNR to minimize background noise and �ensure clear recordings

35 of 219

Microphone sensitivity

microphone's ability to convert sound pressure to an electric voltage

Image Credit: https://www.dpamicrophones.com

Video

  • Output voltage as function of the input at different Sound Pressure Levels (SPL).
  • Influenced by
    • reactiveness of diaphragm,
    • microphone type (dynamic, condenser),
    • amplification within the mic
  • Measure Output using 1Pa sound pressure with a standard 1kHz sine wave

36 of 219

Microphone sensitivity

microphone's ability to convert sound pressure to an electric voltage

Image Credit: https://www.dpamicrophones.com

Video

Lower Red Curve:

  • Represents the output of a microphone with a sensitivity of 1 mV/Pa.
  • Output at an SPL of 124 dB: approximately 31 mV.

Upper Blue Curve:

  • Represents the output of a microphone with a sensitivity of 40 mV/Pa.
  • Output at an SPL of 124 dB: approximately 1.3 V.

37 of 219

Microphone sensitivity

microphone's ability to convert sound pressure to an electric voltage

Image Credit: https://www.dpamicrophones.com

Video

  • Beware of the difference between Europe and US Specification:
    • Europe:
      • output voltage is microphones sensitivity (mV/Pa)
      • Also denoted as transfer factor (TF)
    • USA:
      • ratio between transfer factor and hypothetical highly efficient mic

How much voltage you get for applying one pascal sound pressure

Assumed to produce �1V / 1Pa

Example: Sennheiser MKE600

The lower the negative value the better the sensitivity

38 of 219

Microphone sensitivity

Select the proper microphone

  • Select a mic with a sensitivity that suits your application�
  • Low volume sound sources need
    • high sensitivity mics
    • or a post output amplification (will include sound distortion)
  • High volume sound sources need
    • Low sensitivity mics
    • As high sensitivity mics could easily overload the pre-amplifier or mixer and produce distortion

Image Credit: https://www.dpamicrophones.com

Video

39 of 219

Signal-to-Noise Ratio

Determines how “clean” the output signal is

  • Measure SNR with typical test sound
    • 1Pa sound pressure with a standard 1kHz sine wave used

Video

40 of 219

Signal-to-Noise Ratio

Real Examples

Sennheiser MKE 600 �Shutgun Microphone

Comica �Traxshot Microphone

aka Field Microphone

41 of 219

Frequency Response

How are different frequencies perceived?

  • indicates the complete frequency range at which the microphone responds
  • perfect frequency response is not necessarily flat.
    • tailored frequency responses for better intelligibility

Reproduces sound with little or no coloration/variation from original sound

Flat response microphone (Shure SM81)

Decreased sensitivity for low frequencies reduce pick up of room noise or vibration

As well as counteracts that build up of bass that can occur when the recording distance is low

Increased sensitivity in upper mid-range add clarity to vocals

Shaped response microphone (Shure KSM42)

42 of 219

Frequency Response

Low-Frequency Roll-off Control

  • Some microphones allow frequency response to be adjusted
  • Most common:
    • Low-frequency roll-off control

adjusted

Low frequency roll-off inactive

Low frequency roll-off active

43 of 219

Frequency Response

How are different frequencies perceived from different directions?

Image Credit: https://www.dpamicrophones.com

  • professional microphones may provide more than one frequency response curve,
  • present how the microphone responds to
    • sound coming from different directions and in
    • different acoustic sound fields.

44 of 219

Directivity

“Control” which directions are recorded

  • Measure sensitivity with different directions of the sound source
    • Each circle represent a different microphone sensitivity value
    • 1Pa sound pressure with a standard 1kHz sine wave used
  • Pattern changes with the frequency!
  • Observe the type of direction for your application
    • Shotgun mics also have a height directivity
    • Two shotguns can be used for interview setting
    • Lavalier mics are mostly omnidirectional!

Video

45 of 219

Directivity

“Control” which directions are recorded

Image Credit: https://www.dpamicrophones.com

Types of directional microphones

46 of 219

Pop Filter

Reduce the distortion from plosives

Image Credit: https://www.audiomentor.com

  • Eliminate ‘popping’ noises when you sing or speak into the microphone
  • But you don’t necessarily need a pop filter
    • Position the microphone a little off-axis to the vocalist’s mouth
    • Especially when experiencing a “dulled” sound ⇒ remove the pop filter

47 of 219

Shotgun vs. Condenser vs. Lavalier

Specifications and Specialities

Image Credit: https://www.dpamicrophones.com

  • All Mics use a capacitor to convert sound into electrical signal
    • Need external power to continually charge the capacitor
  • Condenser mics
    • have large diaphragm ⇒ helps to pick up richer and fuller sound
  • Shotgun mics
    • used for long range pickups
    • Narrow pickup pattern helps to eliminate noise from the surrounding/back of the mic
  • Lavalier mic
    • High freedom of movement as it is fixed to the body
    • Allows recording breathing sounds
    • Correct placement is crucial for high-quality recordings

Sennheiser �MKE 600 �

Shure �KSM42

Rode Lavalier GO

48 of 219

Impact of technical measurement factors

Why microphone specs matter

  • Based on a paper by Dineley at al. 2023
  • Experimental Setup:
    • Speech recordings of 42 healthy volunteers in rooms with low and high reverberation
    • Simultaneous recordings with one budget and two higher-end smartphones and a condenser microphone
  • Can the associated variability in the recorded speech signal be erroneously interpreted as related to a change in health state?

Paper

49 of 219

Impact of technical measurement factors

Why microphone specs matter

Participant characteristics (n = 42)

Standardized differences in timing features

  • Errors bars are 95% confidence intervals
  • Negative differences represent lower feature values in the presence of higher reverberation

Sex

female

23

male

19

Age (years)

median

28

IQR

23-32

English L1

Yes

29

No

13

Height (m)

median

1.70

IQR

1.63-1.80

50 of 219

Impact of technical measurement factors

Why microphone specs matter

  • Positive differences represent higher values in the presence of higher reverberation.
  • Voice quality features are seemingly the most affected by reverberation
  • Amplify concerns about the robustness and validity of these features for health assessments

51 of 219

Challenge: Impact of technical measurement factors

52 of 219

Challenge: Impact of technical measurement factors

Speech analytical pipeline is susceptible to many types of variability

In-the-wild longitudinal monitoring introduces multiple technical, acoustic and human factors into the recording process

Can the associated variability in the recorded speech signal be erroneously interpreted as related to a change in health state?

Recorded the speech of 42 healthy volunteers recorded consecutively in rooms with low and high reverberation

Simultaneous recordings on one budget and two higher-end smartphones and a condenser microphone

53 of 219

Challenge: Impact of technical measurement factors

Participant characteristics (n = 42)

Standardized differences in timing features

Errors bars are 95% confidence intervals

Negative differences represent lower feature values in the presence of higher reverberation

Sex

female

23

male

19

Age (years)

median

28

IQR

23-32

English L1

Yes

29

No

13

Height (m)

median

1.70

IQR

1.63-1.80

54 of 219

Challenge: Impact of technical measurement factors

Speech analytical pipeline is susceptible to many types of variability

Standardized differences in acoustic features

Positive differences represent higher values in the presence of higher reverberation.

Error bars represent 95% confidence intervals.

Voice quality features are seemingly the most affected by reverberation

Amplify concerns about the robustness and validity of these features for health assessments

55 of 219

Improved Recording Pipleline?

High-quality shot-gut mics used

  • Does it work?
  • Simultaneous talking
  • Loudness analyses

56 of 219

Improved Recording Pipleline?

High-quality headsets used

  • Does it work?
  • Simultaneous talking
  • Loudness analyses

57 of 219

Recording can still include errors

Why adjustment of recording level (GAIN) is important

  • Recording Loudness has to be adjusted manually regarding
    • current setup and
    • Individual speaker ⇒ note down GAIN for later comparison

Watch the sound level meter!

Clipping!

58 of 219

The Effects of noise on Acoustic Parameters

Why noise should be avoided

  • Based on a paper by Özseven and Düğenci 2015
  • Experimental Setup:
    • Berlin Database of Emotional Speech
    • Additive Gaussian noise with 4 SNR levels (1dB, 5dB, 10dB and 15dB)
    • Praat for audio analyses (pitch, jitter, shimmer, harmonic noise rate, noise harmony rate, format frequencies, and energy density parameters

59 of 219

The Effects of noise on Acoustic Parameters

Why noise should be avoided

60 of 219

The Effects of noise on Acoustic Parameters

Why noise should be avoided

61 of 219

Recording can still include errors

Why sampling rate matters

Sampling is the reduction of a continuous-time signal to a discrete-time signal

  • Sampling frequency is the average number of samples obtained in one second
  • Frequency components whose cycle length (period) is less than 2 sample intervals cannot be reconstructed

Image Credit: https://www.youtube.com/watch?v=Z0EMObqS90U

Range of human hearing:

20 Hz- 20,000 Hz

62 of 219

Recording can still include errors

Why sampling rate matters

Speech Sampling, oriented towards its application:

  • 8000 Hz: Analog telephone, wireless intercom; minor with sibilances (ess sounds like eff (/s/, /f/))
  • 11,025 Hz: Lower-quality PCM, MPEG audio and for audio analysis of subwoofer band passes
  • 16,000 Hz: Wideband telephone speech (VoIP)
  • 22,050 Hz: Lower-quality PCM, MPEG audio and for audio analysis of low frequency energy
  • 32,000 Hz: miniDV, DAT and NICAM audio, high-quality digital wireless microphones
  • 44,100 Hz: AudioCD, MPEG-1 audio, MP3,
  • 48,000 Hz: Professional digital video equipment
  • 96,000 Hz: DVD-Audio, Blu-ray and HD DVD

Human Speech intelligibility is between 300 and 3,400Hz

Video

63 of 219

Recording can still include errors

Why recording format and bitrate matters

  • Audio File formats used to store audio data on computer
    • Specific bit layout of the audio data to reduce file size
    • Uncompressed, compressed (lossless or lossy)

Image Credit: https://filesconverter.com/best-audio-format

64 of 219

Recording can still include errors

Why recording format and bitrate matters

I. Siegert et al., (2016). Measuring the impact of audio compression on the spectral quality of speech data. ESSV, 2016

65 of 219

Compression and acoustic parameters

Why low bitrates influence acoustic measurement

  • Based on a paper by Fuchs and Maxwell, Interspeech 2016
  • Experimental Setup:
    • DyViS database (selected speakers)
    • MP3 codec with 7 compression rates (16, 32, 56, 96, 128, 256, 320 kbps)
    • Praat for audio analyses (f0, PDQ, pitch range, and pitch level)

Pitch dynamism quotient�

(PDQ)⇒

Paper

66 of 219

Compression and emotion recognition

Why recording format and bitrate matters

Speech Emotion Recognition Experiments�

  • Combination of Analysis-by-Synthesis and psychoacoustic modelling�
  • Comparison with two single-operation codecs:�
    • MP3 (Psychoacoustic modelling)
    • OPUS (Hybrid mode)
    • AMR-WB (Analysis-by-Synthesis)

Paper

  • Based on a paper by Siegert et al., Specom 2016

67 of 219

Recording Data:

Towards mobile health application

Speech analytical pipeline is susceptible to many types of variability

68 of 219

Remarks and Questions

Link to survey on speech data acquisition

69 of 219

From

Speech Features to

Speech Measures

  • Existing speech features
  • Methodological shift from features to measures
  • Validation framework of speech measures
  • Case study: FDA breakthrough designation

70 of 219

Reproducibility Challenge

As a community we need to build-up reproducible and reliable clinical evidence

Challenge: Speech analysis has to be reliable at a level that is acceptable for clinical decision making

Issues

  • Absence of standardised collection and analytics protocols
  • Need for more clinically meaningful / interpretable systems
  • Lack of replication and verification analysis
  • Lack of standard definition of core features
  • Many features have poor test-retest reliability
  • Focus on large multivariate representation for use with machine learning

71 of 219

Reproducibility Challenge

Harmonisation across studies

Challenges

  • Core methodological choices often not reported in paper
  • This can include feature names, feature definitions, and key hyperparameter setting associate with their extraction
  • Feature definitions vary between toolboxes
  • Leading toolboxes can be difficult to change from default settings
  • Lack of evidence for many widely used features and popular representations

72 of 219

Reproducibility Challenge

As a community we need to build-up reproducible and reliable clinical evidence

Reproducibility and harmonisation

  • Present a minimal speech-health feature set
  • Features can be used individually or in combinations in hypothesis testing.
  • They can be used as a baseline multivariate space in machine learning or prediction modelling

Starting point: Add additional features and improve reliability through hypothesis driven research

73 of 219

Reproducibility Challenge

Minimal speech-health feature set

Inclusion criteria:

  • The features could not be specific to a single elicitation prompt or strategy
  • ‘Basic’ representation; i.e., not heavily reliant on another feature or not capturing interactions between features
  • Minimal reliance on third-party software; i.e., ASR required before extraction

74 of 219

Minimal Speech Health Data Set

Timing/Fluency

Speed: speed with which speech is performed

Breakdown: the pauses and silences that disrupt the flow of speech

Repair fluency: hesitations, repetitions, and reformulations

75 of 219

Minimal Speech Health Data Set

Timing/Fluency

Phonation ratio: phonation time divided by duration

Speech rate: Number of linguistic units divided by duration

Articulation Rate: Number of linguistic units divided by phonation time

Pause Rate: Number of Pause divided by duration

Mean Pause Duration: Mean duration of all pauses longer than 300ms

76 of 219

Minimal Speech Health Data Set

Timing/Fluency

Extraction:

Our extraction utilises the code originally developed for L2 fluency, with slight modifications regarding efficiency and error catching.

This code uses an intensity threshold to automatically identify pause boundaries,

Syllables identified using a combination of intensity thresholds and voicing information.

Notes:

Other methods, such as forced alignment, can be used to extract these features, this would mean they are reliant on additional third-party software.

Repair fluency properties difficult to extract without ASR

77 of 219

Minimal Speech Health Data Set

Speech production subsystems and representative acoustic features

  • Intensity (Loudness)
  • Fundamental Frequency (Pitch)
  • Harmonic to Noise Ratio
  • Spectral Slope
  • Spectral Tilt:
  • Cepstral Peak Prominence
  • Spectral Moments
  • Formants

78 of 219

Minimal Speech Health Data Set

Respiration

Process of moving air in and out of the lungs

Stability: Speech production requires stable air pressure throughout our vocal tract

Power: When shouting or speaker longer sentences, we need a bigger inhalation

79 of 219

Minimal Speech Health Data Set

Respiration

Features:

  • Intensity (Mean): Mean loudness of a speech signal
  • Intensity (Range): Mean range of loudness values in a speech signal

Extraction:

Utilises Praat Sound: To Intensity… function:

Intensity calculated as the root-mean-square amplitude of the provided audio signal:

Note

Other respiratory measures (E.g., the detection of respiratory events within a speech signal) require third-party software in their extraction

80 of 219

Minimal Speech Health Data Set

Phonation

Production of sound at the level of the vocal folds

Voiced Speech: produced when vocal folds are vibrating

  • Harmonic speech sounds; e.g., /a/, /e/, /i/, /o/, /u/

Unvoiced Speech: vocal folds are lax and open

  • Plosive, noise like speech sounds; e.g., /f/, /s/, /t/

Related perceptual properties include: pitch and vocal quality (E.g., hoarseness, breathiness, roughness)

81 of 219

Minimal Speech Health Data Set

Phonation

Features:

  • Fundamental Frequency: Rate of vocal fold vibration (pitch)
  • Harmonic to Noise Ratio: Degree of acoustic periodicity
  • Spectral Slope: Ratio of voiced energy between 10-1000Hz over 1000-4000Hz
  • Spectral Tilt: Linear slope of voiced energy distribution between 100-5000Hz
  • Cepstral Peak Prominence: Stability of vocal fold vibration

82 of 219

Minimal Speech Health Data Set

Phonation

Fundamental frequency (F0) vs. Pitch?

  • F0 is the rate of vibration of the vocal folds.
  • Rate of vocal fold vibration is associated with the perceptual characteristics of pitch

Features:

  • Mean Pitch value over a recording
  • Mean pitch sigma over a recording
    • Pitch sigma is the pitch standard deviation converted to semitones

Extraction

  • Our extraction utilises the code originally presented developed for standardisation in F0 extraction, with slight modifications regarding efficiency and error catching.

83 of 219

Minimal Speech Health Data Set

Phonation

Harmonics-to-noise ratio (HNR)

  • Additive noise in the voice signal arising from turbulent airflow generated during phonation due to inadequate vocal fold closure or aperiodic vocal fold vibration
  • It relates to the perceptual qualities of roughness and breathiness

Extraction

  • Extracted using the Praat: Sound: To Harmonicity (cc)... function
  • Default settings

84 of 219

Minimal Speech Health Data Set

Phonation

Long-time average spectrum (LTAS)

  • Provides spectral information averaged over time
  • The LTAS captures information on both the glottal source spectrum and the resonant characteristics of the vocal tract
  • Extracted using the Praat Sound: To Ltas (pitch-corrected)... function to minimse the effect of F0 and harmonics

Spectral Slope

  • Ratio of energy in a spectra between 10-1000Hz over 1000-4000Hz
  • Intensity of the lower harmonics compared to the intensity of the higher harmonics
  • Extracted using the Praat Get Slope function on the LTAS

Spectral Tilt

  • Linear slope of energy distribution between 100-5000Hz
  • Amplitude of the fundamental frequency compared to that of higher-frequency harmonics
  • Extracted using the Praat Report spectral tilt function on the LTAS

85 of 219

Minimal Speech Health Data Set

Phonation

Cepstral Peak Prominence

  • Stability of vocal fold vibration
  • The amplitude of the cepstral peak, relative to a regression line through the cepstrum
  • This peak is a function of the level of overall noise in a voice signal
  • Used as an objective measure of breathiness and overall dysphonia

Extraction:

86 of 219

Minimal Speech Health Data Set

Phonation

Jitter and Shimmer

Note that we have not included two of the better-known and widely used phonation measures in jitter and shimmer.

These features are open to errors due to differing sound pressure levels and phonetic content between and within individuals

Should be considered unreliable for analysis of voice pathology

Also subject to measurement errors due to recording conditions

87 of 219

Minimal Speech Health Data Set

Articulatory

Producing unique speech sounds by changing the shape of the vocal tract

Different muscle formations changes the point of constriction in the vocal tract

88 of 219

Minimal Speech Health Data Set

Articulatory

Formants

Formants are the resonances of the vocal tract

The defines consonant and vowel perception

  • Formant Frequency is the centre frequency of a particular resonance.
  • Formant Bandwidth the difference in frequency between the two points either side of a formant frequency where the frequency amplitude has dropped by 3 dB

Extraction

Praat: Sound: To Formant (burg)... function which utilises linear predictive coding

89 of 219

Minimal Speech Health Data Set

Articulatory

Spectral features characterise the speech spectrum.

Typical spectral features are high dimensional representations that capture all the information contained in speech, which could include confounds

90 of 219

Minimal Speech Health Data Set

Articulatory

Use the four spectral moments to characterise the spectrum.

Extraction:

Extracted from a Praat Spectrogram object using the relevant function above

91 of 219

Minimal Speech Health Data Set

Code being tested before initial release

Code is implemented in Parselmouth so it can run in Python

  • Currently being beta tested
  • Release later this year

Please get in touch if you would like early access

92 of 219

Not linked to theoretical construct

Reliability and generalizability not individually characterized

Different implementations by different groups

Designed to measure specific construct

Psychometric properties characterized; normative data

Standardized for comparability

Selected to improve model accuracy

Clinical meaningfulness?

Liss, Julie, and Visar Berisha. "Operationalizing Clinical Speech Analytics: Moving From Features to Measures for Real-World Clinical Impact." Journal of Speech, Language, and Hearing Research (2024): 1-7.

Allen, Mary J., and Wendy M. Yen. Introduction to measurement theory. Waveland Press, 2001.

Features

Measures

Moving from speech features to speech measures

93 of 219

5.

  • 110 participants with ALS
  • 7632 speech sessions
    • All collected remotely using the participant’s own device

  • Speech measure:
    • Articulatory Precision
      • algorithmically calculated
    • Clinically Validated
      • perceptual ratings
      • ALSFRS-R speech subscore

Stegmann, G.M., Hahn, S., Liss, J., Shefner, J., Rutkove, S., Shelton, K., Duncan, C.J. and Berisha, V., 2020. Early detection and tracking of bulbar changes in ALS via frequent and remote speech analysis. Nature npj Digital Medicine, 3(1), pp.1-

ALS@Home and ALS - FTD studies:

Amyotrophic Lateral Sclerosis

94 of 219

Witt, Silke M., and Steve J. Young. "Phone-level pronunciation scoring and assessment for interactive language learning." Speech communication 30.2-3 (2000): 95-108.

Mathad, Vikram C., et al. "The Impact of Forced-Alignment Errors on Automatic Pronunciation Evaluation." Interspeech. 2021.

  • ALS impacts respiratory support, phonation, articulation, resonance, and prosody.

  • Focus on articulatory precision: a range of subjective impressions one experiences when they hear speech, from crisp and clear to slurred and mumbled.

  • Physiologically, articulatory precision is mediated by the integrity of articulatory structures (e.g., lips, tongue, soft palate) as they shape and control the flow of voiced and unvoiced energy through the vocal tract.

  • Acoustically, speech with high articulatory precision contains vowels and consonants with high acoustic and perceptual distinctiveness.

Measuring the right constructs

95 of 219

Witt, Silke M., and Steve J. Young. "Phone-level pronunciation scoring and assessment for interactive language learning." Speech communication 30.2-3 (2000): 95-108.

Mathad, Vikram C., et al. "The Impact of Forced-Alignment Errors on Automatic Pronunciation Evaluation." Interspeech. 2021.

How it works?

Acoustic model for likelihood estimation

Acoustic model for alignment

Speech + Transcript

Likelihood ratio computation

  • Elicitation: Phonemically balanced sentences, randomly presented from session to session.

  • Data acquisition: Participants allowed to use their own iPhone or Android devices (minimum version required).

  • Objective measure: An algorithm for assessing phoneme integrity from the recorded speech.

Operationalizing the construct

96 of 219

Mathad, Vikram C., et al. "The Impact of Forced-Alignment Errors on Automatic Pronunciation Evaluation." Interspeech. 2021.

Stegmann, G., Charles, S., Liss, J., Shefner, J., Rutkove, S., & Berisha, V. (2023). A speech-based prognostic model for dysarthria progression in ALS. Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration, 1-6.

N=110 (40F); tracked longitudinally for up to 1 year

Analysis

Result

Correlation with clinical ratings of articulatory precision

r = .90

Test-retest reliability

ICC = .97

Correlation with ALSFRS-R speech

r = .82

Separation of severe and nonsevere speech

ROC AUC = .97

Correlation between articulatory precision and ALSFRS-R speech

subscale longitudinal slopes

r = .37

Correlation between articulatory precision and ALSFRS-R bulbar

subscale longitudinal slopes

r = .41

Longitudinal change

Slope = -.004

How good does the aligner have to be?

TLDR: MFA forced aligner is fine

Validating the algorithm

97 of 219

Establishing clinical meaningfulness: why is the objective measure meaningful to patients?

ALS impacts the speech production mechanism. This has a negative impact on a patient’s communicative participation and quality of life (Börjesson, 2021; Borrie et al, 2022). This is because others have a difficult time understanding them and there is increased listener effort to understand them (Stipancic et al., 2018). The standard way to measure the impact on speech is the ALSFRS-R speech subscale, a 5-point scale that is part of the clinical standard. We devised a new measure for object assessment of precision of articulation. This tool has high correlation with perceptual measures of articulatory precision, establishing construct validity. It has high test-retest reliability, making it appropriate for longitudinal monitoring (Stegmann et al., 2020; Stegmann et al., 2023). It is more sensitive than the ALSFRS-R speech subscale to longitudinal change in a patient’s speech (Stegmann et al., 2024).

98 of 219

Stegmann, G., Krantsevich, C., Liss, J., Charles, S., Bartlett, M., Shefner, J., Rutkove, S., Kawabata, K., Talkar, T. and Berisha, V., 2024. Automated speech analytics in ALS: higher sensitivity of digital articulatory precision over the ALSFRS-R. Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration, pp.1-9.

Establishing clinical utility: improving clinical research

Reducing sample size requirements in clinical trials

Designing new trials specifically focused on speech outcomes

99 of 219

  • Correlation between predicted and observed was between r = 0.7 (waiting period of 90 days) and 0.96 (waiting period of 30 days)

  • Clinical utility:
    • Predicting loss of intelligibility
    • End of life planning
    • Clinical trial enrollment/ stratification

Stegmann, G., Charles, S., Liss, J., Shefner, J., Rutkove, S., & Berisha, V. (2023). A speech-based prognostic model for dysarthria progression in ALS. Amyotrophic Lateral Sclerosis and Frontotemporal Degeneration, 1-6.

Establishing clinical utility: A prognostic model for ALS

100 of 219

FDA and Clinical Trial Outcome

March 2023: FDA Breakthrough Device Designation

“Medical devices receive breakthrough designation from the FDA when they provide for more effective treatment or diagnosis of life-threatening or irreversibly debilitating diseases or conditions. The goal of the program is to provide patients and healthcare providers with timely access to groundbreaking technologies by expediting the approval process.”

February 2023: “The impact of pridopidine on speech measures was notable, likely due to its S1R mechanism of action. Speech is a highly clinically relevant endpoint in ALS studies, and more than 80 percent of ALS patients become speech impaired, which significantly impacts their quality of life.”

101 of 219

From features to measures

With a subtitle if you need it

When using the template, try to not move the headlines up and down. It will add a level or professional design if the top margin of your headline is not moving around between slide transitions.

102 of 219

Remarks and Questions

103 of 219

Afternoon Session

104 of 219

Recap of Morning Session

  1. Design of speech elicitation task:��→ Impact of clinical conditions to speech production - conceptualization, formulation, articulation��→ Considerations for speech task design - model’s goal, deficits, patient limitation, demographic factors…��Speech task landscape - speech degree of freedom vs. cognitive load��→ Clinical examples - MCI, pulmonary disease, Parkinson’s disease…�

105 of 219

Recap of Morning Session

2. Speech data acquisition�� Considerations for microphone specs � - Frequency response, sensitivity, SNR, directionality, microphone type, etc.�� → Differences in measurement across microphone brands�� Noise / post-processing technique that impact speech features compression, etc.)�� → Variability in data acquisition pipeline�

106 of 219

Recap of Morning Session

3. Speech features and measures�� Limitations in existing speech features - reproducibility, clinical evidence, � implementation, etc.� → Minimal Speech Health Data Set - Timing/frequency, speech production � subsystems, respiration, phonation, articulatory

� → Moving from speech features to speech measures - finding right construct,� operationalize the construct, validate the algorithm, establish clinical meaning� � → FDA breakthrough - An ALS example�

107 of 219

Learning Generalizable & Interpretable

�Speech Representations &�Clinical Models

  • Develop speech measures using ML; Analytical validation
  • Develop clinical ML models using speech measures; Clinical validation

108 of 219

From Speech Data to Clinical Label: Commonly-seen Pipeline

  1. Speech feature: High-dimensional; Don’t address a specific clinical theory�
  2. ML model: tuned to optimize prediction accuracy.

109 of 219

A Revised Pipeline from Speech Data to Clinical Label

  1. Speech measures: linked to clinical construct; interpretability & reliability guaranteed�
  2. Clinical ML model: developed on validated measures; generalizable & interpretable

110 of 219

Design of Speech Measures

Traditional ways to obtain speech measures�� Perceptual rating (use a scale to rate coordination, speed, voice quality, etc.)�� Manual transcription (speech intelligibility, speech rate, etc.)��→ Signal processing methods (F0, jitter, shimmer, cepstral peak prominence, etc.)

111 of 219

Design of Speech Measures

We have speech data and some type of labels…��How to design an ML algorithm to derive speech measure?��- Utilize domain knowledge of speech production, semantic similarity in language.��- Predict clinically-validated measures

Traditional ways to obtain speech measures:�1) Perceptual rating�2) Manual annotation�3) Signal processing

112 of 219

Design New Speech Measures

Using domain knowledge of speech production

Derive measures using probability outputs from acoustic models

  1. Acoustic models: speech signal → linguistic unit mapping (e.g. phoneme)�
  2. Probability output reflects speech production quality�
  3. Can be trained with healthy speech data

113 of 219

Design New Speech Measures

Using domain knowledge of speech production

  1. Classical approach:�- Goodness of pronunciation (GoP) (Witt and Young, 2000)�- Computer assisted language learning (CALL)�
  2. Clinical speech analytics: Specific objectives�- Hypernasality → Nasality in speech production�- Articulation problem → Precision of consonant production�- Speech disorder in children → Manner and place of articulation

114 of 219

Design New Speech Measures

Example 1: Objective Hypernasality Measure (OHM) - Algorithm

  1. Linguistic units:�- Nasal consonants (NC)�- Oral consonants (OC)�- Nasal vowels (NV)�- Oral vowels (OV)�
  2. Probability output �→ Objective Hypernasality Measure (OHM)

Task: �Assess hypernasality in speakers with cleft palate (Mathad et al. 2021)

Training data: Healthy speech

115 of 219

Design New Speech Measures

Example 1: Objective Hypernasality Measure (OHM) - Algorithm

Visualization of OHM:

-> Sentence: “Buy baby a bib”�(No nasal consonant & vowels)�

-> Nasal consonant & vowel detected in hypernasal speech (shown in red).�

Can we use the OHM now?�NO, we need validation.

116 of 219

A Measure’s Reliability and Validity

117 of 219

Learnable Speech Measures

Example 1: Objective Hypernasality Measure (OHM) - Validation

  1. Internal reliability: consistency of OHM measured from different speech materials consistent?�
  2. OHM-rater reliability: consistent with human ratings?�
  3. Cross-corpora performance (External validity): OHM generalizable to different datasets?

118 of 219

Design New Speech Measures

Example 2: Objective Articulation Measure (OAM) - Algorithm

  1. Consonant vowel segments → Classify consonants�
  2. Probability output → Objective Articulation Measure (OAM)

Task: �Evaluation articulation in speakers with dysarthria, cleft lips/palate (Mathad et al. 2023)

119 of 219

Design New Speech Measures

Example 2: Objective Articulation Measure (OAM) - Algorithm

Visualization (CNN Salience Map): ��Trained model is focused on consonants and vowel onsets when classifying the consonants��→ Confirm the model is properly utilizing the inputs.

120 of 219

Design New Speech Measures

Example 2: Objective Articulation Measure (OAM) - Validation

Validation:� �- Compare OAM with perceptual rating of clinical speech

- Compare OAM with existing measures that also measure articulation (GOP)

121 of 219

Design New Speech Measures

NLP: Cosine similarity as a measure

�Speech can be transcribed into text ��→ Natural language processing (NLP) to derive speech measures

122 of 219

Design New Speech Measures

Example 3: Using cosine similarity to assess thought disorder - Algorithm

Task: Assess thought disorder (Bilgrami et al. 2022)

Approach: �1) Use pre-trained BERT model to extract sentence-level embeddings2) Find the minimum semantic coherence based on cosine similarity between sentences.

123 of 219

Design New Speech Measures

Example 3: Using cosine similarity to assess thought disorder - Validation

Validation:��- Construct validity: Compare the minimum semantic coherence with 8 clinical measures �

124 of 219

Learnable Speech Measures

NLP: Cosine similarity as a measure

Cosine similarity between text embeddings → Correlate to clinical measures

�1. Word2Vec → Semantic similarity between flow of ideas → Thought disorder

2. Sentence embedding between adjacent utterance pairs → Coherence measure → Schizophrenia

125 of 219

Predict Validated Measures

Learnable speech measures

  1. Examples of valid measures: �- Perceptual ratings of intelligibility, hypernasality, vocal quality, etc.�- OHM and OAM�- and many others introduced in existing studies…�
  2. Build ML models to predict these validated measures

126 of 219

Predict Validated Measures

Example 1: Information bottleneck - Algorithm & Validation

  1. Intermediate DNN layer → Predict perceptual ratings (Tu et al. 2017)�- Predict nasality, vocal quality, articulatory precision, prosody�
  2. Benefit: �- All features used for clinical label prediction links to clinical construct.�- Interpretable to most clinicians work with pathological speech

W1: Measure prediction;

W2: Predict diagnostic labels.

y: Severity of dysarthric speech

Task: Assess severity of dysarthric speech

127 of 219

Predict Validated Measures

Example 1: Information bottleneck - Algorithm & Validation

Validation: Compare predicted measures with ground truth perceptual labels

128 of 219

Predict Validated Measures

Example 2: Information bottleneck - Algorithm

  1. Intermediate DNN layer → Predict articulatory measures (Xu et al. 2023) �- OAM → Consonant-vowel transition�- GOP → Articulatory precision�- OHM → Hypernasality�- Cepstral peak prominence → Vocal quality

Binary Output → Presence of dysarthria

Obtaining these measures �DO NOT require manual annotation

129 of 219

Predict Validated Measures

Example 2: Information bottleneck - Validation

Validation:�- Compare selected measures with ground truth values

- Evaluate if the selected measures are truly useful for dysarthria detection

Large absolute SHAP value �→ More contribution to decision making

130 of 219

Clinical ML for Predicting Clinical Variables

An ideal clinical ML model: ��1) Interpretable inputs

2) Interpretable & generalizable model

3) Produce reliable and clinically meaningful outputs

How to design a clinical ML model based on speech measures?

131 of 219

Clinical ML for Predicting Clinical Variables

An ideal pipeline of clinical ML model: ��1) Interpretable inputs → Adopt interpretable speech measures

2) Interpretable & generalizable model → Simple ML model, low-dimensional inputs�

3) Produce reliable and clinically meaningful outputs → Perform clinical validation

132 of 219

Clinical ML for Predicting Clinical Variables

Example 1: Assess verbal memory in mental illnesss (Holmlund et al., 2020)

  1. Speech Task: Story recall
  2. Clinical ML input:�Common Word Types & Word Mover’s Distance (WMD)�computed from ASR transcriptions�
  3. Clinical ML Model: �Least square linear regression model�
  4. Clinical output: �Clinician rated accuracy of recalls ()�
  5. Clinical objective:�Perform a self-administered ambulatory verbal memory test with automated scoring (schizophrenia)

133 of 219

Clinical ML for Predicting Clinical Variables

Example 1: Assess verbal memory in mental illnesss (Holmlund et al., 2020) �- Clinical validation

  1. Model interpretation: Regression coefficient�→ Common word types = 0.15; WMD = -0.54�
  2. Clinical validation: �A. Predicted scores similar to human raters?→ Pearson correlation between the two ratings��B. Is the automated data processing pipeline comparable to manual processing?�→ Compare scores: ASR vs. human transcription .

134 of 219

Clinical ML for Predicting Clinical Variables

Example 2: Longitudinal Tracking of Forced Vital Capacity Using Speech Acoustics (Stegmann et al. , 2021)

  1. Speech Task: Maximum phonation�
  2. Clinical ML input: maximum phonation time, age, height�
  3. Clinical ML Model: Linear model�
  4. Clinical output: Forced Vital Capacity (FVC)�
  5. Clinical objective: Longitudinal monitoring of FVC

135 of 219

Clinical ML for Predicting Clinical Variables

Example 2: Longitudinal Tracking of Forced Vital Capacity Using Speech Acoustics (Stegmann et al. , 2021) - Clinical Validation

  • Clinical validity: comparison between observed FVC and predicted FVC (r=0.80)

136 of 219

Clinical ML for Predicting Clinical Variables

Example 2: Longitudinal Tracking of Forced Vital Capacity Using Speech Acoustics (Stegmann et al. , 2021) - Clinical Validation

  • Each subjects provide multiple speech samples → Test-retest reliability of FVC���(ICC: Intra-class correlation; SEM: within-person standard deviation; CV: coefficient of variation)

137 of 219

Clinical ML for Predicting Clinical Variables

Example 2: Longitudinal Tracking of Forced Vital Capacity Using Speech Acoustics (Stegmann et al. , 2021) - Clinical Validation

  • Can the predicted FVC satisfy the clinical objective? → Longitudinal tracking��Compare longitudinal trajectories between observed and predicted FVC��

138 of 219

Learning Generalizable & Interpretable Speech Representations & Clinical models

Suggestions

  • Developing clinical speech ML with existing datasets:

Thoroughly understand data and labels. �- Ground truth measure → Analytical validity �- Repeated tests → Test-retest reliability�- Same subject, measured on different words/sentence → Internal reliability�- Other validated measures that measure the similar construct → Construct validity�- Measures made by multiple raters → Inter-rater reliability�etc. �

139 of 219

Learning Generalizable & Interpretable Speech Representations

Suggestions

  • Building a new clinical datasets to develop clinical speech ML:�- With the given resource, optimize the data acquisition protocol to collect sufficient data for reliability & validity evaluation.�- e.g. different clinical measures, ratings from multiple raters, repeated tests, etc.

140 of 219

Learning Generalizable & Interpretable Speech Representations

Suggestions

  • Continual development of new speech measures

- Lots of measures can be designed without disordered data

- Allow more clinical constructs to be utilized in clinical speech analytics → A more comprehensive diagnosing process

- More feature options in clinical model development → Improve analytic performance

- etc.

141 of 219

Design of Clinical ML for Diagnostic Decision

Example: (Sauder et al. 2017)

1. Interpretable inputs: ��→ Smoothed cepstral peak prominence (CPPS): describe breathiness of voice

�2. Model: Logistic regression �

Disorder status ~ coeff1 + coeff2 * Gender + coefff3 * CPP

3. Clinical output: Presence of voice disorder

�4. Interpretation: Regression coefficient �

142 of 219

Expectation from Regulatory Bodies

A perspective from FDA

Endpoint: what happens to people in the clinical trial.��Clinical endpoint: directly measure what matters most to people, �e.g. occurrence of disease, improvement in symptoms, function better.�→ Occurence of disease�→ Severity score, depression score, mental state score �→ Change of score reflects improvement/worsening of disease�

143 of 219

Here is a text slide

With a subtitle if you need it

When using the template, try to not move the headlines up and down. It will add a level or professional design if the top margin of your headline is not moving around between slide transitions.

144 of 219

Remarks and Questions

145 of 219

Ethical, Privacy and Security Considerations

  • Presented by Nina (60 mins)
  • Experimental designs for robust machine learning
  • Data bias, model discrimination, model fairness
  • Privacy, attacks, misuse
  • Participatory Research
  • Actions supporting responsible development of clinical speech analytics

146 of 219

Overview

Responsible Clinical

Speech Analytics

Participant Factors

Privacy�Attacks�Misuse

Bias�Discrimination Fairness

Access Equity

Sustainability

Robust Experimental Designs

147 of 219

Robust

Experimental Designs

148 of 219

Generalizable AI is Ethical AI

Ghasemzadeh et al. (2024)

The experimental design strategy used in machine learning studies influences:

  • the unbiased estimation of model performance
  • the minimum sample size required to demonstrate a statistically significant outcome (i.e., power)
  • the probability that correct measures are included in the final model after feature selection

The experimental design strategy is particularly important in clinical datasets.

149 of 219

Data Splitting Experimental Designs

Ghasemzadeh et al. (2024)

  1. Single holdout B) k-fold (k = 3)

C) Train-validation-test D) nested k-fold cross validation (k = 3)

150 of 219

Impact of Experimental Design on Clinical Data

Ghasemzadeh et al. (2024)

Voice Dataset

153 females with phonotraumatic vocal hyperfunction and 136 female controls

6 voice measures: F0, cepstral peak prominence, sound pressure level, H1-H2, voicing duration, and phonatory rest duration.

8 distributional characteristics from data on each day: mean, median, standard deviation, interquartile range, 5th percentile, 95th percentile, skewness, and kurtosis

Each participant provided one 48-dimension datapoint to the study.

151 of 219

Impact of Experimental Design on Clinical Data

Ghasemzadeh et al. (2024)

152 of 219

Impact of Experimental Design on Clinical Data

Ghasemzadeh et al. (2024)

153 of 219

Statistical Comparison with Simulated Data

Ghasemzadeh et al. (2024)

154 of 219

Statistical Power of Nested k-fold CV

Ghasemzadeh et al. (2024)

155 of 219

Statistical Power of Nested k-fold CV

Ghasemzadeh et al. (2024)

156 of 219

Statistical Confidence

Ghasemzadeh et al. (2024)

157 of 219

Nested k-fold cross validation

Ghasemzadeh et al. (2024)

Nested k-fold cross validation

158 of 219

Bias in Clinical Speech Analytics

159 of 219

The Presence of Bias in Clinical Datasets

Berisha and Liss (2024)

Published clinical AI models are demographically biased, with 71% of training data coming from the states of California, Massachusetts, and New York (Kaushal et al (2020; JAMA).

This is problematic, and emphasizes the importance of global equity, because health-related changes in speech can be language specific (Garcia et al., 2023; Brain).

Within-language dialect differences are also known to impact error rates in speech machine learning applications (Wassink et al., 2022).

Understanding of sociolinguistic variation is biased towards western, Indo European languages (Adli & Guy, 2022).

160 of 219

The Risks of Bias in Clinical Speech Analytics

Chen et al. (2023)

Bias in training datasets can lead to skewed AI model outputs, affecting clinical decision-making.

Smaller clinical datasets often fail to represent minority groups adequately, amplifying potential biases in model performance.

Disparities in model performance can result in unequal healthcare outcomes, particularly in underrepresented populations.

Enhancing data diversity and inclusion in the data acquisition phase is critical to developing fair and effective clinical AI systems.

161 of 219

Biases in Data: Research Constructs

Mehrabi et al. (2021)

Measurement bias: face validity of features

Omitted variable bias: an important variable is left out

Representation bias: non-representative sampling omits entire subgroups or

outgroups

Sampling bias: non-uniform sampling of subgroups

Aggregation bias: group level averaging that obscures differences in subgroups

162 of 219

User Biases: Social Constructs

Mehrabi et al. (2021)

Historical bias: an already existing bias perpetuates through to the data

Population bias: dataset shift from development to deployment due to sampling

Temporal bias: dataset shift from development to deployment due to social evolution

Self-selection bias: tool works better for people motivated/have the resources to appear in research

Content production bias: linguistic diversity and diversity of speech register between development speech and deployment speech impacts the data

163 of 219

Algorithmic Biases

Mehrabi et al. (2021)

Algorithmic bias: bias is a result of functions, regularizations, or statistically biased estimators.

Evaluation bias: inappropriate or disproportionate benchmarks for evaluation

164 of 219

Case Example: Evaluation Bias

Benway et al. (2024)

165 of 219

Case Example: Evaluation Bias

Benway et al. (2022)

166 of 219

Case Example: Uncovering Population Bias

Benway and Preston (2024)

Prospective clinical validation can help detect biases and distributional shifts.

167 of 219

Clinical Therapeutics Require Clinical Trials

Healthcare and allied health practitioners cannot ethically use AI therapies without efficacy evidence.

Is there evidence that a given AI therapy is valid and reliable for clinical use?

Is there evidence that a given AI therapy is poised to work without bias in real clinical populations with speakers it has never encountered before?

168 of 219

Clinical Therapeutics Require Clinical Trials

Benway and Preston, (2024)

169 of 219

Discrimination in Clinical Speech Analytics

170 of 219

Risks of Discrimination

Obermeyer et al. (2019)

Clinical speech analytic systems may inadvertently perpetuate or amplify existing biases, leading to discriminatory practices in healthcare.

Discriminatory biases can affect diagnosis accuracy, treatment recommendations, and patient trust in healthcare systems and clinical AI.

171 of 219

Direct Discrimination

Obermeyer et al. (2019)

A non-favorable outcome for an individual is because of a protected attribute

Legally-protected attributes will vary by country/jurisdiction but generally might include:

Ethnicity/Race

Religion

Age

Sex/Gender/Sexual Orientation

Disability

172 of 219

Indirect Discrimination

Systemic Discrimination

Obermeyer et al. (2019)

Individual cases are handled according to non-protected attributes that serve as proxy factors for protected attributes

Discrimination against an individual or subgroup is perpetuated by policies and customs.

Obermeyer et al. (2019)

173 of 219

Fairness in Clinical Speech Analytics

174 of 219

Risks of Unfair Clinical Speech Analytics

Binns (2020)

Unfairness in clinical speech analytics refers to systems that offer unequal benefits or burdens to different groups without justified reasons.

Systems perceived as unfair can violate anti-discrimination laws and lead to legal consequences for healthcare providers and clinical tech developers.

175 of 219

Fairness Awareness or Unawareness

Mehrabi et al. (2021)

Fairness through awareness:

Individuals who are similar on a particular task should receive similar predictions across different subgroupings of protected attributes

Fairness through unawareness:

An algorithm is fair as long as protected information is not represented in the decision-making process

However, unawareness raises the opportunity for indirect discrimination that cannot be easily monitored or detected.

176 of 219

Quantifying Fairness

Mehrabi et al. (2021)

Equalized Odds

True positives and false positives rates should be equal across different subgroupings of protected attributes

Equal Opportunity

True positive rates should be equal across different subgroupings of protected attributes

Demographic/Statistical Parity

The likelihood of a positive outcome should be the same across different subgroupings of protected attributes

177 of 219

Quantifying Fairness

Mehrabi et al. (2021)

Treatment Equality

False positives:false negative ratio should be equal across different subgroupings of protected attributes

Test Fairness

For any predicted probability score S, there is an equal probability of belonging to the protected class across different subgroupings of protected attributes

Counterfactual Fairness

A decision is fair if the decision is the same regardless of which protected attribute subgroup the individual belonged to

178 of 219

Case Example: Fairness

Benway et al. (2024a)

A linear mixed model fit on the combined test dataset indicated that neither the fixed effects of age nor sex nor the age-sex interaction significantly influenced classifier performance.

179 of 219

Privacy, Attacks and Misuse

180 of 219

Privacy

Berisha & Liss (2024)

If speech can be used as a biomarker, and the recorded speech of individuals is public, then individuals are effectively leaving a publicly accessible, longitudinal trail of data that can be used to make claims about one’s personal health or emotional status.

Such claims could impact health privacy of private individuals or, for example, support the spread of disinformation in the case of political figures.

Privacy is therefore paramount.

181 of 219

Data Breaches

Liu et al. (2015)

Clinical speech analytic databases themselves could be compromised, potentially exposing individually identifiable protected health information.

Individual jurisdictions will have regulations related to data privacy and data breaches. For example:

  • Health Insurance Portability and Accountability Act (HIPAA) in USA
  • General Data Protection Regulation (GDPR) in Europe
  • Personal Data Protection Act (PDPA) in Singapore

182 of 219

Attacks on Algorithms

Finlayson et al. (2019)

Adversarial attacks pose a significant risk to clinical machine learning models by exploiting vulnerabilities that can alter outputs.

Potential attacks include data poisoning, model inversion, and adversarial examples, each posing unique threats to model reliability and data security.

Compromised models could lead to misdiagnoses, inappropriate treatment suggestions, patient harm, and erosion of trust in clinical speech analytics.

Ensuring robust cybersecurity measures for clinical machine learning models is vital due to the potential for attacks to compromise patient care.

183 of 219

Malicious Misuse

Price et al. (2019)

Analytic systems can be manipulated to generate false claims or diagnoses, especially when used without robust oversight or ethical guidelines.

Using personal data without consent to develop clinical analytics can violate privacy rights and lead to ethical and legal repercussions.

Misrepresentation of data can harm careers, influence treatment plans, and erode trust in healthcare systems.

184 of 219

Misuse through Premature Deployment

Berisha and Liss (2024); Yawer et al. (2023)

Insufficient validation of clinical speech analytics runs the risk of overoptimistic lab testing and clinically inaccurate output in production.

Example: the Cigna StressWaves Test

  • Deployed without public efficacy data
  • Poor test-retest reliability (Figure 1 of Yawer et al., right) in independent validation with 60 adults.
  • Poor agreement with an existing measure of stress.

Models that are deployed prematurely may cause harm to patients through inappropriate treatment, wasted resources, increased anxiety, or false reassurance.

185 of 219

Equitable Access

186 of 219

Environmental Impact of Computation

Bender, Gebru, et al. (2021)

The training of large AI models has an environmental impact due to the energy resources required to support computation.

This environmental impact is most likely to impact climate-vulnerable communities.

Because climate-vulnerable communities are also associated with social factors like lower income and minority languages, the communities bearing the brunt of AI environmental impact are the least likely to benefit from AI advancements.

187 of 219

AI Dataset Microtask Economies

Casilli (2024), Rowe (2023), Munn (2024)

AI-ready datasets must be curated, labeled, and validated.

Who is labeling these data and how fairly are workers being paid?

There is scientific and journalistic evidence that low-wage workers are involved in these tasks.

These low-wage workers also come from the least likely communities to benefit from the technology they are helping to train.

188 of 219

Patient and Public Involvement

189 of 219

Community Based Participatory Research

Active involvement of patients and the public in clinical speech analytics development ensures the technologies address real-world needs and ethical considerations.

Community involvement in clinical speech analytics development can improve the relevance and acceptance of technologies, leading to better adoption of the technology and better trust in the technology.

190 of 219

AI 4 Social Good Framework

Bondi et al. (2021)

Are the communities being designed “for” rather than having a seat at the table?

Does the development of clinical speech analytics mirror the needs and preferences of specific communities with clinical speech needs?

Is this engagement reciprocal and part of the process from the beginning?

How was the community involved in the design, development, and deployment of clinical AI?

191 of 219

Guiding Principles for Participatory AI4SG

Bondi et al. (2021)

How are impacted communities identified and how are they represented in this process? Who represents historically marginalized groups?

How are plans laid out for maintaining and sustaining this work in the long-term, and how would the partnership be ended?

What kind of compensation are interested parties receiving for their time and input? Does that compensation respect them as partners in the process?

How are viewpoints from many different groups understood and incorporated?

What specific concerns are raised during deliberation with interested parties, and how are these addressed?

192 of 219

Determining Capability Sets

Bondi et al. (2021)

What capabilities to interested parties want to achieve?

What are the priorities for the most vulnerable in the community?

Will any of the priorities pay dividends in other important areas (occupational, social, etc)?

Do the values of the project match the values of the community?

193 of 219

Evaluating AI4SG

Bondi et al. (2021)

How does the clinical speech analytics system affect the capabilities prioritized by the community?

Are any capabilities negatively affected by the clinical speech analytic system?

What should the role of the AI researcher be?

194 of 219

Case Study (KCL)

Automating coding of Five Minute Speech Samples

E-Risk Study

195 of 219

Case Study (KCL)

Automating coding of Five Minute Speech Samples

Mothers’ speech samples

When the children were 10 years old, mothers were asked to speak about them to the researcher

“For the next 5 minutes, I would like you to describe [child] to me; what is [child] like?”

This speech was recorded

196 of 219

Case Study (KCL)

Automating coding of Five Minute Speech Samples

Coding mothers’ emotions towards child

197 of 219

Case Study (KCL)

Automating coding of Five Minute Speech Samples

Predicting later mental health issues

198 of 219

Case Study (KCL)

Automating coding of Five Minute Speech Samples

Conducted PPI sessions with parents and young carers

Sessions included the brief outline of the study

Had employed a near-future science fiction who wrote to dystopian stories based our research

These were also read to participants in our workshops

199 of 219

Case Study (KCL)

Automating coding of Five Minute Speech Samples

Parents Workshop

Data & privacy

Will the AI of the future be able to de-anonymise us? What happens to anonymity then?

Consent as a continuous process

Research directions may change over time

The assessment

Parent-blaming attitudes

How can you capture parenting across the lifespan?

Lack of expertise of, and consistency between, clinicians

Discrimination against neurodivergent parents

Key contextual issue: waitlists and lack of NHS support available, health inequities

Predicting mental health outcomes

‘Gaming the system’

Self-fulfilling prophecies

The AI output: reductionism of interpreting health statistics

Limitations of predicting health outcomes and of diagnosis

Issues with the method more than use of AI

200 of 219

Case Study (KCL)

Automating coding of Five Minute Speech Samples

Parents Workshop

Data & privacy

Will the AI of the future be able to de-anonymise us? What happens to anonymity then?

Consent as a continuous process

Research directions may change over time

The assessment

Parent-blaming attitudes

How can you capture parenting across the lifespan?

Lack of expertise of, and consistency between, clinicians

Discrimination against neurodivergent parents

Key contextual issue: waitlists and lack of NHS support available, health inequities

Predicting mental health outcomes

‘Gaming the system’

Self-fulfilling prophecies

The AI output: reductionism of interpreting health statistics

Limitations of predicting health outcomes and of diagnosis

Issues with the method more than use of AI

201 of 219

Case Study (KCL)

Automating coding of Five Minute Speech Samples

Young Carers Workshop

What if parents don’t tell the truth?

Gaming the system’

Lying

May not know their child

Not all cultures discuss mental health

Preference for doctors over AI

Can respond to emotion

More human

Doctors have specialised training

Can take context & tone into account

Risks of using the tool

Self-fulfilling prophecy

Parents treat children differently

202 of 219

Case Study (KCL)

Automating coding of Five Minute Speech Samples

Common Themes

Inability to account for other life events (positive or negative)

Dangers of ‘gaming the system’ and/or self-fulfilling prophecies

Concerns about money & access

Potential positive applications of AI:

  • Synthesising up-to-date research and information for patients on clinicians e.g., on autism
  • Reducing waiting times. Could AI be used to streamline dyslexia or autism assessments – e.g., Part 1 of a two-part assessment?
  • Screening people to flag very high-risk cases

203 of 219

Time-Ordered Actions Toward Responsible Clinical Speech Analytics

204 of 219

Establish Participatory Partnerships

Early in the process

Enlist community members ahead of the development of clinical speech analytics to inform the technological design and technological fidelity for:

  • Motivation for tool development
  • Acceptability of the use case/capabilities to be enhanced by the tool
  • Convenience in use
  • Accessibility of the tool

This is most ernest when done during the proposal writing process.

205 of 219

Design Speech Tasks

Factors for consideration

  1. The goal of model? �- e.g. classification, diagnosis, changes, etc.�
  2. What aspects of speech production could be impacted? �- e.g. Conceptualization, Formulation, Articulation�
  3. What are the participant-related limitation? �-e.g. Vision, hearing, cognition, etc.�
  4. How are the spoken responded be processed?�-e.g. speech recognition, spectral-temporal measures, etc.

206 of 219

Minimal Speech Health Data Set

Group

Feature

Description

unit

Timing/Fluency

Phonation Ratio

Phonation time divided by duration

-

Speaking Rate

Number of syllables divided by duration

syl/sec

Articulation Rate

Number of linguistic units divided by phonation time

syl/sec

Pause Rate

Number of Pause (> 300ms) divided by duration

pause/sec

Mean Pause Duration

Mean duration of all pauses (> 300ms)

sec

Respiratory

Intensity

Root-mean-square amplitude of the provided audio signal

dB

Intensity

Ratio of max to min intensity

-

Phontation

Pitch

Rate of vibration of vocal folds

Hz

Pitch Sigma

Standard deviation of pitch in semitones

Semitone

Harmonic to Noise Ratio

Degree of acoustic periodicity

dB

Cepstral Peak Prominence

Stability of vocal fold vibration

dB

Spectral Slope

Ratio of energy in a spectra between 10-1000Hz over 1000-4000Hz

-

Spectral Tilt

Linear slope of energy distribution between 100-5000Hz

-

Articulatory

Formant Frequency (1st and 2nd)

Centre frequency of vocal tract resonant peaks

Hz

Formant Bandwidths (1st and 2nd)

Bandwidth of vocal tract resonant peak

Hz

Spectral Gravity

Spectral centroid

Hz

Spectral Standard Deviation

Spread of frequencies around the centroid

Hz

Spectral Skewness

Symmetry of frequencies around the centroid

Hz

Spectral Kurtosis

Flatness of the spectrum the centroid

Hz

207 of 219

Select Speech Measures

Factors for consideration

  1. Clinical constructs that link to the disease�
  2. Types of data available, �e.g. speech recordings, text, etc.�
  3. Types of ground truth measures available,�e.g. perceptual ratings, valid articulatory measures, temporal measures, etc.�
  4. Kinds of validation need to be performed�e.g. test-retest reliability, inter-rater reliability, construct validity, etc.

208 of 219

Obtain Informed Consent/Assent

Beauchamp & Childress (2019)

Informed consent is a cornerstone of ethical research and clinical practice, ensuring that participants are fully aware of the nature and risks of the data collection and analysis process.

Obtaining consent is not only an ethical obligation but also a legal requirement under frameworks like GDPR in Europe and HIPAA in the U.S., protecting patient rights and data privacy.

Effective consent processes involve clear communication about how speech data will be used, stored, and shared, ensuring participants' understanding and agreement.

209 of 219

Acquire Speech Data

Recording Device and Environment Bias

  • Recording devices can cause subtle changes in the measurement
    • Same holds true for data storage (w.r.t compression/filtering)
  • Properly select devices and test with participant group in mind
    • Reading task for children? / Screen-Task for color-blind people
  • Log all peculiarities/incidents during data acquisition
    • Possible data anomalies can otherwise not be traced later
    • If possible do a video recording (just for observation, not for analyses)
  • Environment can influence data acquisition:

whisperroom.com Ingo Siegert, Uni Magdeburg Univ. Maryland Medical Center

210 of 219

Transparently Report Dataset and Biases

Datasheets for Datasets Model Cards for Model Reporting Replicability Cards

Gebru et al. (2018)

Mitchell et al. (2018)

Kapoor & Narayanan (2023)

211 of 219

Develop Generalizable Models

Factors for consideration

  1. Amount of speech data available for training�
  2. Data augmentation that does not affect characteristics of clinical speech �e.g. speed augmentation is not suitable when speech rates are analysed�
  3. Number of speech measures selected to construct input features�
  4. Type of clinical validations need to be performed�e.g. relationships with other established clinical labels, reliability in longitudinal tracking, etc.

212 of 219

Demonstrate Interpretability of Model/Measures

  1. The type of ML model being used�
  2. What the model parameters can explain�e.g. feature importance, decision making mechanism �
  3. What kind of visualization can be performed to illustrate the speech measures / feature importance / predicted label?�e.g. severity of speech vs. measure, SHAP values, longitudinal trajectory…

213 of 219

Use Robust Experimental Designs

Ghasemzadeh et al. (2024)

Increase the chances that speech measures and hyperparameters are optimized.

Reduce the chances that reported model performance will change with a different dataset split.

Link to paper ->

214 of 219

Justify Model Improvements

Benway et al. (2024a)

Are model improvements different from chance?

Identifying this helps streamline the development of parsimonious models that do not waste computational or environmental resources.

215 of 219

Search for evidence of biases

What biases are present?

Techniques like hypothesis testing and linear mixed modeling can help identify biases in algorithms before clinical deployment.

Recall previous Case Example from �Benway et al. (2024a)

216 of 219

Clinical Therapeutics Require Clinical Trials

Prospective clinical trials for therapeutic technologies will help identify sources of bias.

Multiple clinical trials will happen in sequence, each focusing on a unique factor of deployment.

217 of 219

Post-deployment Monitoring

Reanalysis of Benway et al. (2024b)

Post-deployment monitoring and continuous model updating are essential to maintain fairness over time, addressing new data or changing population dynamics.

Lab Data Prospectively Collected Data

Balanced Accuracy

Count

218 of 219

Measuring AI Impact

Stahl et al. (2023)

A repository of AI Impact Assessments associated with the systematic review of Stahl and colleagues is available at this QR code →

https://www.zotero.org/groups/4042832/ai_impact_assessments

219 of 219

Remarks and Questions