1 of 20

Legal context biases listeners toward hearing voice pairs as more similar

Vincent Hughes, Carmen Llamas, and Thomas Kettig

NWAV50

San Jose, CA

October 2022

{vincent.hughes|carmen.llamas|thomas.kettig}@york.ac.uk

2 of 20

https://sites.google.com/york.ac.uk/humans-machines/

Joe Cutting

Daniel Slawson

Humans & Machines: Novel methods for assessing speaker recognition performance

(AH/T012978/1)

2

3 of 20

Speaker comparison

  • Forensic phoneticians express the strength of voice evidence via a Likelihood Ratio (LR):

Similarity – “Prosecution hypothesis”

Typicality – “Defense hypothesis”

 

 

 

 

3

4 of 20

Speaker matching

  • ASR systems compare these measures using available data

We want to test how this might work with human listeners

 

Variable A

4

5 of 20

RQs for project:

  1. Human performance vs. ASR performance

(2) Effect of listener group/biases on human performance

(3) Effect of sample type on human performance/ASR performance

(4) Effect of contextual information on human performance

5

6 of 20

Method: Immersive jury game

  • Participants encounter pairs of sound samples
    • Typicality rating elicited for first stimulus
    • Similarity and sameness ratings elicited after second stimulus presented
    • First gave us self-declared accent familiarity ratings / demographic info

  • Two levels discussed here
    • Tutorial level without context
    • Jury/legal context

6

7 of 20

Stimuli

  • Samples from Standard Southern British English (DyViS) &

Newcastle and Middlesbrough men (TUULS)

  • Forensically-realistic quality
      • First sample = landline phone quality (actual or noise/filter added)
      • Second sample = HQ, taken from mock police interviews
      • Short (10-11s)
  • Tagged for accentedness and voice quality
  • Normed for guilt/suspiciousness of sample content

7

8 of 20

Stimuli

  • 120 pairs created
    • 30 SSBE pairs (15 DS, 15 SS)
    • 30 Middlesbrough pairs (15 DS, 15 SS)
    • 30 Newcastle pairs (15 DS, 15 SS)
    • 30 mixed Middlesbrough/Newcastle pairs (30 DS)
  • Distributed into 15 blocks containing 8 pairs each (5 DS, 3 SS)
  • Stimuli within blocks internally randomized
  • 1 block presented per level, counterbalanced
  • 896 participants

8

9 of 20

9

10 of 20

10

11 of 20

11

12 of 20

12

13 of 20

13

14 of 20

14

15 of 20

Sameness ratings vs. similarity ratings

15

16 of 20

Similarity ratings:

Level effects by stimulus accent

Different-speaker pairs

Same-speaker pairs

High accentedness

Mid/Nc cross-accent pairs

Low accentedness

16

17 of 20

Typicality ratings by Northern accent familiarity and accentendess

17

18 of 20

Conclusions

  • On both sides of the similarity/typicality likelihood ratio equation, potential consequences for jury decisions and forensic applications
  • Potential of game-based immersive experimental methods vs. ‘vanilla’ Qualtrics-type ones usually used

18

19 of 20

Thank you!

Questions?

{vincent.hughes|carmen.llamas|thomas.kettig}@york.ac.uk

@VinceH_Forensic @TKettig

19

20 of 20

20