1 of 37

Zero Resource Learning �from Spoken Audio

David Harwath1,2 Timothy J. Hazen1 James Glass2

1MIT Lincoln Laboratory

2MIT Computer Science and Artificial Intelligence Laboratory

This work was sponsored by the Department of Defense under Air Force Contract FA8721-05-C-0002. Opinions, interpretations,

conclusions, and recommendations are those of the authors and are not necessarily endorsed by the United States Government.

MIT Lincoln Laboratory

MIT Lincoln Laboratory

Zero Resource Analysis-1

SANE Workshop 10/24/2012

2 of 37

Outline

  • Introduction
  • Acoustic Pattern Discovery
  • Latent Modeling of Topics and Words
  • Summary and Future Directions

MIT Lincoln Laboratory

Zero Resource Analysis-2

SANE Workshop 10/24/2012

3 of 37

Introduction

  • Current state-of-the-art speech recognition relies on:
    • Large transcribed audio corpora for training of acoustic models
    • Prior knowledge of the words and their phonetic pronunciations
    • Large text corpora for training of language models

  • Zero resource learning from spoken audio data:
    • No transcriptions, annotations or prior knowledge of language
    • Unsupervised learning from spoken audio only

  • Ultimate goal is completely unsupervised learning of:
    • Acoustic phonetic units
    • Sub-word structures (syllables, morphs, etc)
    • Lexical dictionaries (words and pronunciations)
    • Higher level information (syntax and semantics)

  • How do we get there?

MIT Lincoln Laboratory

Zero Resource Analysis-3

SANE Workshop 10/24/2012

4 of 37

Previous Work

  • Acoustic pattern discovery:
    • A. Park and J. Glass, “Unsupervised pattern discovery in speech,” IEEE TASLP, 2008.
    • A. Jansen, et al, “Towards spoken term discovery at scale with zero resources,” Proc. Interspeech, 2010.�
  • Unsupervised learning of acoustic-phonetic inventories:
    • H. Gish, et al, “Unsupervised training of an HMM-based speech recognition system for topic classification,” Proc. Interspeech, 2009.
    • C. Lee and J. Glass, "A nonparametric Bayesian approach to acoustic model discovery," Proc. ACL, 2012.

  • Learning words from phoneme strings:
    • M. Johnson, “Unsupervised word segmentation for Sesotho using adaptor grammars,” Proc. ACL SIG on Computational Morphology and Phonology, 2008.
    • S. Goldwater, et al, “A Bayesian framework for word segmentation: exploring the effects of context,” Cognition, 2009.

  • Higher level learning:
    • M. Drezde, et al, “NLP on spoken documents without ASR,” EMNLP, 2010.
    • T. Hazen, et al, “Topic modeling for spoken documents using only phonetic information,” Proc. ASRU, 2011.

MIT Lincoln Laboratory

Zero Resource Analysis-4

SANE Workshop 10/24/2012

5 of 37

Near-Term Application Areas

  • Useful near-term applications are possible even if unsupervised learning from speech is far from perfect

  • Some useful applications that are feasible in the short term are:
    • Query-by-example spoken term detection
    • Query-by-example audio document retrieval
    • Spoken document clustering
    • Spoken document summarization
    • Spoken corpus summarization

MIT Lincoln Laboratory

Zero Resource Analysis-5

SANE Workshop 10/24/2012

6 of 37

Outline

  • Introduction
  • Acoustic Pattern Discovery
  • Latent Modeling of Topics and Words
  • Summary and Future Directions

MIT Lincoln Laboratory

Zero Resource Analysis-6

SANE Workshop 10/24/2012

7 of 37

Segmental Dynamic Time Warping

(Park and Glass, 2008)

MIT Lincoln Laboratory

Zero Resource Analysis-7

SANE Workshop 10/24/2012

8 of 37

Unsupervised Pattern Discovery

  • Unsupervised S-DTW pattern discovery was a success, but…
    • Direct comparison of MFCCs used in initial Park and Glass work
    • Speaker dependency was a limiting issue

  • Model-based features can help reduce speaker dependence
    • Self-Organizing Units (Gish, et al, Interspeech, 2006)
    • GMM-UBM over MFCCs (Zhang and Glass, ICASSP, 2010)
    • Deep Boltzmann Machine posteriorgrams (Zhang and Glass, ICASSP, 2012)
    • Bayesian nonparametric segmentation (Lee and Glass, ACL, 2012)

MIT Lincoln Laboratory

Zero Resource Analysis-8

SANE Workshop 10/24/2012

9 of 37

SOU Posteriorgrams

  • This work uses BBN’s SOU-based feature representation

  • SOU recognizer generates lattices of SOU hypotheses

  • SOU lattice converted to frame-based posteriorgram

SOU

Index

Frame Index

MIT Lincoln Laboratory

Zero Resource Analysis-9

SANE Workshop 10/24/2012

10 of 37

Posteriorgram Similarity Matrix

  • Example utterance with repeated word:

“…education computers and education…”

  • Figure shows self similarity matrix using posteriorgrams instead of MFCCs

MIT Lincoln Laboratory

Zero Resource Analysis-10

SANE Workshop 10/24/2012

11 of 37

Segmental DTW Pattern Discovery

  • S-DTW search can be slow and is O(N2) over N utterances
  • Fast approximate algorithms are needed:
    • Image filtering then Hough transform (Jansen, et al, Interspeech, 2010)
    • LSH indexing of utterance data (Jansen & Van Durme, Interspeech, 2012)

Original Similarity Matrix

Filtered Similarity Matrix

MIT Lincoln Laboratory

Zero Resource Analysis-11

SANE Workshop 10/24/2012

12 of 37

Document-Link Structure

MIT Lincoln Laboratory

Zero Resource Analysis-12

SANE Workshop 10/24/2012

13 of 37

Document-Link Structure

MIT Lincoln Laboratory

Zero Resource Analysis-13

SANE Workshop 10/24/2012

14 of 37

Document-Link Structure

MIT Lincoln Laboratory

Zero Resource Analysis-14

SANE Workshop 10/24/2012

15 of 37

Experimental Conditions

  • Goal: Completely unsupervised analysis of an audio corpus

  • Experimental corpus: Fisher English
    • Ten-minute phone calls recorded at 8 kHz
    • Conversational telephone speech, wide variety of speakers
    • Topic-prompted conversations across 40 prompts
    • Experiments run on subsets of varying sizes (60 to 389 calls)

  • Posteriorgram feature representation for all utterances
    • Created using BBN Self-Organizing Unit (SOU) system
    • Unsupervised acoustic model training
    • Trained on 60hr independent set of Fisher data

MIT Lincoln Laboratory

Zero Resource Analysis-15

SANE Workshop 10/24/2012

16 of 37

Corpus Analysis & Summarization

  • Example application: Automatic summarization of an audio corpus
  • PLSA-based summarization of Fisher corpus
    • Example below uses text transcripts of 5850 Fisher conversations
    • Good correspondence between PLSA topics and Fisher Topics
    • Can we do something similar using only the audio

(from Hazen & Richardson, SLT Workshop, 2012)

Summaries of Ranked PLSA Topics Matching Fisher Topic

  1. dog, cats, pet, animals, fish, bird, feed, puppy, cute, cage Pets (.900)
  2. minimum wage, pay, jobs, five fifteen and hour, paid, making, tips Minimum Wage (.864)
  3. sports, football, basketball, baseball, game, team, watching, hockey Sports on TV (.849)
  4. airport security, plane, fly, september eleventh, flight, airplane, flown Airport Security (.523)

September 11th (.351)

  • show, watched, survivor, reality t.v.,reality shows, bachelor Reality TV (.901)

. . .

. . .

. . .

  1. back and change, back in time and change, regret, time travel Time Travel (.767)

. . .

. . .

. . .

  1. drugs, drug test, company, hair, marijuana, invasion of privacy Drug Testing (.516)
  2. linguistics, o.k., study, phone number, u. penn, speech recognition None

MIT Lincoln Laboratory

Zero Resource Analysis-16

SANE Workshop 10/24/2012

17 of 37

Topics and Match Intervals

  • Documents contain intervals identified by S-DTW
  • Characterize documents and by the outgoing links associated with all intervals in the document
  • Represent documents as bags-of-links vectors

Link

Document

MIT Lincoln Laboratory

Zero Resource Analysis-17

SANE Workshop 10/24/2012

18 of 37

Document Link Structure

  • Documents can be compared using cosine similarity of link-based feature vectors
  • Document similarity matrix can be viewed as a connected graph between documents
  • Example shows graph over 389 Fisher conversations

MIT Lincoln Laboratory

Zero Resource Analysis-18

SANE Workshop 10/24/2012

19 of 37

Document Link Structure

  • Highlighted conversation on “Minimum Wage” topic
  • Strongly connected to other “Minimum Wage” conversations

MIT Lincoln Laboratory

Zero Resource Analysis-19

SANE Workshop 10/24/2012

20 of 37

Document Links

  • Highlighted conversation on “Personal Habits”
  • Strongly connected to some “Personal Habits” conversations
  • …but also connected to conversations on “Smoking” and “Food”

What can we learn about the topical content of the data from this graph?

MIT Lincoln Laboratory

Zero Resource Analysis-20

SANE Workshop 10/24/2012

21 of 37

Outline

  • Introduction
  • Acoustic Pattern Discovery
  • Latent Modeling of Topics and Words
  • Summary and Future Directions

MIT Lincoln Laboratory

Zero Resource Analysis-21

SANE Workshop 10/24/2012

22 of 37

A Graphical Model

  • A latent model for discovering topics from the link features of a graph.
  • Can be learned using standard probabilistic latent semantic analysis (PLSA) training techniques

= Observed Variable

= Latent Variable

Latent topic

MIT Lincoln Laboratory

Zero Resource Analysis-22

SANE Workshop 10/24/2012

23 of 37

Summarizing The Topics

  • Once model is learned, audio intervals (i.e., nodes of graph) can be used to summarize the latent topics.
  • For each latent topic, extract top N intervals of audio ranked by a weighted pointwise mutual information measure

MIT Lincoln Laboratory

Zero Resource Analysis-23

SANE Workshop 10/24/2012

24 of 37

Pilot Study

  • Pilot study performed to assess feasibility of approach
  • Experimental data:
    • 60 Fisher conversations
    • 10 conversations from 6 different topics
  • Automatic summarization of pilot data set
    • Latent model trained using 6 topics
    • Topic summarized by ten most representative audio intervals
  • In this talk, intervals are represented by the reference words matching the time interval
    • In practice word identities are unknown and the user listens to the audio summary
  • Compute to map latent topics to true topics

MIT Lincoln Laboratory

Zero Resource Analysis-24

SANE Workshop 10/24/2012

25 of 37

Results

Data Set Summary

  • Topic 1 (Minimum Wage, 99.7%) : minimum_wage minimum_wage minimum_wage minimum_wage minimum_wage minimum_wage minimum_wage minimum_wage minimum_wage minimum_wage

  • Topic 2 (Education, 99.9%) : think_computers computers of_computers computer computer computers computers computers computers of_computers

  • Topic 3 (Illness, 37.3%, Corporate Conduct 32.2%) : exactly um country exactly countries um countries exactly exactly

  • Topic 4 (Holidays 83.1%) : holidays holiday holiday_is holidays the_holidays holidays holiday the_holidays a_holiday own_holiday

  • Topic 5 (Anonymous Benefactor 55.3%) : money situations situations the_more_money_you friend educational four_years situation situations make_money

  • Topic 6 (Anonymous Benefactor 45.3%, Corporate Conduct 55.2%) : weather_friends friends friends friends friends some_friends friends kind_of_friends to_happen major_you_know

MIT Lincoln Laboratory

Zero Resource Analysis-25

SANE Workshop 10/24/2012

26 of 37

A Second Graphical Model

  • A doubly stochastic latent model for discovering both words and topics from the graph link structure
  • Can be learned using an Expectation Maximization (EM) style algorithm

Document

Interval

Links

MIT Lincoln Laboratory

Zero Resource Analysis-26

SANE Workshop 10/24/2012

27 of 37

EM Update Equations

M-Step:

E-Step:

MIT Lincoln Laboratory

Zero Resource Analysis-27

SANE Workshop 10/24/2012

28 of 37

Summarizing The Topics

  • For each latent topic, choose top N latent pseudo-words ranked by a weighted pointwise mutual information measure

  • For each latent pseudo-word, extract the interval of audio which maximizes

MIT Lincoln Laboratory

Zero Resource Analysis-28

SANE Workshop 10/24/2012

29 of 37

Results

Topic 1 (Minimum Wage 76.4%) : minimum wage, you, the economy, money, you know, actually, yeah I, right, minimum wage jobs, it’ll be interesting, you know people, economy

Topic 2 (Education 85.5%) : think computers, something that’s, more and, computers, education, you ah, technical ah, know the computerized, ah and, conditioning, and, information

Topic 3 (Corporate Conduct 46.9%, Illness 42.7%) : sicker, c.e.o., stock market, exactly, without the, country, every sick, this guy, in uh in, that um, greedy, enron

Topic 4 (Holidays 76.9%) : I really like, holidays, own holiday, holiday, equality, favorite holiday, you like, considerate, and, the key, keys, this

Topic 5 (Anonymous Benefactor 76.4%) : uh-huh, friend, ‘em twenty, know, maybe ah, I’ve seen it done, that and all, uh, best friend, every day, increased, lazier and

Topic 0 (Anonymous Benefactor 51.8%) : people who, weather friends, situations, no I, the lottery, and, don’t even know who, benefactor, economy, very you know, now um, to happen

Text transcripts of extracted audio intervals

MIT Lincoln Laboratory

Zero Resource Analysis-29

SANE Workshop 10/24/2012

30 of 37

Results

Topic 1 (Minimum Wage 76.4%) : minimum wage, you, the economy, money, you know, actually, yeah I, right, minimum wage jobs, it’ll be interesting, you know people, economy

Topic 2 (Education 85.5%) : think computers, something that’s, more and, computers, education, you ah, technical ah, know the computerized, ah and, conditioning, and, information

Topic 3 (Corporate Conduct 46.9%, Illness 42.7%) : sicker, c.e.o., stock market, exactly, without the, country, every sick, this guy, in uh in, that um, greedy, enron

Topic 4 (Holidays 76.9%) : I really like, holidays, own holiday, holiday, equality, favorite holiday, you like, considerate, and, the key, keys, this

Topic 5 (Anonymous Benefactor 76.4%) : uh-huh, friend, ‘em twenty, know, maybe ah, I’ve seen it done, that and all, uh, best friend, every day, increased, lazier and

Topic 0 (Anonymous Benefactor 51.8%) : people who, weather friends, situations, no I, the lottery, and, don’t even know who, benefactor, economy, very you know, now um, to happen

Text transcripts of extracted audio intervals

MIT Lincoln Laboratory

Zero Resource Analysis-30

SANE Workshop 10/24/2012

31 of 37

Results

Topic 1 (Minimum Wage 76.4%) : minimum wage, you, the economy, money, you know, actually, yeah I, right, minimum wage jobs, it’ll be interesting, you know people, economy

Topic 2 (Education 85.5%) : think computers, something that’s, more and, computers, education, you ah, technical ah, know the computerized, ah and, conditioning, and, information

Topic 3 (Corporate Conduct 46.9%, Illness 42.7%) : sicker, c.e.o., stock market, exactly, without the, country, every sick, this guy, in uh in, that um, greedy, enron

Topic 4 (Holidays 76.9%) : I really like, holidays, own holiday, holiday, equality, favorite holiday, you like, considerate, and, the key, keys, this

Topic 5 (Anonymous Benefactor 76.4%) : uh-huh, friend, ‘em twenty, know, maybe ah, I’ve seen it done, that and all, uh, best friend, every day, increased, lazier and

Topic 0 (Anonymous Benefactor 51.8%) : people who, weather friends, situations, no I, the lottery, and, don’t even know who, benefactor, economy, very you know, now um, to happen

Text transcripts of extracted audio intervals

MIT Lincoln Laboratory

Zero Resource Analysis-31

SANE Workshop 10/24/2012

32 of 37

Information Theoretic Evaluation Metrics

I(Z;T)

Mutual Information

H(T|Z)

Missed Information

H(Z|T)

False Information

H(T)

Entropy of True Topics

H(Z)

Entropy of Latent Topics

H(Z|T)+H(T|Z)

Total Erroneous Information

Erroneous Information Ratio

Normalized Mutual Information

MIT Lincoln Laboratory

Zero Resource Analysis-32

SANE Workshop 10/24/2012

33 of 37

Topic Discovery Evaluation

  • Experimental data:
    • 60 Fisher conversations
    • 10 conversations from 6 different topics
  • Evaluation assessing quality of topic discovery:
    • Normalized mutual information (NMI)
    • Erroneous information ratio (ERI)
  • Baseline systems:
    • PLSA based on text transcripts
    • Hard agglomerative clustering using link structure

Model

NMI

EIR

Text PLSA

0.895

0.210

Hard Clustering

0.529

0.880

Model 1

0.529

0.922

Model 2

0.541

0.909

MIT Lincoln Laboratory

Zero Resource Analysis-33

SANE Workshop 10/24/2012

34 of 37

Outline

  • Introduction
  • Acoustic Pattern Discovery
  • Latent Modeling of Topics and Words
  • Summary and Future Directions

MIT Lincoln Laboratory

Zero Resource Analysis-34

SANE Workshop 10/24/2012

35 of 37

Discussion

  • Successes
    • Achieves completely unsupervised joint learning of pseudo-words and topics from nothing but telephone audio
    • Strong overlap with true topic labels
    • Many topically relevant words in summaries

  • Room for improvement
    • How to choose a good initialization
    • Precision/Recall of match intervals
    • How to deal with speaker dependent stop words

MIT Lincoln Laboratory

Zero Resource Analysis-35

SANE Workshop 10/24/2012

36 of 37

Future Directions

  • Optimizations enabling the use of larger data sets
    • LSH-based dotplot computation on GPU architecture
    • Parallelize EM estimation on GPU architecture

  • New unsupervised acoustic models (Lee and Glass, 2012), (Zhang and Glass, 2012)

  • Lexicon learning for low-supervision ASR systems

MIT Lincoln Laboratory

Zero Resource Analysis-36

SANE Workshop 10/24/2012

37 of 37

Future Directions

  • Bayesian Reformulation

  • Integrate with Lee & Glass system to jointly infer phones, words, and topics all from raw acoustics

MIT Lincoln Laboratory

Zero Resource Analysis-37

SANE Workshop 10/24/2012