1 of 51

Machine Learning for

Natural Language Processing

Fundamentals 2

2 of 51

demo

explosion.ai/demos/displacy

spaCy

3 of 51

How many words in a language?

4 of 51

Size of a language

Words in the dictionary

10^5

Unique tokens in full corpus

10^6

5 of 51

What is the distribution of words?

How do the frequencies vary?

6 of 51

Zipf’s law

“The frequency of any word is inversely proportional to its rank in the frequency table.”

7 of 51

Zipf’s law

“The frequency of any word is inversely proportional to its rank in the frequency table.”

8 of 51

What are corpora?

9 of 51

Corpora

corpus > document > paragraph > sentence > phrase > word > char

10 of 51

Corpora

corpus > document > paragraph > sentence > phrase > word > char

? > query > ?

11 of 51

Corpora

corpus > document > paragraph > sentence > phrase > word > char

? > tweet > ?

12 of 51

Annotated corpus

Human labour

13 of 51

Parallel corpora

The Rosetta stone

14 of 51

Tokenisation

Sentences and words

15 of 51

Morphology

Stemmers and lemmatisers

16 of 51

Grammar

Syntax and agreement

17 of 51

Syntax

Parsing

18 of 51

Canonicalisation

A double-edged sword

19 of 51

When to preprocess?

20 of 51

Language Model

Probability distribution over sequences of words

21 of 51

n-grams

Context

22 of 51

n-grams:

word level

Text

“Is the color of milk and fresh snow, the color produced by the combination of all the colors of the visible spectrum.”

n=2

“Is the”, “the color”, “color of”, “of milk”, “milk and”, “and fresh”, “fresh snow”, “snow ,”, “, the”, “the color”...

23 of 51

n-grams:

char level

Text

“Is the color of milk and fresh snow, the color produced by the combination of all the colors of the visible spectrum.”

n=3

“@Is”, “Is_”, “s_t”, “_th”, “the”, “he_”, “e_c”, “_co”, “col”, “olo”, “lor”, “or_”, “r_o”, “_of”, “of_”, “f_m”...

24 of 51

demo

github.com/SignalN/language

language.ngrams

25 of 51

What is the difference between n-grams and a language model?

26 of 51

NER

Named entities

27 of 51

demo

cloud.google.com/natural-language

Google Cloud Natural Language API

28 of 51

Co-references

Anaphora, pronouns...

29 of 51

Polysemy

It’s not what you think it means.

30 of 51

How do we measure success?

31 of 51

BLEU

Exact match

Fuzzy

32 of 51

demo

rajpurkar.github.io/SQuAD-explorer

SQuAD

33 of 51

Beyond English

34 of 51

What % is English?

% of users? % of content?

35 of 51

How many languages does the average person speak?

36 of 51

1.5

37 of 51

Language codes

ISO

38 of 51

Script

Alphabets+

39 of 51

Is English unique?

How is it easier? How is it harder?

40 of 51

Is Armenian unique?

How is it easier? How is it harder?

41 of 51

Linguistic Typology

42 of 51

Multilingual problems

Identification, translation, transliteration

43 of 51

Multimodal problems

language + x

44 of 51

text + struct

=>

numeric

Ali Baba, YouTube, Amazon

product recommendations

45 of 51

image

=>

struct | text

cs.stanford.edu/.../clevr

CLEVR

46 of 51

47 of 51

image

=>

text

@picdescbot

image captioning

48 of 51

For next time...

49 of 51

Can you find some errors that spaCy and Google make?

What types of errors? Why is it happening?

50 of 51

Why is deep learning useful?

Does it work for language?

51 of 51

deeplanguageclass.github.io

next: deep learning