Machine Learning for
Natural Language Processing
Fundamentals 2
demo
explosion.ai/demos/displacy
spaCy
How many words in a language?
Size of a language
Words in the dictionary
10^5
Unique tokens in full corpus
10^6
What is the distribution of words?
How do the frequencies vary?
Zipf’s law
“The frequency of any word is inversely proportional to its rank in the frequency table.”
Zipf’s law
“The frequency of any word is inversely proportional to its rank in the frequency table.”
What are corpora?
Corpora
corpus > document > paragraph > sentence > phrase > word > char
Corpora
corpus > document > paragraph > sentence > phrase > word > char
? > query > ?
Corpora
corpus > document > paragraph > sentence > phrase > word > char
? > tweet > ?
Annotated corpus
Human labour
Parallel corpora
The Rosetta stone
Tokenisation
Sentences and words
Morphology
Stemmers and lemmatisers
Grammar
Syntax and agreement
Syntax
Parsing
Canonicalisation
A double-edged sword
When to preprocess?
Language Model
Probability distribution over sequences of words
n-grams
Context
n-grams:
word level
Text
“Is the color of milk and fresh snow, the color produced by the combination of all the colors of the visible spectrum.”
n=2
“Is the”, “the color”, “color of”, “of milk”, “milk and”, “and fresh”, “fresh snow”, “snow ,”, “, the”, “the color”...
n-grams:
char level
Text
“Is the color of milk and fresh snow, the color produced by the combination of all the colors of the visible spectrum.”
n=3
“@Is”, “Is_”, “s_t”, “_th”, “the”, “he_”, “e_c”, “_co”, “col”, “olo”, “lor”, “or_”, “r_o”, “_of”, “of_”, “f_m”...
demo
github.com/SignalN/language
language.ngrams
What is the difference between n-grams and a language model?
NER
Named entities
demo
cloud.google.com/natural-language
Google Cloud Natural Language API
Co-references
Anaphora, pronouns...
Polysemy
It’s not what you think it means.
How do we measure success?
BLEU
Exact match
Fuzzy
demo
rajpurkar.github.io/SQuAD-explorer
SQuAD
Beyond English
What % is English?
% of users? % of content?
How many languages does the average person speak?
1.5
Language codes
ISO
Script
Alphabets+
Is English unique?
How is it easier? How is it harder?
Is Armenian unique?
How is it easier? How is it harder?
Linguistic Typology
Multilingual problems
Identification, translation, transliteration
Multimodal problems
language + x
text + struct
=>
numeric
Ali Baba, YouTube, Amazon
product recommendations
image
=>
struct | text
cs.stanford.edu/.../clevr
CLEVR
image
=>
text
@picdescbot
image captioning
For next time...
Can you find some errors that spaCy and Google make?
What types of errors? Why is it happening?
Why is deep learning useful?
Does it work for language?
deeplanguageclass.github.io
next: deep learning