1 of 61

Text representation as �Vector Semantics

2 of 61

Introduction

  • Einat Minkov, senior lecturer at IS dep.
    • 2003-08: PhD at Carnegie Mellon University, Language Technologies Institution (LTI).
    • Joined Univ. of Haifa in 2010
    • Univ. of Edinburgh 2018-19
    • Google Research 2020, Nokia Research 2010, �MS Research 2006
    • Also teach: Intro. to AI / Text mining (advanced)
    • Research interests:
      • Natural Language Processing (NLP)
      • Information extraction (IE) / semantics
      • Relational inference
      • and more.

3 of 61

Information Retrieval (IR)

Goal: rank documents in a corpus (Web) by the match with the query.

How to quantify this match?

The vector space method:

Translate query/document into vectors in a vector space

one vector for each document

How to convert text into a vector?

Measure distance between vectors (documents)

How to measure vector distance?

4 of 61

A bag of words representation

I love this movie! It's sweet, but with satirical humor. The dialogue is great and the adventure scenes are fun… It manages to be whimsical and romantic while laughing at the conventions of the fairy tale genre. I would recommend it to just about anyone. I've seen it several times, and I'm always happy to see it again whenever I have a friend who hasn't seen it yet.

5 of 61

6 of 61

A bag of words representation

great

2

dialogue

1

recommend

1

it

6

I

4

...

...

7 of 61

The Vector-Space Model (IR)

  • Assume t distinct terms remain after preprocessing; call them the vocabulary.
  • These terms form a vector space.

Dimension = t = |vocabulary|

  • Each term i in a document or query j �is given a real-valued weight, wij.
  • Both documents and queries are expressed as �t-dimensional vectors:

dj = (w1j, w2j, …, wtj)

7

8 of 61

Graphic Representation

Example:

D1 = 2T1 + 3T2 + 5T3

D2 = 3T1 + 7T2 + T3

Q = 0T1 + 0T2 + 2T3

8

T3

T1

T2

D1 = 2T1+ 3T2 + 5T3

D2 = 3T1 + 7T2 + T3

Q = 0T1 + 0T2 + 2T3

7

3

2

5

  • Is D1 or D2 more similar to Q?
  • How to measure the degree of similarity? Distance? Angle? Projection?

9 of 61

Cosine Similarity Measure

  • Cosine similarity measures the cosine of the angle between two vectors.
  • Inner product normalized by the vector lengths.

9

D1 = 2T1 + 3T2 + 5T3 CosSim(D1 , Q) = 10 / √(4+9+25)(0+0+4) = 0.81

D2 = 3T1 + 7T2 + 1T3 CosSim(D2 , Q) = 2 / √(9+49+1)(0+0+4) = 0.13

Q = 0T1 + 0T2 + 2T3

θ2

t3

t1

t2

D1

D2

Q

θ1

D1 is 6 times better than D2 using cosine similarity but only 5 times better using inner product.

CosSim(dj, q) =

10 of 61

Example: Cosine similarity amongst 3 documents

term

SaS

PaP

WH

affection

115

58

20

jealous

10

7

11

gossip

2

0

6

wuthering

0

0

38

  • How similar are the novels �SaS: Sense and Sensibility, �PaP: Pride and Prejudice, and WH: Wuthering Heights?��(the first two – authored by Jane Austin, and the latter – by Emily Bronte)

Term frequencies (counts)

11 of 61

Example: Cosine similarity amongst 3 documents

  • Log frequency weighting

term

SaS

PaP

WH

affection

3.06

2.76

2.30

jealous

2.00

1.85

2.04

gossip

1.30

0

1.78

wuthering

0

0

2.58

cos(SaS,PaP)

0.789 × 0.832 + 0.515 × 0.555 + 0.335 × 0.0 + 0.0 × 0.0

0.94

cos(SaS,WH)0.79

cos(PaP,WH) 0.69

12 of 61

Question

  • How would you use this model for text classification, or clustering?

13 of 61

The Problem

  • synonymy: many ways to refer to the same object, e.g. car �and automobile -- leads to poor recall
  • polysemy: most words have more than one distinct meaning, e.g., model, python, chip -- leads to poor precision

auto

engine

bonnet

tyres

lorry

boot

car

emissions

hood

make

model

trunk

make

hidden

Markov

model

emissions

normalize

Synonymy

Will have small cosine

but are related

Polysemy

Will have large cosine

but not truly related

14 of 61

What do words mean?

  • First thought: look in a dictionary

  • http://www.oed.com/

15 of 61

Words, Lemmas, Senses, Definitions

sense

lemma

definition

16 of 61

Lemma pepper

Sense 1: spice from pepper plant

Sense 2: the pepper plant itself

Sense 3: another similar plant (Jamaican pepper)

Sense 4: another plant with peppercorns (California pepper)

Sense 5: capsicum (i.e. chili, paprika, bell pepper, etc)

17 of 61

Name these items

18 of 61

Superordinate Basic Subordinate

chair office chair

piano chair rocking chair

furniture lamp torchiere

desk lamp

table end table coffee table

19 of 61

Distributional similarity �(Ludwig Wittgenstein):

"The meaning of a word is its use in the language"

A data-driven approach!

20 of 61

What does ongchoi mean?

Suppose you see these sentences:

  • Ong choi is delicious sautéed with garlic.
  • Ong choi is superb over rice
  • Ong choi leaves with salty sauces
  • And you've also seen these:
  • …spinach sautéed with garlic over rice
  • Chard stems and leaves are delicious
  • Collard greens and other salty leafy greens
  • Conclusion:
    • Ongchoi is a leafy green like spinach, chard, or collard greens

21 of 61

Ong choi: Ipomoea aquatica �"Water Spinach"

Yamaguchi, Wikimedia Commons, public domain

22 of 61

Let's define words by their usages

  • In particular, words are defined by their environments (the words around them)

  • If A and B have almost identical environments, they are probably synonyms.

23 of 61

word-word matrix�(or "term-context matrix")

  • Two words are similar in meaning if their context vectors are similar

23

24 of 61

24

large

data

computer

apricot

1

0

0

digital

0

1

2

information

1

6

1

Which pair of words is more similar?

cosine(apricot,information) =

cosine(digital,information) =

cosine(apricot,digital) =

25 of 61

Visualizing cosines �(angles)

26 of 61

Ask humans how similar two words are

word1

word2

similarity

vanish

disappear

9.8

behave

obey

7.3

belief

impression

5.95

muscle

bone

3.65

modest

flexible

0.98

hole

agreement

0.3

SimLex-999 dataset (Hill et al., 2015)

27 of 61

Vector Semantics, part II:

�Dimensionality reduction & �Word embeddings

28 of 61

  • Freq-based word vectors are
    • long (length |V|= 20,000 to 50,000)
    • sparse (most elements are zero)

29 of 61

Alternative: dense vectors

  • vectors which are
    • short (length 50-1000)
    • dense (most elements are non-zero)

29

30 of 61

31 of 61

Singular Value Decomposition (SVD)

  • �Example: apply SVD to extract dimensions of meaning (topics) from documents.
    • consider documents d1-d4

Doc1 and doc2 similar 🡪 `computers’�Doc3 and doc4 similar 🡪 `automotive

(the term ‘the’ is not really related to any topic)

  • SVD will help up extract these dimensions, which we refer to as computers and automotive. it won‘t, however, give us nice human readable names for them.

32 of 61

SVD Example

SVD decomposes a matrix A into the product of three specially formed matrices (mysteriously named) U, S and V

S – describes the relative strength of features

U – relationship between terms and features

Vt – relationship between features and documents

33 of 61

Dimensionality Reduction

  • Singular Value Decomposition

{A}={U}{S}{V}T

  • Dimension Reduction

{~A}~={~U}{~S}{~V}T

34 of 61

35 of 61

(Pre-trained) dense word embeddings

  • Glove (Pennington, Socher, Manning)
  • http://nlp.stanford.edu/projects/glove/

36 of 61

Word2Vec and Word embeddings

37 of 61

Word2Vec: Model I

38 of 61

39 of 61

Can you compute Softmax?

40 of 61

Can you compute Softmax?

41 of 61

Word2Vec: Model II

42 of 61

“ … who passes the … “

43 of 61

Skip-Gram Training Data

  • Training sentence:
  • ... lemon, a tablespoon of apricot jam a pinch ...
  • c1 c2 target c3 c4

43

9/13/23

Asssume context words are those in +/- 2 word window

44 of 61

Basic idea behind skip-gram embeddings

from an input word w(t) in a document

construct hidden layer that “encodes” that word

So that the hidden layer will predict likely nearby words w(t-K), …, w(t+K)

final step of this prediction is a softmax over lots of outputs

45 of 61

Basic idea behind skip-gram embeddings

Training data:

positive examples are pairs of words w(t), w(t+j) that co-occur

Training data:

negative examples are samples of pairs of words w(t), w(t+j) that don’t co-occur

You want to train over a very large corpus (100M words+) and hundreds+ dimensions

46 of 61

Skip-Gram Training

  • Training sentence:
  • ... lemon, a tablespoon of apricot jam a pinch ...
  • c1 c2 t c3 c4

46

9/13/23

  • For each positive example, we'll create k negative examples.
  • Using noise words
    • Any random word that isn't t

47 of 61

Skip-Gram Training

  • Training sentence:
  • ... lemon, a tablespoon of apricot jam a pinch ...
  • c1 c2 t c3 c4

47

9/13/23

k=2

48 of 61

49 of 61

Negative sampling

increase the probability for less frequent words and decrease the probability for more frequent words.

50 of 61

Learning the classifier

  • Iterative process.
  • We’ll start with 0 or random weights
  • Then adjust the word weights to
    • make the positive pairs more likely
    • and the negative pairs less likely

51 of 61

52 of 61

Evaluating embeddings

  • Compare to human scores on word similarity-type tasks:
  • WordSim-353 (Finkelstein et al., 2002)
  • SimLex-999 (Hill et al., 2015)
  • Stanford Contextual Word Similarity (SCWS) dataset (Huang et al., 2012)
  • TOEFL dataset: Levied is closest in meaning to: imposed, believed, requested, correlated

53 of 61

Ask humans how similar two words are

word1

word2

similarity

vanish

disappear

9.8

behave

obey

7.3

belief

impression

5.95

muscle

bone

3.65

modest

flexible

0.98

hole

agreement

0.3

SimLex-999 dataset (Hill et al., 2015)

54 of 61

Words as vectors of context words

  • Each word = a vector
  • Similar words are "nearby in space"

55 of 61

Results from word2vec

https://www.tensorflow.org/versions/r0.7/tutorials/word2vec/index.html

56 of 61

Analogy: Embeddings capture relational meaning!

  •  

56

57 of 61

58 of 61

59 of 61

Word2Vec vs. �Distributional Semantic Models

  • “ it generally makes no difference whatsoever whether you use word embeddings or distributional methods -- what matters is that you tune your hyperparameters and employ the appropriate pre-processing and post-processing steps.
  • Recent papers from Jurafsky's group [5] [6] echo these findings and show that SVD -- not SGNS -- is often the preferred choice when you care about accurate word representations.”

60 of 61

Question

  • How would you use word vectors (Counts or word embeddings) for document classification?

61 of 61

Conclusion