1 of 28

ME 5990: Introduction to Machine Learning

Attention Model Discussion

2 of 28

Outline

  • Seq2Seq encoder-decoder
  • Attention

3 of 28

Universal Grammar

  • Noam Chomsky
  • The basic postulate of UG is that there are innate constraints on what the grammar of a possible human language could be.

Noam Chomsky in 2017�from Wikipedia

Fun fact: legend says Prof Chomsky is the most cited living individual for publications between 1980-1992�(Totaled 3874 citations)

4 of 28

Review: Autoencoder

  • Encoder: transform the input (human-interpretable) information into a “code”. The code is a highly-abstracted that contains the input information.
  • Decoder: regenerate the human-interpretable information from the code. A good encoder-decoder model will let the regenerated information be similar to the input

5 of 28

Review: recurrent network

  • For simple RNN, there are two transitions:

6 of 28

Review: recurrent network

  • A recurrent network is efficient for sequential information

 

 

 

7 of 28

Seq-2-seq

  • The encoding and decoding
    • For instance, we want to train translation: input ABC -> output WXYZ

Target: W X Y Z <EOS>

Loss

Encoder

Decoder

8 of 28

Seq-2-seq

  • We can use a similar idea to encode sequential information using recurrent models

9 of 28

Seq-2-seq

  • RNN with two languages

10 of 28

Seq-2-seq

  • Training the encoder and decoder:

11 of 28

Seq-2-seq

  • RNN:

12 of 28

Gated Recurrent Unit

  • The hidden layer is not necessarily to be “fully connected”, it can be a compound of various functions as a “cell”
  • The cell is referred as gated recurrent unit
  • It can be:

13 of 28

Seq-2-seq

  • Implementation

https://pytorch.org/tutorials/intermediate/seq2seq_translation_tutorial.html

  • GRU: gated recurrent unit,
  • Embedding: a way to encode vocabulary based on input words exist

14 of 28

Seq-2-seq

  • Decoder

15 of 28

Attention

  • Problem of seq-2-seq:
  • We assume at the end of each sentence, the hidden vector carries a meaning of the sentence, regardless of the language

16 of 28

Attention

  • The concept

17 of 28

Attention

  • The attention output and input

18 of 28

Attention

19 of 28

Attention

20 of 28

Attention

21 of 28

Attention

22 of 28

Attention

23 of 28

Attention

24 of 28

Attention

25 of 28

Attention

26 of 28

Attention

  • Recurrent GRU cell can be FC, but also can be LSTM
    • C state: long-term memory
    • H state: short-term memory

C State

H State

27 of 28

Attention

28 of 28

Summary

  • All recurrent model’s long term memory fades away
  • Attention give “weights” to the hidden states, so the hidden states appears to “pay attention to” several sequence of outputs
  • Apply to sequences where the outputs order may be different from inputs order