1 of 37

Neural Networks with Memory

Memory Networks, End-to-End Memory Networks, Condensed Memory Networks, Neural Turing Machine & Differential Neural Computer

Aaditya Prakash

Nov 5, 2018��

2 of 37

Why?

Siri, play me a good song

Sorry, I couldn’t find a ‘good song’ in your music.

Because AI is good only with context !

3 of 37

Neural Network - limitation

  1. Not all problems can be mapped to y = f(x)
  2. Need to remember external context
  3. Need to know where to look for in the context�
  4. Need to know what to look for in the context�
  5. Need to know how to reason using external context�
  6. Need to handle potentially changing external context

Text source: Sumit Chopra FAIR, DLSS 2016

4 of 37

Why not RNN ?

  1. Inherent temporal structure (not all knowledge is temporal)�
  2. Very slow (because of sequential learning, not embarrassingly parallel)�
  3. Cannot scale to large network (overhead in terms of learning)�
  4. Prone to “vanishing/exploding gradient” (multiply gradient many times)�
  5. Back-propagation in time is hard (always needs truncation and clipping)�

5 of 37

Reasoning is hard

Bilbo travelled to the cave. Gollum dropped the ring there. Bilbo took the ring.�Bilbo went back to the Shire. Bilbo left the ring there. Frodo got the ring.�Frodo journeyed to Mount-Doom. Frodo dropped the ring there. Sauron died.�Frodo went back to the Shire. Bilbo travelled to the Grey-havens. The End.

Where is the ring?

Where is Bilbo now?

Where is Frodo now?

6 of 37

Reasoning is hard

Bilbo travelled to the cave. Gollum dropped the ring there. Bilbo took the ring.�Bilbo went back to the Shire. Bilbo left the ring there. Frodo got the ring.�Frodo journeyed to Mount-Doom. Frodo dropped the ring there. Sauron died.�Frodo went back to the Shire. Bilbo travelled to the Grey-havens. The End.

Where is the ring? A: Mount-Doom

Where is Bilbo now? A: Grey-havens

Where is Frodo now? A: Shire

7 of 37

Memory Networks

I

G

O

R

Convert the input (x) to internal representation (aka feature space)� - learn embeddings (bow/gru/lstm), use pre-trained vectors (glove/w2v)

Make a memory state (u), write initial data to the memory, update when necessary

Generate the output state, which is combination of memory state and input�

Convert the output (O) into response as desired by the model

8 of 37

Memory Networks

Figure: Saina Sukhbaatar

9 of 37

Memory Networks

Figure: Saina Sukhbaatar

John is in the playground.

Bob is in the office.

John picked up the football.

Bob went to the kitchen.

bABI tasks ---

Where is the football?

Where was Bob before the kitchen?

Where is the football? A:playground

Where was Bob before the kitchen? A:office

10 of 37

Hard Attention (Strong supervision)

Supervision

(hard attention)

Supervision

(hard attention)

11 of 37

Soft Attention (Weak supervision)

Supervision

(soft attention)

Supervision

(soft attention)

12 of 37

Soft Attention (Weak supervision)

Supervision

(soft attention)

Supervision

(soft attention)

Now applicable to more problems where it is hard to get strong supervision

13 of 37

End to End Memory Networks

14 of 37

End to End Memory Networks

15 of 37

One hop is just not sufficient

16 of 37

Condensed Memory Network [ C-MemNN ]

MemNN

17 of 37

Condensed Memory Network [ C-MemNN ]

MemNN

C-MemNN

18 of 37

Condensed Memory Network [ C-MemNN ]

A-MemNN

MemNN

C-MemNN

C-MemNN

19 of 37

Condensed Memory Network [ C-MemNN ]

A-MemNN

C-MemNN

C-MemNN

Condensation of memory state - continuous hierarchy.

20 of 37

Neural Turing Machine

NTM can learn basic algorithms like sorting.

21 of 37

Neural Turing Machine

Sharp functions made smooth. Now can be trained with ‘backpropagation’

22 of 37

Neural Turing Machine

Sharp functions made smooth. Now can be trained with ‘backpropagation’

23 of 37

Neural Turing Machine - Selective Attention

  • Focus on the parts of the memory the network will read and write to:� ‘An introspective attention model’�
  • Use the controller outputs to parameterise a distribution (weights) over the rows (slots) in the memory matrix�
  • Weights are defined by two main attention mechanism:
    • One based on content
    • One based on location

Source: Alex Graves

24 of 37

Neural Turing Machines - Memory

25 of 37

Neural Turing Machines - Memory

26 of 37

Neural Turing Machines - Memory

Forget gate !!

Déjà vu

27 of 37

Neural Turing Machines - Memory

Forget gate !!

Déjà vu

28 of 37

Neural Turing Machines - Memory

29 of 37

Neural Turing Machines - Memory

The key vector kt, & key strength βt, are used to perform content-based addressing of the memory matrix Mt.The resulting content-based weighting is interpolated with the weighting from the previous time step based on the value of the interpolation gate gt.The shift weighting st, determines whether & by how much the weighting is rotated. Depending on γt,the weighting is sharpened and used for memory access.

30 of 37

NTM - Copy

- A simple task of ‘copying’ the input back can be hard, when the time period becomes very long.

- NTM can learn basic algorithms like copy, loop, sort, associative recall, N-gram inference.

31 of 37

NTM - Mult Copy

Copying the same sequence many times.

32 of 37

Differential Neural Computer

Architecture of DNC

Architecture of NTM

33 of 37

Differential Neural Computer

Architecture of DNC

  • NTM was able to retrieve memories in order of their index but not in order in which they were written
  • Preserving temporal order is necessary for many tasks (e.g. sequence of instructions) and appears to play an important role in human cognition
  • DNC tries to iterate through memories in the order they were written (or rewritten)
  • A precedence weighting (pt) keeps track of which locations were most recently written to:

34 of 37

Differential Neural Computer

  • RNNs are great at processing sequences: text, audio, time-series…
  • But graph-structured data is even more general: maps, molecules, parse-trees, knowledge graphs, social networks…
  • DNC is able to interpret and answer questions about graphs, even if presented in sequential form
  • Started with explicit graphs, but ultimately interested in implicit graphs: relations in natural language, scene parsing, agent’s surroundings…

35 of 37

How DNC Reads Graphs

36 of 37

Differential Neural Computer

37 of 37

Relevant papers

Source materials (also good for tutorials)

  1. J. Weston, S. Chopra, A. Bordes. Memory Networks. ICLR 2015 (and arXiv:1410.3916).
  2. S. Sukhbaatar, A. Szlam, J. Weston, R. Fergus. End-To-End Memory Networks. NIPS 2015
  3. J. Weston, et al. Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
  4. A. Bordes, N. Usunier, S. Chopra, J. Weston. Large-scale Simple Question Answering with Memory Networks
  5. Alex Graves, et al Neural Turing Machine. Nature 2015
  6. Prakash, et al. Condensed Memory Networks. AAAI 2017
  7. Alex Graves, et al Differentiable Neural Computers. Nature 2017