1 of 30

Convolutional Layers in Neural Machine Translation Architectures

Pratik Gujjar

October 23, 2017

1

2 of 30

Sequence learning in RNNs

2

3 of 30

Sequence learning in CNNs

3

4 of 30

  • Compared to RNNs, convolutions create representations for fixed size contexts.
  • Hierarchical stacking of layers leads to dense representations with large but controllable length of dependencies. O(n/k)
  • RNNs maintain a hidden state over the entire sequence; O(n)
  • Computations can be parallelized over every element in the sequence.
  • Convolutional networks have historically performed poorly over recurrent networks.

4

5 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

Recurrent Continuous Translation Models

Nal Kalchbrenner and Phil Blunsom

University of Oxford

5

6 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

  • One of the first successful NMT efforts.
  • Distinct phrase pairs often sharing significant linguistic similarities, do not share statistical weight in the earlier estimations of their transition probabilities.
  • Estimation is sparse or skewed for rare or unseen phrase pairs.
  • Continuous representations over phrase translation and alignment models.

6

7 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

  • They present two RCTMs modelled with conditional generation and translation units.
  • In both cases the convolutional layers are used to generate representations for phrases in a sentences from representations for words in the sentence.
  • Advantage: Lack of latent alignment segmentations and sparsity.
  • Perplexity of the model is more than 43% lower that the IBM Model 2 (Brown et.al., 1993; Dyer et.al., 2013)
  • Model is highly sensitive to word position and order.

7

8 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

Recurrent Language Model

8

9 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

Convolutional Sentence Model

  • (Ki )2 ≤ i ≤ r is a sequence of convolutional kernels.
  • Each row of Ki is a vector of i weights.
  • Given a source sentence of length k, the CSM convolves successively with the sentence matrix. Ee →e ∈ Rq x 1

9

10 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

RCTM I

  • Length of the target sentence is predicted by the RLM itself. By its architecture is biased towards shorter sentences.
  • e constrains all target words equally; target words may depend on some source words more than others

10

11 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

RCTM II

  • Estimate the length m of the target sentence independently.
  • Construct a representation for the n-grams in e
  • Build an m-length representation from e through the inverted truncated CGM.

11

12 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

Experiments

1A low perplexity indicates the probability distribution is good at predicting the sample

  • Trained on WMT’13 data. 25403 English words and 34831 French words.
  • Test sets: WMT News Test sets- 2009, 2010, 2011, 2012.
  • 256 dimensional word vectors. Adagrad optimizer; 15 hours of training on 3 multicore CPUs. Randomly initialized weights.
  • Experiment 1 reports 1perplexities of the model with respect to reference translations.
  • Experiment 2 and 3 test the sensitivity of RCTM II to linguistic aspects of the source sentences.
  • Experiment 4 tests rescoring performances of RCTM I and RCTM II.

12

13 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

13

14 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

14

15 of 30

Kalchbrenner and Blunsom, Recurrent Continuous Translation Models, EMNLP 2013

15

16 of 30

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

Convolutional Sequence to Sequence Learning

Jonas Gehring, Michael Auli, David Grangier, Denis Yarats and Yann N. Dauphin

Facebook AI Research

16

17 of 30

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

  • Current models use Bidirectional RNNs after Sutskever et.al., 2014. Soft attention (Bahdanau et.al., 2014; Luong et.al.,) interfaces are standard.
  • CNNs though have fixed size contexts, offer controllable context sizes and huge parallelization benefits.
  • Multilayer CNN provides a shorter path to capture long-range dependencies compared to the chain-structure in RNNs. O(n/k) vs O(n)
  • Fixing number of non-linearities eases learning.
  • Bardbury et.al., 2016 introduce recurrent pooling between convolutions and Kalchbrenner et.al., 2016 do translation without attention. No improvements over the state of the art.

17

18 of 30

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

  • Gated convolutions (Meng et.al., 2015)
  • Most convolutional approaches still rely on a recurrent decoder. (Gehring et.al., 2016)
  • This model is fully convolutional.
  • Attention in every decoder layer.
  • Gated linear units (Dauphin et.al., 2016) as the non-linearities.
  • Performance over the state of the art with an order of magnitude faster speed on CPU and GPU hardware. WMT’14 En-Fr and En-De.

18

19 of 30

Position Embeddings

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

  • Elements x = (x1,…, xm) in distributional space as w = (w1,..., wm).
  • wj is a column in an embedding matrix.
  • A sense of ‘order’ by embedding the absolute position p = (p1,..., pm).
  • Input element representation e = (w1 + p1, …, wm + pm)

19

20 of 30

Convolutional Block Structure

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

  • For block L:
    • Decoder output:
    • Encoder output:
  • Each convolutional kernel takes X- k input elements embedded in d dimensions and maps them to a single output Y in 2d dimensions
  • Gated Linear Units
    • Gates control which inputs of the context are relevant
    • is pointwise multiplication

20

21 of 30

Convolutional Block Structure

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

  • Output of convolutional layers match input length in the encoders.
    • Ensured by zero padding at the edges.
  • Deep convolutions enabled by residual connection (Oord et.al., 2016)

  • A distribution over T next possible target elements is estimated:

21

22 of 30

Multi-step Attention

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

  • Every stage of the decoder attends to separate sections; gi is an embedding of the previous target element.

  • Conditional input ci is a weighted sum of both the encoder output and input element embeddings. ej provides point information about a specific input element.

22

23 of 30

Multi-step Attention

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

  • Every decoder layer has access to its previous layer’s attention history.
  • The model can account for what has been attended to. In a recurrent attention scheme, this information cannot survive the many non-linearities at every step in the sequence.
  • Model performs multiple attention ‘hops’ per time step.
  • Convolutional architecture allows batching of attention computations for each layer individually.

23

24 of 30

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

24

25 of 30

Source: https://norman3.github.io/papers/docs/fairseq.html

25

26 of 30

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

Source: A novel approach to neural machine translation

https://code.facebook.com/posts/1978007565818999/a-novel-approach-to-neural-machine-translation/

26

27 of 30

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

27

28 of 30

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

28

29 of 30

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

29

30 of 30

Gehring et.al., Convolutional Sequence to Sequence Learning, arXiv July 2017

30