1 of 47

CSCI-SHU 376: Natural Language Processing

Hua Shen

2026-09-23

Fall 2026

Lecture 7: Sequence-to-Sequence Modeling

​

​

2 of 47

Today’s Plan

  • Machine Translation
  • Evaluation
  • Neural Machine Translation (NMT)
  • Attention

​

​

3 of 47

Translation

  • Translation: Use computers to translate from one language to another
  • Information access: e.g., instructions from the web
  • Aid human translators: computer-aided translation

​

​

​

​

4 of 47

Translation Examples

  • Easy example:
  • More challenging examples:

5 of 47

Machine Translation

 

6 of 47

Machine Translation is challenging

  • Single words may be replaced with multi-word phrases

​

​

​

  • Reordering of phrases

​

​

  • Contextual dependence

7 of 47

Vauquois Pyramid

  • Hierarchy of concepts and distances between them in different languages
  • Lowest level: individual words / characters
  • Higher levels: syntax, semantics
  • Interlingua: Generic language-agnostic representation of meaning

8 of 47

Today’s Plan

  • Machine Translation
  • Evaluation
  • Neural Machine Translation (NMT)
  • Attention

​

​

9 of 47

Evaluation metrics

  • Two main criteria:
    • Adequacy: Translation should adequately reflect the linguistic content
    • Fluency: Translation should be fluent text in the target language

10 of 47

Evaluation metrics

  • Manual evaluation: ask a native speaker to verify the translation
    • Most accurate, but expensive

​

  • Automated evaluation metrics:
    • Compare system hypothesis with reference translations
    • BiLingual Evaluation Understudy (BLEU)
      • Modified n-gram precision

​

11 of 47

Evaluation metrics (BLEU)

  • To avoid log0, all precisions are smoothed
  • Each n-gram in reference can be used at most once
    • E.g., Hypothesis: to to to to to and Reference: to be or not to be should not get a unigram precision of 1
  • BLEU-k: average of BLEU scores computing using 1-gram through k-gram
  • Avoid short translations favor:
    • Multiply score with a brevity penalty for translations shorter than reference

12 of 47

Evaluation metrics (BLEU)

  • Correlates with human judgements

​

13 of 47

Today’s Plan

  • Machine Translation
  • Evaluation
  • IBM Model

​

​

14 of 47

Training Data

  • Training machine translation system relies parallel corpora (bilingual)

​

  • Not easily available for many low-resource languages

​

15 of 47

Today’s Plan

  • Machine Translation
  • Evaluation
  • Neural Machine Translation (NMT)
  • Attention

​

​

16 of 47

Neural Machine Translation

  • A single neural network is used to translate from source to target language
  • Encoder-Decoder Model, or Sequence-to-Sequence model
  • Encoder: Convert source sentence (input) into a vector representation
  • Decoder: Convert encoding into a sentence in target language (output)

​

​

​

17 of 47

Recall: RNNs

18 of 47

The Encoder-Decoder (Or Seq2seq) Model

  • Encoder-decoder Structure
  • The Encoder is a RNN to read the input sentence
  • The Decoder is another RNN to generate output word by word

​

19 of 47

Encoder

20 of 47

Encoder

21 of 47

Encoder

22 of 47

Encoder

23 of 47

Decoder

24 of 47

Decoder

25 of 47

Decoder

26 of 47

Decoder

27 of 47

Seq2seq Training

 

28 of 47

Encoder-Decoder Training

29 of 47

Greedy Decoding

  • Greedy Decoding: Compute argmax (over entire vocab) at every step

30 of 47

Beam Search

  • At every step, keep track of the k most probable partial translations
  • Score of each hypothesis = log probability of sequence so far

​

​

  • Not guaranteed to be optimal, but more efficient than exhaustive search

31 of 47

Beam Search

  • Beam size K=2

32 of 47

Beam Search

  • Beam size K=2

33 of 47

Beam Search

  • Beam size K=2

34 of 47

Beam Search (Backtrack)

  • Beam size K=2

35 of 47

Progress on Neural Machine Translation

36 of 47

Progress on Neural Machine Translation

37 of 47

Today’s Plan

  • Machine Translation
  • Evaluation
  • Neural Machine Translation (NMT)
  • Attention

​

​

38 of 47

Seq2seq: Bottleneck

  • A single encoder for all information
  • Vanishing Gradients
  • Model may overfit to training sequences

38

39 of 47

Alignment

39

40 of 47

Attention

  • Solve the bottleneck problem
  • At each decoding step, focusing on specific parts of input
    • Attention Score
    • Attention Weight
    • Attention Output

​

​

40

41 of 47

Attention Score

41

42 of 47

Attention Weight

  • Softmax over Attention Scores

​

​

42

43 of 47

Attention Output

  • Weighted Sum of vectors

​

​

43

44 of 47

Summary: Compute Attention

44

45 of 47

Summary: Understanding Attention Intuitions

45

46 of 47

Attention does improve MT

46

47 of 47

Visualizing Attention

47