Large Language Models
Lecture 7
Introduction to Transformer: Self-Attention and Multi-Head Attention
Krishnendu Ghosh
Recap: Attention
Recap: Attention
Recap: Attention
Recap: Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
Self-Attention
From Self-Attention to Transformers
• We will talk about a class of models for processing sequences that does not use recurrent connections but instead relies entirely on attention and will build up towards a class of models called Transformers.
• To address a few key limitations, we need to add certain elements:
1. Positional encoding addresses lack of sequence information
2. Multi-headed attention allows querying multiple positions at each layer
3. Adding nonlinearities so far, each successive layer is linear in the previous one
4. Masked decoding how to prevent attention lookups into the future?
Positional Encoding: Motivation
Sinusoidal Positional Encoding
Multi-Head Attention
Multi-Head Attention
Multi-Head Attention
Multi-Head Attention
Multi-Head Attention
Multi-Head Attention
Multi-Head Attention
Self-Attention (In Encoder)
Self-Attention (In Encoder)
Self-Attention (In Encoder)
Self-Attention (In Encoder)
Self-Attention (In Encoder)
Self-Attention (In Encoder)
Self-Attention (In Encoder)
Self-Attention (In Encoder)
Multi-Head Self-Attention
Multi-Head Self-Attention
Self-Attention Is “Linear”
Position-wise
Feed-Forward Networks
Position-wise
Feed-Forward Networks
Self-attention can see the future!
Self-attention can see the future!
Masked Attention
Example Tokens
Self-Attention
Self-Attention
Multi-Head Attention
Cross-Attention
Cross-Attention
Masked Attention
Masked Attention