Transformers
Discussion Mini Lecture 10
Why is Attention All You Need?
CS 189/289A, Fall 2025 @ UC Berkeley
Sara Pohland
Concepts Covered
Overview of a Transformer
How do we represent our data?
Image
The ducks were swimming near the bank.
Text
Naïve: Vector of Pixels
…
…
Naïve: Vector of Characters
…
…
T
s
w
.
Better: Word Parts
The ducks were swimming near the bank.
Better: Patches
Token Embeddings
The
Token Embeddings
How do we represent our data?
Image
The ducks were swimming near the bank.
Text
Token Embeddings
The
Token Embeddings
What do we lose in this representation?
Context is important – in both image and text processing.
What are we looking at here?
bank
What type of bank are we talking about?
How do we capture context?
How do we capture context?
How do we learn appropriate attention weights?
learned attention weight
learned context matrix
How to Model Attention
Capturing Similarity via Inner Product
feature 1
feature 2
Token Keys
Token Queries
*Assume vectors have unit norm.
*
Capturing Similarity via Inner Product
Token Keys
Token Queries
We’ll review these calculations in Problem 1 on the Disc. 10 worksheet.
Token Similarities
Transforming Similarity into Probability
This is not guaranteed using the inner product for similarity, but we can define
attention matrix
Ensuring Unit Variance Distributions
We’ll discuss this scaling in Problem 2 on the Disc. 10 worksheet.
The Self Attention Layer
– attention matrix
– value (context) matrix
– key matrix
– query matrix
The Self Attention Layer
Transformers
Discussion Mini Lecture 10
Contributors: Sara Pohland
Additional Resources