1 of 17

Transformers

Discussion Mini Lecture 10

Why is Attention All You Need?

CS 189/289A, Fall 2025 @ UC Berkeley

Sara Pohland

2 of 17

  1. Transformer Architecture
  2. Modeling Attention

Concepts Covered

3 of 17

Overview of a Transformer

  1. Transformer Architecture
  2. Modeling Attention

4 of 17

How do we represent our data?

Image

The ducks were swimming near the bank.

Text

Naïve: Vector of Pixels

Naïve: Vector of Characters

T

s

w

.

Better: Word Parts

The ducks were swimming near the bank.

Better: Patches

Token Embeddings

The

 

 

Token Embeddings

 

 

5 of 17

How do we represent our data?

Image

The ducks were swimming near the bank.

Text

Token Embeddings

The

 

 

Token Embeddings

 

 

 

 

 

 

6 of 17

What do we lose in this representation?

Context is important – in both image and text processing.

What are we looking at here?

bank

What type of bank are we talking about?

7 of 17

How do we capture context?

 

 

 

 

 

 

 

 

8 of 17

How do we capture context?

 

 

How do we learn appropriate attention weights?

 

learned attention weight

 

 

learned context matrix

9 of 17

How to Model Attention

  1. Transformer Architecture
  2. Modeling Attention

10 of 17

Capturing Similarity via Inner Product

feature 1

feature 2

 

 

 

 

 

 

 

 

Token Keys

 

 

Token Queries

 

 

*Assume vectors have unit norm.

 

 

*

11 of 17

Capturing Similarity via Inner Product

Token Keys

 

 

Token Queries

 

 

We’ll review these calculations in Problem 1 on the Disc. 10 worksheet.

Token Similarities

 

 

 

12 of 17

Transforming Similarity into Probability

 

 

 

This is not guaranteed using the inner product for similarity, but we can define

 

 

attention matrix

 

13 of 17

Ensuring Unit Variance Distributions

 

 

 

We’ll discuss this scaling in Problem 2 on the Disc. 10 worksheet.

14 of 17

The Self Attention Layer

 

 

– attention matrix

 

– value (context) matrix

 

 

– key matrix

– query matrix

15 of 17

The Self Attention Layer

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

16 of 17

Transformers

Discussion Mini Lecture 10

Contributors: Sara Pohland

17 of 17

Additional Resources

  1. Transformers