1 of 83

Introduction to Multimodal AI

2026 Spring Semester | Hankuk University of Foreign Studies

Week 2: Fundamentals (1)

​

Division of Language & AI

seungtaek.choi@hufs.ac.kr

Seungtaek Choi

2 of 83

2

1

Unimodal Representation (Text)

2

Unimodal Representation (Image)

3

Unimodal Representation (Audio)

3 of 83

1 - Unimodal Representation (Text)

3

4 of 83

왜 텍스트를 숫자로 바꿔야 하는가?

1 - Unimodal Representation (Text)

  1. 모든 신경망 연산은 실수(float) 텐서 위에서 동작합니다.

4

5 of 83

왜 텍스트를 숫자로 바꿔야 하는가?

1 - Unimodal Representation (Text)

  • 컴퓨터에서 텍스트가 어떻게 저장되는가?

5

단순 문자 인코딩은 의미(semantic)를 담지 못합니다.

6 of 83

왜 텍스트를 숫자로 바꿔야 하는가?

1 - Unimodal Representation (Text)

  • 단어를 의미를 나타낼 수 있는 벡터로 바꿀 수 있다.

6

7 of 83

텍스트 처리 파이프라인

1 - Unimodal Representation (Text)

  • 텍스트 = a sequence of words

7

8 of 83

텍스트 처리 파이프라인

1 - Unimodal Representation (Text)

  • Look up embedding matrix (vocabulary)

8

9 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

  • Word vs. Character vs. Sub-word

9

fields 등 못 본 단어는 Out-of-Vocabulary (OOV) 발생

모든 문자를 봤지만 단어의 의미 파악이 어려움

vocabulary 크기와 expression power 간의 균형

10 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

  • It’s the problem of optimizing trade-off between vocab size and sequence length

10

11 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

  • Sub-word tokenization 이 사실상 표준.

11

12 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

  • Byte Pair Encoding (BPE)
    1. 가장 자주 등장하는 문자 쌍을 반복적으로 합쳐 vocab을 구성

12

13 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

  • Byte Pair Encoding (BPE) - Training Step 1
    • 모든 단어는 개별 문자의 나열로 시작

13

Initial Vocabulary:

{l, o, w, e, r, n, s, t, i, d}

Training Corpus:

  • low low low low low
  • lower lower
  • newest newest newest newest newest newest
  • wider wider wider

14 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

  • Byte Pair Encoding (BPE) - Training Step 2
    • 가장 빈번한 인접쌍 합치기

14

Frequency of Pair at Step 2:

  • (l, o) x 7번
  • (o, w) x 7번
  • (e, s) x 6번
  • …

Training Corpus:

  • low low low low low
  • lower lower
  • newest newest newest newest newest newest
  • wider wider wider

Vocabulary at Step 2:

{l, o, w, e, r, n, s, t, i, d, lo}

At Step 1: low = “l” + “o” + “w”

At Step 2: low = “lo” + “w”

15 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

  • Byte Pair Encoding (BPE) - Training Step 3
    • 원하는 vocabulary 크기에 도달할 때까지 이 과정을 반복

15

Frequency of Pair at Step 2:

  • (lo, w) x 7번
  • (n, e) x 6번
  • (e, s) x 6번
  • …

Training Corpus:

  • low low low low low
  • lower lower
  • newest newest newest newest newest newest
  • wider wider wider

Vocabulary at Step 2:

{l, o, w, e, r, n, s, t, i, d, lo, low}

At Step 1: low = “l” + “o” + “w”

At Step 2: low = “lo” + “w”

At Step 3: low = “low”

16 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

  • Byte Pair Encoding (BPE) - Training Step 4
    • 여러번 반복하면 의미있는 sub-word가 형성

16

Training Corpus:

  • low low low low low
  • lower lower
  • newest newest newest newest newest newest
  • wider wider wider

Final Vocabulary

{l, o, w, e, r, n, s, t, i, d, lo, low, es, est, er, …}

lower = low + er

newest = n + ew + est

…

17 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

  • Byte Pair Encoding (BPE) - Example of Llama

17

18 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

  • Byte Pair Encoding (BPE) - Example of Llama (vocab size: 32000)

18

19 of 83

Step 1: Tokenization

1 - Unimodal Representation (Text)

19

20 of 83

Step 2: Embedding

1 - Unimodal Representation (Text)

  • Embedding layer = Lookup table
    • Embedding matrix = (Vocab size x Embedding dim.)
    • n-th row of the embedding matrix = the embedding vector of n-th word

20

21 of 83

Step 2: Embedding

1 - Unimodal Representation (Text)

  • Word embedding vector?
    • 비슷한 의미(semantic) → 가까운 벡터(vector)
    • → word analogy: vec(“king”) - vec(“man”) + vec(“woman”) = vec(“queen”)

21

22 of 83

Step 2: Embedding

1 - Unimodal Representation (Text)

  • How to obtain such embeddings?
    • The most well-known example: word2vec skip-gram.
    • 중심 단어가 주어졌을 때 주변 단어를 예측하는 objective로 학습.

22

23 of 83

Step 2: Embedding

1 - Unimodal Representation (Text)

  • How to obtain such embeddings? – skip-gram.
    • text corpus 를 훑으면서 중심 단어와 주변 단어로 데이터 만듬.

23

24 of 83

Step 2: Embedding

1 - Unimodal Representation (Text)

  • How to obtain such embeddings? – skip-gram.
    • 유사한 문맥에 등장하는 단어들은 유사한 벡터를 갖게 된다.

24

학습되는 부분

25 of 83

Step 2: Embedding

1 - Unimodal Representation (Text)

  • Word2vec = 정적 (static) 임베딩
    • 고정된 벡터를 사용하기 때문에 빠르고 효율적이지만, 문맥에 따라 달라질 수 있는 의미를 반영하지 못 함.
    • → Contextual Embedding

25

26 of 83

Step 2: Embedding

1 - Unimodal Representation (Text)

  • What is contextual embedding?
    • Same word + different context = different meaning → different embedding (vector)

26

27 of 83

Step 2: Embedding

1 - Unimodal Representation (Text)

  • What is contextual embedding?
    • Language models for word representation (RNN-based)

27

28 of 83

Step 3: Language Models

1 - Unimodal Representation (Text)

  • What is language modeling?
    • Measuring “how likely is a sentence to appear in a language?”

28

29 of 83

Step 3: Language Models

1 - Unimodal Representation (Text)

  • What is language modeling?
    • Decomposing into smaller parts
      1. P(y1, y2, …, yn) = P(y2 | y1) * P(y3 | y1, y2) * … * P(yn | y1, y2, …, yn-1)

29

30 of 83

Step 3: Language Models

1 - Unimodal Representation (Text)

  • What is language modeling?
    • Once we have such a language model, we can use it to generate text.

30

31 of 83

Step 3: Language Models

1 - Unimodal Representation (Text)

  • How do neural networks language modeling?
    • text → tokens → sequence of embeddings → neural networks → next token prediction

31

32 of 83

Step 3: Language Models

1 - Unimodal Representation (Text)

  • How do neural networks language modeling?
    • text → tokens → sequence of embeddings → neural networks → next token prediction

32

33 of 83

Step 3: Language Models

1 - Unimodal Representation (Text)

  • How do neural networks language modeling?
    • next token prediction objective (loss)

33

34 of 83

Step 3: Language Models

1 - Unimodal Representation (Text)

  • Which architecture do neural networks language modeling?
    • Recently, the Transformer architecture is the de-facto standard.

34

35 of 83

Step 3: Language Models

1 - Unimodal Representation (Text)

  • Which architecture do neural networks language modeling?
    • The self-attention is the building block of Transformer architecture.
    • self-attention = 같은 sequence 내의 모든 토큰들끼리 서로 관계를 학습

35

36 of 83

Step 3: Language Models

1 - Unimodal Representation (Text)

  • Special case: BERT (Bidirectional Encoder Representations from Transformers)
    • … will be covered later in more detail (regarding CLIP, …)
    • 주로 마지막 레이어의 [CLS] 토큰 representation이 sentence 의 대표 representation으로 활용됨.

36

37 of 83

Summary

1 - Unimodal Representation (Text)

  • Text Representation
    • Input: text = sequence of tokens
    • Inside: sequence of tokens → sequence of embedding vectors → sequence of hidden representations
    • Output: prediction (next token, classification, …) or hidden representation(s)

37

38 of 83

2 - Unimodal Representation (Image)

38

39 of 83

컴퓨터는 이미지를 어떻게 이해하는가?

2 - Unimodal Representation (Image)

  • Images are numbers
    • A grayscale image is just a matrix of numbers [0, 255]

39

40 of 83

컴퓨터는 이미지를 어떻게 이해하는가?

2 - Unimodal Representation (Image)

  • Images are numbers
    • An RGB image is also a tensor of numbers, i.e., (3 x 1080 x 1080)

40

41 of 83

컴퓨터는 이미지를 어떻게 이해하는가?

2 - Unimodal Representation (Image)

  • It’s same with text processing.
    • Input image → pixel representation → … → prediction

41

Lincoln

​

Washington

​

Jefferson

​

Obama

​

Trump

0.8

​

0.05

​

0.05

​

0.01

​

0.09

Input Image

Pixel Representation

Prediction

42 of 83

컴퓨터는 이미지를 어떻게 이해하는가?

2 - Unimodal Representation (Image)

  • It’s same with text processing.
    • Input image → pixel representation → … → prediction

42

43 of 83

Convolution 이란 무엇인가?

2 - Unimodal Representation (Image)

  • Fully connected neural networks have …
    • no spatial information
    • huge parameters

43

44 of 83

Convolution 이란 무엇인가?

2 - Unimodal Representation (Image)

  • Fully connected neural networks have …
    • no spatial information
    • huge parameters

44

45 of 83

Convolution 이란 무엇인가?

2 - Unimodal Representation (Image)

  • Locally connected neural networks have …
    • spatial information
    • less parameters

45

46 of 83

Convolution 이란 무엇인가?

2 - Unimodal Representation (Image)

  • Convolutional neural networks have …
    • spatial information
    • share parameters

46

47 of 83

Convolution 이란 무엇인가?

2 - Unimodal Representation (Image)

  • Convolutional operation uses spatial structure
    • process an image patch

47

48 of 83

Convolution 이란 무엇인가?

2 - Unimodal Representation (Image)

  • Convolutional operation uses spatial structure
    • a patch in upper layer = a single neuron in subsequent layer
    • use a sliding window to define connections

48

49 of 83

Convolution 이란 무엇인가?

2 - Unimodal Representation (Image)

  • The Convolutional operation is an element-wise multiplication by a matrix (kernel)
    • It enables translation invariance (이동불변성).

49

Image

Kernel

Feature Map

50 of 83

Convolutional Neural Networks

2 - Unimodal Representation (Image)

  • Image → Conv Layers → Pool Layers → Prediction

50

51 of 83

Convolutional Neural Networks

2 - Unimodal Representation (Image)

  • Many well-known vision models are based on the CNN architecture.
    • VGG-16, ResNet-50, EfficientNet, …

51

52 of 83

Receptive Field

2 - Unimodal Representation (Image)

  • Receptive field = 수용 범위
    • 출력 뉴런 하나를 만들 때 참조하는 입력 이미지의 영역
    • 층이 깊어질 수록 receptive field 가 커짐 (선형적) = 더 큰 패턴 인식

52

53 of 83

Receptive Field

2 - Unimodal Representation (Image)

  • In the case of convolutional language models

53

54 of 83

Receptive Field

2 - Unimodal Representation (Image)

  • 한계: 층 수에 비례해 선형적으로만 증가
  • = 이미지 전체 맥락 (ex. 멀리 있는 물체 간 관계) 을 파악하려면 매우 많은 층이 필요

54

55 of 83

CNN → Vision Transformer (ViT)

2 - Unimodal Representation (Image)

  • Vision Transformer (ViT)
    • 핵심 아이디어: 이미지를 패치(patch)로 쪼개서 텍스트 토큰처럼 처리

55

56 of 83

Vision Transformer (ViT)

2 - Unimodal Representation (Image)

  • 이미지를 패치로 쪼개기
    • (224, 224, 3) 크기의 이미지를 16 x 16 픽셀의 패치로 쪼개면 총 (224 / 16) x (224 / 16) = 196 개.

56

57 of 83

Vision Transformer (ViT)

2 - Unimodal Representation (Image)

  • 이미지 패치들을 평탄화하기
    • 각 패치: (16, 16, 3) → flatten → 768차원 벡터
    • 현재 196 x 768 크기.

57

58 of 83

Vision Transformer (ViT)

2 - Unimodal Representation (Image)

  • 패치로부터 임베딩 만들기
    • 768 차원 벡터를 고정된 d 차원 벡터로 변경
    • 현재 196 x d

58

59 of 83

Vision Transformer (ViT)

2 - Unimodal Representation (Image)

  • 분류 토큰 (CLS) 추가하기
    • CLS 토큰은 학습 가능한(learnable) 벡터
    • 현재 (196 + 1) x d.

59

60 of 83

Vision Transformer (ViT)

2 - Unimodal Representation (Image)

  • 위치 임베딩 (Positional Embedding) 벡터 추가하기
    • Transformer 는 순서 정보가 없음 → 위치 정보를 별도로 더해줘야 함
    • ex. “이 패치는 이미지의 3번째 행 5번째 열에 있다.”

60

61 of 83

Vision Transformer (ViT)

2 - Unimodal Representation (Image)

  • 위치 임베딩 (Positional Embedding) 벡터 추가하기
    • Transformer 는 순서 정보가 없음 → 위치 정보를 별도로 더해줘야 함
    • ex. “이 패치는 이미지의 3번째 행 5번째 열에 있다.”

61

62 of 83

Vision Transformer (ViT)

2 - Unimodal Representation (Image)

  • 이후로는 stacked self-attention …

62

63 of 83

Vision Transformer (ViT)

2 - Unimodal Representation (Image)

  • ViT models have become a new standard …

63

64 of 83

CNN vs. ViT

2 - Unimodal Representation (Image)

  • ​

64

​

CNN

ViT

귀납적 편향 (Inductive Bias)

강함 (Locality, Translation Invariance, …)

약함 (데이터로 학습)

전역 관계 (Global Relations)

깊은 층 필요

모든 층에서 가능

필요 데이터

✅ 소량으로도 잘 동작

❌ 일정량 이상에서 유의미

텍스트 통합

별도 설계 필요

✅ 동일 구조 공유

대표 모델

ResNet, EfficientNet, …

ViT-B/16, MAE, …

65 of 83

왜 ViT가 멀티모달 AI의 핵심이 되었는가?

2 - Unimodal Representation (Image)

  • ViT는 Text Transformer와 동일한 구조를 공유한다.
    • 이미지 패치 벡터 시퀀스를 LLM의 입력 토큰으로 그대로 사용 가능
    • 텍스트-이미지 간 Cross-Attention도 자연스럽게 연결 가능

65

66 of 83

3 - Unimodal Representation (Audio)

66

67 of 83

소리를 어떻게 숫자로 표현하는가?

3 - Unimodal Representation (Audio)

  • 오디오 = 시간에 따른 공기 압력의 변화
    • Waveform (파형) = 1차원 시계열 신호

67

68 of 83

소리를 어떻게 숫자로 표현하는가?

3 - Unimodal Representation (Audio)

  • Sampling rate
    • ex) 16,000 Hz = 1초에 16,000개의 숫자값으로 기록
    • → 1초 오디오 = shape (16000, ) 배열

68

69 of 83

소리를 어떻게 숫자로 표현하는가?

3 - Unimodal Representation (Audio)

69

70 of 83

소리를 어떻게 숫자로 표현하는가?

3 - Unimodal Representation (Audio)

70

71 of 83

Mel Spectrogram (1): 왜 파형 그대로 쓰지 않는가?

3 - Unimodal Representation (Audio)

  • Raw waveform 의 문제점
    • 1초 오디오 = 16,000개 값 → 시퀀스가 너무 김
    • 30초 오디오 = 480,000개 값 (Transformer에 비현실적)
    • “어떤 주파수가 언제 나타나는지" 파형에서 바로 알 수 없음

71

72 of 83

Mel Spectrogram (2): Spectrogram 이란?

3 - Unimodal Representation (Audio)

  • Spectrogram: 오디오의 2D 표현
    • x축: 시간 (time)
    • y축: 주파수 (frequency)
    • 색상/밝기: 해당 시간·주파수에서의 진폭(Amplitude) — scalar 값 하나
  • → "어떤 주파수가 언제 얼마나 강한가"를 이미지처럼 표현

72

73 of 83

Mel Spectrogram (3): Fourier Transform & FFT

3 - Unimodal Representation (Audio)

  • Fourier Transform
    • 시간 영역 신호 → 주파수 영역으로 변환
    • “이 소리에는 어떤 주파수가 얼마나 강한가?”
  • FFT = Fast Fourier Transform: O(n^2) → O(n log n)
  • 한계: 시간 정보가 사라진다.
    • 오디오 전체에 FFT 한 번 → “그 주파수가 언제 등장했는지” 알 수 없음.

73

74 of 83

Mel Spectrogram (4): STFT → Spectrogram 생성

3 - Unimodal Representation (Audio)

  • Short-Time Fourier Transform (STFT)
    • 짧은 구간(윈도우)으로 나눠서 각각 FFT

74

75 of 83

Mel Spectrogram (5): Mel Filter Bank 적용

3 - Unimodal Representation (Audio)

  • Filter의 동작
    • 담당 주파수 범위 (예: 500~1500 Hz)의 bin 값들을 숫자 하나로 합산
    • 삼각형 모양 = 중심 주파수에 가까울수록 많이 반영, 양 끝은 0
  • Mel Filter Bank
    • Spectrogram: y축 = Frequency (Hz)
    • Mel Filter Bank 적용 후: y축 = Mel bin (ex. 80개)

75

76 of 83

Mel Spectrogram (5): Mel Filter Bank 적용

3 - Unimodal Representation (Audio)

  • 왜 불균등하게 배치하는가?
    • 낮은 주파수: 삼각형 좁음 → Mel bin이 촘촘 (사람 귀가 민감)
    • 높은 주파수: 삼각형 넓음 → Mel bin이 성김 (사람 귀가 둔감)
  • DIY: 가청주파수 테스트 사이트

​

​

76

77 of 83

Mel Spectrogram (6): Log 스케일 → 최종 입력

3 - Unimodal Representation (Audio)

  • Log 스케일 변환
    • Mel filter bank 출력값의 범위가 매우 크기 때문에 log를 취해 압축
    • 큰 값의 차이는 줄이고, 작은 값의 차이는 더 잘 보이게 함
    • 사람의 소리 크기 인식도 대체로 로그적 경향을 보임
      • 10 dB 증가 → 대략 2배 크게 느낌

​

​

77

78 of 83

Mel Spectrogram (7): Summary

3 - Unimodal Representation (Audio)

  • 전체 흐름
    • Waveform → STFT → Power Spectrogram (T, 201) → Mel Filterbank → Mel Spectrogram (T, 80) → Log-Mel Spectrogram (T, 80) → 모델 입력
  • + Log-Mel Spectrogram 은 “소리의 2D 이미지” → 이미지 처리 기술을 오디오에도 적용 가능

​

​

78

79 of 83

Mel Spectrogram (7): Summary

3 - Unimodal Representation (Audio)

  1. 추가자료: https://www.youtube.com/watch?v=eKSmEPAEr2U

79

80 of 83

OpenAI – Whisper

3 - Unimodal Representation (Audio)

  • 대규모 다국어 음성 인식 모델 (2022)
    • 입력: 30초 길이의 오디오 → Log-mel spectrogram
    • 출력: 텍스트 토큰

80

81 of 83

OpenAI – Whisper

3 - Unimodal Representation (Audio)

81

82 of 83

Today’s Summary

3 - Unimodal Representation (Audio)

  1. 어떤 모달리티든 Transformer에 넣을 수 있는 vector sequence로 변환하면 된다.

82

​

Text

Image

Audio

Raw Data

문자 시퀀스

픽셀 행렬 (H x W x 3)

파형 (1D sequence)

전처리

Tokenization

정규화/리사이즈

Mel Spectrogram

토큰/패치

Sub-word token

16 x 16 pixel patch

25ms 프레임

최종 형태

seq of 768 벡터

seq of 768 벡터

seq x 80 행렬

핵심 모델

BERT, GPT, …

ViT, CLIP-ViT, …

Whisper, …

83 of 83

Thank you!

Any questions?

83