Introduction to Multimodal AI
2026 Spring Semester | Hankuk University of Foreign Studies
Week 2: Fundamentals (1)
Division of Language & AI
seungtaek.choi@hufs.ac.kr
Seungtaek Choi
2
1
Unimodal Representation (Text)
2
Unimodal Representation (Image)
3
Unimodal Representation (Audio)
1 - Unimodal Representation (Text)
3
왜 텍스트를 숫자로 바꿔야 하는가?
1 - Unimodal Representation (Text)
4
왜 텍스트를 숫자로 바꿔야 하는가?
1 - Unimodal Representation (Text)
5
단순 문자 인코딩은 의미(semantic)를 담지 못합니다.
왜 텍스트를 숫자로 바꿔야 하는가?
1 - Unimodal Representation (Text)
6
텍스트 처리 파이프라인
1 - Unimodal Representation (Text)
7
텍스트 처리 파이프라인
1 - Unimodal Representation (Text)
8
Step 1: Tokenization
1 - Unimodal Representation (Text)
9
fields 등 못 본 단어는 Out-of-Vocabulary (OOV) 발생
모든 문자를 봤지만 단어의 의미 파악이 어려움
vocabulary 크기와 expression power 간의 균형
Step 1: Tokenization
1 - Unimodal Representation (Text)
10
Step 1: Tokenization
1 - Unimodal Representation (Text)
11
Step 1: Tokenization
1 - Unimodal Representation (Text)
12
Step 1: Tokenization
1 - Unimodal Representation (Text)
13
Initial Vocabulary:
{l, o, w, e, r, n, s, t, i, d}
Training Corpus:
Step 1: Tokenization
1 - Unimodal Representation (Text)
14
Frequency of Pair at Step 2:
Training Corpus:
Vocabulary at Step 2:
{l, o, w, e, r, n, s, t, i, d, lo}
At Step 1: low = “l” + “o” + “w”
At Step 2: low = “lo” + “w”
Step 1: Tokenization
1 - Unimodal Representation (Text)
15
Frequency of Pair at Step 2:
Training Corpus:
Vocabulary at Step 2:
{l, o, w, e, r, n, s, t, i, d, lo, low}
At Step 1: low = “l” + “o” + “w”
At Step 2: low = “lo” + “w”
At Step 3: low = “low”
Step 1: Tokenization
1 - Unimodal Representation (Text)
16
Training Corpus:
Final Vocabulary
{l, o, w, e, r, n, s, t, i, d, lo, low, es, est, er, …}
lower = low + er
newest = n + ew + est
…
Step 1: Tokenization
1 - Unimodal Representation (Text)
17
Step 1: Tokenization
1 - Unimodal Representation (Text)
18
Step 1: Tokenization
1 - Unimodal Representation (Text)
19
Step 2: Embedding
1 - Unimodal Representation (Text)
20
Step 2: Embedding
1 - Unimodal Representation (Text)
21
Step 2: Embedding
1 - Unimodal Representation (Text)
22
Step 2: Embedding
1 - Unimodal Representation (Text)
23
Step 2: Embedding
1 - Unimodal Representation (Text)
24
학습되는 부분
Step 2: Embedding
1 - Unimodal Representation (Text)
25
Step 2: Embedding
1 - Unimodal Representation (Text)
26
Step 2: Embedding
1 - Unimodal Representation (Text)
27
Step 3: Language Models
1 - Unimodal Representation (Text)
28
Step 3: Language Models
1 - Unimodal Representation (Text)
29
Step 3: Language Models
1 - Unimodal Representation (Text)
30
Step 3: Language Models
1 - Unimodal Representation (Text)
31
Step 3: Language Models
1 - Unimodal Representation (Text)
32
Step 3: Language Models
1 - Unimodal Representation (Text)
33
Step 3: Language Models
1 - Unimodal Representation (Text)
34
Step 3: Language Models
1 - Unimodal Representation (Text)
35
Step 3: Language Models
1 - Unimodal Representation (Text)
36
Summary
1 - Unimodal Representation (Text)
37
2 - Unimodal Representation (Image)
38
컴퓨터는 이미지를 어떻게 이해하는가?
2 - Unimodal Representation (Image)
39
컴퓨터는 이미지를 어떻게 이해하는가?
2 - Unimodal Representation (Image)
40
컴퓨터는 이미지를 어떻게 이해하는가?
2 - Unimodal Representation (Image)
41
Lincoln
Washington
Jefferson
Obama
Trump
0.8
0.05
0.05
0.01
0.09
Input Image
Pixel Representation
Prediction
컴퓨터는 이미지를 어떻게 이해하는가?
2 - Unimodal Representation (Image)
42
Convolution 이란 무엇인가?
2 - Unimodal Representation (Image)
43
Convolution 이란 무엇인가?
2 - Unimodal Representation (Image)
44
Convolution 이란 무엇인가?
2 - Unimodal Representation (Image)
45
Convolution 이란 무엇인가?
2 - Unimodal Representation (Image)
46
Convolution 이란 무엇인가?
2 - Unimodal Representation (Image)
47
Convolution 이란 무엇인가?
2 - Unimodal Representation (Image)
48
Convolution 이란 무엇인가?
2 - Unimodal Representation (Image)
49
Image
Kernel
Feature Map
Convolutional Neural Networks
2 - Unimodal Representation (Image)
50
Convolutional Neural Networks
2 - Unimodal Representation (Image)
51
Receptive Field
2 - Unimodal Representation (Image)
52
Receptive Field
2 - Unimodal Representation (Image)
53
Receptive Field
2 - Unimodal Representation (Image)
54
CNN → Vision Transformer (ViT)
2 - Unimodal Representation (Image)
55
Vision Transformer (ViT)
2 - Unimodal Representation (Image)
56
Vision Transformer (ViT)
2 - Unimodal Representation (Image)
57
Vision Transformer (ViT)
2 - Unimodal Representation (Image)
58
Vision Transformer (ViT)
2 - Unimodal Representation (Image)
59
Vision Transformer (ViT)
2 - Unimodal Representation (Image)
60
Vision Transformer (ViT)
2 - Unimodal Representation (Image)
61
Vision Transformer (ViT)
2 - Unimodal Representation (Image)
62
Vision Transformer (ViT)
2 - Unimodal Representation (Image)
63
CNN vs. ViT
2 - Unimodal Representation (Image)
[2010.11929] An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale �[2111.06377] Masked Autoencoders Are Scalable Vision Learners
64
| CNN | ViT |
귀납적 편향 (Inductive Bias) | 강함 (Locality, Translation Invariance, …) | 약함 (데이터로 학습) |
전역 관계 (Global Relations) | 깊은 층 필요 | 모든 층에서 가능 |
필요 데이터 | ✅ 소량으로도 잘 동작 | ❌ 일정량 이상에서 유의미 |
텍스트 통합 | 별도 설계 필요 | ✅ 동일 구조 공유 |
대표 모델 | ResNet, EfficientNet, … | ViT-B/16, MAE, … |
왜 ViT가 멀티모달 AI의 핵심이 되었는가?
2 - Unimodal Representation (Image)
65
3 - Unimodal Representation (Audio)
66
소리를 어떻게 숫자로 표현하는가?
3 - Unimodal Representation (Audio)
67
소리를 어떻게 숫자로 표현하는가?
3 - Unimodal Representation (Audio)
68
소리를 어떻게 숫자로 표현하는가?
3 - Unimodal Representation (Audio)
69
소리를 어떻게 숫자로 표현하는가?
3 - Unimodal Representation (Audio)
70
Mel Spectrogram (1): 왜 파형 그대로 쓰지 않는가?
3 - Unimodal Representation (Audio)
71
Mel Spectrogram (2): Spectrogram 이란?
3 - Unimodal Representation (Audio)
72
Mel Spectrogram (3): Fourier Transform & FFT
3 - Unimodal Representation (Audio)
73
Mel Spectrogram (4): STFT → Spectrogram 생성
3 - Unimodal Representation (Audio)
74
Mel Spectrogram (5): Mel Filter Bank 적용
3 - Unimodal Representation (Audio)
75
Mel Spectrogram (5): Mel Filter Bank 적용
3 - Unimodal Representation (Audio)
76
Mel Spectrogram (6): Log 스케일 → 최종 입력
3 - Unimodal Representation (Audio)
77
Mel Spectrogram (7): Summary
3 - Unimodal Representation (Audio)
78
Mel Spectrogram (7): Summary
3 - Unimodal Representation (Audio)
79
OpenAI – Whisper
3 - Unimodal Representation (Audio)
80
OpenAI – Whisper
3 - Unimodal Representation (Audio)
81
Today’s Summary
3 - Unimodal Representation (Audio)
82
| Text | Image | Audio |
Raw Data | 문자 시퀀스 | 픽셀 행렬 (H x W x 3) | 파형 (1D sequence) |
전처리 | Tokenization | 정규화/리사이즈 | Mel Spectrogram |
토큰/패치 | Sub-word token | 16 x 16 pixel patch | 25ms 프레임 |
최종 형태 | seq of 768 벡터 | seq of 768 벡터 | seq x 80 행렬 |
핵심 모델 | BERT, GPT, … | ViT, CLIP-ViT, … | Whisper, … |
Thank you!
Any questions?
83