οΏ½NeuronMM: High-Performance οΏ½Matrix MultiplicationοΏ½for LLM Inference on AWS TrainiumοΏ½οΏ½
Dinghong Song
University of California, Merced
PASA Lab
PASA Lab
Introduction
2
PASA Lab
Introduction
3
PASA Lab
Introduction
for the tensor engine
4
PASA Lab
Introduction
5
PASA Lab
Motivation
6
PASA Lab
Motivation
7
PASA Lab
Motivation
8
PASA Lab
NeuronMM
9
PASA Lab
Block-Aligned SVD
10
PASA Lab
Block-Aligned SVD
11
PASA Lab
LoRA Fine-Tuning
12
PASA Lab
XUV NKI Kernel
13
PASA Lab
MLP NKI Kernel
14
PASA Lab
Evaluation
15
PASA Lab
Evaluation of πππ kernel
16
PASA Lab
Evaluation of πππ kernel
17
PASA Lab
Evaluation of πππ kernel
18
PASA Lab
Evaluation of MLP Kernel
19
PASA Lab
Impact of Block Size on Kernel Performance
HBM and degrade performance.
20
PASA Lab
Case Study
21
PASA Lab
Conclusion
22
PASA Lab
οΏ½οΏ½οΏ½Thank you!
Dinghong Song
University of California, Merced
PASA Lab
PASA Lab