1 of 81

Large Language Models

Lecture 15

Parameter Efficient Fine-Tuning (PEFT)

Krishnendu Ghosh

2 of 81

Transfer Learning Before LLM Era

3 of 81

Transfer Learning Before LLM Era

4 of 81

Transfer Learning Before LLM Era

5 of 81

Transfer Learning Before LLM Era

6 of 81

Downsides of In-context Learning

7 of 81

Full Fine-tuning in LLM Challenging?

8 of 81

Parameter Efficient Fine Tuning

9 of 81

PEFT Advantages

10 of 81

PEFT Techniques

11 of 81

(Soft) Prompt Tuning (Lester et al. 2021)

12 of 81

(Soft) Prompt Tuning:

Multi-Task Serving

13 of 81

(Soft) Prompt Tuning:

Multi-Task Serving

14 of 81

(Soft) Prompt Tuning

15 of 81

(Soft) Prompt Tuning

16 of 81

(Soft) Prompt Tuning

17 of 81

PEFT Techniques

18 of 81

Prefix Tuning (Li & Liang 2021)

19 of 81

Prefix Tuning (Li & Liang 2021)

20 of 81

Prefix Tuning (Li & Liang 2021)

21 of 81

PEFT Techniques

22 of 81

Adapters (Houlsby et al 2019)

23 of 81

Adapters

24 of 81

Adapters

25 of 81

PEFT Techniques

26 of 81

Low Rank Composition

27 of 81

Intrinsic Dimensionality (ID)

28 of 81

Structure-Aware Intrinsic Dimension

29 of 81

Structure-Aware Intrinsic Dimension

30 of 81

Structure-Aware Intrinsic Dimension

31 of 81

Low Rank Adaptation (LoRA)

32 of 81

Low Rank Adaptation (LoRA)

33 of 81

Low Rank Adaptation (LoRA)

34 of 81

Low Rank Adaptation (LoRA)

35 of 81

LoRA: Effect on Weight Matrices

36 of 81

LoRA: Effect on Performance

37 of 81

LoRA Weights Initialization

38 of 81

Extensions of LoRA

39 of 81

PEFT Techniques

40 of 81

LLM Sizes

41 of 81

LLM Sizes

42 of 81

LLM Sizes

43 of 81

LLM Inference

44 of 81

LLM Inference

45 of 81

LLM Inference

46 of 81

Cost Effective Inference

47 of 81

Model Compression

48 of 81

Quantization: Problem with LLMs

49 of 81

Quantization: Numerical Values Representation

50 of 81

Quantizing FP32 to INT8

51 of 81

Dequantizing INT8 to FP32

52 of 81

53 of 81

Post Training Quantization (PTQ)

54 of 81

Post Training Quantization (PTQ)

55 of 81

Post Training Quantization (PTQ)

56 of 81

PTQ: LLM.int8() [Dettmers et al., 2022]

57 of 81

PTQ: LLM.int8() [Dettmers et al., 2022]

58 of 81

PTQ: LLM.int8() [Dettmers et al., 2022]

59 of 81

QLoRA [Dettmers et al. 2023]

60 of 81

QLoRA [Dettmers et al. 2023]

61 of 81

QLoRA [Dettmers et al. 2023]

62 of 81

QLoRA [Dettmers et al. 2023]

63 of 81

QLoRA [Dettmers et al. 2023]

64 of 81

QLoRA [Dettmers et al. 2023]

65 of 81

QLoRA [Dettmers et al. 2023]

66 of 81

QLoRA [Dettmers et al. 2023]

67 of 81

QLoRA [Dettmers et al. 2023]

68 of 81

QLoRA [Dettmers et al. 2023]

69 of 81

Pruning

70 of 81

Magnitude Pruning

[Han et al. 2015, See et al. 2016]

71 of 81

Wanda [Sun et al. 2023]

72 of 81

Wanda [Sun et al. 2023]

73 of 81

Unstructured Pruning

74 of 81

Structured Pruning

75 of 81

Structured Pruning [Xia et al. 2022]

76 of 81

Distillation [Hinton et al 2015]

77 of 81

Distillation [Hinton et al 2015]

78 of 81

Distillation [Hinton et al 2015]

79 of 81

Sequence Level Distillation

[Kim et al. 2016]

80 of 81

Self-Instruct [Wang et al. 2023]

81 of 81

Source: https://lcs2-iitd.github.io/ELL881-AIL821-2401/lectures/