1 of 39

Adaptive Budget Allocation for �Parameter-Efficient Fine Tuning��Qingru Zhang, Minshuo Chen, Alexander Bukharin, Pengcheg He, Yu Cheng, Weizhu Chen, Tuo Zhao ��ICLR 2023

Presented by: Meghna Kalra

Electrical and Computer Engineering

2 of 39

Introduction to Language model and fine tuning

  • Pretrained language model (PLM)

- Model trained on a vast corpus of text to understand language patterns/contexts

- Example – ChatGPT , Bert.

  • Fine tuning

- Adjusting a pre-trained model to specialize/adapt in a particular task.

  • Challenges of Fine tuning

- Memory cost (separate copy for each fine – tuned task )

- Computational cost (resource intensive , billions of parameters)

Fine Tuning all the parameter uniformly is not efficient !

3 of 39

Parameter-Efficient Fine-Tuning Methods

  •  

4 of 39

Low-Rank Adaptation (LoRA)

  •  

 

Instead of changing the entire weight matrix, only two smaller matrices are

adjusted !

 

5 of 39

Why Low rank works ?

The core function/behavior learned by the model resides in

a much lower-dimensional space than number of its parameters.

 

 

 

 

 

 

 

 

 

Applications: signal processing, optimization, computer algebra, machine learning ……..

  • Low rank approximation of an image
  • Full rank = 256, even with rank = 5,

we can identify Joseph Fourier.

Madathil, Baburaj, et al. "Tensor low rank modeling and its applications in signal processing.

6 of 39

Limitation of LoRA

  •  

Fine tuning FFN modules achieve better performance than self-attention.

Weight matrices in top layers are more important than those in bottom layers.

7 of 39

How can we allocate the parameter budget adaptively according to importance of modules to improve the performance of parameter-efficient fine-tuning ?

8 of 39

ADALoRA – Adaptive Low-rank Update

  • The core idea is to represent the incremental updates to the pre-trained weight matrices by using low rank approximation of Singular Value Decomposition (SVD).

 

Matrix dimensions

Rank

Computation cost

 

 

 

 

 

 

 

 

 

 

 

 

 

9 of 39

ADALoRA – Adaptive Low-rank Update

 

Structured Pruning in LoRA:

Adalora's Approach:

  • It prunes elements doublet-wise, which can permanently remove trainable parameters.
  • Once pruned, these elements cannot contribute to learning, potentially eliminating important model capacity.

  • It maintains the singular vectors even when singular values are pruned.
  • This allows the possibility of reactivating\recovering pruned components.

10 of 39

Notations

  •  

11 of 39

Importance – Aware Rank Allocation

  •  

12 of 39

 

  •  

13 of 39

Equation for global budget scheduler

  •  

 

14 of 39

Designing sensitivity-based “importance scores”

  •  

15 of 39

Refining Sensitivity for Robustness

  •  

16 of 39

AdaLoRA Algorithm

 

 

 

 

 

17 of 39

Experimental results

  • Natural language understanding - AdaLoRA applied to DeBERTaV3-base, evaluated on GLUE (General Language Understanding Evaluation) benchmark.
  • Compared with the following Baseline algorithms:
    • Full Fine-Tuning
      • All parameters updated
    • BitFit
      • Only fine-tunes bias vectors – parameter efficient method
    • Adapter Tuning
      • Inserts two-layer adapters between Transformer blocks
        • Houlsby Adapter: Between self-attention and FFN
        • Pfeiffer Adapter: After FFN and LayerNorm, more efficient
    • LoRA (Low-Rank Adaptation)
      • Incremental updates with two small matrices

18 of 39

AdaLoRA outperforms or equals existing methods across most datasets and budget levels

19 of 39

Comparison to low-rank parameterization.

 

 

Pruning LoRA doublet-wise to conduct the rank allocation. In this case, the doublets are zeroed out entirely, permanently removing the pruned parameters

AdaLoRA outperforms or equals existing methods across all low rank parameterizations

Importance scores:

20 of 39

Budget Distribution

AdaLoRA - Resulting rank of each incremental matrix

LoRA

This validates that the proposed importance metric in AdaLoRA focuses on crucial modules and layers.

21 of 39

Conclusion

  •  

22 of 39

 

23 of 39

IncreLoRA: Incremental Parameter Allocation Method for Parameter-Efficient Fine-tuning��Feiyu Zhang, Liangzhi Li, Junhao Chen, Zhouqiang Jiang, Bowen Wang, and Yiming Qian��

Presented by: Seonho Kim

Electrical and Computer Engineering

24 of 39

Recap

  • LoRA : gives the equal rank to each module matrix.

- Importance of each matrix is ignored.

- needs to give more rank to more important matrix.

  • AdaLoRA : mitigate these limitations using SVD form and pruning with importance scores.

- upper bound on the rank for each matrix is determined by the initial rank (the sum of ranks of total matrices decreases over iteration).

- cannot give more rank on a matrix than the initial rank even if the matrix is very important.

25 of 39

Basic idea of IncreLoRA

  • Increase the total rank automatically by giving more rank to important matrices without upper bound on rank.

26 of 39

Basic idea of IncreLoRA

  • Increase the total rank automatically by giving more rank to important matrices without upper bound on rank.

27 of 39

Methodology

 

  • Formulation

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

 

Pretrained

Weights

 

28 of 39

Methodology

  • SVD form

 

 

 

 

 

 

 

 

 

29 of 39

Methodology

  •  
  • Parameter Update

 

  • Parameter updates are applied to all linear layers.

30 of 39

Methodology

  • Incremental Parameter Allocation

 

SS with an exponential moving average helps to mitigate the effect of variability in the gradients across training iterations.

 

31 of 39

Methodology

  • Incremental Parameter Allocation

 

32 of 39

Methodology

  • Advanced Learning

 

33 of 39

Methodology

 

 

34 of 39

Experiments

Transformer : DeBERTaV3-base

Dataset : GLUE

  • Observation : IncreLoRA at the same budget level is likely to occupy more ranks to FFN modules, which have four times parameters than other modules.

35 of 39

Experiments

  • Different Budget Levels

  • IncreLoRA at the lowest budget outperforms the all performances of LoRA in considered budgets.

36 of 39

Experiments

  • Ablation Experiment

 

37 of 39

Experiments

  • Ablation Experiment

 

38 of 39

Conclusion

  • IncreLoRA improve the parameter efficiency compared with AdaLoRA.
  • IncreLoRA incrementally assign trainable parameters while improving training effect and stability using advance learning.
  • IncreLoRA outperforms the existing methods such as LoRA and AdaLoRA.

39 of 39

Remaining Questions

  • Is the Importance Scoring necessary?

-> 1. It depends on hyperparameters.

2. increase the computational cost.

3. Ignore the importance of singular values.

  • Is the Frobenius norm for orthogonalization the best choice?

  • The regularization to promote the low-rank structure has not been explored yet.