Resource-Efficient and Cross-Modal Learning Toward Foundation Models
1
Dr. Huck Yang�Amazon Alexa Speech
Section 1 �(overview and advances of resource-efficient learning)�1hr 20 min
Dr. Pin-Yu Chen�IBM Research
Section 2 �(cross-modal reprogramming for foundation models)
50 min
Dr. Shalini Ghosh
Amazon Alexa Speech
Section 3
(cross-modal learning for speech recognition)�40 min
7 min Spotlight Featured Talks
2
Mr. Srijith Radhakrishnan�KAUST
Practices in Neural Adapters Designs and Cross-Modal
10:15 to 10:30 am
Prof. Marcel Worring �Universiteit van Amsterdam
GPT for Multi-Modal Processing
Backup Video
Dr. Chunyang Wu
Meta AI
Prompting LLM for ASR
12:20 to 12:30 pm
Overview: from Parameter-Efficient Learning (PEL) to Multimodal Adaptation
Dr. Huck Yang
Amazon Alexa ASR
9:05 am to 9:35 am: from Theory to Neural Modules (A)
9:35 am to 10:15 am: Advanced PEL Topics (B)
3
Tutorial 3 - Part 1A
In this Tutorial …
4
Parameter Efficient Learning (e.g., Reprograms & Prompts)
Huck Yang 2023
Pre-Trained
Model�
P
Trainable �Features
P
(e.g., Adapters, BitFit, LoRA)
Memory
Tutorial 3 - Part 1A
Pre-Tutorial Surveys (1/3) - Topics
5
Tutorial 3 - Part 1A
Pre-Tutorial Surveys (2/3) - Background
6
Tutorial 3 - Part 1A
Pre-Tutorial Surveys (3/3) - Insights
7
Tutorial 3 - Part 1A
Background: an eagle eye from 2021 to 2023
8
🦅
Tutorial 3 - Part 1A
Outline (mostly new)
9
Tutorial 3 - Part 1A
NOT to be covered in the Session but in Interspeech
Materials in previous editions or to be presented at Interspeech�
10
Tutorial 3 - Part 1A
What is Parameter-Efficient Learning? (1/2)
11
Tutorial 3 - Part 1A
What is Parameter-Efficient Learning? (2/3)
Connection to “Theoretical Background” of Neural Net (NN) based Approximation
Universal Approximation Theorem (AR Barren 1993)
12
w: weights
Idea of hypercube
s: activation function
x: intput
Tutorial 3 - Part 1A
[Recap] What is Parameter-Efficient Learning? (3/3)
13
(Qi et al. 2020 and Fan 2020 et al.)
data and its label pairs of (x, y)
Frozen Pre-trained Model �(e.g., PLM or PAM)
Parameter-Efficient Adaptation
Tutorial 3 - Part 1A
[Recap] Pre-Trainings to Prompting / Reprogramming (1/2)
Say we have a “pre-training model” and only few training samples...
14
It is hard to fine-tune a large pretrained model (source domain) with few target domain samples. �Due to: (1) Small training data, (2) Domain mismatching, (3) Smoothness.��
New test data are often outliers sampling from a (low-resource) target domain.
pre-trained decision boundary
fine-tuned transfer learning
Illustration drafted by huck yang 2021 ICML
Tutorial 3 - Part 1A
[Recap] From Pre-Training to Model Reprogramming (2/2)
The representation power of the pre-training model is good but …��“Could we also use the established decision boundaries from pre-trains?”
15
source decision boundaries
class a
class b
class c
class d
Mismatching domains, label numbers, ...
Target Data Classes 1�Target Data Classes 2
Source Data Classes 1�Source Data Classes 2
Illustration drafted by huck yang ICML 2021 Voice2Series
Trainable* ?
P
Tutorial 3 - Part 1A
[Recap] Cross-Modal Reprogram from Speech to Time Series
16
Reprogram Layer
xt
(a)
Pretrained AM
(b)
xt’
Label Mapping
(c)
ys
yt
Target �(e.g., ECG)
Reprogrammed �(e.g., 16k sampling rate speech)
source output
target output
R
Tutorial 3 - Part 1A
[Recap] Reprogram = Trainable Input + Label Mapping
17
nine
Ys
no
Label Mapping
Yt
source class probabilities
target class probabilities
ne
Reprogram Layer
trainable
noise θ
(prompt)
Xt
Lithuanian “ne”
Pretrained or Frozen AM
Xt’
reprogrammed
as English “no”
R
Hao Yen et al, Interspeech 2023, (Best Student Paper Candidate) �Code: SpeechReprogramSimularity in Arxiv 2021
Tutorial 3 - Part 1A
[Recap] Audio Examples and Similar Mapping
18
Original
Reprogrammed
Dysarthric Mandarin��一 [yi] (one)
Lithuanian��vienas (one)
Hao Yen et al, Interspeech 2023, (Best Student Paper Candidate) �Code: SpeechReprogramSimularity in Arxiv 2021
Tutorial 3 - Part 1A
[Recap] SpeechPrompt Tuning on Generative Spoken LM
19
Tutorial 3 - Part 1A
[Recap] Theoretical Analysis for Reprogram / Prompt
20
This results suggest that reprogram / prompt can perform better (lower risk) when the source model has a lower source loss and smaller representation loss.
Presented by c.-h. huck yang et al. ICML 2021
Tutorial 3 - Part 1A
Architecture Introductions
21
Tutorial 3 - Part 1A
Saying … a massive stacked-transformer model
22
Adding Trainable Input or Tokens
Reprogram & Prompt
1
2
3
4
Tuning Key, Query, and Value
Prompt* & Adapter
Adding Trainable Latent Variables
Adapter* & Reprogram & Prompt
Output Mapping or Verbalizer
Reprogram* & Prompt (* for mainly)
Tutorial 3 - Part 1A
Reprogramming vs. Prompting: Any Difference?
23
Tutorial 3 - Part 1A
Recent Works for Speech and Acoustic Modeling
24
Conformer ASR Reprogramming* [ICASSP 23]
Speech to Music Reprogramming [ICASSP 23]
Whisper Token Reprogramming [Interspeech 23]
SpeechGen [Arxiv 23]
TTS Accent Adaptation Reprogramming [Interspeech 23]
SpeechPromptV2 [Arxiv 23]
Whisper Prompting [Interspeech 23]
Multilingual Spoken Command [Interspeech 23; Arxiv 21]
Tutorial 3 - Part 1A
Take Home Message 1: Who should work on PEL?
25
Tutorial 3 - Part 1A
Generic Concept of “Modular Deep Learning” (1/2)
26
Tutorial 3 - Part 1B
Generic Concept of “Modular Deep Learning” (2/2)
Modular Domain Adaptation for Conformer-Based Streaming ASR [Interspeech 2023]
27
Tutorial 3 - Part 1B
Recent Development of Neural (Residual) Adapter
28
Tutorial 3 - Part 1B
Convolution Adapter for Vision
29
Residual adapter in a ResNet (adapter parameters are in blue) [Rebuffi et al., 2017]
Tutorial 3 - Part 1B
Adapter for NLP (1/2)
30
Tutorial 3 - Part 1B
NLP Adapter (2/2)
31
Tutorial 3 - Part 1B
Adapter in Speech Processing (1/2)
32
Tutorial 3 - Part 1B
Adapter in Speech Processing (2/2)
33
Tutorial 3 - Part 1B
Efficient On-Device Learning via Feature Reprogramming (1/5)
34
Tutorial 3 - Part 1B
Efficient On-Device Learning via Feature Reprogramming (2/5)
35
Tutorial 3 - Part 1B
Efficient On-Device Learning via Feature Reprogramming (3/5)
36
Tutorial 3 - Part 1B
Efficient On-Device Learning via Feature Reprogramming (4/5)
37
Tutorial 3 - Part 1B
Efficient On-Device Learning via Feature Reprogramming (5/5)
38
How to Mix Pre-Training Languages?
Can PEL work for SSL Speech Models over SOTA?
Tutorial 3 - Part 1B
[Recap] How to Estimate Pre-trained Speech Models? (1/3)
39
Zih-Ching Chen, C.-H. Huck Yang et al, to Appear Interspeech 2023
a joint work with National Taiwan University, Google, and Georgia Tech
Targeted
Speech Data
Zero-Shot
Encoded Features
👑 Best Layer
👑 Best Model
Score �Assessments
Step 1
Pretrained AM
Pre-trained �Speech Models (SMs)
Pretrained SM
Step 2
Step 3
Step 4
Tutorial 3 - Part 1B
[Recap] Estimate Pre-trained Speech Models? (2/3)
40
Image Source from Zih-Ching Chen
C.-H. H. Yang et al. ICML 21
K. You et al. ICML 21
Zih-Ching Chen et al, to Appear Interspeech 2023
Tutorial 3 - Part 1B
[Recap] Estimate Pre-trained Speech Models? (3/3)
41
Zih-Ching Chen et al, to Appear Interspeech 2023
Tutorial 3 - Part 1B
42
Tutorial 3 - Part 1B
Fun Photos with Prof. Chin-Hui Lee before Interspeech 23
Got a Forced Landing at Gander on the Way of a Direct Flight to Ireland
43
Overview: from Parameter-Efficient Learning (PEL) to Multimodal Adaptation
Dr. Huck Yang
Amazon Alexa ASR Science
9:05 am to 9:35 am: from Theory to Neural Modules
9:35 am to 10:15 am: Advanced PEL Topics
44
Tutorial 3 - Part 1B
BitFit: Bias-Only Fine-Tuning (1/2)
45
Tutorial 3 - Part 1A
BitFit: Bias-Only Fine-Tuning (2/2)
46
Q and GELU
Low-Rank Adaptation (LoRA) Background (1/2)
47
Hu, Edward J., et al. "Lora: Low-rank adaptation of large language models." arXiv preprint arXiv:2106.09685 (2021).
�
🔥
🔥
🔥
Tutorial 3 - Part 1A
Low-Rank Adaptation (LoRA) Background (2/2)
48
Hu, Edward J., et al. "Lora: Low-rank adaptation of large language models." arXiv preprint arXiv:2106.09685 (2021).
�
Tutorial 3 - Part 1A
Limitations of Low-Rank Adaptation (LoRA)
49
rank
Tutorial 3 - Part 1A
LoRA Advances – Adaptive Rank Selection (1/3)
50
Zhang, Qingru, et al. "Adaptive budget allocation for parameter-efficient fine-tuning." ICLR 2023.
P
Q
Pretrained Weight k
x
h
Tutorial 3 - Part 1A
LoRA Advances – Adaptive Rank Selection (2/3)
51
Zhang, Qingru, et al. "Adaptive budget allocation for parameter-efficient fine-tuning." ICLR 2023.
P
Q
Pretrained Weight k
x
h
Tutorial 3 - Part 1A
LoRA Advances – Adaptive Rank Selection (3/3)
52
Zhang, Qingru, et al. "Adaptive budget allocation for parameter-efficient fine-tuning." ICLR 2023.
P
Q
Pretrained Weights
x
h
Tutorial 3 - Part 1A
LoRA Advances – ReLoRA Pre-Training (1/2)
53
Lialin, Vladislav, et al. "Stack More Layers Differently: High-Rank Training Through Low-Rank Updates." arXiv preprint arXiv:2307.05695 (2023).
Tutorial 3 - Part 1A
LoRA Advances – ReLoRA Pre-Training (2/2)
54
Lialin, Vladislav, et al. "Stack More Layers Differently: High-Rank Training Through Low-Rank Updates." arXiv preprint arXiv:2307.05695 (2023).
Tutorial 3 - Part 1A
In-Context Learning Basics (1/3)
55
How can a deployed model can still learn from input?�
Image Source from Xie et al. 2022
Tutorial 3 - Part 1A
In-Context Learning in Theories (2/3)
56
Theory
Methods
Kazuki Irie, Róbert Csordás, and Jürgen Schmidhuber ICML 2022
D. Da et al. 2022
Tutorial 3 - Part 1A
In-Context Learning in Theories (3/3)
57
Tutorial 3 - Part 1A
CoT: Chains of Thoughts (1/3)
58
Tutorial 3 - Part 1A
Chains of Thoughts Performance (2/3)
59
Tutorial 3 - Part 1A
Chains of Thoughts Performance (3/3)
60
ICASSP 2024 Special Session of In-Context Learning
61
(New) Multi-Modal Weights Merging
An Empirical Study of Multimodal Model Merging, Arxiv 2023
62
Conclusion
63
Related Works Presented in this Talk
64
Acknowledgments (1/2)
65
Acknowledgments (2/2)
66
More References
67
Hu, Edward J., et al. "Lora: Low-rank adaptation of large language models." arXiv preprint arXiv:2106.09685 (2021).
�
Eger, Steffen, and Yannik Benz. "From Hero to Z\'eroe: A Benchmark of Low-Level Adversarial Attacks.” AACL2020.
Zhang, Qingru, et al. "Adaptive budget allocation for parameter-efficient fine-tuning." ICLR 2023.
Lialin, Vladislav, et al. "Stack More Layers Differently: High-Rank Training Through Low-Rank Updates." arXiv preprint arXiv:2307.05695 (2023).
Wang, Xuezhi, Haohan Wang, and Diyi Yang. "Measure and improve robustness in nlp models: A survey.” NAACL 2022.
Yu, Yu, Abdul Rafae Khan, and Jia Xu. "Measuring Robustness for NLP." COLING 2022.
Smorga's Board. ” Frequently Misspelled Word List for Dyslexia.”, https://www.teacherspayteachers.com/Product/Frequently-Misspelled-Word-List-for-Dyslexia-5295631
.
Huang, Chengsong, et al. "LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition." arXiv preprint arXiv:2307.13269 (2023).
7 min Spotlight Featured Talks
68
Dr. Chunyang Wu
Meta AI
Prompting LLM for ASR
12:20 to 12:30 pm