1 of 68

Resource-Efficient and Cross-Modal Learning Toward Foundation Models

1

Dr. Huck Yang�Amazon Alexa Speech

Section 1 �(overview and advances of resource-efficient learning)�1hr 20 min

Dr. Pin-Yu Chen�IBM Research

Section 2 �(cross-modal reprogramming for foundation models)

50 min

Dr. Shalini Ghosh

Amazon Alexa Speech

Section 3

(cross-modal learning for speech recognition)�40 min

2 of 68

7 min Spotlight Featured Talks

2

Mr. Srijith Radhakrishnan�KAUST

Practices in Neural Adapters Designs and Cross-Modal

10:15 to 10:30 am

Prof. Marcel Worring �Universiteit van Amsterdam

GPT for Multi-Modal Processing

Backup Video

Dr. Chunyang Wu

Meta AI

Prompting LLM for ASR

12:20 to 12:30 pm

3 of 68

Overview: from Parameter-Efficient Learning (PEL) to Multimodal Adaptation

Dr. Huck Yang

Amazon Alexa ASR

9:05 am to 9:35 am: from Theory to Neural Modules (A)

9:35 am to 10:15 am: Advanced PEL Topics (B)

3

Tutorial 3 - Part 1A

4 of 68

In this Tutorial …

4

Parameter Efficient Learning (e.g., Reprograms & Prompts)

Huck Yang 2023

Pre-Trained

Model�

P

Trainable �Features

P

(e.g., Adapters, BitFit, LoRA)

Memory

Tutorial 3 - Part 1A

5 of 68

Pre-Tutorial Surveys (1/3) - Topics

5

Tutorial 3 - Part 1A

6 of 68

Pre-Tutorial Surveys (2/3) - Background

6

Tutorial 3 - Part 1A

7 of 68

Pre-Tutorial Surveys (3/3) - Insights

7

Tutorial 3 - Part 1A

8 of 68

Background: an eagle eye from 2021 to 2023

  • ICASSP 22: Frozen Model Adaptation
    • Neural Model Reprogramming (Yang et al. ICML 21)
    • Review Prompting and Prefix-Tuning from NLP to Speech
  • Interspeech 23: from Parameter-efficient learning to Multimodal Adaptation
    • Memory-efficient Learning
    • Multimodal Weight Merging

8

🦅

  • ICASSP 23: Auxiliary Objectives, SSM, In-Context Learning
    • Part 1. How Parameter-Efficient Learning works in Theories
    • Part 2. Recent Advances of Parameter Efficient Speech Processing
      • Frozen Models Adaptation
      • Neural Space State Machines (SSM)
    • Part 3. From Parameter-Efficient Adaptation to In-Context Learning

Tutorial 3 - Part 1A

9 of 68

Outline (mostly new)

  • Basic Definition and Literatures of Parameter Efficient Learning (PEL)
    • Connection to the Universal Approximation Theory (5 min)
    • “How to Estimate a best model or layer to tune?” [Interspeech 23]
    • (new) Standard PEL Modules

  • Open Problems and Opportunities in PEL
    • (new) Generalization and Stability
    • “Multi-loss Adapter Training” [Interspeech 23]
    • “Privacy-enhanced Frozen Model Adaptation“ [Interspeech 23]
    • (new) Data and Model Scaling Law of PEL�
  • (new) Memory Efficient Learning
    • Quantization and Serialization
  • (new) Advances in Low-rank adaptation over 1B models
  • (new) Modular Deep Learning
    • Design principles and limitation of current public codebase

9

Tutorial 3 - Part 1A

10 of 68

NOT to be covered in the Session but in Interspeech

Materials in previous editions or to be presented at Interspeech�

  • Specific works on Multilingual and Cross-Modal Learning
    • ASR Conformer Reprogramming [ICASSP 23]
    • Speech or Vision Transformer to Music Reprogramming [ICASSP 23]
    • General Speech Model to Specific Task Reprogramming [Interspeech 23]

  • Other parameter efficient learning architectures or methods
    • Neural State Space Machine (SSM)
      • SSM for Enhancement [Interspeech 23]
      • SSM for Separation [Interspeech 23]
    • Bayesian and Non-Parametric Adaptation (e.g., kNN)�

10

Tutorial 3 - Part 1A

11 of 68

What is Parameter-Efficient Learning? (1/2)

  • Background of Frozen Model Adaptation
    • Related Works
      • Frozen Model Distillation
      • Bayesian Adaptation
    • Frozen Representation Learning
      • Neural Adapter
      • Prompt-Tuning
      • Model Reprogramming
      • k-Nearest Neighbors based DataStore Inference

11

Tutorial 3 - Part 1A

12 of 68

What is Parameter-Efficient Learning? (2/3)

Connection to “Theoretical Background” of Neural Net (NN) based Approximation

Universal Approximation Theorem (AR Barren 1993)

12

  • Any Continuous Function u(x) can based approximated by NN
  • The test error is upper bounded by ɛu

w: weights

Idea of hypercube

s: activation function

x: intput

Tutorial 3 - Part 1A

13 of 68

[Recap] What is Parameter-Efficient Learning? (3/3)

13

(Qi et al. 2020 and Fan 2020 et al.)

data and its label pairs of (x, y)

Frozen Pre-trained Model �(e.g., PLM or PAM)

Parameter-Efficient Adaptation

Tutorial 3 - Part 1A

14 of 68

[Recap] Pre-Trainings to Prompting / Reprogramming (1/2)

Say we have a “pre-training model” and only few training samples...

14

It is hard to fine-tune a large pretrained model (source domain) with few target domain samples. �Due to: (1) Small training data, (2) Domain mismatching, (3) Smoothness.��

New test data are often outliers sampling from a (low-resource) target domain.

pre-trained decision boundary

fine-tuned transfer learning

Illustration drafted by huck yang 2021 ICML

Tutorial 3 - Part 1A

15 of 68

[Recap] From Pre-Training to Model Reprogramming (2/2)

The representation power of the pre-training model is good but …��“Could we also use the established decision boundaries from pre-trains?”

15

source decision boundaries

class a

class b

class c

class d

Mismatching domains, label numbers, ...

Target Data Classes 1�Target Data Classes 2

Source Data Classes 1�Source Data Classes 2

Illustration drafted by huck yang ICML 2021 Voice2Series

Trainable* ?

P

  • 19 out 29 Time Series Classification SOTA in 2021
  • Optimal Transport based Approximation in ICASSP 2022 Tutorial & �Dr. Pin-Yu Chen’s talk

Tutorial 3 - Part 1A

16 of 68

[Recap] Cross-Modal Reprogram from Speech to Time Series

16

  • Schematic illustration of the proposed Voice2Series Neural Reprogramming�(CHH Yang et al. ICML 2021)

Reprogram Layer

xt

(a)

Pretrained AM

(b)

xt

Label Mapping

(c)

ys

yt

Target �(e.g., ECG)

Reprogrammed �(e.g., 16k sampling rate speech)

source output

target output

R

Tutorial 3 - Part 1A

17 of 68

[Recap] Reprogram = Trainable Input + Label Mapping

17

  • Schematic illustration of the proposed reprogramming for multilingual speech

nine

Ys

no

Label Mapping

Yt

source class probabilities

target class probabilities

ne

Reprogram Layer

trainable

noise θ

(prompt)

Xt

Lithuanian “ne”

Pretrained or Frozen AM

Xt’

reprogrammed

as English “no”

R

Hao Yen et al, Interspeech 2023, (Best Student Paper Candidate) �Code: SpeechReprogramSimularity in Arxiv 2021

Tutorial 3 - Part 1A

18 of 68

[Recap] Audio Examples and Similar Mapping

18

  • Label mapping based on cosine similarity
  • We beat wav2vec pre-trained models in Lithuanian, Dysarthric Mandarin, and Arabic Speech.

Original

Reprogrammed

Dysarthric Mandarin��一 [yi] (one)

Lithuanian��vienas (one)

Hao Yen et al, Interspeech 2023, (Best Student Paper Candidate) �Code: SpeechReprogramSimularity in Arxiv 2021

Tutorial 3 - Part 1A

19 of 68

[Recap] SpeechPrompt Tuning on Generative Spoken LM

19

Tutorial 3 - Part 1A

20 of 68

[Recap] Theoretical Analysis for Reprogram / Prompt

  • Population Risk Analysis via Trainable Inputs based on Optimal Transport

20

This results suggest that reprogram / prompt can perform better (lower risk) when the source model has a lower source loss and smaller representation loss.

Presented by c.-h. huck yang et al. ICML 2021

Tutorial 3 - Part 1A

21 of 68

Architecture Introductions

  • Neural Architectures
    • Input-Only Frozen Model Adaption
      • Reprogramming
      • Prompting
    • Latent Space Neural Adapter
      • Convolution Adapter
      • Low-Rank Adaptation (LoRA)

21

Tutorial 3 - Part 1A

22 of 68

Saying … a massive stacked-transformer model

22

Generative Pre-trained Transformer (GPT)

GPT Image Source: GPT-3 An Overview by dzlab

Adding Trainable Input or Tokens

Reprogram & Prompt

1

2

3

4

Tuning Key, Query, and Value

Prompt* & Adapter

Adding Trainable Latent Variables

Adapter* & Reprogram & Prompt

Output Mapping or Verbalizer

Reprogram* & Prompt (* for mainly)

Tutorial 3 - Part 1A

23 of 68

Reprogramming vs. Prompting: Any Difference?

  • In terms of their semantic meaning, “reprogramming” aims to change the behaviors of existing software systems or agents' learning
    • Mainly used in ML, Computer Security, Single Cell Biology �

23

  • "prompting," as a way of illustrating behavior, focuses on injecting information.
    • Mainly used in NLP Literatures but also World-Level Adversarial Reprogramming (ACL 21)

  • In speech processing, the first reprogramming work happened a year earlier than prompting works, but it was still within the similar period.
    • Recommend to cite both when writing a paper discussion injecting trainable input toward models
    • VoiceToSeries (ICML 21); SpeechPrompt & WavPrompt (Interspeech 22)

Tutorial 3 - Part 1A

24 of 68

Recent Works for Speech and Acoustic Modeling

  • Input Information Injection

24

  • Voice2Series [ICML 21]

Conformer ASR Reprogramming* [ICASSP 23]

Speech to Music Reprogramming [ICASSP 23]

Whisper Token Reprogramming [Interspeech 23]

  • SpeechPrompt [Interspeech 22]
  • WavPrompt [Interspeech 22]

SpeechGen [Arxiv 23]

TTS Accent Adaptation Reprogramming [Interspeech 23]

SpeechPromptV2 [Arxiv 23]

Whisper Prompting [Interspeech 23]

Multilingual Spoken Command [Interspeech 23; Arxiv 21]

  • Fairness Reprogramming [NeurIPS 22]
  • Label Space Mapping Design [CVPR 22]
  • LLM for Antibody Sequence Infilling �[ICML 23; Natural Machine Intelligence 23]
  • More Detailed NLP Related works Can be Referred to Mr. Cheng-Han Chiang advised by Prof. Hung-yi Lee (NTU)’s AACL Tutorial and his ICASSP 23 Tutorial

Tutorial 3 - Part 1A

25 of 68

Take Home Message 1: Who should work on PEL?

  • If you have been working on adversarial robustness or speaker vector
    • The Problem Framework is Similar
      • Having a Frozen Deployed Model
      • Trainable Perturbation or Speaker Vector
      • Also, for the Contextual Biasing for ASR Modeling�
  • If you have been working on Bayesian Adaptation
    • Additive Loss or Margin based Approximation is unexplored in PEL
    • The Theoretical Connection on Frozen Model Adaptation �
  • If you have been working on Multi-Loss and Multi-Modal Training
    • Bringing the Representation Power from Other Modals

25

Tutorial 3 - Part 1A

26 of 68

Generic Concept of “Modular Deep Learning” (1/2)

26

Tutorial 3 - Part 1B

27 of 68

Generic Concept of “Modular Deep Learning” (2/2)

27

Tutorial 3 - Part 1B

28 of 68

Recent Development of Neural (Residual) Adapter

  • A Modern Format of Neural (Residual) Adapter
    • Rebuffi, CVPR 2017
    • Introduced the concept of extra tunable weight for ResNet
  • Parameter-efficient learning of NLP tasks
    • Houlsby et al., ICML 2019
    • Highlighted the idea of adding adapters to transformer-based pre-trained models (BERT)
    • Up and down projectors with GELU activation layers
  • Some very closed early works in speech community
    • x-Vector and i-Vector
    • Likelihood Adaptation Methods

28

Tutorial 3 - Part 1B

29 of 68

Convolution Adapter for Vision

  • Rebuffi et al., 2017
  • First introduced the concept of “tunable deep network architecture”
  • Parametrize the ResNet module with convolutional layers

29

Residual adapter in a ResNet (adapter parameters are in blue) [Rebuffi et al., 2017]

Tutorial 3 - Part 1B

30 of 68

Adapter for NLP (1/2)

  • Houlsby et al., 2019
  • Adapters for Pretrained LM
    • After BERT was introduced
    • Houlsby added extra tunable weights to pretrained Transformer models.
  • Proposed a fixed Architecture
    • Down & Up Projections.
    • GELU activation layer.

30

Tutorial 3 - Part 1B

31 of 68

NLP Adapter (2/2)

  • Fu et al., NAACL 2022 (NTU)
  • Proposed token-dependent representation shift for NLP tasks
  • Further reduces the trainable parameters.

31

Tutorial 3 - Part 1B

32 of 68

Adapter in Speech Processing (1/2)

  • Tomanek et al., 2021
  • Proposed to add adapters to ASR models for atypical speech adaptation.

32

Tutorial 3 - Part 1B

33 of 68

Adapter in Speech Processing (2/2)

  • Chen et al., 2022, SLT 2022 (NTU)
  • Evaluated adapters on SSL speech models on multiple downstream speech tasks
  • Chen et al., 2023, ICASSP 2023 (NTU)
  • Proposed adding CNN adapters on the feature extractor for speaker adaptation

33

Tutorial 3 - Part 1B

34 of 68

Efficient On-Device Learning via Feature Reprogramming (1/5)

  • Adapter with Neural Reprogramming and Fixed Activation

34

Tutorial 3 - Part 1B

35 of 68

Efficient On-Device Learning via Feature Reprogramming (2/5)

35

Tutorial 3 - Part 1B

36 of 68

Efficient On-Device Learning via Feature Reprogramming (3/5)

36

Tutorial 3 - Part 1B

37 of 68

Efficient On-Device Learning via Feature Reprogramming (4/5)

37

Tutorial 3 - Part 1B

38 of 68

Efficient On-Device Learning via Feature Reprogramming (5/5)

38

How to Mix Pre-Training Languages?

Can PEL work for SSL Speech Models over SOTA?

Tutorial 3 - Part 1B

39 of 68

[Recap] How to Estimate Pre-trained Speech Models? (1/3)

39

Zih-Ching Chen, C.-H. Huck Yang et al, to Appear Interspeech 2023

a joint work with National Taiwan University, Google, and Georgia Tech

Targeted

Speech Data

Zero-Shot

Encoded Features

👑 Best Layer

👑 Best Model

Score �Assessments

Step 1

Pretrained AM

Pre-trained �Speech Models (SMs)

Pretrained SM

Step 2

Step 3

Step 4

Tutorial 3 - Part 1B

40 of 68

[Recap] Estimate Pre-trained Speech Models? (2/3)

40

Image Source from Zih-Ching Chen

C.-H. H. Yang et al. ICML 21

K. You et al. ICML 21

Zih-Ching Chen et al, to Appear Interspeech 2023

Tutorial 3 - Part 1B

41 of 68

[Recap] Estimate Pre-trained Speech Models? (3/3)

41

  • 17 Layers Conformer RNN-T

  • HuBERT in SUPERB Benchmark

Zih-Ching Chen et al, to Appear Interspeech 2023

Tutorial 3 - Part 1B

42 of 68

  • Website (1 to 3 pages)
    • Bayesian methods for machine learning and deep learning�
    • In-context learning and generative models�
    • Adaptation and few-shot learning for speech and language processing�
    • Theory and parameter efficient tuning for large speech and language models (LLMs)�
    • Multimodal intelligence across audio, text, and vision

42

Tutorial 3 - Part 1B

43 of 68

Fun Photos with Prof. Chin-Hui Lee before Interspeech 23

Got a Forced Landing at Gander on the Way of a Direct Flight to Ireland

43

44 of 68

Overview: from Parameter-Efficient Learning (PEL) to Multimodal Adaptation

Dr. Huck Yang

Amazon Alexa ASR Science

9:05 am to 9:35 am: from Theory to Neural Modules

9:35 am to 10:15 am: Advanced PEL Topics

44

Tutorial 3 - Part 1B

45 of 68

BitFit: Bias-Only Fine-Tuning (1/2)

45

Tutorial 3 - Part 1A

46 of 68

BitFit: Bias-Only Fine-Tuning (2/2)

46

Q and GELU

47 of 68

Low-Rank Adaptation (LoRA) Background (1/2)

47

  • LoRA [1] for frozen model adaptation
    • For pre-trained weights, apply low rank decomposition to the change of weights

    • WB = 0; WA ~ N (0, σ2)

Hu, Edward J., et al. "Lora: Low-rank adaptation of large language models." arXiv preprint arXiv:2106.09685 (2021).

🔥

🔥

🔥

Tutorial 3 - Part 1A

48 of 68

Low-Rank Adaptation (LoRA) Background (2/2)

48

Hu, Edward J., et al. "Lora: Low-rank adaptation of large language models." arXiv preprint arXiv:2106.09685 (2021).

Tutorial 3 - Part 1A

49 of 68

Limitations of Low-Rank Adaptation (LoRA)

  • Can rank search process be automated?
    • Which rank to use?
    • What target modules/layers to insert LoRA?
      • Target modules:
        • Self-attention: Wq, Wk, Wv, Wo
        • Two-layer FFN: Wf1, Wf2

49

rank

Tutorial 3 - Part 1A

50 of 68

LoRA Advances – Adaptive Rank Selection (1/3)

 

50

Zhang, Qingru, et al. "Adaptive budget allocation for parameter-efficient fine-tuning." ICLR 2023.

P

Q

 

Pretrained Weight k

x

h

 

Tutorial 3 - Part 1A

51 of 68

LoRA Advances – Adaptive Rank Selection (2/3)

 

51

Zhang, Qingru, et al. "Adaptive budget allocation for parameter-efficient fine-tuning." ICLR 2023.

P

Q

 

Pretrained Weight k

x

h

 

Tutorial 3 - Part 1A

52 of 68

LoRA Advances – Adaptive Rank Selection (3/3)

 

52

Zhang, Qingru, et al. "Adaptive budget allocation for parameter-efficient fine-tuning." ICLR 2023.

P

Q

 

Pretrained Weights

x

h

 

Tutorial 3 - Part 1A

53 of 68

LoRA Advances – ReLoRA Pre-Training (1/2)

53

Lialin, Vladislav, et al. "Stack More Layers Differently: High-Rank Training Through Low-Rank Updates.arXiv preprint arXiv:2307.05695 (2023).

Tutorial 3 - Part 1A

54 of 68

LoRA Advances – ReLoRA Pre-Training (2/2)

54

Lialin, Vladislav, et al. "Stack More Layers Differently: High-Rank Training Through Low-Rank Updates.arXiv preprint arXiv:2307.05695 (2023).

Tutorial 3 - Part 1A

55 of 68

In-Context Learning Basics (1/3)

55

How can a deployed model can still learn from input?�

  • autoregressive learning
  • recurrent network
  • attention head

Image Source from Xie et al. 2022

Tutorial 3 - Part 1A

56 of 68

In-Context Learning in Theories (2/3)

56

Theory

  • Bayesian View of In-Context Learning
  • Dual-Form of Gradients

Methods

  • Demonstration based Tuning
  • Instruction based Tuning

Kazuki Irie, Róbert Csordás, and Jürgen Schmidhuber ICML 2022

  • The dual form of neural networks revisited: Connecting test time predictions to training patterns via spotlights of attention.

D. Da et al. 2022

  • Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Tutorial 3 - Part 1A

57 of 68

In-Context Learning in Theories (3/3)

57

Tutorial 3 - Part 1A

58 of 68

CoT: Chains of Thoughts (1/3)

58

Tutorial 3 - Part 1A

59 of 68

Chains of Thoughts Performance (2/3)

59

Tutorial 3 - Part 1A

60 of 68

Chains of Thoughts Performance (3/3)

60

61 of 68

ICASSP 2024 Special Session of In-Context Learning

  • Special Session
    • Co-Chairs: Prof. Chao Zhang (THU and Cambridge), Prof. Marco Marco Siniscalchi (NTNU), Dr. Huck Yang (Amazon)
    • Please Share 1 Page of Abstract Before Aug 26th

61

62 of 68

(New) Multi-Modal Weights Merging

62

63 of 68

Conclusion

  • Understand the theory helps to design better parameter efficient learning modules�
  • Multi-losses or Low-rank adaptation could be a promising direction for both adapters, prompting, and reprogramming�
  • In-Context Learning serves a parameter-efficient layer itself as well as new test-time adaptation framework

63

64 of 68

Related Works Presented in this Talk

  1. Differentially Private Adapters for Parameter Efficient Acoustic Modeling, to appear Interspeech 23�
  2. Parameter-Efficient Learning for Text-to-Speech Accent Adaptation, to appear Interspeech 23�
  3. A Parameter-Efficient Learning Approach to Arabic Dialect Identification with Pre-Trained General-Purpose Speech Model, to appear Interspeech 23�
  4. How to Estimate Model Transferability of Pre-Trained Speech Models?, to appear Interspeech 23

64

65 of 68

Acknowledgments (1/2)

  • Amazon Alexa: Dr. Ivan Bulyko, Dr. Andreas Stolcke, Dr. Shalini Ghosh, Dr. Ariya Rastrow and Dr. Björn Hoffmeister�
  • Georgia Tech: Prof. Alex Lerch and Prof. Chin-Hui Lee�
  • Industry: Dr. Pin-Yu Chen (IBM); Dr. Bo Li, Dr. Yu Zhang, Dr. Nanxin Chen, Dr. Heiga Zen and Dr. Tara N. Sainath (Google)�
  • Academics: Prof. Sabato Marco Siniscalchi (NTNU), Prof. Chao Zhang (THU), Prof. Jesper Tegner (KAUST), Prof. Hung-yi Lee (NTU), �Prof. Yu Tsao (Sinica), Prof. Jen-Tzung Chien (NYCU), and Prof. Eng Siong Chng (Nanyang Tech)

65

66 of 68

Acknowledgments (2/2)

  • More Early-Career Collaborators
    • Neural Reprogramming
      • Amy Yung-Ning Hung (now TikTok), Rick Yen (Georgia Tech), Srijith Radhakrishnan (KAUST), Yun-Yun Tsai (Columbia)�
    • Pre-Trained Model Estimation
      • Zih-Ching Chen (National Taiwan University)�
    • Multi-Losses Adapter
      • Chun-Wei Ho (Georgia Tech), Li-Jen Yang (National Yang Ming Chiao Tung University)�
    • Neural Space State Machine
      • Pin-Jui Ku (Georgia Tech) and Chen Chen (Nanyang Tech)�
    • In-Context Learning
      • Chen Chen (Nanyang Tech) and Yuchen Hu (Nanyang Tech)

66

67 of 68

More References

67

Hu, Edward J., et al. "Lora: Low-rank adaptation of large language models." arXiv preprint arXiv:2106.09685 (2021).

Eger, Steffen, and Yannik Benz. "From Hero to Z\'eroe: A Benchmark of Low-Level Adversarial Attacks.” AACL2020.

Zhang, Qingru, et al. "Adaptive budget allocation for parameter-efficient fine-tuning." ICLR 2023.

Lialin, Vladislav, et al. "Stack More Layers Differently: High-Rank Training Through Low-Rank Updates." arXiv preprint arXiv:2307.05695 (2023).

Wang, Xuezhi, Haohan Wang, and Diyi Yang. "Measure and improve robustness in nlp models: A survey.” NAACL 2022.

Yu, Yu, Abdul Rafae Khan, and Jia Xu. "Measuring Robustness for NLP." COLING 2022.

Smorga's Board. ” Frequently Misspelled Word List for Dyslexia.”, https://www.teacherspayteachers.com/Product/Frequently-Misspelled-Word-List-for-Dyslexia-5295631

.

Huang, Chengsong, et al. "LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition." arXiv preprint arXiv:2307.13269 (2023).

68 of 68

7 min Spotlight Featured Talks

68

Dr. Chunyang Wu

Meta AI

Prompting LLM for ASR

12:20 to 12:30 pm