1 of 11

On Robustness of Finetuned Transformer-based NLP Models

Pavan Kalyan Reddy Neerudu1, Subba Reddy Oota2 , Mounika Marreddy3 Venkateswara Rao Kagita1 , Manish Gupta3,4

neerud_951963@student.nitw.ac.in, subba-reddy.oota@inria.fr, mounika.marreddy@research.iiit.ac.in venkat.kagita@nitw.ac.in, gmanish@microsoft.com

1NIT Warangal, India; 2INRIA, Bordeaux, France; 3IIIT Hyderabad, India; 4Microsoft, India

1

2 of 11

Interesting questions about BERT, GPT2, T5

  • Finetuning modifies the representations generated by each layer.
    • While fine-tuning these models, what changes across layers with respect to the pre-trained checkpoints?
  • Robustness to input perturbations
    • How robust are BERT, GPT2, T5 to input perturbations?
    • Is the effect of finetuning consistent across all models for various NLP tasks?
    • Do these models exhibit varying levels of robustness to input text perturbations when finetuned for different NLP tasks?

Pavan Kalyan Reddy Neerudu, Subba Reddy Oota, Mounika Marreddy, Venkateswara Rao Kagita, Manish Gupta. On Robustness of Finetuned Transformer-based NLP Models. EMNLP 2023.

2

3 of 11

Representational Similarity Analysis between pretrained and finetuned models

  •  

Pavan Kalyan Reddy Neerudu, Subba Reddy Oota, Mounika Marreddy, Venkateswara Rao Kagita, Manish Gupta. On Robustness of Finetuned Transformer-based NLP Models. EMNLP 2023.

  •  

3

4 of 11

Text perturbations and robustness tasks

  •  

Pavan Kalyan Reddy Neerudu, Subba Reddy Oota, Mounika Marreddy, Venkateswara Rao Kagita, Manish Gupta. On Robustness of Finetuned Transformer-based NLP Models. EMNLP 2023.

  • Robustness Evaluation Tasks
    • Generative tasks: GPT-2 and T5-base
    • Classification tasks: BERT-base, GPT-2, T5-base encoder.
    • Classification tasks: GLUE
      • Single-sentence tasks (CoLA, SST-2)
      • Similarity and paraphrase tasks (MRPC, STS-B, QQP)
      • Inference tasks (MNLI, QNLI, WNLI, RTE)
    • Generative tasks:
      • Text summarization: XSum
      • Free-form text generation: CommonGen
      • Question generation: SQuAD

4

5 of 11

How does finetuning modify the layers representations for different models?

  • GPT-2 had fewer affected layers, indicating higher semantic stability.
  • CKA values for GPT-2 remained mostly higher than those for BERT
  • CKA/STIR drop from initial to later layers; accuracy increases.
    • Finetuning impacts later layers more.
  • CKA and STIR vary consistently across all three models.
    • 🡺 data characteristics of these particular tasks are better captured by the pretrained representations
    • explains why transfer learning using these models is so successful

Pavan Kalyan Reddy Neerudu, Subba Reddy Oota, Mounika Marreddy, Venkateswara Rao Kagita, Manish Gupta. On Robustness of Finetuned Transformer-based NLP Models. EMNLP 2023.

5

Results of layer-wise CKA/STIR comparisons between pre-trained and finetuned BERT, GPT-2 and T5 on various GLUE tasks. STIR values reported here are STIR(finetuned model|pre-trained model).

6 of 11

How robust are the classification models to perturbations in input text?

  • GPT-2 is the most robust model, followed by T5 and then BERT.
  • GPT-2 may be better suited when input text is noisy or incomplete.
  • Most impact: changing characters, removing nouns and verbs
  • Lowest impact: Bias.

Pavan Kalyan Reddy Neerudu, Subba Reddy Oota, Mounika Marreddy, Venkateswara Rao Kagita, Manish Gupta. On Robustness of Finetuned Transformer-based NLP Models. EMNLP 2023.

6

7 of 11

Is the impact of input text perturbations on finetuned models task-dependent?

Pavan Kalyan Reddy Neerudu, Subba Reddy Oota, Mounika Marreddy, Venkateswara Rao Kagita, Manish Gupta. On Robustness of Finetuned Transformer-based NLP Models. EMNLP 2023.

  • NLI tasks: GPT2 is better for RTE.
  • Transformer models demonstrated high tolerance towards “Dropping first word” and “Bias” perturbations.
  • Impact of text perturbations on finetuned models is task-dependent.
  • Single-sentence tasks
    • CoLA: GPT2 is most robust.
    • Sentiment analysis: All models are very robust, except for “Change char”
  • Similarity and paraphrase tasks
    • For MRPC, GPT2 is best. For STS-B, BERT is best.

7

8 of 11

Is the impact of input text perturbations on finetuned models task-dependent?

Pavan Kalyan Reddy Neerudu, Subba Reddy Oota, Mounika Marreddy, Venkateswara Rao Kagita, Manish Gupta. On Robustness of Finetuned Transformer-based NLP Models. EMNLP 2023.

  • Dropping nouns, verbs, or changing characters have highest impact.

8

9 of 11

Is the impact of perturbations on finetuned models different across layers?

Pavan Kalyan Reddy Neerudu, Subba Reddy Oota, Mounika Marreddy, Venkateswara Rao Kagita, Manish Gupta. On Robustness of Finetuned Transformer-based NLP Models. EMNLP 2023.

  • BERT
  • blue → most affected layer
  • green → second most affected layer
  • orange → third most affected layer
  • Initial and last layers are most sensitive to perturbations.
  • “Add text” affects the lower layers

9

10 of 11

Is the impact of perturbations on finetuned models different across layers?

  • Similarity across models
    • The layers most affected by text perturbations tend to be consistent across different models.
    • Certain shared linguistic features and contextual info are crucial for all these models.
  • Variation across tasks
    • BERT: Later layers are more impacted
    • GPT-2: Initial layers are more impacted
    • T5: Similar to GPT-2, but some middle layers are also affected.

Pavan Kalyan Reddy Neerudu, Subba Reddy Oota, Mounika Marreddy, Venkateswara Rao Kagita, Manish Gupta. On Robustness of Finetuned Transformer-based NLP Models. EMNLP 2023.

  • Swap text
    • Tends to affect multiple layers across different models.
    • Implies that contextual information is distributed and integrated throughout the layers of these models.

10

11 of 11

Conclusion

  • Finetuned vs pretrained models: last layers of the models are more affected than the initial layers when finetuning.
  • GPT-2 exhibits more robust representations than BERT and T5 across multiple types of input perturbation.
  • Models are seen to be most affected by dropping nouns, verbs or changing characters.
  • Certain layers are consistently impacted across different models, indicating the importance of specific linguistic features and contextual information.

11

Pavan Kalyan Reddy Neerudu, Subba Reddy Oota, Mounika Marreddy, Venkateswara Rao Kagita, Manish Gupta. On Robustness of Finetuned Transformer-based NLP Models. EMNLP 2023.