1 of 17

Sentiment Analysis: Fine-Tuning Open-Source Large Language Models

Hasin Zaman, Yanmin Sun

2 of 17

Table of Content

  1. What is Sentiment Analysis
  2. Models
  3. Training
  4. Datasets
  5. Experiments

2

3 of 17

What is Sentiment Analysis?

What is it?

  • NLP (Natural Language Processing) Classification Task
  • Determine the attitude of a text
    • Positive, Negative (Binary)
    • Positive, Negative, Neutral (Tertiary)
    • Number (Spectrum)

Use Cases?

  • Public Opinion Monitoring
  • Threat Detection & Risk Assessment

πŸ˜„

😐

😑

3

4 of 17

Models

Encoder Models

  • DeBERTaV3
    • Microsoft
    • 2021
  • ERNIE2
    • BAIDU
    • 2019
  • FlanT5
    • Google
    • 2022

Lexicographic

  • VADAR

Zero Shot Classification

  • BART

Highlighted Encoder Block

Vaswani, A. (2017). Attention is all you need. Advances in Neural Information Processing Systems.

4

5 of 17

Models

GLUE Benchmark

SuperGLUE Benchmark

5

6 of 17

Training

Pre-Training

  • Linguistic Understanding
  • Unsupervised
  • Starting point

Fine-Tuning

  • Specialized training for downstream task
  • Supervised
  • Labeled data
  • End point

6

7 of 17

Datasets

Distribution

  • Syntactic Distribution
    • Structure of datapoints
  • Semantic Distribution
    • Meaning of datapoints
    • Culture changes the meaning
      • Western
      • Radical ideologies
    • Platform
      • Social media
      • Product reviews
      • News articles

Hierarchy of Syntactic Granularity Levels

7

8 of 17

Datasets

Stanford Sentiment Treebank (SST2)

  • Movie review
  • Sentences and phrases
  • 49,676 samples
  • Larger dataset

Multiple perspective Question Answer (MPQA3.0)

  • News articles
  • Phrase level sentiment
  • Contains 2317 samples

TweetEval

  • Tweets
  • Document level
  • Contains 31,289 samples

Finance Phrase Bank (FPB)

  • Financial news
  • 4845 English articles headlines
  • positive, neutral, and negative sentiment
  • One of the smaller datasets used

8

9 of 17

Datasets

IMDB

  • Movie reviews
  • Document level
  • 49,969 samples

Histogram of Datapoint Length of Different Granularity Levels

9

10 of 17

Baseline Benchmark

Goal

  • Compare all LLM models to each other
  • Compare against traditional solution
  • Provides context for other experiments

Results

  • Fine-tuned models perform the best
  • DeBERTa > ERNIE2 > FlanT5
  • Zero-shot ~5-10% compared to fine-tuned
  • Machine learning methods significantly outperforms traditional method
  • MPQA is the hardest dataset

10

11 of 17

Minimum Cardinality

Goal

  • Smallest dataset required before plateau

Results

  • Performance plateaus after 600-1100 labeled samples
  • Smaller models need fewer fine-tuning samples
  • Smaller models shows large accuracy variation initially
  • Smaller models learn faster

11

12 of 17

Effect of Noise

Goal

  • How does noise/incorrectly labeled data affect models

Results

  • Models degrade after 30%-40% noisy data
  • Larger models handle noise better overall
  • FlanT5 > DeBERTaV3 > ERNIE2

12

13 of 17

Out-of-Distribution

Goal

  • How does training-application distribution affect models?
  • Most open-source datasets are created from commercial sources
    • Movie & product reviews
    • Social media
  • OSINT applications have different semantic compared to commercial applications

13

14 of 17

Out-of-Distribution

DeBERTaV3

ERNIE2

FlanT5

14

15 of 17

Out-of-Distribution

Granularity Impact

  • IMDB vs. SST2
  • Going from high granularity to low granularity has less loss
    • 2%-3% loss from ideal
    • SST2 (Train) -> IMDB (Test)
  • Going from low granularity to high granularity has more loss
    • 4%-6% loss from ideal
    • IMDB (Train) -> SST2 (Test)

DeBERTaV3

ERNIE2

FlanT5

15

16 of 17

Conclusion

  • Smaller models are:
    • More susceptible to noise
    • Train faster with less data
  • Larger Models are:
    • More susceptible to noise
    • Train slower to be well fine-tuned
  • Minimize semantic domain shift when possible
  • Try to go down in granularity if the ideal syntax isn’t possible

16

17 of 17