1 of 16

Artificial Intelligence on Social Media (AISoMe): FIRE 2023 Track

Soham Poddar, IIT Kharagpur

Moumita Basu, Amity University Kolkata

Kripabandhu Ghosh, IISER Kolkata

Saptarshi Ghosh, IIT Kharagpur

2 of 16

AISoMe 2023

A multi-label classification problem on tweets in the healthcare domain

Specifically, classify tweets according to anti-vaccine opinions expressed

3 of 16

Stances towards vaccines

Pro-vax

support and promote the benefits of vaccines

Anti-vax

believe that vaccines do more harm than good

4 of 16

Different people have different Anti-vax concerns

Vaccines shouldn’t be mandatory

Vaccines are for money-making

Vaccines are unnecessary

5 of 16

The training dataset: CAVES

  • CAVES: “Concerns About Vaccines with Explanations and Summaries”�
  • A large dataset of 9,921 Anti-Vax tweets labelled with 12 concerns about vaccines in a multi-label setting [Poddar et al., SIGIR 2022]

Conspiracy

Country

Ineffective

Ingredients

Mandatory

Pharma

Political

Religious

Rushed

Side-effect

Unnecessary

None

6 of 16

The classes / labels in CAVES dataset

7 of 16

Examples of tweets and labels

STOP TAKING TOXIC VAX and expose COVID hoax and murders with morphine and ventillators. there is No covid!

The reason insurance companies won't pay out if you experience the inevitable adverse reactions, including death is because it is an "Experimental Vaccine"

ingredients

rushed

side-effect

unnecessary

conspiracy

8 of 16

AISoMe 2023 evaluation details

Test dataset:

  • 486 tweets labeled with the same 12 classes
  • About COVID vaccines as well as other vaccines

Task: Each tweet in the test set has to be assigned to one or more of the 12 classes (anti-vaccine concerns).

Metric: Macro-F1 of all classes

9 of 16

Submitted runs

  • 19 teams participated
  • 48 runs were submitted
  • Fine-tuned pre-trained transformer based models such as BERT, CT-BERT, RoBERTa
  • LLM-based models such as GPT 3.5 and GPT2LMheadmodel
  • Multinomial Naïve Bayes and Support Vector Machines, Multi-Output Classifiers

  • Fine-tuned CT-BERT models outperformed other models (Macro F1: 0.71 )

10 of 16

Best-performing runs of top 10 teams

Results of all runs available in the overview paper of the track

11 of 16

Analysis on Result

Fine-tuned CT-BERT model performs best and used by top two teams

  • Pre-trained Contextualized Representations: pre-trained on a large corpus of Covid related tweets posted during January–April 2020, on the topic of COVID-19

  • Transfer Learning : knowledge gained from pre-training on a large dataset is transferred to a smaller dataset for fine-tuning.

12 of 16

Analysis on Result

  • Bidirectional Context Understanding: processes text bidirectionally, considering both left and right context for each word

  • Attention Mechanism: beneficial for capturing long-range dependencies and contextual information

  • Parameter Fine-Tuning: helps the model specialize and optimize its performance for the task at hand

13 of 16

Conclusion

  • A challenging multi-label classification task�
  • The CAVES dataset contains explanations and summaries as well - can be used for explainable classification and summarization tasks�
  • Thanks to all participating times and the FIRE conference organizing committee

14 of 16

References

S Poddar, AM Samad, R Mukherjee, N Ganguly, S Ghosh. “CAVES: A dataset to facilitate explainable classification and summarization of concerns towards COVID vaccines.” In Proceedings of ACM SIGIR Conference on Research and Development in Information Retrieval. Vol 45, 2022

15 of 16

Thank you! �Questions?

16 of 16

The training dataset: CAVES