1 of 45

Contextualized Topic Models with Commonsense Knowledge

CPSC 532G - Final Project Presentation

Felipe González-Pizarro, Raymond Li

2 of 45

3 of 45

Corpus-Level Analysis

  • In this course, we have studied multi-document summarization for analyzing a collection of documents

  • An alternative solution: Topic Modeling
    • Statistical approach for extracting topics from large text corpora.
    • Can be viewed as a form of extreme summarization
    • Example: In news articles, the results might include a list of topics related to politics, economy, sports, etc.

4 of 45

Latent Dirichlet Allocation (LDA)

  • Generative model that induce latent topical structure from the observed documents
  • Core Idea:
    • Documents are represented by a few topics
    • Topics are represented by a few keywords

5 of 45

Generative Process of a Document

  • Topic distributions are sampled from a dirichlet prior
  • For each word of the words:
    • Sample topic
    • Sample word

[1] Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent dirichlet allocation. Journal of machine Learning research, 3(Jan), 993-1022.

6 of 45

Neural Topic Models (NTM)

  • Variational Autoencoders (VAE) trains an encoder network mapping a document to an approximate posterior distribution
  • Example: LDA with Product of Experts (ProdLDA)
    • Approximate prior with a softmax-normal distribution
    • Topical words can be sampled with weights of the FC layer on the decoder

7 of 45

Variational Autoencoder

  • Variational autoencoders learn the parameters of a probability distribution representing the data

  • We can sample from the distribution and generate new input data samples

8 of 45

Variational Autoencoder as a Topic Model

  • Input: Documents represented as a Bag of Words (BOW)
  • The encoder samples the topic document representation (hidden representation) from the learned parameters of the distribution
  • The top-words of a topic are obtained by the weight matrix that reconstruct the BOW

[1] Akash Srivastava and Charles Sutton. 2017. Autoencoding variational inference for topic models, ICLR,2017

9 of 45

Contextualized Topic Models (CTM)

  • It is based on variational autoencoders
  • Instead of BoW representation, CTM maps contextualized document embeddings onto a continuous latent space
  • Missing: Global (corpus-level) semantic relationships between words

Contextualized Embeddings (e.g, SBERT)

[1] Bianchi, F., Terragni, S., Hovy, D., Nozza, D., & Fersini, E. (2021, April). Cross-lingual Contextualized Topic Models with Zero-shot Learning. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 1676-1683).

10 of 45

Clustering-Based Topic Models

  • Rather than approximating the topic distribution from document-level context, we can also approximate ’s from the global semantic relationships between all words
  • Then each is represented as the cluster centroid, where is the reciprocal distance (similarity) in latent space.

[1] Zihan Zhang, Meng Fang, Ling Chen, and Mohammad Reza Namazi Rad. 2022. Is Neural Topic Modelling Better than Clustering? An Empirical Study on Clustering with Contextual Embeddings for Topics. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 3886–3893, Seattle, United States. Association for Computational Linguistics.

11 of 45

Topic modeling: poor quality of topics

  • Results are not always understandable and useful [1]:
    • incoherent or noisy topics [2,3]
    • misaligned with domain expert understanding of the corpus (e.g., repetitive topics) [2]

[1] Harrando, I., & Troncy, R. (2021, December). Discovering Interpretable Topics by Leveraging Common Sense Knowledge. In Proceedings of the 11th on Knowledge Capture Conference (pp. 265-268).

[2]Smith, A., Kumar, V., Boyd-Graber, J., Seppi, K., & Findlater, L. (2018, March). Closing the loop: User-centered design and evaluation of a human-in-the-loop topic modeling system. In 23rd International Conference on Intelligent User Interfaces (pp. 293-304). ACM.

[3] Wang, J., Zhao, C., Xiang, J., & Uchino, K. (2019). Interactive Topic Model with Enhanced Interpretability. In IUI Workshops.

12 of 45

Topic modeling and Commonsense

  • Topic modeling algorithms focus on co-occurrences of terms
    • "You shall know a word by the company it keeps" [1]

  • As a result, they can not capture relations between words that are not explicitly present in the training data [2]

  • Injecting external knowledge such as commonsense might mitigate this problem [2]

[1] Firth, J. R. (1957). A synopsis of linguistic theory, 1930-1955. Studies in linguistic analysis

[2] Ismail Harrando and Raphaël Troncy. 2021. Discovering Interpretable Topics by Leveraging Common Sense Knowledge. In Proceedings of the 11th on Knowledge Capture Conference (K-CAP '21). Association for Computing Machinery, New York, NY, USA, 265–268. DOI:https://doi.org/10.1145/3460210.3493586

13 of 45

Why do NLP Models Need Commonsense?

Natural language is ...

* Based on slides from Vered Shwartz’s Talk: Incorporating Commonsense Reasoning into NLP Models (July, 2022)

Ambiguous

Under-Specified

Social

14 of 45

What is Commonsense?

*Based on Introductory Tutorial on Commonsense Reasoning. Maarten Sap, Vered Shwartz, Antoine Bosselut, Dan Roth, and Yejin Choi. ACL 2020.

The basic level of practical knowledge and reasoning concerning everyday situations and events that are commonly shared among most people.

15 of 45

Explore Techniques to integrate Commonsense Knowledge into Neural Topic Models

Motivation

Explore Techniques to integrate Commonsense Knowledge

into Neural Topic Models

16 of 45

Explore Techniques to integrate Commonsense Knowledge into Neural Topic Models

To the best of our knowledge, there have been no attempts on integrating relational knowledge into neural topic models.

Motivation

Explore Techniques to integrate Commonsense Knowledge

into Neural Topic Models

17 of 45

  1. Explore techniques on integrating commonsense knowledge with neural topic models
  2. Visualize the latent space of CTM to understand how topic representations are distributed and evaluate our intermediate results
  3. Evaluate our approaches on a wide-range of automatic evaluation metrics

Project Plan

18 of 45

[1] Bianchi, F., Terragni, S., & Hovy, D. (2021, August). Pre-training is a Hot Topic: Contextualized Document Embeddings Improve Topic Coherence. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers) (pp. 759-766).

[2] Srivastava, A., & Sutton, C. (2016). Autoencoding Variational Inference For Topic Models, ICLR 2017

Chicago

  • Baseline Model:
    • Contextualized Topic Model (CTM) [1]: which extends ProdLDA[2]
  • Corpora
    • 20Newsgroups: 18,173 documents, 2,000 vocab
    • Wiki20K: 20,000 documents, 2,000 vocab
    • Tweets2011: 2,471 documents, 5,098 vocab
  • Commonsense Knowledge:
    • ConceptNet: ~1.5M nodes, 34 types of relations
    • Atomic: ~300K nodes Mainly focuses on events, causes, and effects.

Settings

19 of 45

  • Incorporate Commonsense into CTM
    • Incorporated into the contextualized embeddings (e.g., using COMET embeddings)
  • Clustering
    • Cluster commonsense-aware embeddings for all concepts in the corpus

Knowledge Incorporation Approaches

20 of 45

ConceptNet Numberbatch

  • Embeddings have semi-structured, common sense knowledge from ConceptNet, giving them a way to learn about words that isn't just observing them in context

  • Numberbatch is built using an ensemble that combines data from ConceptNet, word2vec, GloVe, and OpenSubtitles 2016 [1]

[1] Speer, R., Chin, J., & Havasi, C. (2017, February). Conceptnet 5.5: An open multilingual graph of general knowledge. In Thirty-first AAAI conference on artificial intelligence.

21 of 45

COMET

  • COMmonsEnse Transformers (COMET) [1] learns to generate rich and diverse commonsense descriptions in natural language

[1] Bosselut, A., Rashkin, H., Sap, M., Malaviya, C., elikyilmaz, A., & Choi, Y. (2019, July). COMET: Commonsense Transformers for Automatic Knowledge Graph Construction. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics (pp. 4762-4779).

22 of 45

Incorporate Commonsense into CTM

Contextualized Embeddings

  • We encode documents keywords by using NumberBatch/COMET
  • Commonsense document embedding: Average/Max keywords- embeddings
  • We concatenate commonsense-based embeddings to SBERT embeddings

New Document representations

SBERT

NumberBatch

SBERT

COMET

Reimers, N., & Gurevych, I. (2019, November). Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP) (pp. 3982-3992).

23 of 45

Clustering-Based Approach with Commonsense Knowledge

  • Extracting ConceptNet nodes from the input documents using CoCo-Ex [1]
    • Match phrases with a dictionary of normalized ConceptNet nodes
    • Filter extractions by calculating the similarity between the ConceptNet node and the extracted phrase.
  • Cluster the extracted concepts using Numberbatch Embeddings
    • Label each topical cluster based on distance from centroid

[1] Becker, Maria, Katharina Korfhage, and Anette Frank. "COCO-EX: A tool for linking concepts from texts to ConceptNet." Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations. 2021.

24 of 45

Evaluation metrics

  • Topic Coherence: Topic descriptors (e.g., keywords) must share some level of semantic relatedness
    • (Co-occurrence based) Normalized Pointwise Mutual Information (NPMI) [1]
    • (Co-occurrence based) Cv [5]
    • (Semantic based) External word embeddings topic coherence (WECO) [2]
  • Topic segregation: Topics should have little lexical/semantic overlap between them
    • Topic diversity (TD) [3]
    • Inversed Rank-Biased Overlap (I-RBO) [4]

[1] Jey Han Lau, David Newman, and Timothy Baldwin. 2014. Machine reading tea leaves: Automatically evaluating topic coherence and topic model quality. In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, pages 530–539.

[2] Ran Ding, Ramesh Nallapati, and Bing Xiang. 2018. Coherence-aware neural topic modeling. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 830–836, Brussels, Belgium. Association for Computational Linguistics.

[3] Adji B Dieng, Francisco JR Ruiz, and David M Blei. 2020. Topic modeling in embedding spaces. Transactions of the Association for Computational Linguistics, 8:439–453

[4] Federico Bianchi, Silvia Terragni, and Dirk Hovy. 2021a. Pre-training is a hot topic: Contextualized document embeddings improve topic coherence. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 2: Short Papers), pages 759–766, Online. Association for Computational Linguistics

[5] Röder, M., Both, A., & Hinneburg, A. (2015, February). Exploring the space of topic coherence measures. In Proceedings of the eighth ACM international conference on Web search and data mining (pp. 399-408).

25 of 45

Final Results

  • We average results for each metric over 30 runs of each model for Tweets2011 and over 20 runs for each model for 20NewsGroups and Wiki20K

  • We have also evaluated other hyperparameters such as different number of topics (i.e., 50, 75) → will be included in the report.

26 of 45

Neural Topic Model: CTM

  • Injecting commonsense-based embeddings into a Neural Topic Model can increase the coherence (NPMI, Cv) and diversity (TD) of resulting topics in short-text datasets (i.e., Tweets2011)

27 of 45

Clustering

  • Clustering-based methods can largely increase the diversity of the resulting topics. However, they must be used with caution (decrease in coherence)

28 of 45

NTMs vs Clustering

  • Topics produced by NTM can better capture the co-occurrence relationship between words in the document (as evident by NPMI).
  • The topics produced through clustering can better capture the semantic relationship between words in the corpus (as evident by WECO)
  • Our work provide motivation to better integrate semantic relationships into NTM or vice versa.

29 of 45

Ablation study

Embeddings

30 of 45

Ablation study

  • CLIP Text embeddings might have higher performance than SBERT embeddings in short-text datasets

31 of 45

Exploring other configurations

  • CLIP Text embeddings might have higher performance than SBERT embeddings in short-text datasets.
  • Commonsense embeddings can result in higher topic diversity (TD)

32 of 45

Visualizations

Sievert, C., & Shirley, K. (2014, June). LDAvis: A method for visualizing and interpreting topics. In Proceedings of the workshop on interactive language learning, visualization, and interfaces (pp. 63-70).

We used an interactive topic modeling visualization tool to interpret topics and analyze the quality of our intermediate results.

33 of 45

Visualizations

Sievert, C., & Shirley, K. (2014, June). LDAvis: A method for visualizing and interpreting topics. In Proceedings of the workshop on interactive language learning, visualization, and interfaces (pp. 63-70).

20 NewsGroups - SBERT

20 NewsGroups - CTM+ConceptNet

34 of 45

Lessons learned

  • Gained a better understanding of:
    • Advantages/disadvantages between LDA-based topic models (NTMs) and clustering-based topic models.
    • Evaluation metrics and their limitations (e.g., measuring topic diversity only on top ten keywords is not very helpful).
    • Variational Autoencoders and Variational Inferences
  • We developed skills to :
    • Improve our algorithms in terms of space and time-complexity. E.g., Currently, we evaluate topic models 5x faster!
    • Write modular code. We can easily apply our pipeline to any dataset
  • Increase our critical thinking skills:
    • Formulate better research questions before diving into coding.

35 of 45

Reflections

Was the project successful?

  • Yes! We have completed nearly all the items initially planned.

Strengths:

    • Robust evaluation: We consider three full datasets, several topics' coherence and diversity metrics. We have also evaluated our models several times, under different settings (e.g., # of topics)
    • Our current pipeline make easy to expand this work

Weakness:

    • Incomplete integration of Clustering and NTMs (work-in-progress)
      • Injecting corpus level semantic information into NTMs (E.g. Align the CTM decoder p(w|z) with probability computed from clustering )
      • We underestimated the efforts of changing the inner-workings of CTM.

36 of 45

Contributions and Takeaways

  • Systematically experimented with commonsense embeddings (COMET and NumberBatch) as a viable solution to incorporate commonsense knowledge into a neural topic model (NTMs)
    • We find evidence that commonsense based embeddings can increase the quality of resulting topics, specially in short-text datasets
  • Explored clustering techniques as a potential alternative to LDA-inspired methods (e.g. NTM)
  • Our results provided motivation to incorporate corpus-level semantic relationships between words into NTMs

37 of 45

Future Work

  • Consider other topics' quality metrics (e.g., Topics' coverage, topic's granularity)
    • Manual annotations can help!

38 of 45

Future Work

  • Consider other topics' quality metrics (e.g., Topics' coverage, topic's granularity)
    • Manual annotations can help!
  • For each document, generate precise commonsense inferences (e.g., COMET).
    • This might be an alternative to commonsense-based embeddings.

39 of 45

Future Work

  • Consider other topics' quality metrics (e.g., Topics' coverage, topic's granularity)
    • Manual annotations can help!
  • For each document, generate precise commonsense inferences (e.g., COMET).
    • This might be an alternative to commonsense-based embeddings.
  • Integration between CTM and clustering:
    • Align output topic-word distribution from CTM with the distribution produced by the clustering model using KL divergence.
      • Resolve converge problems during training (work-in-progress)
    • Alternative: Use weighted similarity between intra-topic word distribution as an additional constraint in CTM. Intuition.
      • Penalize semantically distant word pairs from appearing in the same topic-word distribution.

40 of 45

Appendix

41 of 45

Questions?

42 of 45

Contextualized Topic Models with Commonsense Knowledge

CPSC 532G - Final Project Presentation

Felipe González-Pizarro, Raymond Li

43 of 45

Expected Outcomes

  • Gain a deeper understanding of neural topic models (inner-workings, shortcomings)
  • Design a technique to integrate relational knowledge with neural topic models through knowledge injection
  • Conduct automatic evaluation of topic models
  • Answer the question: Could commonsense knowledge benefit topic modeling algorithms? How?

44 of 45

Computational Cost of Posterior

  • Posterior inference over the hidden variable is intractable
  • Two common approaches
    • MCMC (e.g. Gibbs Sampling)
    • Variational methods (e.g. Mean Field Approximation)

[1] Blei, D. M., Ng, A. Y., & Jordan, M. I. (2003). Latent dirichlet allocation. Journal of machine Learning research, 3(Jan), 993-1022.

45 of 45

Topic Keywords Coverage (20Newsgroup)