1 of 63

Trustworthy Generative AI:

The Hybridization of Large Language Models with NeuroSymbolic AI��Manas Gaur�manas@umbc.edu || https://manasgaur.github.io/

2 of 63

Knowledge-infused Learning

Manas Gaur

Amit P. Sheth

3 of 63

Key Focus Areas

Achieve Consistency in Gen AI by NeuroSymbolic means

Customization and Compactness of Gen AI for High-Stake Decision-Making

Beyond RAG: Context Addition, Metacognition, and Knowledge Gap Minimization

Achieve Grounding in Gen AI by NeuroSymbolic Means

4 of 63

Why the need for Trustworthiness in Generative AI

5 of 63

Nearly Impossible to Explain or Reason Generative Answers

Prompt Injections can leak data

Context Windows are and will remain limited

Bias in Large Language Models that Supervised Learning cannot reduce

Reliability Issues: Different Large Language Models Yield Different Outcomes

Inconsistency in Prompts for Completeness in Outcomes

6 of 63

#Grounding #Intructability #Alignment

#Explainability #Intrepretability #Safety

#Causality #Attribution #Abstraction

#Analogy

#Reliability

#Consistency

More Challenges for the Generative AI

7 of 63

Knowledge

8 of 63

Benchmarking Example Generative AI

WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions,

In BlackboxNLP @ EMNLP 2024

9 of 63

Wellness Dimension

10 of 63

Wellness Dimension

Wellness Dimension Definitions and Questionnaire

https://store.samhsa.gov/sites/default/files/sma16-4958.pdf

11 of 63

MultiWD and

WellXplain Datasets

Content worth 4000 users

6 Wellness Dimensions

  • Physical
  • Intellectual
  • Vocational
  • Social
  • Spiritual
  • Emotional

Clinical expert explanations

12 of 63

13 of 63

Robustness in Prediction

Maximum Attention Overlap over Explanations

14 of 63

Input

Attention Matrix

Explanation

Prediction

Input

Attention Matrix

Explanation

Prediction

Definitions In-Context Learning

Design 1

Design 2

15 of 63

Input

Attention Matrix

Explanation

Prediction

Questionnaire

Workflow-based

In-Context Learning

Design 4

Input

Attention Matrix

Explanation

Prediction

Chain of Thoughts with Definitions

Design 3

16 of 63

Outcomes

Design Choices

Robustness using SVD Attention Rank

(lower is better)

Attention Overlap (Jaccard Similarity between Attention and Explanations)

Design 1

General-purpose LLMs

91.0 +/- 8.0

0.0

Design 1

Domain-specific LLMs

50.0 +/- 4.0

0.03 +/- 0.05

Design 2

General-purpose LLMs

25 +/- 2.0

0.28 +/- 0.15

Design 2

Domain-specific LLMs

18 +/- 1.0

0.35 +/- 0.05

Design 3

General-purpose LLMs

38.0 +/- 3.0

0.20 +/- 0.12

Design 3

Domain-specific LLMs

22.0 +/- 2.0

0.33 +/- 0.02

Design 4

General-purpose LLMs

15 +/- 1.0

0.65 +/- 0.005

Design 4

Domain-specific LLMs

9 +/- 2.0

0.68 +/- 0.002

PMC- LLAMA

Meditron

17 of 63

Hybridized Architectures: NeuroSymbolic AI

18 of 63

NeuroSymbolic AI

19 of 63

20 of 63

NeuroSymbolic AI for Consistency

Bonagiri, Vamshi Krishna, Sreeram Vennam, Priyanshul Govil, Ponnurangam Kumaraguru, and Manas Gaur. "SaGE: Evaluating Moral Consistency in Large Language Models." In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation LREC-COLING 2024.

21 of 63

Semantic Consistency: the ability to make consistent decisions in semantically equivalent contexts. i.e, Semantically equivalent questions should yield semantically equivalent answers

Claim: LLMs are not semantically consistent, and can give contradictory answers to paraphrased questions

22 of 63

NeuroSymbolic Empirical Analysis

23 of 63

Semantic Graph-driven Consistent LLM Training (SaGE)

Moral Consistency Corpus (MCC)

24 of 63

BLEURT

BLEU

ROUGE-L

BERTScore

SaGE (LLAMA 3)

GPT-4

25 of 63

NeuroSymbolic AI for Grounding and Instructibility

26 of 63

Grounding

A successful AI teammate requires several cognitive capacities including situation assessment, task behavior, language comprehension and generation , and knowledge gap resolution processes. Grounding enables agents with different capabilities to communicate.

Bajaj, Goonmeet, Valerie L. Shalin, Srinivasan Parthasarathy, and Amit Sheth. "Grounding From an AI and Cognitive Science Lens." IEEE Intelligent Systems 39, no. 2 (2024): 66-71.

27 of 63

Knowledge Gap Resolution

  • Bajaj, Goonmeet, Bortik Bandyopadhyay, Daniel Schmidt, Pranav Maneriker, Christopher Myers, and Srinivasan Parthasarathy. "Understanding knowledge gaps in visual question answering: Implications for gap identification and testing." In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 386-387. 2020.

28 of 63

Language Gap

29 of 63

  • Problem:
    • Real-time mental health monitoring is difficult with traditional models due to evolving language.
    • A neurosymbolic approach integrating neural networks with knowledge bases to adapt to changing language.
  • Importance and Impact
    • Enables timely insights during crises (e.g., COVID-19), supporting public health and policy-making.
    • Achieves greater accuracy and adaptability in mental health analysis, outperforming traditional models.

Real-Time Mental Health Analysis

Khandelwal, V., Gaur, M., Kursuncu, U., Shalin, V. L., & Sheth, A. P. (2024). “A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19.” Proceedings of the IEEE Big Data Conference 2024.

30 of 63

NeuroSymbolic Architecture

This Knowledge-Infused Learning based Neurosymbolic architecture consists of three components: B1 (Semantic Gap Management) gathers and filters data, B2 (Metadata Scoring) generates classification labels via semantic mapping, and B3 (Adaptive Classifier Training) uses metadata-enhanced data for accurate labeling. Drawing on 12 billion tweets, 2.5 million Reddit posts, 700,000 news articles, and multiple knowledge bases like DAO and SNOMED-CT, this setup supports real-time mental health sentiment analysis.

31 of 63

NeuroSymbolic Architecture

75%

CN

Pk

Decoder

Encoder

posts

Concept Classes

Autoencoder

  • Representation Learners
  • Representation Modulators

32 of 63

32

Posts-by-Posts Self-Attention Matrix

Concept-by-Concept Self-Attention Matrix

Linear Projection between two embedding spaces

Sylvester Equation

Simoncini et al. SIAM’16

Semantic Encoding and Decoding

33 of 63

33

Machine Learning Model

Sylvester Equation

Simoncini et al. SIAM’16

Input Layer

Feed Forward Neural Network

δ : Tunable Parameter for Knowledge Infusion

(1-δ) : Forces Knowledge Infusion

Allow model interpretability

Semantic Encoding and Decoding

34 of 63

34

Really struggling with my bisexuality which is causing chaos in my relationship with a girl. Being a fan of LGBTQ community, I am equal to worthless for her. I’m now starting to get drunk because I can’t cope with the obsessive, intrusive thoughts, and need to get out of my head.

Don’t want to live anymore. Sexually assault, ignorant family members and my never ending loneliness brights up my path to death.

I do have a potential to live a decent life but not with people who abandon me. Hopelessness and feelings of betrayal have turned my nights to days. I am developing insomnia because of my restlessness. I just can’t take it anymore. Been abandoned yet again by someone I cared about. I've been diagnosed with borderline for a while, and I’m just going to isolate myself and sleep forever.

Really struggling with my bisexuality which is causing chaos in my relationship with a girl. Being a fan of LGBTQ community, I am equal to worthless for her. I’m now starting to get drunk because I can’t cope with the obsessive, intrusive thoughts, and need to get out of my head.

Don’t want to live anymore. Sexually assault, ignorant family members and my never ending loneliness brights up my path to death.

I do have a potential to live a decent life but not with people who abandon me. Hopelessness and feelings of betrayal have turned my nights to days. I am developing insomnia because of my restlessness. I just can’t take it anymore. Been abandoned yet again by someone I cared about. I've been diagnosed with borderline for a while, and I’m just going to isolate myself and sleep forever.

δ = 1.0 (No Knowledge)

δ = 0.84 (16% knowledge)

Interpretability with Semi-Deep Infusion

Really struggling with my bisexuality which is causing chaos in my relationship with a girl. Being a fan of LGBTQ community, I am equal to worthless for her. I’m now starting to get drunk because I can’t cope with the obsessive, intrusive thoughts, and need to get out of my head.

Don’t want to live anymore. Sexually assault, ignorant family members and my never ending loneliness brights up my path to death.

I do have a potential to live a decent life but not with people who abandon me. Hopelessness and feelings of betrayal have turned my nights to days. I am developing insomnia because of my restlessness. I just can’t take it anymore. Been abandoned yet again by someone I cared about. I've been diagnosed with borderline for a while, and I’m just going to isolate myself and sleep forever.

δ = 0.71 (29% knowledge)

Expert Evaluation Agreement: 84%

Really struggling with my bisexuality which is causing chaos in my relationship with a girl. Being a fan of LGBTQ community, I am equal to worthless for her. I’m now starting to get drunk because I can’t cope with the obsessive, intrusive thoughts, and need to get out of my head.

Don’t want to live anymore. Sexually assault, ignorant family members and my never ending loneliness brights up my path to death.

I do have a potential to live a decent life but not with people who abandon me. Hopelessness and feelings of betrayal have turned my nights to days. I am developing insomnia because of my restlessness. I just can’t take it anymore. Been abandoned yet again by someone I cared about. I've been diagnosed with borderline for a while, and I’m just going to isolate myself and sleep forever.

δ = 0.66 (34% knowledge)

35 of 63

Results

The tables compare model performance for mental health classification across Precision, Recall, and F1-Score. The left table shows traditional models’ results with and without the Neurosymbolic approach, while the right table contrasts the Neurosymbolic model with state-of-the-art LLMs like LLama, Phi, and Mistral.

The Neurosymbolic model consistently outperforms both traditional models and state-of-the-art LLMs, achieving higher performance metrics and adaptability in mental health sentiment classification.

36 of 63

Language Gap : Pay attention to Important Domain Concepts

37 of 63

ProcesS knowledge-infused cross Attention (PSAT)

38 of 63

39 of 63

40 of 63

Reasoning Gap by Improving Retrievers

41 of 63

42 of 63

43 of 63

44 of 63

Truthful QA: 817 Questions on Conspiracies, Supernatural, Paranormal, Fictions …

Natural Questions: 19000 questions

Integrated Mental Health Dataset: 10 Datasets and >100K queries

Mean Average Precision

45 of 63

Reasoning �Gap

Procedural

Describe the process by which hair cells transduce mechanical energy from sound waves into electrical signals?

  1. What are hair cells?
  2. What is mechanical energy?
  3. What are sound waves?
  4. What are electrical signals?

Comparative

How do neutrinos differ from other subatomic particles, and why are they considered potential candidates for dark matter?

  1. What are neutrinos?
  2. What are subatomic particles?
  3. What is dark matter?
  4. What characteristics do particles need to be considered candidates for dark matter?

46 of 63

Instructible while maintaining Grounding

47 of 63

48 of 63

49 of 63

Grounding

Work in Progress

50 of 63

51 of 63

Zhang, Yue, Yuntian He, Saket Gurukar, and Srinivasan Parthasarathy. "HeteroMILE: a Multi-Level Graph Representation Learning Framework for Heterogeneous Graphs." arXiv preprint arXiv:2404.00816 (2024).

52 of 63

Safety and Robustness Enforced Next Steps

with Generative AI

53 of 63

Implicit Personalization in Domain-specific Conversations

Workflows-guided Safe Decision Making

Explainable and Attributable Reasoning on Narratives

Grounded Debiasing using Rules and Domain Knowledge (No additional Data)

Social Welfare Optimized Ensemble of LLMs

CoPilots for Health, Manufacturing

54 of 63

Putting it All together

  • Safety is a function of Grounding, Instructibility, and Alignment
  • Grounding is well connected with the notion of the Knowledge Gap
  • Metacognition, Path-based Retrievers can help address Reasoning Gap in Gen AI
  • Robustness in Gen AI requires
    • Uncertainty Quantification using Graph Theory:
    • Measuring and Optimizing with Attention Fidelity
    • Knowledge infusion (e.g., Definitions)
  • Reservoir Networks, Neural Graph elicitation and Graph Coarsening are following best methods for customizing and lossless compactness in Gen AI models

55 of 63

Thank You for Your Attention

56 of 63

[Not Covered] Publications intersecting Neurosymbolic AI and Generative AI

Personalized Response Generation using Dynamic Knowledge Retrieval and Persona-Adaptive Queries

https://ojs.aaai.org/index.php/AAAI-SS/article/view/31203

Unboxing Occupational Bias: Grounded Debiasing LLMs with US Labor Data

To be published in the AAAI Fall Symposium on AI Trustworthiness and Risk Assessment for Challenged Contexts

Building trustworthy NeuroSymbolic AI Systems: Consistency, reliability, explainability, and safety

https://onlinelibrary.wiley.com/doi/pdf/10.1002/aaai.12149

Localintel: Generating organizational threat intelligence from global and local cyber knowledge

In AAAI AICS Workshop 2024 (https://arxiv.org/pdf/2401.10036)

Neurosymbolic Artificial Intelligence (Why, What, and How)

https://www.computer.org/csdl/magazine/ex/2023/03/10148662/1NVf9V0YKze

57 of 63

Knowledge-guided Machine Unlearning

Vision

The vision focuses on multimodal, knowledge graph-focused, controlled, targeted unlearning in publicly available LLMs and their respective quantized versions.

Objectives: Can LLMs

  1. Can LLMs unlearn private and sensitive data while still provide reasonable outcomes and explanations?
  2. Can LLMs unlearn ambiguous and outdated protocols, while still retaining fresh information or protocols?

Success:

  • Change in Protocols should be reflected in Unlearnt Model.
  • Degree of Unlearning
  • Time to regain the unlearnt information
  • Reduced Bias and Increased Safety

PhD. Student: Shaswati Saha

58 of 63

Knowledge-driven and Explainable Model Pipelining

Vision

The vision focuses on dynamical and knowledge-driven pipelining of Models for task-specific explainable recommendation and generation

Objectives:

  1. Can we dynamically create performance enhanced and efficient model pipelines for dynamic tasks?
  2. Can we enhance ensemble performance with domain knowledge and guidelines?

Success:

  1. Resource Constrained Model Pipelining
  2. Dynamic Ensembling (Tree, Sequential, Graph) specific to task
  3. Better Explanations and Safety

Industry Work: HPE

59 of 63

59

60 of 63

Context-Oriented Bias Indicator and Assessment Score

60

61 of 63

61

62 of 63

62

63 of 63

63