Trustworthy Generative AI:
The Hybridization of Large Language Models with NeuroSymbolic AI��Manas Gaur�manas@umbc.edu || https://manasgaur.github.io/
Knowledge-infused Learning
Manas Gaur
Amit P. Sheth
Key Focus Areas
Achieve Consistency in Gen AI by NeuroSymbolic means
Customization and Compactness of Gen AI for High-Stake Decision-Making
Beyond RAG: Context Addition, Metacognition, and Knowledge Gap Minimization
Achieve Grounding in Gen AI by NeuroSymbolic Means
Why the need for Trustworthiness in Generative AI
Nearly Impossible to Explain or Reason Generative Answers
Prompt Injections can leak data
Context Windows are and will remain limited
Bias in Large Language Models that Supervised Learning cannot reduce
Reliability Issues: Different Large Language Models Yield Different Outcomes
Inconsistency in Prompts for Completeness in Outcomes
#Grounding #Intructability #Alignment
#Explainability #Intrepretability #Safety
#Causality #Attribution #Abstraction
#Analogy
#Reliability
#Consistency
More Challenges for the Generative AI
Knowledge
Benchmarking Example Generative AI
WellDunn: On the Robustness and Explainability of Language Models and Large Language Models in Identifying Wellness Dimensions,
In BlackboxNLP @ EMNLP 2024
Wellness Dimension
Wellness Dimension
Wellness Dimension Definitions and Questionnaire
https://store.samhsa.gov/sites/default/files/sma16-4958.pdf
MultiWD and
WellXplain Datasets
Content worth 4000 users
6 Wellness Dimensions
Clinical expert explanations
Robustness in Prediction
Maximum Attention Overlap over Explanations
Input
Attention Matrix
Explanation
Prediction
Input
Attention Matrix
Explanation
Prediction
Definitions In-Context Learning
Design 1
Design 2
Input
Attention Matrix
Explanation
Prediction
Questionnaire
Workflow-based
In-Context Learning
Design 4
Input
Attention Matrix
Explanation
Prediction
Chain of Thoughts with Definitions
Design 3
Outcomes
Design Choices | | Robustness using SVD Attention Rank (lower is better) | Attention Overlap (Jaccard Similarity between Attention and Explanations) |
Design 1 | General-purpose LLMs | 91.0 +/- 8.0 | 0.0 |
Design 1 | Domain-specific LLMs | 50.0 +/- 4.0 | 0.03 +/- 0.05 |
Design 2 | General-purpose LLMs | 25 +/- 2.0 | 0.28 +/- 0.15 |
Design 2 | Domain-specific LLMs | 18 +/- 1.0 | 0.35 +/- 0.05 |
Design 3 | General-purpose LLMs | 38.0 +/- 3.0 | 0.20 +/- 0.12 |
Design 3 | Domain-specific LLMs | 22.0 +/- 2.0 | 0.33 +/- 0.02 |
Design 4 | General-purpose LLMs | 15 +/- 1.0 | 0.65 +/- 0.005 |
Design 4 | Domain-specific LLMs | 9 +/- 2.0 | 0.68 +/- 0.002 |
PMC- LLAMA
Meditron
Hybridized Architectures: NeuroSymbolic AI
NeuroSymbolic AI
NeuroSymbolic AI for Consistency
Bonagiri, Vamshi Krishna, Sreeram Vennam, Priyanshul Govil, Ponnurangam Kumaraguru, and Manas Gaur. "SaGE: Evaluating Moral Consistency in Large Language Models." In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation LREC-COLING 2024.
Semantic Consistency: the ability to make consistent decisions in semantically equivalent contexts. i.e, Semantically equivalent questions should yield semantically equivalent answers
Claim: LLMs are not semantically consistent, and can give contradictory answers to paraphrased questions
NeuroSymbolic Empirical Analysis
Semantic Graph-driven Consistent LLM Training (SaGE)
Moral Consistency Corpus (MCC)
BLEURT
BLEU
ROUGE-L
BERTScore
SaGE (LLAMA 3)
GPT-4
NeuroSymbolic AI for Grounding and Instructibility
Grounding
A successful AI teammate requires several cognitive capacities including situation assessment, task behavior, language comprehension and generation , and knowledge gap resolution processes. Grounding enables agents with different capabilities to communicate.
Bajaj, Goonmeet, Valerie L. Shalin, Srinivasan Parthasarathy, and Amit Sheth. "Grounding From an AI and Cognitive Science Lens." IEEE Intelligent Systems 39, no. 2 (2024): 66-71.
Knowledge Gap Resolution
Language Gap
Real-Time Mental Health Analysis
Khandelwal, V., Gaur, M., Kursuncu, U., Shalin, V. L., & Sheth, A. P. (2024). “A Domain-Agnostic Neurosymbolic Approach for Big Social Data Analysis: Evaluating Mental Health Sentiment on Social Media during COVID-19.” Proceedings of the IEEE Big Data Conference 2024.
NeuroSymbolic Architecture
This Knowledge-Infused Learning based Neurosymbolic architecture consists of three components: B1 (Semantic Gap Management) gathers and filters data, B2 (Metadata Scoring) generates classification labels via semantic mapping, and B3 (Adaptive Classifier Training) uses metadata-enhanced data for accurate labeling. Drawing on 12 billion tweets, 2.5 million Reddit posts, 700,000 news articles, and multiple knowledge bases like DAO and SNOMED-CT, this setup supports real-time mental health sentiment analysis.
NeuroSymbolic Architecture
75%
CN
Pk
Decoder
Encoder
posts
Concept Classes
Autoencoder
32
Posts-by-Posts Self-Attention Matrix
Concept-by-Concept Self-Attention Matrix
Linear Projection between two embedding spaces
Sylvester Equation
Simoncini et al. SIAM’16
Semantic Encoding and Decoding
33
Machine Learning Model
Sylvester Equation
Simoncini et al. SIAM’16
Input Layer
Feed Forward Neural Network
δ : Tunable Parameter for Knowledge Infusion
(1-δ) : Forces Knowledge Infusion
Allow model interpretability
Semantic Encoding and Decoding
34
Really struggling with my bisexuality which is causing chaos in my relationship with a girl. Being a fan of LGBTQ community, I am equal to worthless for her. I’m now starting to get drunk because I can’t cope with the obsessive, intrusive thoughts, and need to get out of my head.
Don’t want to live anymore. Sexually assault, ignorant family members and my never ending loneliness brights up my path to death.
I do have a potential to live a decent life but not with people who abandon me. Hopelessness and feelings of betrayal have turned my nights to days. I am developing insomnia because of my restlessness. I just can’t take it anymore. Been abandoned yet again by someone I cared about. I've been diagnosed with borderline for a while, and I’m just going to isolate myself and sleep forever.
Really struggling with my bisexuality which is causing chaos in my relationship with a girl. Being a fan of LGBTQ community, I am equal to worthless for her. I’m now starting to get drunk because I can’t cope with the obsessive, intrusive thoughts, and need to get out of my head.
Don’t want to live anymore. Sexually assault, ignorant family members and my never ending loneliness brights up my path to death.
I do have a potential to live a decent life but not with people who abandon me. Hopelessness and feelings of betrayal have turned my nights to days. I am developing insomnia because of my restlessness. I just can’t take it anymore. Been abandoned yet again by someone I cared about. I've been diagnosed with borderline for a while, and I’m just going to isolate myself and sleep forever.
δ = 1.0 (No Knowledge)
δ = 0.84 (16% knowledge)
Interpretability with Semi-Deep Infusion
Really struggling with my bisexuality which is causing chaos in my relationship with a girl. Being a fan of LGBTQ community, I am equal to worthless for her. I’m now starting to get drunk because I can’t cope with the obsessive, intrusive thoughts, and need to get out of my head.
Don’t want to live anymore. Sexually assault, ignorant family members and my never ending loneliness brights up my path to death.
I do have a potential to live a decent life but not with people who abandon me. Hopelessness and feelings of betrayal have turned my nights to days. I am developing insomnia because of my restlessness. I just can’t take it anymore. Been abandoned yet again by someone I cared about. I've been diagnosed with borderline for a while, and I’m just going to isolate myself and sleep forever.
δ = 0.71 (29% knowledge)
Expert Evaluation Agreement: 84%
Really struggling with my bisexuality which is causing chaos in my relationship with a girl. Being a fan of LGBTQ community, I am equal to worthless for her. I’m now starting to get drunk because I can’t cope with the obsessive, intrusive thoughts, and need to get out of my head.
Don’t want to live anymore. Sexually assault, ignorant family members and my never ending loneliness brights up my path to death.
I do have a potential to live a decent life but not with people who abandon me. Hopelessness and feelings of betrayal have turned my nights to days. I am developing insomnia because of my restlessness. I just can’t take it anymore. Been abandoned yet again by someone I cared about. I've been diagnosed with borderline for a while, and I’m just going to isolate myself and sleep forever.
δ = 0.66 (34% knowledge)
Results
The tables compare model performance for mental health classification across Precision, Recall, and F1-Score. The left table shows traditional models’ results with and without the Neurosymbolic approach, while the right table contrasts the Neurosymbolic model with state-of-the-art LLMs like LLama, Phi, and Mistral.
The Neurosymbolic model consistently outperforms both traditional models and state-of-the-art LLMs, achieving higher performance metrics and adaptability in mental health sentiment classification.
Language Gap : Pay attention to Important Domain Concepts
ProcesS knowledge-infused cross Attention (PSAT)
Reasoning Gap by Improving Retrievers
Truthful QA: 817 Questions on Conspiracies, Supernatural, Paranormal, Fictions …
Natural Questions: 19000 questions
Integrated Mental Health Dataset: 10 Datasets and >100K queries
Mean Average Precision
Reasoning �Gap
Procedural
Describe the process by which hair cells transduce mechanical energy from sound waves into electrical signals?
Comparative
How do neutrinos differ from other subatomic particles, and why are they considered potential candidates for dark matter?
Instructible while maintaining Grounding
Grounding
Work in Progress
Zhang, Yue, Yuntian He, Saket Gurukar, and Srinivasan Parthasarathy. "HeteroMILE: a Multi-Level Graph Representation Learning Framework for Heterogeneous Graphs." arXiv preprint arXiv:2404.00816 (2024).
Safety and Robustness Enforced Next Steps
with Generative AI
Implicit Personalization in Domain-specific Conversations
Workflows-guided Safe Decision Making
Explainable and Attributable Reasoning on Narratives
Grounded Debiasing using Rules and Domain Knowledge (No additional Data)
Social Welfare Optimized Ensemble of LLMs
CoPilots for Health, Manufacturing
Putting it All together
Thank You for Your Attention
[Not Covered] Publications intersecting Neurosymbolic AI and Generative AI
Personalized Response Generation using Dynamic Knowledge Retrieval and Persona-Adaptive Queries
https://ojs.aaai.org/index.php/AAAI-SS/article/view/31203
Unboxing Occupational Bias: Grounded Debiasing LLMs with US Labor Data
To be published in the AAAI Fall Symposium on AI Trustworthiness and Risk Assessment for Challenged Contexts
Building trustworthy NeuroSymbolic AI Systems: Consistency, reliability, explainability, and safety
https://onlinelibrary.wiley.com/doi/pdf/10.1002/aaai.12149
Localintel: Generating organizational threat intelligence from global and local cyber knowledge
In AAAI AICS Workshop 2024 (https://arxiv.org/pdf/2401.10036)
Neurosymbolic Artificial Intelligence (Why, What, and How)
https://www.computer.org/csdl/magazine/ex/2023/03/10148662/1NVf9V0YKze
Knowledge-guided Machine Unlearning
Vision
The vision focuses on multimodal, knowledge graph-focused, controlled, targeted unlearning in publicly available LLMs and their respective quantized versions.
Objectives: Can LLMs
Success:
PhD. Student: Shaswati Saha
Knowledge-driven and Explainable Model Pipelining
Vision
The vision focuses on dynamical and knowledge-driven pipelining of Models for task-specific explainable recommendation and generation
Objectives:
Success:
Industry Work: HPE
59
Context-Oriented Bias Indicator and Assessment Score
60
61
62
63