Yantian Zha
Ph.D. Dissertation Defense
Committee Members:
Prof. Subbarao Kambhampati, Chair and Advisor
Prof. Baoxin Li
Prof. Siddharth Srivastava
Dr. Jianjun Wang
Perceiving, Acting, Planning, and Self-Explaining:
A Cognitive Quartet with Four Neural Networks
+
Thesis Website
2
3
Perceiving
Component
Acting
Component
Planning
Component
Self-Explaining
Component
Perceiving
4
[1] Action Genome: Actions as Composition of Spatio-temporal Scene Graphs, Ji, Jingwei, et al., CVPR 2020
[1]
Acting
5
Planning
6
preconditions:� at(x, table)� ∧ empty(gripper)�effects:� ~at(x, table) ∧� ~empty(gripper)
Self-Explaining
7
Self-Explaining
8
A Successful case
L1
L2
Self-Explaining
9
Self-Explanation:
object-location > object-color
A Successful case
L1
L2
Self-Explaining
10
Self-Explanation:
object-location > object-color
A Successful case
L1
L2
L1
L2
Self-Explaining
11
Self-Explanation:
object-location > object-color
A Successful case
Unsuccessful
L1
L2
L1
L2
Self-Explaining
12
Self-Explanation:
object-location > object-color
A Successful case
The successful and unsuccessful cases are only different in color assignment
Unsuccessful
L1
L2
L1
L2
Self-Explaining
13
Self-Explanation:
object-location > object-color
A Successful case
The successful and unsuccessful cases are only different in color assignment
Self-Explanation:
object-color > object-location
Unsuccessful
L1
L2
L1
Self-Explaining
14
Self-Explanation:
object-location > object-color
A Successful case
The successful and unsuccessful cases are only different in color assignment
Self-Explanation:
object-color > object-location
Unsuccessful
Successful
L1
L2
L1
L2
L2
L1
Self-Explaining
15
A Coupling of the Four
16
A Coupling of the Four
17
A Coupling of the Four
Their couplings motivate a stronger cognitive intelligence
Allows cognitive agents to learn to accomplish more complex tasks better and more efficiently
18
19
Subconscious Level
Metacognition Level
Conscious Level
Planning
Perceiving
Acting
Self-Explaining
Background Knowledge
Zha et al, PRDA, AAAI-WS, 18
Zha et al, Affordance-Aware-Imitation-Learning, IROS, 21
Zha et al, Self-Explanation-guided-RLfD, AAAI-RLG Workshop, 22
Cognitive Agent Learning
Model
Zha et al, Distr2Vec, AAMAS, 18
Kulkarni et al., Explicable Planning, AAMAS, 19
Zha et al, Self-Explanation-guided-RLfD, Under-Review, 21
Zha et al, Distr2Vec, AAMAS, 18
Zhuo et al., RNNPlanner, TIST, 20
Zha et al., Distr2Vec, AAMAS, 18
Zhuo et al., RNNPlanner, TIST, 20
Kulkarni et al., Explicable Planning, AAMAS, 19
20
Coupling the Perception and Acting
Part I
Contrastively Learning Visual Attention as Affordance Cues from Demonstrations for Robotic Grasping
Yantian Zha, Siddhant Bhambri, Lin Guan
IROS 2021�
21
Subconscious Level
Metacognition Level
Conscious Level
Perceiving
Acting
Cognitive Agent Learning
Zha et al, Affordance-Aware-Imitation-Learning, IROS, 21
Affordance
22
action
object
effect
Grasping Affordance
23
action
object
effect
Grasping Affordance
24
action
object
effect
Grasping Affordance
25
action
object
effect
Grasping Affordance Modeling
26
Clusters of approaching orientations [1]
Clusters of grasp pre-shapes [2]
[1] de Granville C et al. Learning grasp affordances through human demonstration. In Proceedings of ICDL 2006.
[2] Detry, Renaud, et al. "Learning a dictionary of prototypical grasp-predicting parts from grasping experience." ICRA, 2013.
27
Baxter Robot
Image from: https://www.active8robots.com/baxter-manufacturing-robot-commitment/
Merging the learning of affordance-cues and policy imitation
28
Insights from Cognitive Psychology
29
Merging the learning of affordance-cues and policy imitation
30
31
Contrastive Learning for Affordance Discovery
32
body-grasp
handle-left-right-grasp
handle-front-back-grasp
1
2
3
4
5
6
7
8
33
The learning of Affordance Embeddings Guides the learning of Observation Embeddings
Simultaneously Learning Affordance-Cues and Grasping Policy
34
Attention Model
Policy Model
Encoder
…
…
Simultaneously Learning Affordance-Cues and Grasping Policy
35
Attention Model
Policy Model
Encoder
Optimize it with the coupled triplet loss (need to make two other copies of the framework)
…
…
Simultaneously Learning Affordance-Cues and Grasping Policy
36
Attention Model
Policy Model
Encoder
Optimize it with our coupled triplet loss (need to make two other copies of the framework)
Final Loss
…
…
Video: Simultaneously Predicting Affordance-Cues and Actions
37
Full Model:
The affordance-cue highlights mug body and robot grasps the mug by body
Full Model:
The affordance-cue is shifted to the handle and robot grasps the mug by handle
Panda Robot
Image from: https://wiredworkers.io/our-cobots/cobot-franka-emika-panda/
38
The same encoder-decoder without contrastive learning:
The model seem to always focus on handles of different mugs
Siamese Encoder + Normal Triplet Loss:
The model seem to always focus on the whole mug for different mugs
Experimental Results
39
[20] J. Morton and M. J. Kochenderfer, “Simultaneous policy learning and latent state inference for imitating driver behavior,” in ITSC, IEEE, 2017
Takeaways
40
Takeaways
41
Takeaways
42
43
Learning a Shallow Domain Model for Planning
Part II
Recognizing Plans by Learning Embeddings from Observed Action Distributions
Yantian Zha, Yikang Li, Sriram Gopalakrishnan, Baoxin Li, Subbarao Kambhampati�AAMAS 18 and IJCAI/AAMAS/ICML Joint Workshop 18
Discovering Underlying Plans Based on Shallow Models
Hankz Hankui Zhuo, Yantian Zha, Subbarao Kambhampati, Xin Tian
ACM Transactions on Intelligent Systems and Technology, Volume 11, Issue 2, March 2020
44
Subconscious Level
Metacognition Level
Conscious Level
Planning
Perceiving
Background Knowledge
Cognitive Agent Learning
Model
Zha et al, Distr2Vec, AAMAS, 18
Zha et al, Distr2Vec, AAMAS, 18
Zhuo et al., RNNPlanner, TIST, 20
Zha et al., Distr2Vec, AAMAS, 18
Zhuo et al., RNNPlanner, TIST, 20
Shallow Planning Domain Models
45
46
47
The plan traces are sequences of symbolic activities
They can be viewed as sentences in NLP
Can we borrow some NLP techniques to solve our planning or plan recognition problems?
Using Shallow Models for Planning
48
Word Embedding Learning
A Problem
49
50
51
Learn distr2vec
52
PER: Perception Error Rate; Distr2Vec: Our model
RBM: Resampling deterministic action sequences from distribution sequences (Baseline-1)
NM: Directly take the actions of the highest recognition confidences from distribution sequences (Baseline-2)
Takeaways
53
54
Coupling the Perceiving and Planning
Part III
Plan-Recognition-Driven Attention Modeling for Visual Recognition�Yantian Zha, Yikang Li, Tianshu Yu, Subbarao Kambhampati, Baoxin Li
AAAI-PAIR Workshop, 2019
55
Subconscious Level
Metacognition Level
Conscious Level
Planning
Perceiving
Background Knowledge
Zha et al, PRDA, AAAI-WS, 18
Cognitive Agent Learning
Zha et al, Distr2Vec, AAMAS, 18
Zhuo et al., RNNPlanner, TIST, 20
Model
Zha et al., Distr2Vec, AAMAS, 18
Zhuo et al., RNNPlanner, TIST, 20
Zha et al, Distr2Vec, AAMAS, 18
Where Would You Focus On?
56
Plan Recognition Driven Attention
57
Your mental procedure is perhaps this:
58
Focusing on the moving policemen
Estimating their plans
Also paying attention to regions where they may go to
Your mental procedure is perhaps this:
59
Focusing on the moving policemen
Estimating their plans
Also paying attention to regions where they may go to
What state-of-the-art attention models can do
What we are trying to do
Why using plans to drive attention is hard?
60
How to do plan recognition?
61
Focusing on the moving policemen
Recognize their plans
Also paying attention to regions where they may go to
We use a shallow-model based plan recognition approach
62
[1] Tian, X., Zhuo, H.H. and Kambhampati, S., 2016, May. Discovering underlying plans based on distributed representations of actions. In AAMAS 16
[2] Zha, Y., Li, Y., Gopalakrishnan, S., Li, B. and Kambhampati, S., 2018, July. Recognizing plans by learning embeddings from observed action distributions. In AAMAS 18.
Extracting an AMP (action) from two video frames
63
Plan Recognition (UDUP) | Plan Recognition (our work) |
| |
| |
Extracting an AMP (action) from two video frames
64
…
Video Stream
Frame-Level Feature
Frame Difference
Optical Flow
Clustering
…
AMP Features Learning
Extracting an AMP (action) from two video frames
65
…
Video Stream
Frame-Level Feature
Frame Difference
Optical Flow
Clustering
…
…
AMP Features Learning
Plan Corpora Collection
Extracting an AMP (action) from two video frames
66
…
Video Stream
Frame-Level Feature
Frame Difference
Optical Flow
Clustering
…
…
Obtain an AMP for two consecutive frames:
AMP Features Learning
Plan Corpora Collection
Extracting an AMP (action) from two video frames
67
AMP Features Learning
Plan Corpora Collection
…
Video Stream
Frame-Level Feature
Frame Difference
Optical Flow
Clustering
…
…
Obtain an AMP for two consecutive frames:
Temporarily merging if they have the same order w.r.t probabilities
68
Observation
A glimpse as a visual observation
69
Observation
A glimpse as a visual observation
Plan Recognition via Distr2Vec
70
Observation
How each AMP would shift states in a glimpse to the next state
A glimpse as a visual observation
Plan Recognition via Distr2Vec
71
Observation
Where would be important in future!
A glimpse as a visual observation
Plan Recognition via Distr2Vec
PDN in Event Recognition (ER-PDN)
72
Plan Recognition Driven Attention
Pretrained!
Recognized plans pf M groups at step t
Extracted M glimpses at step t
Shallow-Planning Model
PDN in Event Recognition (ER-PDN)
73
Plan Recognition Driven Attention
Pretrained!
Shallow-Planning Model
Shallow-Planning Model
Evaluation
74
Plan Recognition Driven Attention
Girdhar, Rohit, and Deva Ramanan. "Attentional pooling for action recognition." In Advances in Neural Information Processing Systems, pp. 34-45. 2017.
Evaluation-Visualization
75
Plan Recognition Driven Attention
Takeaways
76
77
Self-Explanation Supported Learning
Part IV
Learning from Ambiguous Demonstrations with Self-Explanation Guided Reinforcement Learning
Yantian Zha, Lin Guan, Subbarao Kambhampati
AAAI-22 Workshop on Reinforcement Learning in Games
78
Subconscious Level
Metacognition Level
Conscious Level
Perceiving
Acting
Self-Explaining
Background Knowledge
Cognitive Agent Learning
Zha et al, Self-Explanation-guided-RLfD, AAAI-RLG Workshop, 22
Zha et al, Self-Explanation-guided-RLfD, Under-Review, 21
Motivation
79
Motivation
80
Learning Guidance
Motivation – Learning from Demonstrations
81
robot
humans'
Settings and Overview
82
Settings and Overview
83
RL
from
demonstrations
Can be ambiguous in complex tasks
Settings and Overview
84
RL
from
demonstrations
samples
Interact with an Environment
Can be ambiguous in complex tasks
Settings and Overview
85
RL
from
demonstrations
samples
failure ones
successful ones
Interact with an Environment
Can be ambiguous in complex tasks
Settings and Overview
86
RL
from
demonstrations
samples
failure ones
successful ones
Self-Explain
human-related background
knowledge
Interact with an Environment
Can be ambiguous in complex tasks
Settings and Overview
Robots take human advice
87
RL
from
demonstrations
samples
failure ones
successful ones
Self-Explain
human-related background
knowledge
Can be ambiguous in complex tasks
Guide:
State-Augmentation and/or
Reward Augmentation
Self-Explanation
Interact with an Environment
88
What is a Self-Explanation?
89
90
Image from: https://www.theconstructsim.com/fetch-robot-simulation-3/
91
Image from: https://www.theconstructsim.com/fetch-robot-simulation-3/
92
93
94
95
96
This demo is made by my co-author Lin Guan
97
Blue curves: Baseline RLfD models: TD3fD, SACfD, and SQLfD
Aqua Curves: RLfD + SE; RL learning only uses self-explanation to augment states
Pink Curves: RLfD + SE; RL learning only uses self-explanation to augment rewards
Red curves: RLfD + Self-Explainer (SE); RL learning uses self-explanation to augment states and rewards
Gold Curves: Baseline GAN-Inverse-RL Model
Gray Curves: RLfD + SE; RL learning does not use task rewards, only self-explanation augmented rewards
Uses Our Self-Explainer
Takeaways
98
Takeaways
99
When Cognitive Agents take Advice from Humans
100
When Cognitive Agents take Advice from Humans
101
We suggest this!
102
Cognitive Learning with Humans in the Loop
Looking up at the Blue Sky
Symbols as a Lingua Franca for Bridging Human-AI Chasm for Explainable and Advisable AI Systems
Kambhampati, Sarath Sreedharan, Mudit Verma, Yantian Zha, Lin Guan
AAAI Blue Sky Paper 2022.
When Robots need Humans – What could be Communication Channels?
103
When Robots need Humans – What could be Communication Channels?
104
Symbolic Interface for Human-AI Systems
105
Symbolic Interface in our Framework
106
The symbolic interface can be used to take advices from humans or give explanations for their decision to humans
Challenges
107
The challenge of approximating explanations constrained by the symbolic interface
The challenge of assembling the symbolic interface itself
The challenge of effectively taking information from the symbolic interface
Challenges – Advice Incorporation
108
The challenge of approximating explanations constrained by the symbolic interface
The challenge of assembling the symbolic interface itself
The challenge of effectively taking information from the symbolic interface
Advice Incorporation – Self-Explanation
109
110
Conclusion and Future Directions
Part V
Conclusion – What I have Addressed
111
Conclusion – What I have Addressed
112
Conclusion – What I have Addressed
113
Looking Forward..
114
Perception-Acting Coupling
Self-Explanation Guided Learning
Perception-Planning Coupling
Challenges
115
116
Perception-Acting Coupling
Self-Explanation Guided Learning
Perception-Planning Coupling
Self-Explanation-guided Learning for Tasks that Require Explicit Knowledge
117
Advice Incorporation – Self-Explanation
118
Self-Explanation-guided Learning for Tasks that Require Explicit Knowledge
119
120
Perception-Acting Coupling
Self-Explanation Guided Learning
Perception-Planning Coupling
Self-Explanation-guided Learning for Tasks that Require Tacit Knowledge
121
Self-Explanation-guided Learning for Tasks that Require Tacit Knowledge
122
I helped Dr. Anagha Kulkarni on her Explainable Planning paper by making the following demo
123
Explicability as Minimizing Distance from Expected Behavior
Anagha Kulkarni, Yantian Zha, Tathagata Chakraborti, Satya Gautam Vadlamudi, Yu Zhang and Subbarao Kambhampati
AAMAS 2019
124
Publication
125
Awards
126
Thank you!
127
Prof. Subbarao Kambhampati
Prof. Baoxin Li
Prof. Siddharth Srivastava
Dr. Jianjun Wang
Prof. Tony Zhang
Prof. Yezhou Yang
Prof. Heni Ben Amor