1 of 127

Yantian Zha

Ph.D. Dissertation Defense

Committee Members:

Prof. Subbarao Kambhampati, Chair and Advisor

Prof. Baoxin Li

Prof. Siddharth Srivastava

Dr. Jianjun Wang

Perceiving, Acting, Planning, and Self-Explaining:

A Cognitive Quartet with Four Neural Networks

+

Thesis Website

2 of 127

2

3 of 127

3

Perceiving

Component

Acting

Component

Planning

Component

Self-Explaining

Component

4 of 127

Perceiving

  • Understanding the real-world
  • Providing the most useful information for higher-level cognition components

4

[1] Action Genome: Actions as Composition of Spatio-temporal Scene Graphs, Ji, Jingwei, et al., CVPR 2020

[1]

5 of 127

Acting

  • A robot needs to react to its perception by taking primitive actions to change the world

​

5

6 of 127

Planning

  • With a model of the world, robot can plan to achieve long-term goals
  • Searching a solution from the model with initial and goal states

6

preconditions:�    at(x, table)�    ∧ empty(gripper)�effects:�    ~at(x, table) ∧� ~empty(gripper)

7 of 127

Self-Explaining

  • Self-explaining experiences about how a problem was solved, why a mistake was made, why others made it, etc.
  • Metacognition: learners be aware and reflective of their own learning

​

7

8 of 127

Self-Explaining

  • Self-explaining experiences about how a problem was solved, why a mistake was made, why others made it, etc.
  • Metacognition: learners be aware and reflective of their own learning

​

8

A Successful case

L1

L2

9 of 127

Self-Explaining

  • Self-explaining experiences about how a problem was solved, why a mistake was made, why others made it, etc.
  • Metacognition: learners be aware and reflective of their own learning

​

9

Self-Explanation:

object-location > object-color

A Successful case

L1

L2

10 of 127

Self-Explaining

  • Self-explaining experiences about how a problem was solved, why a mistake was made, why others made it, etc.
  • Metacognition: learners be aware and reflective of their own learning

​

10

Self-Explanation:

object-location > object-color

A Successful case

L1

L2

L1

L2

11 of 127

Self-Explaining

  • Self-explaining experiences about how a problem was solved, why a mistake was made, why others made it, etc.
  • Metacognition: learners be aware and reflective of their own learning

​

11

Self-Explanation:

object-location > object-color

A Successful case

Unsuccessful

L1

L2

L1

L2

12 of 127

Self-Explaining

  • Self-explaining experiences about how a problem was solved, why a mistake was made, why others made it, etc.
  • Metacognition: learners be aware and reflective of their own learning

​

12

Self-Explanation:

object-location > object-color

A Successful case

The successful and unsuccessful cases are only different in color assignment

Unsuccessful

L1

L2

L1

L2

13 of 127

Self-Explaining

  • Self-explaining experiences about how a problem was solved, why a mistake was made, why others made it, etc.
  • Metacognition: learners be aware and reflective of their own learning

​

13

Self-Explanation:

object-location > object-color

A Successful case

The successful and unsuccessful cases are only different in color assignment

Self-Explanation:

object-color > object-location

Unsuccessful

L1

L2

L1

14 of 127

Self-Explaining

  • Self-explaining experiences about how a problem was solved, why a mistake was made, why others made it, etc.
  • Metacognition: learners be aware and reflective of their own learning

​

14

Self-Explanation:

object-location > object-color

A Successful case

The successful and unsuccessful cases are only different in color assignment

Self-Explanation:

object-color > object-location

Unsuccessful

Successful

L1

L2

L1

L2

L2

L1

15 of 127

Self-Explaining

  • Self-explaining experiences about how a problem is solved, why a mistake is made, why others made it, etc.
  • Metacognition: learners be aware and reflective of their own learning

​

  • Self-explaining, as a kind of metacognition, helps the robot gain insights from what it has experienced so far
  • Its future learning of other cognitive components can then be guided and improved

15

16 of 127

A Coupling of the Four

  1. Perception-Acting Coupling

​

  • Perception-Planning Coupling

​

  • Self-Explaining – Metacognition – monitors and improves other cognitive functions, e.g. 1 and 2

​

​

16

17 of 127

A Coupling of the Four

  1. Perception-Acting Coupling
    1. Can robots learn tacit perception knowledge from demonstrations to guide the learning of acting?
  2. Perception-Planning Coupling
    • Can we use the recognized plans of humans to guide visual recognition tasks?
  3. Self-Explaining to guide learning
    • Can we make cognitive agents more reflective?

17

18 of 127

A Coupling of the Four

  1. Perception-Acting Coupling
    1. Can robots learn tacit perception knowledge from demonstrations to guide the learning of acting?
  2. Perception-Planning Coupling
    • Can we use the recognized plans of humans to guide visual monitoring tasks?
  3. Self-Explaining to guide learning
    • Can we make cognitive agents more reflective?

Their couplings motivate a stronger cognitive intelligence

Allows cognitive agents to learn to accomplish more complex tasks better and more efficiently

18

19 of 127

19

Subconscious Level

Metacognition Level

Conscious Level

Planning

Perceiving

Acting

Self-Explaining

Background Knowledge

Zha et al, PRDA, AAAI-WS, 18

Zha et al, Affordance-Aware-Imitation-Learning, IROS, 21

Zha et al, Self-Explanation-guided-RLfD, AAAI-RLG Workshop, 22

Cognitive Agent Learning

Model

Zha et al, Distr2Vec, AAMAS, 18

Kulkarni et al., Explicable Planning, AAMAS, 19

Zha et al, Self-Explanation-guided-RLfD, Under-Review, 21

Zha et al, Distr2Vec, AAMAS, 18

Zhuo et al., RNNPlanner, TIST, 20

​

Zha et al., Distr2Vec, AAMAS, 18

Zhuo et al., RNNPlanner, TIST, 20

Kulkarni et al., Explicable Planning, AAMAS, 19

​

20 of 127

20

Coupling the Perception and Acting

Part I

Contrastively Learning Visual Attention as Affordance Cues from Demonstrations for Robotic Grasping

Yantian Zha, Siddhant Bhambri, Lin Guan

IROS 2021�

21 of 127

21

Subconscious Level

Metacognition Level

Conscious Level

Perceiving

Acting

Cognitive Agent Learning

Zha et al, Affordance-Aware-Imitation-Learning, IROS, 21

22 of 127

Affordance

22

action

object

effect

23 of 127

Grasping Affordance

23

action

object

effect

 

24 of 127

Grasping Affordance

24

action

object

effect

25 of 127

Grasping Affordance

25

action

object

effect

26 of 127

Grasping Affordance Modeling

  •  

26

Clusters of approaching orientations [1]

Clusters of grasp pre-shapes [2]

[1] de Granville C et al. Learning grasp affordances through human demonstration. In Proceedings of ICDL 2006.

27 of 127

27

Baxter Robot

Image from: https://www.active8robots.com/baxter-manufacturing-robot-commitment/

28 of 127

Merging the learning of affordance-cues and policy imitation

  • We open the direction of integrating learning affordance from demonstrations and deep learning based imitation learning in an end-to-end deep learning framework
  • This enables their learning to support each other in a tightly coupled way

​

  • Objective: Robot imitates not only expert behavior, but also the affordance tacit knowledge of the teacher

​

28

29 of 127

Insights from Cognitive Psychology

  • There is a close association between affordance and attention

​

  • Humans tend to focus on object parts that are permissible to accomplish a task

​

29

30 of 127

Merging the learning of affordance-cues and policy imitation

  • We initialize the direction of integrating learning affordance from demonstrations and policy imitation learning in an end-to-end deep learning framework
  • Objective: Robot imitates not only expert behavior, but also the affordance tacit knowledge of the teacher

​

30

31 of 127

31

32 of 127

Contrastive Learning for Affordance Discovery

  • Contrastive learning is a machine learning technique used to learn the general features of a dataset without labels by teaching the model which data points are similar or different

​

​

32

body-grasp

handle-left-right-grasp

handle-front-back-grasp

1

2

3

4

5

6

7

8

33 of 127

33

The learning of Affordance Embeddings Guides the learning of Observation Embeddings

34 of 127

Simultaneously Learning Affordance-Cues and Grasping Policy

34

Attention Model

Policy Model

 

 

 

Encoder

 

 

 

 

…

…

35 of 127

Simultaneously Learning Affordance-Cues and Grasping Policy

35

Attention Model

Policy Model

 

 

 

Encoder

Optimize it with the coupled triplet loss (need to make two other copies of the framework)

 

 

 

 

 

…

…

36 of 127

Simultaneously Learning Affordance-Cues and Grasping Policy

36

Attention Model

Policy Model

 

 

 

Encoder

Optimize it with our coupled triplet loss (need to make two other copies of the framework)

 

Final Loss

 

 

 

 

…

…

37 of 127

Video: Simultaneously Predicting Affordance-Cues and Actions

37

Full Model:

The affordance-cue highlights mug body and robot grasps the mug by body

Full Model:

The affordance-cue is shifted to the handle and robot grasps the mug by handle

Panda Robot

Image from: https://wiredworkers.io/our-cobots/cobot-franka-emika-panda/

38 of 127

38

The same encoder-decoder without contrastive learning:

The model seem to always focus on handles of different mugs

Siamese Encoder + Normal Triplet Loss:

The model seem to always focus on the whole mug for different mugs

39 of 127

Experimental Results

39

[20] J. Morton and M. J. Kochenderfer, “Simultaneous policy learning and latent state inference for imitating driver behavior,” in ITSC, IEEE, 2017

40 of 127

Takeaways

  • Learning affordances from demonstrations has the benefit of avoiding labor-intensive affordance label collection
  • It would be promising to replace the external motion planner in such works with a learnable policy neural network

40

41 of 127

Takeaways

  • Learning affordances from demonstrations has the benefits of avoiding labor-intensive affordance label collections
  • It would be promising to replace the external motion planner in such works with a learnable policy neural network
  • We combine two directions together – affordance-cue learning from demonstrations and deep policy imitation learning – in an end-to-end deep learning framework
  • It would be interesting to consider multiple levels of interactions that drive richer affordance-cues. It would be also interesting to apply our coupled triplet loss to other robot learning problems

​

41

42 of 127

Takeaways

  • Learning affordances from demonstrations has the benefits of avoiding labor-intensive affordance label collections
  • It would be promising to replace the external motion planner in such works with a learnable policy neural network
  • We combines two directions together – affordance-cue learning from demonstrations and deep policy imitation learning – in an end-to-end deep learning framework
  • It would be interesting to consider multiple levels of interactions that drive richer affordance-cues. It would be also interesting to apply our coupled triplet loss to other robot learning problems

​

  • While the affordance-cue learning made the perception more “conscious”, its learning is “model-free” that does not support “long-term thinking”

42

43 of 127

43

Learning a Shallow Domain Model for Planning

Part II

Recognizing Plans by Learning Embeddings from Observed Action Distributions

Yantian Zha, Yikang Li, Sriram Gopalakrishnan, Baoxin Li, Subbarao Kambhampati�AAMAS 18 and IJCAI/AAMAS/ICML Joint Workshop 18

​

Discovering Underlying Plans Based on Shallow Models

Hankz Hankui Zhuo, Yantian Zha, Subbarao Kambhampati, Xin Tian

ACM Transactions on Intelligent Systems and Technology, Volume 11, Issue 2, March 2020

44 of 127

44

Subconscious Level

Metacognition Level

Conscious Level

Planning

Perceiving

Background Knowledge

Cognitive Agent Learning

Model

Zha et al, Distr2Vec, AAMAS, 18

Zha et al, Distr2Vec, AAMAS, 18

Zhuo et al., RNNPlanner, TIST, 20

​

Zha et al., Distr2Vec, AAMAS, 18

Zhuo et al., RNNPlanner, TIST, 20

45 of 127

Shallow Planning Domain Models

45

46 of 127

46

47 of 127

47

The plan traces are sequences of symbolic activities

They can be viewed as sentences in NLP

Can we borrow some NLP techniques to solve our planning or plan recognition problems?

48 of 127

Using Shallow Models for Planning

48

Word Embedding Learning

49 of 127

A Problem

  • They assume deterministic action recognition

​

  • What if the action recognition has errors?

​

  • Then the plan recognizer has ZERO access to the correct action at those steps

49

50 of 127

50

51 of 127

51

Learn distr2vec

 

 

52 of 127

52

PER: Perception Error Rate; Distr2Vec: Our model

RBM: Resampling deterministic action sequences from distribution sequences (Baseline-1)

NM: Directly take the actions of the highest recognition confidences from distribution sequences (Baseline-2)

53 of 127

Takeaways

  • If recognized actions are from a visual recognizer trained on real world data, the action recognition should have uncertainties
  • We want to bridge the gap between plan recognition and real-world visual recognition
  • We propose Distr2Vec embeddings for learning shallow planning domain models from observed action distributions

​

  • Can we use such a leaned model to support plan recognition to improve perception?

53

54 of 127

54

Coupling the Perceiving and Planning

Part III

Plan-Recognition-Driven Attention Modeling for Visual Recognition�Yantian Zha, Yikang Li, Tianshu Yu, Subbarao Kambhampati, Baoxin Li

AAAI-PAIR Workshop, 2019

55 of 127

55

Subconscious Level

Metacognition Level

Conscious Level

Planning

Perceiving

Background Knowledge

Zha et al, PRDA, AAAI-WS, 18

Cognitive Agent Learning

Zha et al, Distr2Vec, AAMAS, 18

Zhuo et al., RNNPlanner, TIST, 20

​

Model

Zha et al., Distr2Vec, AAMAS, 18

Zhuo et al., RNNPlanner, TIST, 20

Zha et al, Distr2Vec, AAMAS, 18

56 of 127

Where Would You Focus On?

56

57 of 127

Plan Recognition Driven Attention

57

58 of 127

Your mental procedure is perhaps this:

58

Focusing on the moving policemen

Estimating their plans

Also paying attention to regions where they may go to

59 of 127

Your mental procedure is perhaps this:

59

Focusing on the moving policemen

Estimating their plans

Also paying attention to regions where they may go to

What state-of-the-art attention models can do

What we are trying to do

60 of 127

Why using plans to drive attention is hard?

  • Task plans are typically discrete symbolic actions

​

  • Attention maps and video frames are pixel-like values in continuous space
    • The same object in a video may have different point of views, scalars, etc.

​

  • How to connect them together?

​

​

60

61 of 127

How to do plan recognition?

61

Focusing on the moving policemen

Recognize their plans

Also paying attention to regions where they may go to

62 of 127

We use a shallow-model based plan recognition approach

  • Although plan recognition is hard, there is a line of works based on using pre-defined models

​

  • For vision data, identifying high level actions and a model for those actions is difficult

​

  • We use a shallow model based approach1,2: learning a shallow domain model via learning affinities from plan traces, and do plan recognition based on that model

​

  • Observation at each time step are action distributions, and we use the learned distribution affinities to do plan recognition (the work Distr2Vec2)

62

[1] Tian, X., Zhuo, H.H. and Kambhampati, S., 2016, May. Discovering underlying plans based on distributed representations of actions. In AAMAS 16

[2] Zha, Y., Li, Y., Gopalakrishnan, S., Li, B. and Kambhampati, S., 2018, July. Recognizing plans by learning embeddings from observed action distributions. In AAMAS 18.

63 of 127

Extracting an AMP (action) from two video frames

63

Plan Recognition (UDUP)

Plan Recognition (our work)

​

​

​

​

64 of 127

Extracting an AMP (action) from two video frames

64

…

Video Stream

Frame-Level Feature

Frame Difference

Optical Flow

Clustering

 

…

AMP Features Learning

65 of 127

Extracting an AMP (action) from two video frames

65

…

Video Stream

Frame-Level Feature

Frame Difference

Optical Flow

Clustering

 

…

…

 

 

 

 

 

AMP Features Learning

Plan Corpora Collection

66 of 127

Extracting an AMP (action) from two video frames

66

…

Video Stream

Frame-Level Feature

Frame Difference

Optical Flow

Clustering

 

…

…

 

 

 

 

 

Obtain an AMP for two consecutive frames:

 

AMP Features Learning

Plan Corpora Collection

67 of 127

Extracting an AMP (action) from two video frames

67

AMP Features Learning

Plan Corpora Collection

…

Video Stream

Frame-Level Feature

Frame Difference

Optical Flow

Clustering

 

…

…

 

 

 

 

 

Obtain an AMP for two consecutive frames:

 

 

 

Temporarily merging if they have the same order w.r.t probabilities

68 of 127

 

68

 

Observation

A glimpse as a visual observation

69 of 127

69

 

 

Observation

A glimpse as a visual observation

Plan Recognition via Distr2Vec

 

70 of 127

 

70

 

 

Observation

 

 

 

How each AMP would shift states in a glimpse to the next state

A glimpse as a visual observation

Plan Recognition via Distr2Vec

71 of 127

 

71

 

 

Observation

 

 

 

 

 

Where would be important in future!

A glimpse as a visual observation

Plan Recognition via Distr2Vec

72 of 127

PDN in Event Recognition (ER-PDN)

72

Plan Recognition Driven Attention

Pretrained!

Recognized plans pf M groups at step t

Extracted M glimpses at step t

Shallow-Planning Model

73 of 127

PDN in Event Recognition (ER-PDN)

73

Plan Recognition Driven Attention

Pretrained!

Shallow-Planning Model

Shallow-Planning Model

74 of 127

Evaluation

74

Plan Recognition Driven Attention

Girdhar, Rohit, and Deva Ramanan. "Attentional pooling for action recognition." In Advances in Neural Information Processing Systems, pp. 34-45. 2017.

75 of 127

Evaluation-Visualization

75

Plan Recognition Driven Attention

 

 

 

 

 

 

76 of 127

Takeaways

  • Intuitively, improving perception ability also improves planning abilities

​

  • We use plan recognition to improve visual recognition (perception) tasks, via modeling Plan-Recognition-Driven-Attention

​

  • “Think ahead to better see the present”

​

  • Can we use even higher-level cognitive modeling (e.g. self-explaining) to improve perception and lower-level cognitive functions?

​

​

​

​

​

​

​

76

77 of 127

77

Self-Explanation Supported Learning

Part IV

Learning from Ambiguous Demonstrations with Self-Explanation Guided Reinforcement Learning

Yantian Zha, Lin Guan, Subbarao Kambhampati

AAAI-22 Workshop on Reinforcement Learning in Games

78 of 127

78

Subconscious Level

Metacognition Level

Conscious Level

Perceiving

Acting

Self-Explaining

Background Knowledge

Cognitive Agent Learning

Zha et al, Self-Explanation-guided-RLfD, AAAI-RLG Workshop, 22

Zha et al, Self-Explanation-guided-RLfD, Under-Review, 21

79 of 127

Motivation

  • We want an agent to be reflective of its own learning

​

  • The agent should be able to gain insights from its own and others’ experiences to guide its future learning

​

  • Most of its own experiences might be failure ones while others’ experiences could be successful ones

​

  • What makes them successful or not?

​

​

​

79

80 of 127

Motivation

  • We want an agent to be reflective

​

  • The agent should be able to gain insights from its own and others’ experiences to guide its future learning

​

  • Most of its own experiences might be failure ones while others’ experiences are successful ones

​

  • What makes them successful or not?

​

​

​

80

Learning Guidance

81 of 127

Motivation – Learning from Demonstrations

  • We want an agent to be reflective

​

  • The agent should be able to gain insights from its own and others’ (successful) experiences to guide its future learning

​

  • It makes robot easier to learn from humans if robot can “think” at a similar level to humans by having some human-relevant background knowledge

​

​

​

81

robot

humans'

82 of 127

Settings and Overview

  • Robots take human advice

82

83 of 127

Settings and Overview

  • Robots take human advice

83

RL

from

demonstrations

Can be ambiguous in complex tasks

84 of 127

Settings and Overview

  • Robots take human advice

84

RL

from

demonstrations

samples

Interact with an Environment

Can be ambiguous in complex tasks

85 of 127

Settings and Overview

  • Robots take human advice

85

RL

from

demonstrations

samples

failure ones

successful ones

Interact with an Environment

Can be ambiguous in complex tasks

86 of 127

Settings and Overview

  • Robots take human advice

86

RL

from

demonstrations

samples

failure ones

successful ones

Self-Explain

human-related background

knowledge

Interact with an Environment

Can be ambiguous in complex tasks

87 of 127

Settings and Overview

Robots take human advice

87

RL

from

demonstrations

samples

failure ones

successful ones

Self-Explain

human-related background

knowledge

Can be ambiguous in complex tasks

Guide:

State-Augmentation and/or

Reward Augmentation

Self-Explanation

Interact with an Environment

88 of 127

88

What is a Self-Explanation?

89 of 127

89

90 of 127

90

Image from: https://www.theconstructsim.com/fetch-robot-simulation-3/

91 of 127

91

Image from: https://www.theconstructsim.com/fetch-robot-simulation-3/

92 of 127

92

93 of 127

93

94 of 127

94

95 of 127

95

96 of 127

96

This demo is made by my co-author Lin Guan

97 of 127

97

Blue curves: Baseline RLfD models: TD3fD, SACfD, and SQLfD

Aqua Curves: RLfD + SE; RL learning only uses self-explanation to augment states

Pink Curves: RLfD + SE; RL learning only uses self-explanation to augment rewards

Red curves: RLfD + Self-Explainer (SE); RL learning uses self-explanation to augment states and rewards

Gold Curves: Baseline GAN-Inverse-RL Model

Gray Curves: RLfD + SE; RL learning does not use task rewards, only self-explanation augmented rewards

Uses Our Self-Explainer

98 of 127

Takeaways

  • We open the direction of learning to self-explain the past experiences (which includes demonstrated ones) to guide the control-level learning

​

  • Using self-explanation for learning is beneficial for complex settings, e.g., environments that easily cause ambiguous demonstrations

​

98

99 of 127

Takeaways

  • We open the direction of learning to self-explain the past experiences (which includes demonstrated ones) to guide the control-level learning

​

  • Using self-explanation for learning is beneficial for complex settings, e.g., environments that easily cause ambiguous demonstrations

​

  • It would be interesting to explore how to use our self-explaining mechanism to help other robot learning problems

​

  • It would also be interesting to think about if there are better ways to endow agents the self-explaining abilities

99

100 of 127

When Cognitive Agents take Advice from Humans

  • Robots can take humans’ advice directly from raw data
    • Contrastively Learning Visual Attention as Affordance Cues from Demonstrations for Robotic Grasping

​

  • Robots can also take humans’ advice through the lens of human-related background knowledge (e.g. human-understandable predicates as in my self-explanation work)
    • Learning from Ambiguous Demonstrations with Self-Explanation Guided Reinforcement Learning

​

​

​

100

101 of 127

When Cognitive Agents take Advice from Humans

  • Robots can take humans’ advice directly from raw data
    • Contrastively Learning Visual Attention as Affordance Cues from Demonstrations for Robotic Grasping

​

  • Robots can also take humans’ advice through the lens of human-related background knowledge (e.g. human-understandable predicates as in my self-explanation work)
    • Learning from Ambiguous Demonstrations with Self-Explanation Guided Reinforcement Learning

​

​

​

101

We suggest this!

102 of 127

102

Cognitive Learning with Humans in the Loop

Looking up at the Blue Sky

Symbols as a Lingua Franca for Bridging Human-AI Chasm for Explainable and Advisable AI Systems

Kambhampati, Sarath Sreedharan, Mudit Verma, Yantian Zha, Lin Guan

AAAI Blue Sky Paper 2022.

103 of 127

When Robots need Humans – What could be Communication Channels?

  1. Asking humans to figure out how to understand robots’ own (internal) representations

​

  • Share the raw data (attention images, videos, trajectories) with humans

​

  • Share human-understandable concepts with humans

103

104 of 127

When Robots need Humans – What could be Communication Channels?

  1. Asking humans to figure out how to understand robots’ own (internal) representations
    • Human-AI systems should be designed for the benefits of humans
  2. Share the raw data (attention images, videos, trajectories) with humans
    • Causes high cognitive load for humans
    • Scaling-up issues for human-interaction tasks that involve both tacit and explicit task knowledge
  3. Share human-understandable concepts with humans
    • Humans themselves develop symbols for efficient human-human communications

104

105 of 127

Symbolic Interface for Human-AI Systems

  • While it might be good for AI/Robots to have their own learned knowledge

​

  • It would be beneficial for their interactions with humans if AI/Robots can also develop local symbolic representations that are understandable to humans in the loop

​

  • The symbolic interface can be used to take advices from humans or give explanations to humans

105

106 of 127

Symbolic Interface in our Framework

106

The symbolic interface can be used to take advices from humans or give explanations for their decision to humans

107 of 127

Challenges

107

The challenge of approximating explanations constrained by the symbolic interface

The challenge of assembling the symbolic interface itself

The challenge of effectively taking information from the symbolic interface

108 of 127

Challenges – Advice Incorporation

108

The challenge of approximating explanations constrained by the symbolic interface

The challenge of assembling the symbolic interface itself

The challenge of effectively taking information from the symbolic interface

109 of 127

Advice Incorporation – Self-Explanation

109

110 of 127

110

Conclusion and Future Directions

Part V

111 of 127

Conclusion – What I have Addressed

  • Perception-Acting Coupling
    • Affordance-Aware-Imitation-Learning: Combines the directions of learning affordances from demonstrations and deep imitation learning together

​

111

112 of 127

Conclusion – What I have Addressed

  • Perception-Acting Coupling
    • Affordance-Aware-Imitation-Learning: Combines the directions of learning affordances from demonstrations and deep imitation learning together
  • Perception-Planning Coupling
    • Learning a shallow-domain model for planning / plan recognition
      • RNNPlanner and Distr2Vec: Use Neural Networks to learn action affinity models that allow searching for the best actions. RNNPlanner takes deterministic action sequences and Distr2Vec takes observed action distribution sequences that expresses visual recognition uncertainties
    • Plan-Recognition-Driven-Attention: It is intuitive to improve perception to improve high-level planning. We instead use plan recognition (long-term thinking) to improve perception

​

112

113 of 127

Conclusion – What I have Addressed

  • Perception-Acting Coupling
    • Affordance-Aware-Imitation-Learning: Combines the directions of learning affordances from demonstrations and deep imitation learning together
  • Perception-Planning Coupling
    • Learning a shallow-domain model for planning / plan recognition
      • RNNPlanner and Distr2Vec: Use Neural Networks to learn action affinity models that allow searching for the best actions. RNNPlanner takes deterministic action sequences and Distr2Vec takes observed action distribution sequences that expresses visual recognition uncertainties
    • Plan-Recognition-Driven-Attention: It is intuitive to improve perception to improve high-level planning. We instead use plan recognition (long-term thinking) to improve perception
  • Self-explanation supported learning
    • Self-Explanation-for-RLfD: Agents with metacognitive abilities are reflective of themselves. Self-explaining one’s past experiences and gain insights to guide its future learning is a type of metacognitions. Can we use Deep Neural Networks to approximate such behavior?

​

113

114 of 127

Looking Forward..

114

Perception-Acting Coupling

Self-Explanation Guided Learning

Perception-Planning Coupling

115 of 127

Challenges

  • The challenge of defining the form of teachers' advice and formulating its incorporation -- either at the level of symbolic representations or raw data

​

  • The challenge of designing the self-explanation mechanism that can be tightly coupled into robots' learning algorithms

​

  • The challenge of using self-explanation to improve lower-level cognitive components, like perception, planning and performance components, as a whole based on deep neural networks

115

116 of 127

116

Perception-Acting Coupling

Self-Explanation Guided Learning

Perception-Planning Coupling

117 of 127

Self-Explanation-guided Learning for Tasks that Require Explicit Knowledge

  • In my current self-explanation work, we introduce background knowledge that humans can understand (i.e. a symbolic interface)

​

  • This helps robots to self-explain human advice from humans’ perspectives

117

118 of 127

Advice Incorporation – Self-Explanation

118

119 of 127

Self-Explanation-guided Learning for Tasks that Require Explicit Knowledge

  • However, self-explainer is still model-free and needs to be learned

​

  • That is, the prediction of self-explanation hypothesis does not rely on rigorous formats of action models

​

  • Can we provide some shallow-domain planning models to help the learning of self-explainer? (Note that the model is still imperfect)
  • This should further improve the learning efficiency and generalizability of the robot system

119

120 of 127

120

Perception-Acting Coupling

Self-Explanation Guided Learning

Perception-Planning Coupling

121 of 127

Self-Explanation-guided Learning for Tasks that Require Tacit Knowledge

  • In some tasks like learning to grasp, cognitive learners essentially learn more tacit knowledge than explicit knowledge
  • Maintaining a symbolic interface unnecessarily makes the learning more complex

​

​

121

122 of 127

Self-Explanation-guided Learning for Tasks that Require Tacit Knowledge

  • However, self-explaining could still be helpful

​

  • For example, non-professional robot users might not be able to provide demonstrations that cover all REQUIRED tacit affordance knowledge

​

  • Exploration is needed

​

  • Self-explaining could providing guiding signals for exploration

122

123 of 127

 

I helped Dr. Anagha Kulkarni on her Explainable Planning paper by making the following demo

123

Explicability as Minimizing Distance from Expected Behavior

Anagha Kulkarni, Yantian Zha, Tathagata Chakraborti, Satya Gautam Vadlamudi, Yu Zhang and Subbarao Kambhampati

AAMAS 2019

124 of 127

 

124

125 of 127

Publication

  • Symbols as a Lingua Franca for Bridging Human-AI Chasm for Explainable and Advisable AI Systems, Subbarao Kambhampati, Sarath Sreedharan, Mudit Verma, Yantian Zha, Lin Guan AAAI Blue Sky Paper, 2022.
  • Learning from Ambiguous Demonstrations with Self-Explanation Guided Reinforcement Learning, Yantian Zha, Lin Guan, Subbarao Kambhampati, AAAI-22 Workshop on Reinforcement Learning in Games 2022.
  • Contrastively Learning Visual Attention as Affordance Cues from Demonstrations for Robotic Grasping, Yantian Zha, Siddhant Bhambri, and Lin Guan, The IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2021
  • Symbols as a Lingua Franca for Bridging Human-AI Chasm for Explainable and Advisable AI Systems, Subbarao Kambhampati, Sarath Sreedharan, Mudit Verma, Yantian Zha, Lin Guan, AAAI 2022 (Blue Sky Track)
  • Plan-Recognition-Driven Attention Modeling for Visual Recognition, Yantian Zha, Yikang Li, Tianshu Yu, Subbarao Kambhampati and Baoxin Li, AAAI 2019 Workshop on Plan, Activity, and Intent Recognition (PAIR)
  • Recognizing plans by learning embeddings from observed action distributions, Yantian Zha, Yikang Li, Sriram Gopalakrishnan, Baoxin Li, and Subbarao Kambhampati In Proceedings of International Conference on Autonomous Agents and Multiagent Systems (AAMAS), Extended Abstract, 2018, and IJCAI/AAMAS/ICML Joint Workshop 2018
  • Discovering Underlying Plans Based on Shallow Models Hankz Hankui Zhuo, Yantian Zha, Subbarao Kambhampati, and Xin Tian In Proceedings of ACM Transactions on Intelligent Systems and Technology (ACM-TIST), Volume 11, Issue 2, March 2020
  • Explicability as Minimizing Distance from Expected Behavior Anagha Kulkarni, Yantian Zha, Tathagata Chakraborti, Satya Gautam Vadlamudi, Yu Zhang and Subbarao Kambhampati In Proceedings of International Conference on Autonomous Agents and Multiagent Systems (AAMAS), Extended Abstract, 2019.

125

126 of 127

Awards

  • Maryland Robotics Center Postdoctoral Fellowship at the Institute for Systems Research (ISR), University of Maryland, July 2022 – July 2023

​

  • CIDSE Doctoral Fellowship at Arizona State University, 2021

126

127 of 127

Thank you!

127

Prof. Subbarao Kambhampati

Prof. Baoxin Li

Prof. Siddharth Srivastava

Dr. Jianjun Wang

​

Prof. Tony Zhang

Prof. Yezhou Yang

Prof. Heni Ben Amor