1 of 31

Do We Mean It When �We Say We Want Robots to Engage?

Jesse Thomason

http://glamor.rocks/

2 of 31

Human–Robot Dialogue

“While most robot learning works focus on unidirectional communication from humans to robots, having robots that not only understand what human users say but also initiate and engage in conversations can help them learn and deploy more effectively in human environments.”

3 of 31

One Instruction, One Scene, Multiple Policies

4 of 31

One Instruction, One Scene, Multiple Policies

​

5 of 31

One Instruction, One Scene, Multiple Policies

​

6 of 31

What Does “Engagement” Look Like for People?

7 of 31

Okay, but when should we add friction?

8 of 31

SAFE: Sequential Conformal Failure Prediction

9 of 31

Insight from NLP: Test in a neighborhood around inputs

10 of 31

Contrast Sets for Evaluating Language-Guided RoboNLP

Go right past the couch and stop at the table.

Go left past the couch and stop at the table.

11 of 31

Contrast Sets for Evaluating Language-Guided RoboNLP

12 of 31

Estimate Policy Metrics with Less Experimenter Labor!

13 of 31

Transfer insight: Probe in a neighborhood around inputs

14 of 31

15 of 31

SAFECAST improves risk control with cost on autonomy

16 of 31

We’ve made progress on when to slow down; now what?

  • Given suspected future failure, we could fall back to requesting human teleoperation
  • Do we need a full demo?
  • Is the high fidelity control of teleoperation even needed, or would a “nudge” do the trick?

Really cool ongoing work from Harshitha Rajaprakash

17 of 31

We’ve made progress on when to slow down; now what?

  • What if we could reverse these perturbations?
  • Given suspected future failure, predict what needs to change to undo that failure
  • Could we ask operators to rearrange the scene? Clarify their language-based intent?

Really cool ongoing work from Harshitha Rajaprakash

18 of 31

Can we give commands and corrections in speech at all?

19 of 31

Modern LLMs are so much worse at speech vs text input

20 of 31

Speech Recognition with Environmental Cues

…

“Wash the dirty pug in the sink”

21 of 31

Speech Recognition with Environmental Cues

22 of 31

Speech Recognition with Environmental Cues

23 of 31

Call to Action: Speech and Interaction-first Robotics

  • Typing to robots will always be silly
  • Speaking to robots is currently also silly, but it shouldn’t be
  • Socially and physically embodied speech recognition is hard
  • Human-human dialogue and interaction involves so much more than just the speech signal

24 of 31

Human-Robot Spoken Dialogue Systems

25 of 31

Human-Robot Building Game

Gaze

Gesture

Speech

Motion

26 of 31

Cooperative Building Games

Player A Goal

Player B Goal

Reachable Space

Gaze

Gesture

Speech

Motion

Implicit requests

Planning

27 of 31

Cooperative Building Games

Player A Goal

Player B Goal

Reachable Space

Competition�Ready??

28 of 31

Cooperative Transportation Games

This way a little

Hang on, lift up

Speech

Intonation

Push

Pull

Shear forces

Rotation

Planning

Stress

Facial expression

29 of 31

Thanks to the GLAMOR ✨ Lab PhD students and alumni

Jesse Zhang

Lee Kezar

Tejas Srinivasan

Leticia Pinto-Alva

Ting-Yun Chang

Abrar Anwar

Ishika Singh

Wang Zhu

Harshitha Rajaprakash

Zizhao Hu

Yuliang Cai

Jie Cai

Wei Yang

Dhananjay Ashok

Gabriela Pinto

30 of 31

GLAMOR robotics MS students applying for Fall’27 PhDs!

Mousumi Das

Aditeya Prajapati

SWAP, SAFECAST, Sign Language, Lip Reading

SWAP, RL, best paper award @ SIGDIAL’26

31 of 31

Do We Mean It When �We Say We Want Robots to Engage?

Jesse Thomason

http://glamor.rocks/