1 of 12

When do neuromorphic sensors outperform cameras?

Learning from dynamic features

Cornelia Fermüller

University of Maryland, College Park

joint work with D. Deniz, E. Ros, and F. Barranco

University of Granada, Spain

2 of 12

Prediction of dexterous actions

Motivation: Human Robot Collaboration

  • Capturing motion dynamics is crucial

especially when carrying out different actions

with the same object��

  • Appearance-based features can lead to

incorrect classifications

dynamic features need to be reconstructed from frames

3 of 12

Manipulation Action dataset

  • Manipulation Action Dataset [FER18]
    • RGB Video
    • 5 objects each of 5 actions, 5 actors���
  • Event Manipulation Action Dataset [DEN23]
    • Event-based counterpart of MAD [FER18]
    • 5 objects each of 5 actions, 5 actors

[FER18] C. Fermüller, F. Wang, Y. Yang, K. Zampogiannis, Y. Zhang, F. Barranco, and M. Pfeiffer, “Prediction of manipulation actions,” IJCV 2018

[DEN23] D. Deniz, E. Ros, C. Fermüller, and F. Barranco, “Event-based vision for early prediction of manipulation actions,” submitted to PAMI.

4 of 12

Manipulation Actions

Videos

Events

scoop with a spoon

Videos

Events

shaking a cup

5 of 12

Deep Learning Architectures

LRCN

3DConv-Net

Time Surface

Time Surface

[MA18] N. Ma, X. Zhang, et al., “Shufflenet v2: Practical guidelines for efficient cnn architecture design,” ECCV, 2018.

[TAN19] M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” ICML, 2019.

[TAN19]

[MA19]

6 of 12

Deep Learning Architectures

Mobilenet Transformer

Time Surface

[SAN18] M. Sandler, A. Howard, M. Zhu, et al. “Mobilenetv2: Inverted residuals and linear bottlenecks,” in CVPR, 2018.

[SAN18]

7 of 12

Evaluation on DVS Gestures

Architecture

Accuracy

RG-CNN [BI20]

97.20%

TORE [BAL21]

96.20%

EvT [SAB22]

96.20%

Shufflenet v2 3D

94.31%

Mobilenet Transformer

96.21%

EfficientNet GRU

97.34%

TABLE I: Accuracy of models on DVS 128 Gestures

[BI20] Y. Bi, A. Chadha, A. Abbas, et al. , “Graph-based spatio-temporal feature learning for neuromorphic visión sensing,” IEEE Trans. on Image Processing, 2020.

[BAL21] R. Baldwin, R. Liu, M. Almatrafi, V. Asari, and K. Hirakawa, “Time-ordered recent event (tore) volumes for event cameras,” arXiv:2103.06108, 2021.

[SAB22] A. Sabater, L. Montesano, and A. C. Murillo, “Event transformer. asparse-aware solution for efficient event data processing,” CVPR, 2022.

8 of 12

Event-based Action Recognition

Time Surfaces

Inference

(DL network)

T = 0.25 s

Sponge squeeze 🡪 5%

Sponge flip 🡪 26.4%

Sponge wash 🡪 7.6%

Sponge wipe 🡪 9.5%

Sponge scratch 🡪 11%

T = 1 s

Sponge squeeze 🡪 4%

Sponge flip 🡪 44.9%

Sponge wash 🡪 9.2%

Sponge wipe 🡪 15.2%

Sponge scratch 🡪 10%

T = 2 s

Sponge squeeze 🡪 0%

Sponge flip 🡪 1.7%

Sponge wash 🡪 1.1%

Sponge wipe 🡪 92.8%

Sponge scratch 🡪 3%

Confidence

Confidence

Confidence

Inference with EfficientNet GRU

9 of 12

Video-based Action Recognition

Inference

(DL network)

T = 0.25 s

Sponge squeeze 🡪 5%

Sponge flip 🡪 3%

Sponge wash 🡪 3%

Sponge wipe 🡪 4.1%

Sponge scratch 🡪 4%

T = 1 s

Sponge squeeze 🡪 8%

Sponge flip 🡪 35.6%

Sponge wash 🡪 18.4%

Sponge wipe 🡪 21.6%

Sponge scratch 🡪 12%

T = 2 s

Sponge squeeze 🡪 6%

Sponge flip 🡪 22.8%

Sponge wash 🡪 26.1%

Sponge wipe 🡪 28.5%

Sponge scratch 🡪 15%

Confidence

Confidence

Confidence

Inference with EfficientNet GRU

10 of 12

Video-based Action Recognition

DL architecture

Video-based

Event-based

Accuracy

Accuracy

Mobilenet Transformer

90.02 ± 4.40

96.94 ± 2.50

EfficientNet GRU

79.02 ± 7.07

97.63 ± 0.94

TABLE II: 4-Fold Cross Validation Manipulation Action classification

Event-based Action Recognition

T = 2 s

Sponge squeeze 🡪 0%

Sponge flip 🡪 1.7%

Sponge wash 🡪 1.1%

Sponge wipe 🡪 92.8%

Sponge scratch 🡪 3%

T = 2 s

Sponge squeeze 🡪 6%

Sponge flip 🡪 22.8%

Sponge wash 🡪 26.1%

Sponge wipe 🡪 28.5%

Sponge scratch 🡪 15%

Confidence

Confidence

11 of 12

t-SNE projection of features

a) Events

b) Videos

12 of 12

Conclusion

  • Event-driven solutions offer real-time operation and low latency

  • Despite their struggles with accuracy, they excel in manipulation recognition outperforming video-based approaches

  • Event-driven neural networks effectively learn spatio-temporal features

  • The integration of hand tracking demonstrates the scene- and object-agnostic nature of event-driven approaches, more focused on dynamics�