When do neuromorphic sensors outperform cameras?
Learning from dynamic features
Cornelia Fermüller
University of Maryland, College Park
joint work with D. Deniz, E. Ros, and F. Barranco
University of Granada, Spain
Prediction of dexterous actions
Motivation: Human Robot Collaboration
especially when carrying out different actions
with the same object��
incorrect classifications
dynamic features need to be reconstructed from frames
Manipulation Action dataset
[FER18] C. Fermüller, F. Wang, Y. Yang, K. Zampogiannis, Y. Zhang, F. Barranco, and M. Pfeiffer, “Prediction of manipulation actions,” IJCV 2018
[DEN23] D. Deniz, E. Ros, C. Fermüller, and F. Barranco, “Event-based vision for early prediction of manipulation actions,” submitted to PAMI.
Manipulation Actions
Videos
Events
scoop with a spoon
Videos
Events
shaking a cup
Deep Learning Architectures
LRCN
3DConv-Net
Time Surface
Time Surface
[MA18] N. Ma, X. Zhang, et al., “Shufflenet v2: Practical guidelines for efficient cnn architecture design,” ECCV, 2018.
[TAN19] M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” ICML, 2019.
[TAN19]
[MA19]
Deep Learning Architectures
Mobilenet Transformer
Time Surface
[SAN18] M. Sandler, A. Howard, M. Zhu, et al. “Mobilenetv2: Inverted residuals and linear bottlenecks,” in CVPR, 2018.
[SAN18]
Evaluation on DVS Gestures
Architecture | Accuracy |
RG-CNN [BI20] | 97.20% |
TORE [BAL21] | 96.20% |
EvT [SAB22] | 96.20% |
Shufflenet v2 3D | 94.31% |
Mobilenet Transformer | 96.21% |
EfficientNet GRU | 97.34% |
TABLE I: Accuracy of models on DVS 128 Gestures
[BI20] Y. Bi, A. Chadha, A. Abbas, et al. , “Graph-based spatio-temporal feature learning for neuromorphic visión sensing,” IEEE Trans. on Image Processing, 2020.
[BAL21] R. Baldwin, R. Liu, M. Almatrafi, V. Asari, and K. Hirakawa, “Time-ordered recent event (tore) volumes for event cameras,” arXiv:2103.06108, 2021.
[SAB22] A. Sabater, L. Montesano, and A. C. Murillo, “Event transformer. asparse-aware solution for efficient event data processing,” CVPR, 2022.
Event-based Action Recognition
Time Surfaces
Inference
(DL network)
T = 0.25 s
Sponge squeeze 🡪 5%
Sponge flip 🡪 26.4%
Sponge wash 🡪 7.6%
Sponge wipe 🡪 9.5%
Sponge scratch 🡪 11%
T = 1 s
Sponge squeeze 🡪 4%
Sponge flip 🡪 44.9%
Sponge wash 🡪 9.2%
Sponge wipe 🡪 15.2%
Sponge scratch 🡪 10%
T = 2 s
Sponge squeeze 🡪 0%
Sponge flip 🡪 1.7%
Sponge wash 🡪 1.1%
Sponge wipe 🡪 92.8%
Sponge scratch 🡪 3%
Confidence
Confidence
Confidence
Inference with EfficientNet GRU
Video-based Action Recognition
Inference
(DL network)
T = 0.25 s
Sponge squeeze 🡪 5%
Sponge flip 🡪 3%
Sponge wash 🡪 3%
Sponge wipe 🡪 4.1%
Sponge scratch 🡪 4%
T = 1 s
Sponge squeeze 🡪 8%
Sponge flip 🡪 35.6%
Sponge wash 🡪 18.4%
Sponge wipe 🡪 21.6%
Sponge scratch 🡪 12%
T = 2 s
Sponge squeeze 🡪 6%
Sponge flip 🡪 22.8%
Sponge wash 🡪 26.1%
Sponge wipe 🡪 28.5%
Sponge scratch 🡪 15%
Confidence
Confidence
Confidence
Inference with EfficientNet GRU
Video-based Action Recognition
DL architecture | Video-based | Event-based |
Accuracy | Accuracy | |
Mobilenet Transformer | 90.02 ± 4.40 | 96.94 ± 2.50 |
EfficientNet GRU | 79.02 ± 7.07 | 97.63 ± 0.94 |
TABLE II: 4-Fold Cross Validation Manipulation Action classification
Event-based Action Recognition
T = 2 s
Sponge squeeze 🡪 0%
Sponge flip 🡪 1.7%
Sponge wash 🡪 1.1%
Sponge wipe 🡪 92.8%
Sponge scratch 🡪 3%
T = 2 s
Sponge squeeze 🡪 6%
Sponge flip 🡪 22.8%
Sponge wash 🡪 26.1%
Sponge wipe 🡪 28.5%
Sponge scratch 🡪 15%
Confidence
Confidence
t-SNE projection of features
a) Events
b) Videos
Conclusion