1 of 1

CHARM3R: Towards Unseen Camera Height Robust Monocular 3D DetectorAbhinav Kumar1, Yuliang Guo2, Zhihao Zhang1, Xinyu Huang2, Liu Ren2, Xiaoming Liu1

1Michigan State University (MSU), 2Bosch Center for AI, Bosch Research North America

Detection at Unseen Ego Height is Hard

Support

Code

Regressed Depth

Depth Extrapolation Behavior

  • Neural Networks struggle with out of domain data [3].
  • Ego height changes induce projective transformations which CNNs and ViTs do not handle [2,4].
  • Projective equivariant backbones inapplicable since they handle one projective transformation.
  • Projective transformations do not interpolate linearly.
  • Disentagled learning infeasible from single height data.

Why is this problem hard?

CARLA Results

Conclusion

CHARM3R Pipeline

  • Mathematically prove and empirically observe contrasting trend in regression and ground-based depth estimates for Mono3D.
  • Average these estimate to generalize detectors to unseen heights

[1] Lu et al, Geometry uncertainty projection network for Mono3D, ICCV 21

[2] Kumar et al, DEVIANT: Depth Equivariant Network, ECCV 22

[3] Xu et al, How neural networks extrapolate, ICLR 21

[4] Sarkar et al, Shadows do not lie and lines don’t bend, CVPR 24

References:

  • These depths have contrasting extrapolation behavior.

1

2

3

4

5

6

Ground Depth

Comparison with Augmentation

  • Unseen ego camera heights drops performance by 40 AP points.
  • CHARM3R improves AP by 45% at unseen heights.
  • Augmentation fails at unseen heights.

Main Results

  • Train on Single Car Height Data
  • Test on Bots, Cars and Truck Height Data
  • Average both these depth estimates within the network.