1 of 43

RoRD : Rotation Robust Descriptors and Orthographic Views for Local Feature Matching

Udit Singh Parihar

Accepted IROS 2021

2 of 43

Problem Statement

  • Developing local descriptors invariant to high viewpoints
  • Use in relative pose estimation
  • Front-end pipelines in SLAM system
  • Image Retrievals

3 of 43

Introduction

Feature Correspondences

RORD (ours)

SIFT

✔️

x

4 of 43

Contributions

  • Proposed deep learnt descriptors which works under extreme viewpoints
  • Used self-supervised training to generate huge training data with simple homographic transformations
  • Achieves good generalisation across datasets and perform state of art in variety of tasks
  • Proposed dataset comprising of images from high change in camera viewpoint
  • Dual headed architecture, where one head is rotation invariant and other is illumination invariant

5 of 43

Correspondences from RoRD

  • Results from extreme viewpoints
  • Works in both Indoor and outdoor

6 of 43

Improvements

  • MMA in HPatches Dataset
  • Pose estimation on DiverseView Dataset
  • Use of rotation robust descriptors in Visual Place Recognition Task

7 of 43

Supervised Training

  • Generates correspondences from monocular images by applying SfM
  • Data labelling doesn’t require human labelling
  • Inherit biases from SfM data pipeline which uses classical methods

SfM Data

8 of 43

Self Supervised Training

  • Training on primitive synthetic shapes and deploying on real world scenarios

Homographic Transformations

9 of 43

Orthographic Views

  • Intermediate representation of scene as orthographic views
  • Orthographic views improves image retrieval, feature matching and planning
  • Ability to calculates orthographic views during 6 DoF camera motion

10 of 43

Pipeline

  • Orthographic view rectifies the perspectivity
  • Planar patches extraction and surface normal calculation

Orthographic view

11 of 43

Pipeline

Ensemble Architecture

  • Dual headed architecture for illumination and rotation invariance
  • D2Net model is supervised using 3D depth and pose information
  • RoRD head is trained using self-supervised rotational homographies

12 of 43

Orthographic View Generation

  • Depth information along with desired ROI for 3D scene representation

  • Virtual Camera anti parallel to surface normal

13 of 43

Ortographic to Perspective Matching

  • Inverse homography to project correspondences from orthographic view to perspective view

14 of 43

IPM for Autonomous Driving

  • Challenge of matching images from front and rear camera
  • Discriminative road patches

15 of 43

VPR for Autonomous Driving

  • Image pairs with maximum inliers are considered a match

16 of 43

Network Architecture

  • VGG-16 common backbone architecture for better efficency
  • Last layer is fine-tune for D2-Net and RoRd
  • D2-Net head is train using SfM data while RoRD is trained using Homography data
  • Feature correspondences are calculated independently for each head
  • RANSAC based geometric verification for geometric verification

17 of 43

Loss Function

  • Related via rotation homography
  • Triplet margin loss, with margin
  • p(c) is euclidean distance between descriptors
  • Negative distance is calculated via hardest negative

18 of 43

Training Data

  • Training data from phototourim and Oxford Robot Car Dataset

Homographic Transformations

19 of 43

Results

20 of 43

Datasets and Tasks

  • MMA on Hpatches Dataset
  • VPR on Oxford RobotCar Dataset
  • Pose Estimation Error on DiverseView Dataset

21 of 43

HPatches Dataset

  • Extended Hpatches Dataset to scenes having high rotation
  • Evaluated for Mean matching accuracy, by using ground truth homography
  • D2 Net
  • RoRD

22 of 43

Qualitative Results (MMA)

  • Comparison on Standard and extended HPatches Dataset

23 of 43

Oxford RobotCar Dataset

  • D2-Net with Orthographic View
  • RoRD with Orthographic View

24 of 43

VPR Results with Front and Rear Camera

  • Testing VPR results on the sequence not seen during training

Video

25 of 43

Recall for VPR

  • Obtain twice the recall compared to second best on challenging Oxford RobotCar Dataset

26 of 43

DiverseView Dataset

  • Orthographic views in tandem with RoRD gives best correspondences

27 of 43

Pose Estimation Results

Indoor

Video

28 of 43

Pose Estimation Results

Outdoor

Video

29 of 43

Ablation Study for Pose Estimation

  • Rotation and Translation error

30 of 43

Opposite View Loop Closures in SLAM

  • Global feature matching for VPR and local feature matching for transformations

31 of 43

Transformation Estimation using Rotation Invariant Descriptors

32 of 43

Video Results of Loop Closure in Lab Dataset

33 of 43

Pose Graph Optimization on Lab Dataset

  • Extended RTABMAP to work with scenes where robot revisits the scene

34 of 43

Code and Dataset

  • Code and Dataset are publicly available
  • https://github.com/UditSinghParihar/RoRD

35 of 43

Topological Mapping for Manhattan-like Repetitive Environments

Sai Shubodh Puligilla *, Satyajit Tourani *, Tushar Vaidya *,

Udit Singh Parihar *, Ravi Kiran Sarvadevabhatla and K. Madhava Krishna

*Denotes authors with equal contribution

Accepted to International Conference on Robotics and Automation 2020

36 of 43

Problem Formulation

  1. Loop closure constraints are proposed by MLP and computed using ICP.
  2. Manhattan constraints are obtained from orthogonal relations between nodes in Manhattan graph.

37 of 43

Pose Graph Optimization Pipeline

38 of 43

RESULTS

  1. Nodes of the pose graph are labelled based on the topological labels by CNN to create a topological graph
  1. A highly distorted unoptimized pose graph with labelled topologies is recovered using an intermediate Manhattan Graph.
  2. A loop detection event between two Manhattan Nodes in the Manhattan Graph is shown in bold color. This results in both Manhattan and/or Loop constraints.

39 of 43

RECOVERED TRAJECTORIES

Top row shows unoptimized trajectories, middle row shows trajectories recovered using our pipeline and last row shows ground truth trajectories

40 of 43

Improving RTABMAP with Topological Mapping

41 of 43

Benchmarking RTABMAP

  • Evaluation of RTABMAP on warehouse dataset leads to detection of many False Positive loop closure constraints, due to repeating corridors.
  • Incorporation of Topological constraints leads to better ATE.
  • Our topological comparator utilizes geometric structure of the topological representation.

RTABMAP

RTABMAP + �Topological Constraints

4.45

3.36

42 of 43

References

  • RoRD - Rotation Robust Descriptors and Orthographic Views for Local Feature Matching, IEEE IROS 2021
  • Early Bird : Loop Closures from Opposing Viewpoints for Perceptually-Aliased Indoor Environments, VISAPP 2020
  • Topological Mapping for Manhattan-like Repetitive Environments, IEEE ICRA 2020

43 of 43

Thanks