1 of 21

1

Real-time 3D Reconstruction

for Robot Manipulation

Team F14: Rajath Aralikatti, Santhoshini Gongidi

Collaborators:

Bardenious Duisterhof

Jeffrey Ke

Advisor:

Prof. Jeff Ichnowski

2 of 21

Need for 3D Understanding in Manipulation

��Depth and shape awareness

2

Full Scene Understanding

Optimal viewpoints can improve robot policies

3 of 21

Background: 3D reconstruction

3

Camera Calibration +

Sparse Reconstruction

Set of 2D Images

Dense Depth Maps,�Dense 3D Pointclouds

4 of 21

SfM pipeline in COLMAP

  • What’s the problem with a pipelined approach ?
    • Error propagation
      • Bad matches = Bad reconstruction
    • Slow
    • Matching is inherently a 3D problem

4

5 of 21

5

Dust3R:

Unified Model for 3D Vision

Release Date: CVPR 2024

6 of 21

Dust3r

  • Pointmap representation
    • Mapping from each pixel to its corresponding 3D point (X, Y, Z)
    • Image Pixel (H, W, 3) -> 3D Point (H, W, 3)

6

Wang, Shuzhe, et al. "Dust3r: Geometric 3d vision made easy."

7 of 21

Dust3r

  • Using dense 2D-3D correspondences other 3D geometry information can be inferred

7

Wang, Shuzhe, et al. "Dust3r: Geometric 3d vision made easy."

8 of 21

Dust3r

8

Wang, Shuzhe, et al. "Dust3r: Geometric 3d vision made easy."

9 of 21

Dust3r

  • Works on single pair of images
  • What about a collection of images?
    • Solution: Global Alignment

~20 FPS (0.05s)

~0.5 FPS (2s) - 8 views

9

Wang, Shuzhe, et al. "Dust3r: Geometric 3d vision made easy." CVPR 2024

10 of 21

10

Fast3r:

No costly N-view alignment

Coherent multi-view inference

Release Date: CVPR 2025

11 of 21

Fast3r: Multiview Inference

11

Loss:

Yang, Jianing, et al. "Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass."

12 of 21

Fast3r: Trends

  • Global attention removes sequential dependencies
    • Reduces Error Accumulation
  • Local heads lead to preserving details
    • Invariant to training, unlike global heads
  • Extremely engineered to work for 1000 views in a single pass!
    • FlashAttention, Data Sharding, Model Sharding

12

Yang, Jianing, et al. "Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass."

13 of 21

13

VGGT:

State-of-the-art performance with

camera parameters + point tracking

Release Date: CVPR 2025

14 of 21

VGGT: Multiview + Multitask

14

Wang, Jianyuan, et al. "VGGT: Visual Geometry Grounded Transformer."

Rotation(4), �Translation(3)

FoV(3)

Depth Map D

(HxW)

Point Map X

(HxWx3)

Tracking

Features F

(HxWxC)

Loss:

15 of 21

VGGT: Trends

  • Redundant 3D task-loss can improve performance.
  • Camera prediction task -> Most significant gains
  • VGGT’s frame attention <-> Fast3r’s Local head
  • Predict canonical points maps with ground-truth normalization

15

Wang, Jianyuan, et al. "VGGT: Visual Geometry Grounded Transformer."

16 of 21

Summary: 3D Geometric Models

Unified 3D-model recipe:

  • Over-parameterized representations(point maps) -> dense prediction task
  • Global module -> MV Triangulation
  • Local/View-specific module -> Details preserved
  • A large NN with minimal inductive biases.

16

17 of 21

Comparison of 3D Geometric Models

17

COLMAP

Dust3r

Fast3r

VGGT

Pose Est. Accuracy ()

45.2

67.7

72.7

85.3

MVS Inference

Time (↓)

∼ 15s

∼ 7s

∼ 0.2s

∼ 0.2s

Absolute Depth

Recovery

Possible with

  • LiDAR alignment
  • known cameras
  • Know true size of an object

Temporal

Consistency

(Empirical)

-

Good

Bad

Better

18 of 21

18

Our Goal:

Learn better 3D Scene Representation

for Robot Manipulation

19 of 21

3D Geometric Model Tuning for Robotics

  • Literature: Simple Point-MLP encoders are better
  • Goal: Study and enable these pretrained 3D geometric representations to work for robot manipulation

19

VGGT

Adaptor

Manipulation policy

Next Robot State

Current State

20 of 21

Open problems: 3D Geometric Models

20

  • Edge bleeding -> Incorrect grasp positions
  • Bad reference frame -> Occluded views results in bad reconstruction.
  • Trained on perspective images -> Requires correction
  • Trained on opaque object scenes -> Deformable objects ?

21 of 21

21

Thank You!