1 of 34

1

2 of 34

Lifelong Spatial Understanding: Dense Monocular Reconstruction and Hierarchical Scene Graphs for Multi-Session Mapping

Christina Kassab

PhD Candidate

Haedam Oh

PhD Candidate

Maurice Fallon

PI

3 of 34

Dynamic Robot Systems Group

OXFORD ROBOTICS INSTITUTE

INTRODUCTION

LIFELONG SPATIAL UNDERSTANDING

Our Key Research Themes

  • SLAM
  • Semantic Mapping
  • Navigation
  • Reconstruction
  • Change Detection
  • Autonomous Inspection
  • Mapping in Forests (Digiforest EU Project)

3

4 of 34

Lifelong Spatial Understanding

Ground Truth Changes in the 3RScan Dataset�Can we detect these changes automatically?

INTRODUCTION

LIFELONG SPATIAL UNDERSTANDING

Motivation

  • Agents must operate persistently as environments change
  • Static maps degrade – autonomy requires representations that can adapt

Two Main Questions:

  • How can we build accurate maps with minimal sensing?
  • How can we make maps that are robust to environment changes?

5 of 34

Building Accurate Maps with Minimal Sensing

LIFELONG SPATIAL UNDERSTANDING

6 of 34

Feed-forward 3D Reconstruction

Example Output Reconstruction from DUSt3R

MONOCULAR RECONSTRUCTION

LIFELONG SPATIAL UNDERSTANDING

  • Predict poses, depth and dense 3D structure directly from multi-view RGB images
  • Single forward pass

  • How to scale to larger scenes?
    • Fixed number of input views
    • More views → higher memory and compute cost
    • Combine submaps using SLAM
      • MASt3R-SLAM
      • VGGT-SLAM

7 of 34

Feed-forward 3D Reconstruction

Example Output Reconstruction from DUSt3R

MONOCULAR RECONSTRUCTION

LIFELONG SPATIAL UNDERSTANDING

Our Two Main Approaches:

  1. LEXI-SG
    1. Monocular SLAM system + scene graph
    2. Using semantics to inform batch selection

  • ScaRF-SLAM
    • Leverages poses from classical visual SLAM to improve reconstruction quality

8 of 34

LEXI-SG

Monocular 3D Scene Graph Mapping with Room-Guided FeedForward Reconstruction

MONOCULAR RECONSTRUCTION

LIFELONG SPATIAL UNDERSTANDING

Key Concept

Semantic structure (such as rooms) can help guide geometric reconstruction using feed-forward models.

Contributions

  • First dense monocular SLAM system that builds an open-vocabulary 3D scene graph from RGB images alone
  • A vision-only room identification method using DINO
  • A room-based reconstruction and global alignment
  • An open-vocabulary segmentation module that lifts 2D mask tracklets into the scene graph

Christina Kassab

Hyeonjae Gil

(SNU)

ROOMS

TRAJECTORY

OBJECTS

RECONSTRUCTION

dynamic.robots.ox.ac.uk/projects/lexisg/

8

9 of 34

LEXI-SG: System Overview

LIFELONG SPATIAL UNDERSTANDING

MONOCULAR RECONSTRUCTION

Christina Kassab

Hyeonjae Gil

(SNU)

dynamic.robots.ox.ac.uk/projects/lexisg/

9

Room transition detected

Subsample previous batch

Room-Based Reconstruction

Object Segmentation & Tracking

Loop Closure

Optimisation

Transition Edge Estimation

Previous Batch

Current Batch

Transition Pairs

Transition Edges

Depth

&

Poses

Loop Closure Found?

Merge Batches & Recalculate Edges

2D Object Tracking

3D Overlap Check

Per-view features

Merge

Per-object feature

DINO

DINO

MapA

Room A

Room B

Room C

10 of 34

LEXI-SG: Results

Ours

MASt3R-SLAM

VGGT-SLAM

ViSTA-SLAM

Data recorded in an office environment using Aria Gen 1

LIFELONG SPATIAL UNDERSTANDING

MONOCULAR RECONSTRUCTION

Christina Kassab

Hyeonjae Gil

(SNU)

dynamic.robots.ox.ac.uk/projects/lexisg/

10

11 of 34

LEXI-SG: Results

LIFELONG SPATIAL UNDERSTANDING

MONOCULAR RECONSTRUCTION

Christina Kassab

Hyeonjae Gil

(SNU)

dynamic.robots.ox.ac.uk/projects/lexisg/

11

12 of 34

ScaRF-SLAM

Scale-Consistent Reconstruction with Feed-Forward Models and Classical Visual SLAM

MONOCULAR RECONSTRUCTION

LIFELONG SPATIAL UNDERSTANDING

Key Concept

Classical visual SLAM poses can anchor and scale-correct predictions from feed-forward geometric foundation models

Contributions

  • A practical decoupled GFM-based mapping framework that consistently integrates with existing feature-based SLAM systems
  • Efficient scale correction and point cloud fusion for globally consistent reconstruction

Yuhao

Zhang

dynamic.robots.ox.ac.uk/projects/scarf-slam/

12

13 of 34

ScaRF-SLAM: System Overview

LIFELONG SPATIAL UNDERSTANDING

MONOCULAR RECONSTRUCTION

Yuhao

Zhang

dynamic.robots.ox.ac.uk/projects/scarf-slam/

13

Classical vSLAM

(ORB-SLAM or OpenVINS)

Geometric Feed-Fwd Model

(MapAnything, DepthAnything)

Online Map Fusion

Loop closures and poses

Dense point clouds

Images

plus IMU

calibration

calibration

Point

cloud map

14 of 34

ScaRF-SLAM: Results

LIFELONG SPATIAL UNDERSTANDING

MONOCULAR RECONSTRUCTION

Yuhao

Zhang

dynamic.robots.ox.ac.uk/projects/scarf-slam/

14

15 of 34

Robust Mapping in the Presence of Long-Term Semantic Change

LIFELONG SPATIAL UNDERSTANDING

16 of 34

SG Matching

3D Scene Graph Matching and Updating through learned graph matching

Scene Graph 1

Scene Graph 2

table

chair

computer

mouse

table

chair

book

MULTI-SESSION MAPPING

LIFELONG SPATIAL UNDERSTANDING

Key Concept

We extend 3D scene graph frameworks to long-term dynamic environments by updating observations via learned graph matching

Contributions

  • An online hierarchical 3D scene graph pipeline for multi-session scene understanding
  • A multi-modal cross-session scene graph alignment network to establish object correspondences across sessions
  • An ambiguity-aware object correspondence mechanism that resolves whether visually identical objects across sessions represent the same moved object or different object instances
  • Room-identification mechanism that combines visual similarity and scene graph alignment

Mengyuan Yin

16

17 of 34

Early Results

Object-to-object matching on 3RScan

* Indicates privileged baselines

LIFELONG SPATIAL UNDERSTANDING

MULTI-SESSION MAPPING

Mengyuan Yin

17

SGAligner [1]

Full Scene Graph

SGReg [2]

ROMAN [3]

Ours

SGAligner*

SGReg*

ROMAN*

[1] Sarkar et al. SGAligner: 3D Scene Alignment with Scene Graphs. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023.

[2] Liu et al. SG-Reg: Generlizable and Efficient Scene Graph Registration. In IEEE Transactions on Robotics, 2025.

[3] Peterson et al. ROMAN: Open-set Object Map Alignment for Robust View-Invariant Global Localization. In Robotics: Scene and Systems. 2025.

18 of 34

OASIS-Map

Object-Level Change Detection in Multi-Session Mapping using Semantic Correspondence Matching

MULTI-SESSION MAPPING

LIFELONG SPATIAL UNDERSTANDING

Key Concept

A unified spatio-temporal object map that tracks environmental changes across sessions using dense semantic correspondences.

Contributions

  • A unified object-level mapping framework integrating instances, semantics, and 3D geometry for change-aware mapping across sessions
  • A dense semantic correspondence module using DINOv3 features for robust matching
  • Evaluation on indoor and outdoor benchmarks showing robust performance across diverse scenarios.

Haedam OH

18

3D change map

Static

Appear

Previous session

Current session

RGB

Spatio-Temporally Consistent Map

Change

RGB

Change

19 of 34

OASIS-Map: System Overview

LIFELONG SPATIAL UNDERSTANDING

MULTI-SESSION MAPPING

Haedam OH

19

Step 1: Front-end Session Processing

Step 2: Back-end Session Comparison

Geometric Reconstruction

Object Detection & Tracking

Object association &

Change detection

Previous Session

Geometric map

Object map

RGB images

RGB images

Map & Poses

Depth / LiDAR / Learned Depth

Semantic Correspondences

Poses

Front-end

Back-end

Current Session

Input

Patch-patch

Mask-Mask

RGB images

Depth / LiDAR / Learned Depth

Poses

Input

20 of 34

OASIS-Map: Results

Haedam OH

LIFELONG SPATIAL UNDERSTANDING

MULTI-SESSION MAPPING

20

Market

Car park

Previous image

Geometric change

Semantic change

Current image

Static

Disappeared

Appeared

LT Mapper

Ours

Concept-Graphs

Where’s-my-glasses

Ground Truth

Session t1

Session t0

2D Change Detection (Left) in real world scenarios.

  • Market: entire market installed overnight at a previously empty location.

  • Car park: empty car park in the morning, occupied by cars in the afternoon.

3D Change Detection (Right): Car park

  • Different cars occupy the same spatial location, making the association geometrically ambiguous.

21 of 34

Ellison Institute of Technology Site

Construction site monitoring

MULTI-SESSION MAPPING

Multi-session mapping of an active construction site collected across six months using LiDAR, a 360° camera, and Aria Gen 1.

  • Multi-session Mapping (Current) – Build temporally consistent 3D maps by registering LiDAR scans from multiple recording sessions.
  • Change Detection (Current) – Identify structural changes as the construction site evolves.
  • High-Level Scene Understanding (Next steps) – Interpret construction progress by recognizing new structures, semantic changes, and overall site evolution.

LIFELONG SPATIAL UNDERSTANDING

21

22 of 34

Multi-Session Recordings

: structural changes

LIFELONG SPATIAL UNDERSTANDING

MULTI-SESSION MAPPING

*LiDAR maps are shown for visualizations.

22

Dec 25

Dec 25

Feb 26

Jan 26

Combined

23 of 34

Ellison Institute of Technology Site

December

March

  • Two 3D maps reconstructed via SLAM and monocular depth (Depth Anything V3) capture real-world construction site changes over time.
  • Scene changes include roof coverage in March and scaffold removal between visits.
  • We work towards visually understanding spatio-temporal changes beyond geometry, exploiting semantics to reason about the nature of change

MULTI-SESSION MAPPING

LIFELONG SPATIAL UNDERSTANDING

23

Aria Results from Site 1

24 of 34

Ellison Institute of Technology Site

December

January

MULTI-SESSION MAPPING

LIFELONG SPATIAL UNDERSTANDING

24

Aria Results from Site 1

25 of 34

Thank you

dynamic.robots.ox.ac.uk

LIFELONG SPATIAL UNDERSTANDING

25

26 of 34

Instructions:

  1. Use the following template to complete your slide deck by end of day Friday, June 26th.
  2. Please share the deck and viewer access to all assets with Maddy Bowen (Maddybowen@meta.com) and Ben Porter (benjport).
  3. We will add your slides to a main presentation deck. If you need to make last minute adjustments to your slides please email Maddy or Ben directly.

Title slides will be taken care of by Meta and shared prior to summit.

PRESENTATION TITLE ALL CAPS

NAME OF SECTION

27 of 34

Headline + Text

SUBTITLE HERE IN ALL CAPS

NAME OF SECTION

PRESENTATION TITLE ALL CAPS

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Phasellus nec ligula odio. Praesent dolor nunc, mollis eu volutpat eu, tincidunt in diam. Fusce mollis, ante ut elementum consectetur, ante massa congue libero, nec mattis erat dolor eget enim. Nam tincidunt, arcu eu cursus tincidunt, enim sapien facilisis erat, vel lobortis lectus augue a massa. Sed imperdiet dictum leo, id luctus diam vulputate non. In hac habitasse platea dictumst. Nullam dapibus eget ligula at laoreet.

  • Praesent dolor nunc, mollis eu volutpat eu, tincidunt in diam. Fusce mollis, ante ut elementum consectetur
  • Praesent dolor nunc, mollis eu volutpat eu, tincidunt in diam. Fusce mollis, ante ut elementum consectetur
  • Praesent dolor nunc, mollis eu volutpat eu, tincidunt in diam. Fusce mollis, ante ut elementum consectetur

27

28 of 34

This is a divider slide. �Add as a point of emphasis.

PRESENTATION TITLE ALL CAPS

NAME OF SECTION

29 of 34

PRESENTATION TITLE ALL CAPS

Ego-centric data collection for long duration recordings in space and time

Name Here

NAME OF SECTION

New York University, XXX Department

30 of 34

PRESENTATION TITLE ALL CAPS

Ego-centric data collection for long duration recordings in space and time

Name Here

Firstname Lastname�Title/role

Firstname Lastname�Title/role

Firstname Lastname�Title/role

New York University, XXX Department

NAME OF SECTION

31 of 34

SUBTITLE HERE IN ALL CAPS

SUBTITLE HERE ALSO ALL CAPS

NAME OF SECTION

PRESENTATION TITLE ALL CAPS

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Phasellus nec ligula odio. Praesent dolor nunc, mollis �eu volutpat eu, tincidunt in diam. Fusce mollis, ante ut

Lorem ipsum dolor sit amet, consectetur adipiscing elit. Phasellus nec ligula odio. Praesent dolor nunc, mollis �eu volutpat eu, tincidunt in diam. Fusce mollis, ante ut

31

32 of 34

Headline

Lorem ipsum dolor sit amet, consectetur adipiscing elit.

Lorem ipsum dolor sit amet, consectetur adipiscing elit.

Lorem ipsum dolor sit amet, consectetur adipiscing elit.

Subtitle

NAME OF SECTION

PRESENTATION TITLE ALL CAPS

32

33 of 34

Thank you

Contact Details

FUSCE MOLLIS, ELEMENTUM CONSECTETUR, ANTE MASSA �WWW.WEB.COM

PRESENTATION TITLE ALL CAPS

NAME OF SECTION

33

34 of 34

Sources Slide

SUBTITLE HERE CAPS

NAME OF SECTION

PRESENTATION TITLE ALL CAPS

1. Fusce mollis, elementum consectetur, ante massa �www.source.com

2. Fusce mollis, elementum consectetur, ante massa �www.source.com

3. Fusce mollis, elementum consectetur, ante massa �www.source.com

4. Fusce mollis, elementum consectetur, ante massa �www.source.com

5. Fusce mollis, elementum consectetur, ante massa �www.source.com

6. Fusce mollis, elementum consectetur, ante massa �www.source.com

7. Fusce mollis, elementum consectetur, ante massa �www.source.com

8. Fusce mollis, elementum consectetur, ante massa �www.source.com

34