1 of 51

CSE 5524: �3D

1

2 of 51

HW 4 & quizzes

  • HW 3
    • Caution: Please do NOT change anything outside those “your job” part
    • We have changed the GitHub readme to 20x20 to match the code.
    • Due: 4/17/2025

  • HW 4
    • Plan: A lighter homework
    • Due: 4/21/2025

  • Quizzes:
    • Two quizzes to be released this week --- True/False, multiple choices, unlimited tries

3 of 51

Final project presentation: 4/22 & 4/23

  • Presentation: each team 7~8 minutes
  • Detailed requirements will be given – please follow them carefully
  • Timer will be used – otherwise, we will spend a whole day

  • You are required to come to both days, but if you have a reason that you cannot come to one of the day, email me.

4 of 51

Today (44 & 45)

  • Recap
  • Structure from motion
  • Neural radiance filed

4

5 of 51

3D reconstruction

  • Estimate the 3D location of each pixel
  • Estimate the 3D geometry (and corresponding colors) that lead to the observed images

Depth estimation and 3D reconstruction

6 of 51

3D reconstruction

  • For a particular pixel (u, v)
  • Now we can read out its 3D location: X(u, v), Y(u, v), Z(u, v)

7 of 51

Stereo depth estimation

7

=

Il

Ir

D

Z

depth

Focal length Baseline

disparity

u

v

8 of 51

The issues with stereo vision?

  • Can we fully recover the 3D world?

    • May not be “robust”
    • Occlusion is still occlusion

9 of 51

Multi-view Geometry

  • Looking through multiple “cameras” or “eyes”

[Building Rome in a Day, ICCV 2009]

10 of 51

Some history …

  • By watching the motions of a set of points, not all the pixels, we can recognize actions.

  • We can perceive forms/shapes from motions.

  • This phenomenon from Gunnar Johansson (psychophysicist) in 1971 was later mathematically formulated as the problem of structure from motion (SFM) with “rigid” assumptions.

11 of 51

Some history …

  • 3D reconstruction of “rigid objects” by watching key point motions

12 of 51

Some history …

  • A rigid object’s motion can be interpreted as camera’s motion
  • A camera is moving around the object to capture multiple views of it

13 of 51

From single camera to multiple cameras

  • Single camera moving around a rigid object

  • Multiple cameras taking photos at certain poses (translation, rotations) around the rigid object

14 of 51

Structure from Motions (SFM)

  • Key elements:
    • Multiple “views” of the same object
    • Key points from each frame
    • Correspondence of those points: knowing that they are from the same 3D point

  • With this information, we want to
    • Reconstruct the “3D” locations of the key points and hopefully the whole object from 2D images

15 of 51

Questions?

15

16 of 51

Sparse SFM

  • Given a set of M images with overlapping content

  • We want to recover the 3D structure of the scene and the locations of the camera for each image

17 of 51

Problem formulation

  •  

rotation

translation

Intrinsic

 

 

 

 

Which are knowns? Which are unknowns?

18 of 51

Problem formulation

  •  

rotation

translation

Intrinsic

 

 

 

 

19 of 51

Quick glance of camera parameters

 

 

20 of 51

Quick glance of camera parameters

  • Homogeneous and heterogeneous coordinates

heterogeneous

homogeneous

21 of 51

Quick glance of camera parameters

  • Scaling equivalence properties

  • Illustration

heterogeneous

homogeneous

22 of 51

Quick glance of camera parameters

  • Intrinsic parameters

(X, Y, Z)

Matrix K

23 of 51

Quick glance of camera parameters

  • Extrinsic parameters

24 of 51

Quick glance of camera parameters

  • Put together (intrinsic + extrinsic)

25 of 51

Quick glance of camera parameters

  • Put together (intrinsic + extrinsic)

Camera

3D

Intrinsic

Rotation

Translation

Matrix “M

26 of 51

Questions?

26

27 of 51

Problem formulation

  •  

rotation

translation

Intrinsic

 

 

 

 

28 of 51

Simplified version

  • Images: We have 2NM “matched” knowns

  • Points: We have 3N unknowns

  • Cameras: We have M-1 unknown matrices (one as origin)
    • 3(M - 1) for rotation
    • 3(M - 1) -1 for translation
    • 3~4 M for intrinsic

29 of 51

Reprojection error for optimization

  • High-levelly:

    • From 2D knowns, we estimate the unknowns (3D locations + camera parameters)
    • From the estimated unknowns, we can reproject them to 2D
    • We can then compare the “unknown” 2D vs. “estimated” 2D

30 of 51

Reprojection error for optimization

  • High-levelly:

    • From 2D knowns, we estimate the unknowns (3D locations + camera parameters)
    • From the estimated unknowns, we can reproject them to 2D
    • We can then compare the “unknown” 2D vs. “estimated” 2D

31 of 51

Reprojection error for optimization

  • High-levelly:

    • From 2D knowns, we estimate the unknowns (3D locations + camera parameters)
    • From the estimated unknowns, we can reproject them to 2D
    • We can then compare the “unknown” 2D vs. “estimated” 2D

32 of 51

Reprojection error for optimization

  •  

33 of 51

Questions?

33

34 of 51

What is missing beyond the loss function?

  • How do we detect key points and correspond them across images?

  • Interest point detection:
    • Stable under different sets of geometric transformations and illumination changes
    • Stable in the sense of “detectability”, “distinctiveness”, “feature invariance”

35 of 51

Interest point detection and matching

Traditional: SIFT (Scale-invariant feature transform)

36 of 51

Interest point detection and matching

Recent: neural networks

37 of 51

Interest point detection and matching

Points surrounded by image regions with strong spatial changes

Likely will be “re-found” from another views

38 of 51

Interest point detection and matching

  • Need “discriminative” local image features

39 of 51

Interest point detection and matching

  • Need “discriminative” local image features

40 of 51

SFM results

41 of 51

SFM results

42 of 51

SFM: The devil is in the details

  • How to optimize and iteratively correct mistakes …

  • Check section 44.3.4

43 of 51

Today (44 & 45)

  • Recap
  • Structure from motion
  • Neural radiance filed

43

44 of 51

What if we have more images?

45 of 51

What if we have more images?

46 of 51

Can we synthesize images from other views?

47 of 51

Ray

48 of 51

Pixel value = integration along the ray

49 of 51

Pixel value = integration along the ray

  • Volume rendering

r: ray

t: location on the ray

What don’t we know?

50 of 51

Goal

51 of 51

Neural Radiance Fields (NeRF)

  • From multiple images, use reprojection loss to learn the neural network