1 of 25

TextMesh: Generation of Realistic 3D Meshes From Text Prompts

Christina Tsalicoglou, Fabian Manhardt, Alessio Tonioni, Michael Niemeyer, Federico Tombari

ETH Zurich, Google, Technical University of Munich

2023/05/29 陸品潔

2 of 25

Motivation

01

Contents

Contributions

03

Related Work

02

Method

04

Conclusion

06

Experiments

05

3 of 25

Motivation

01

4 of 25

Text-to-image

    • Has recently made huge progress in terms of speed and quality, thanks to the advent of image diffusion models.

5 of 25

Text-to-3D

    • Generate neural radiance fields (NeRFs), making them impractical for most real applications

    • Tend to produce over-saturated models, giving the output a cartoonish looking effect.

6 of 25

Related Work

02

7 of 25

3D Reconstruction with Neural Fields

    • NeRFs :
      • Enabled view synthesis from only image input via volume rendering

    • Recent Methods:
      • Surface + Volume rendering

    • Our goal:
      • Optimize a high-quality mesh and texture from text input

8 of 25

Signed Distance Fields (SDF)

9 of 25

Score Distillation Sampling (SDS)

10 of 25

Contributions

03

11 of 25

Contributions

    • Modify DreamFusion to model radiance in the form of SDF to tailor the model towards mesh extraction.

    • Propose a novel multi-view consistent and mesh conditioned re-texturing, enabling the generation of photorealistic 3D mesh models.

    • Obtained meshes are geometrically of high quality and showcase more natural textures than the current methods.

12 of 25

Method

04

13 of 25

Schematic Overview

14 of 25

Initial Scene Representation

    • Neural Radiance Fields:
      • Ok

    • Signed Distance Fields:
      • ok
      • Ok
      • SDF-based neural field representation can be rendered to the image plane using the same volume rendering technique.

 

15 of 25

Text-to-3D via Score-based Distillation

    • Imagen:
      • ok
      • Ok

    • Difference from DreamFusion:
      • Sample the whole elevation range for the camera, to avoid bleeding artifacts at the model bottom.

16 of 25

Text-to-3D via Score-based Distillation

    • Initial mesh:
      • Extracted from the SDF as the surface at the zero-level set using Marching Cubes.

    • Drawbacks:
      • Miss high frequency details.
      • Show over-saturated (’cartoonish’)

17 of 25

Photorealistic Texturing Using Multi-View Consistent Diffusion

    • Render from 4 viewpoints:
      • Ok
      • Ok

    • Drawbacks:
      • Still exhibit minor misalignment at their intersection as well as on unobserved object parts.

18 of 25

Photorealistic Texturing Using Multi-View Consistent Diffusion

    • Second optimization stage:
      • Ok

      • Use very small guidance weight.

19 of 25

20 of 25

Experiments

05

21 of 25

Qualitative Comparison

22 of 25

Extracting Meshes

23 of 25

3D Consistency

24 of 25

Conclusion

06

25 of 25

    • A novel approach for text-to-3D mesh.

    • Represent the geometry as a distance field which is optimized using SDS.

    • Supervise the texture refinement with a photometric loss on enhanced 2D mesh renderings generated by a depth condition image-to-image diffusion model.

    • Rely on SDS to smooth out transitions within multi-view supervision.