1 of 27

Generative AI:

A New World Simulator to Transform Synthetic Data Generation

06.28.2024

Lily (Xianling) Zhang

2 of 27

Agenda

Massive Data Need in AI Era Across Industries

01

02

03

04

Progression in AI Models' Data Generation Capabilities

Generative World Models as Learned Simulators

Discussion: Interactive Q & A

Presentation Overview

3 of 27

Massive Data Need in AI Era Across Industries

01

4 of 27

Businesses Embracing AI Have Massive Data Needs

Automotive & Robotics

Healthcare

Banking & Finance

Agriculture

Sales & eCommerce

AI is everywhere. Data is the key to scale.

5 of 27

Progression in AI Models' Data Generation Capabilities

02

6 of 27

Growing Data Challenges for Level 1 to 5 Autonomy

From Level 1 to Level 5 autonomy, the data complexity and volume requirements increase significantly across different Operational Design Domains (ODD). More sophisticated AI-based data generation methods are needed.

Source: SAE, Medium

Eyes-off and hands-off driving in certain conditions

7 of 27

Progression in AI Models' Data Generation Capabilities

Level 5

Level 4

Level 2

3D Scene Generation

Level 3

Compositional Generation

LLM Guided Behavior Generation

Agent-Driven Generation

Level 1

2D Style Transfer

Growing Intelligence

The thinner the level is, the less explored it is.

8 of 27

Level 0: Gaming Engine Based Data Rendering

Traditional gaming engine based rendering

Limitation: It requires intensive labor and graphics expertise for 3D asset modeling in Maya and Blender.

Pillar

Agent

Drivable

Area

Waypoints

Limitation: Hand-crafted waypoints are required in the simulator for data generation along a moving trajectory.

9 of 27

Level 1

2D Style Transfer

Level 1 AI models focus on style transfer (image or video) for generating diverse visual features and patterns.

  • [NeurIPS 2017] UNIT, Liu, Ming-Yu et al.
  • [CVPR 2017] Pix2pix, Isola, Phillip, et al.
  • [ICCV 2017] CycleGAN, Zhu, Jun-Yan, et al.
  • [CVPR 2019] StyleGAN, Karras, Tero, et al.

10 of 27

Level 1 (2D): Image to Image Transfer

N Jaipuria, X Zhang, R Bhasin., et al. "Deflating dataset bias." CVPR-W. 2020.

Data augmentation using synthetic images improved performance in parking slot detection, highway lane segmentation, and monodepth estimation..

Sim to Real GAN Translated

Day to Night GAN Translated

Sunny to Cloud GAN Translated

11 of 27

Level 2

3D Scene Generation

Level 2 AI models enable learning-based 3D scene reconstruction and the generation of data with geometry awareness.

  • [ECCV 2020] NeRF, Mildenhall, Ben, et al.
  • [ICCV 2021] WorldSheet, Hu, Ronghang, et al.
  • [CVPR 2022] SIMBAR, Zhang, Xianling, et al.
  • [SIGGRAPH 2023] 3D Gaussian Splatting
  • for Real-Time Radiance Field Rendering. Kerbl, Bernhard, et al

Level 1

Style Transfer

12 of 27

Context: COLMAP – SFM and MVS

3D Scene Reconstruction using KITTI

Restrictions

  • 11 out of 21 KITTI sequences have insufficient feature matching through different viewpoints, which can result in failed mesh reconstruction.
  • The SFM+MVS-based COLMAP approach requires significant computing power and time, taking 5-10 hours for one sequence.

13 of 27

Level 2: Generate 3D Scene from 2D Images

NeRF, Mildenhall et al. ECCV 2020

It implicitly learns the 3D geometry of the scene as a side effect of its primary task of color prediction. Need to retrain per scene.

Worldsheet, Hu, R., et al. ICCV 2021. It has explicit 3D mesh prediction that can generalize across different scenes.

3D Gaussian Splatting, Kerbl, B, et al.

SIGGRAPH 2023.

It is a hybrid method that uses explicit Gaussian parameters to implicitly represent a 3D scene.

14 of 27

Level 2: Generalizable Scene Relighting with 3D Scene Reconstruction

SIMBAR In The Wild: Single Image-based Scene Relighting

Confidential

Novel Relit Results from Racing Car Field

Input Frame

3D Scene Mesh

Input Frame

Relit Images for 4 different sun angles

X Zhang, N Tseng, A Syed, et al. “SIMBAR”, CVPR. 2022.

15 of 27

Level 2

3D Scene

Level 3

Compositional Generation

Level 3 AI models transition from single entity or scene modeling to compositional actor and background modeling.

  • [CVPR 2021] NSG, Ost, Julian, et al.
  • [CVPR 2022] Block-NeRF, Tancik, Matthew, et al.
  • [Arxiv 2024] Street Gaussians, Yan, Yunzhi, et al.
  • [CVPR 2024] Driving Gaussian, Zhou, Xiaoyu, et al.
  • [WACV-W 2024] LANe. Krishnan, Akshay, et al.

Level 1

Style Transfer

16 of 27

Level 3: Driving Scene Generation

[CVPR 2022] Block-NeRF, Tancik, Matthew, et al.

It builds a grid of Block-NeRFs from 2.8 million images to create the largest neural scene representation to date and is capable of rendering an entire neighborhood in San Francisco.

[Arxiv 2024] Street Gaussian, Tancik, Matthew, et al.

It models dynamic urban street scenes from monocular videos. The dynamic urban street is represented as a set of point clouds equipped with semantic logits and 3D Gaussians.

17 of 27

Level 3: Lighting-aware Compositional Generation

A Krishnan, A Raj, X Zhang, et al. “LANe”, WACV-W. 2024.

18 of 27

Level 3: Lighting-aware Composition on Real Data

Lighting unaware

Lighting aware

Lighting unaware

Lighting aware

A Krishnan, A Raj, X Zhang, et al. “LANe”, WACV-W. 2024.

19 of 27

Generative World Model as Learned Simulator

03

20 of 27

Level 4

Level 2

3D Scene

Level 3

Compositionality

LLM Guided Behavior Generation

Level 4 AI models provide enhanced behavior and granular control through language-guided generation of novel scenes, reducing the manual effort required to design actor trajectory and environment settings (lighting, time of day etc.).

  • [Arxiv 2024] DriveDreamer-2. Zhao, Guosheng, et al.
  • [ICLR 2024] MagicDrive. Gao, Ruiyuan, et al.
  • [ICRA 2024] SceneControl. Lu, Jack, et al.

Technical Leap

  • Level 3 methods augment existing real-world data with hand crafted trajectories.
  • Level 4 models takes language as prompt to control actions in the generation process and can generalize to unseen environments.

Level 1

Style Transfer

21 of 27

Level 4: Language Guided Generation using MagicDrive

MagicDrive: Latent interpolation with the same geometric conditions

22 of 27

Level 4: Language Guided Behavior Generation using DriveDreamer

"Rainy day, a person crosses the road in the front of the ego-car."

"Rainy day, the windshield wipers of the truck are continuously clearing the windshield."

"Rainy day / at night, ego-car drives through urban street, surrounded by a flow of vehicles on both sides."

23 of 27

Level 5

Level 4: Behavior

Level 2:

3D Scene

Level 3:

Compositionality

Agent-Driven Generation

Level 5 AI models feature an intelligent agent system capable of autonomously creating and designing virtual worlds with minimal human intervention.

Examples that are important steps towards Level 5 Generation:

  • [ICLR 2024] Copilot4D: Learning Unsupervised World Models for Autonomous Driving via Discrete Diffusion
  • [ICML 2024] LEO: An Embodied Generalist Agent in 3D World
  • SORA: Video generation models as world simulators

Technical Leap

  • Level 4 methods involve human-designed virtual worlds and object trajectories for ego-perspective data capture.
  • Level 5 models require autonomous agent-driven data generation, with the agent possessing perceptive and decision-making capabilities to handle all virtual driving situations independently.

Level 1:

Style Transfer

24 of 27

Level 5: Can Sora be A World Simulator ?

Input

“Change the setting to 1920s”

“Make it go under water”

“Put the video in space with rainbow road”

“Make it have dinosaurs”

“Rewrite the video in pixel art style.”

25 of 27

Level 5: Agent Driven World Generation

SORA still falls short in

  • Adherence to physical laws in complex environment simulation
  • Agent-controlled virtual world generation and exploration
  • Dynamic adaptation with generalization capabilities
  • Future prediction based on past motion

Input: HD Map

Controllable Traffic Generation

Temporally and physically realistic world generation

Traffic Agent

Rendering Agent

26 of 27

List of Contributors

  • Nikita Jaipuria
  • Vidya N. Murali
  • Rohan Bhasin
  • Jinesh Jain
  • Mayar Arafa
  • Punarjay Chakravarty
  • Shubham Shrivastava
  • Sagar Manglani
  • Nathan Tseng
  • Ameerah Syed
  • Alexandra Carlson
  • Akshay Krishnan (GaTech)
  • Amit Raj (GaTech)
  • James Hays (GaTech)

Acknowledgement

27 of 27

Agents-Driven Generation

Lily Xianling Zhang

Technical Project Lead

Machine Learning Research Scientist

Ford Autonomy, Perception & Prediction

Ford Greenfield Labs, Palo Alto, California email: xzhan258@ford.com

04: Interactive Q & A

Level 5

Level 4

Level 2

3D Scene Generation

Level 3

Compositional Generation

LLM Guided Behavior Generation

Level 1

2D Style Transfer