1 of 25

Driving in Bad Weather

Presenter: Zi Wang, Peipei Zhong

Advisor: Srinivasa Narasimhan

MSCV Capstone 2024 Fall

1

1

1

1

2 of 25

Outline

Part 1: Driving in Bad Weather

  • Motivation and Problem Setting
  • Progress and Results
    • Physics-based and Diffusion-based Foggy Image Generation
    • Finetune Faster R-CNN to Improve Detection and tracking
    • Utilize NeRF to Defog the Images

Part 2: Reconstruction of Dense 3D Point Cloud from Images

  • Motivation
  • Progress and Results
    • Method 1: SLAM and Filtering

2

2

2

2

3 of 25

Motivation

  • The autonomous driving industry has seen rapid growth
  • A key challenge: Vehicle perception performance in adverse weather conditions, which can be severely degraded due to limited visibility, low quality of captured images and insufficient training data.
  • Insufficient training data in these conditions often leads to poor performance of perception algorithms.

3

3

3

3

4 of 25

Problem Statement

  • Goal: Improve the performance of the detection model in bad weather
  • Ways to Improve:
    • Synthesize more training data to train or finetune detection models
    • Pre-process the images to de-weather

4

4

4

4

5 of 25

Solutions

Clear images

Bad weather

images

Bad weather images

De-weathered

images

When testing with bad weather images

Clear images

Train

Train

Test

synthesize

De-weathering

Detector

Detector

Detector

Baseline:

Synthesis:

De-weather:

5

5

5

5

6 of 25

Progress and Results

6

6

6

6

7 of 25

Dataset Collection

  • BDD100K
    • A large and diverse driving video dataset, has more than 100,000 videos captured in a wide range of weather conditions, times of day, and urban and suburban environments
    • With ground truth for multiple tasks, like objection detection and segmentation
  • DENSE
    • Included a rich variety of sensor data—such as camera, radar, and lidar
    • Collected under various challenging scenarios like fog, rain, and snow
  • Driving videos
    • All data collected as videos with an iPhone 11, iPhone 14, or Rove R2-4k Dash Cam, with a frame rate of 30 fps
    • Under different weather conditions: rain, snow or fog; day or night
  • Collect data from CARLA by ourselves
    • Under different scenes and with different amounts of fog

7

7

7

7

8 of 25

Physics-based Foggy Image Generation

E: Foggy Image

R: Clear Image

: Atmosphere Light

𝛽: Scattering Coefficient

d: Depth Map

  • Input clear images
  • Use Depth Anything method to generate depth map
  • Generate synthetic foggy images based on atmospheric scattering model
    • Adjust the amount of fog by β

Input Images and Corresponding Foggy Images

8

8

8

8

9 of 25

Diffusion-based Foggy Image Generation

U-Net

Source domain X:

BDD clear day images

Target domain Y:

Carla foggy images

Decoder

Encoder

“Driving in the fog”

Text Encoder

First Stage Connections

Generator G/F

9

9

9

9

10 of 25

Results: Diffusion-based Foggy Image Generation

Source Images

Generated Images

Source Images

Generated Images

10

10

10

10

11 of 25

Training dataset for fine-tuning

(Pretrained on COCO)

Testing dataset

mAP

AP@50

AP@70

mAP_S

mAP_M

mAP_L

None

DENSE(heavy fog)

0.261

0.422

0.277

0.011

0.188

0.367

BDD100k clear images

DENSE(heavy fog)

0.304

0.491

0.327

0.043

0.225

0.412

BDD100k all images

DENSE(heavy fog)

0.314

0.513

0.332

0.043

0.234

0.424

BDD100k generated foggy images

DENSE(heavy fog)

0.342

0.551

0.368

0.070

0.266

0.438

None

Transweather(fog)

0.268

0.431

0.282

0.013

0.194

0.373

BDD100k clear images

Transweather(fog)

0.314

0.500

0.341

0.074

0.238

0.414

BDD100k all images

Transweather(fog)

0.324

0.527

0.346

0.051

0.246

0.432

BDD100k generated foggy images

Transweather(fog)

0.343

0.554

0.364

0.074

0.264

0.447

Finetune Faster R-CNN and test on DENSE dataset’s heavy foggy images

Fine-tuning can significantly improve the performance of Faster R-CNN!

11

11

11

11

12 of 25

Finetune Faster R-CNN to Improve Detection

  • Used real driving dataset (and generated foggy images) to finetune the Faster R-CNN model
  • Test on DENSE dataset’s heavy foggy images (w/ or w/o dedazing)

Ground Truth Finetune on clear images Generated foggy images

12

12

12

12

13 of 25

Application: Object Filtering Based on Tracking

  • Use our fine-tuned Faster R-CNN detector to infer on some unlabeled driving videos
  • Track the trajectories of the vehicles
  • Remove the false positives by filtering out the trajectories whose lengths < 20 frames

Detection

Output

(with FP)

FP removal

13

13

13

13

14 of 25

Visualization of Trajectory Length

  • To further extend the tracking duration:
    • Employ the bounding boxes from the current frame as region proposals to facilitate backward tracking.
    • Based on 'Tracking by Detection' method, which is compatible with any two-stage detector.

Current Frame

bboxes Output

Next Frame’s bboxes as

Region Proposals

Around frame 500, it becomes evident that our fine-tuned detection model enhances the length of trajectories.

14

14

14

14

15 of 25

Utilizing NeRF to Defog the Images

  • Collecting data from CARLA, an autonomous driving simulation platform
  • Generating foggy images
  • Training NeRF with the help of Atmospheric Scattering Model

15

15

15

15

16 of 25

Pipeline of Using NeRF to Defog

L2-Loss

L2-Loss

Depth

Ground Truth

Atmospheric

Scattering

Model

Ray Origin o

Ray Direction d

Neural Radiance

Field

Color c

Density σ

Volume

Rendering

Pixel Color C

Pixel Depth D

Foggy Image

β: Learnable Parameter

16

16

16

16

17 of 25

Result

Physics-based Fog, with Depth Supervision

Defogged Image

Reconstructed Foggy Image

CARLA Foggy Image

17

17

17

17

18 of 25

Result

Physics-based Fog, with Depth Supervision

CARLA

Foggy

Image

Defogged

Image

18

18

18

18

19 of 25

Result

Controlling the amount of fog

CARLA foggy image

beta=0

0.5*beta

beta

2*beta

19

19

19

19

20 of 25

Reconstruction of Dense 3D Point Cloud from Images

Motivation:

  • Corner cases in autonomous driving occur not only in bad weather but also in scenarios with static obstacles.
  • Research scope expanded to 3D reconstruction of road work scenarios.
  • Reconstructed obstacles could be further used in downstream planning tasks to test algorithm performance in avoiding construction zones.
  • Improves automation level of autonomous driving and reduces the need for human intervention.

20

20

20

20

21 of 25

Reconstruction

Image Sequences

Metric3D

Model

Metric Depth

Pipeline

Example

Input Image Sequence

Output Point Cloud

21

21

21

21

22 of 25

Filtering

To get the positions and points of obstacles in road work areas, we process the 3D point cloud and validates points based on projections across multiple 2D image masks. The goal is to filter points in the 3D space that are visible and valid in a specific number of consecutive frames.

SLAM Result

Filtering Based on 2D Segmentation Labels

BEV Visualization

Input Sequence

Segmentation Mask

Before Filtering

After Filtering

22

22

22

22

23 of 25

Result

Geo-alignment

23

(a)

(b)

The camera trajectories (a) before geo-alignment. (b) after geo-alignment.

The red trajectories represent the SLAM poses, while the blue trajectories depict the geo-aligned COLMAP poses.

23

23

23

23

24 of 25

Future Work

Conclusion

  • Incorporating the obstacles with planning task in autonomous driving
    • Inserting the 3D obstacles into simulation platforms and create more road work scenarios
    • Hopefully, we can improve the obstacle-avoidance of planning algorithms
  • Fine-tuning Faster R-CNN for enhanced vehicle detection in adverse weather.
  • Leveraging NeRF for effective defogging to improve perception.
  • Employing 3D reconstruction to enable navigation and planning in construction zones.

24

24

24

24

25 of 25

Thanks for Your Listening!

25

25

25

25