Driving in Bad Weather
Presenter: Zi Wang, Peipei Zhong
Advisor: Srinivasa Narasimhan
MSCV Capstone 2024 Fall
1
1
1
1
Outline
Part 1: Driving in Bad Weather
Part 2: Reconstruction of Dense 3D Point Cloud from Images
2
2
2
2
Motivation
3
3
3
3
Problem Statement
4
4
4
4
Solutions
Clear images
Bad weather
images
Bad weather images
De-weathered
images
When testing with bad weather images
Clear images
Train
Train
Test
synthesize
De-weathering
Detector
Detector
Detector
Baseline:
Synthesis:
De-weather:
5
5
5
5
Progress and Results
6
6
6
6
Dataset Collection
7
7
7
7
Physics-based Foggy Image Generation
E: Foggy Image
R: Clear Image
: Atmosphere Light
𝛽: Scattering Coefficient
d: Depth Map
Input Images and Corresponding Foggy Images
8
8
8
8
Diffusion-based Foggy Image Generation
U-Net
Source domain X:
BDD clear day images
Target domain Y:
Carla foggy images
Decoder
Encoder
“Driving in the fog”
Text Encoder
First Stage Connections
Generator G/F
9
9
9
9
Results: Diffusion-based Foggy Image Generation
Source Images
Generated Images
Source Images
Generated Images
10
10
10
10
Training dataset for fine-tuning (Pretrained on COCO) | Testing dataset | mAP | AP@50 | AP@70 | mAP_S | mAP_M | mAP_L |
None | DENSE(heavy fog) | 0.261 | 0.422 | 0.277 | 0.011 | 0.188 | 0.367 |
BDD100k clear images | DENSE(heavy fog) | 0.304 | 0.491 | 0.327 | 0.043 | 0.225 | 0.412 |
BDD100k all images | DENSE(heavy fog) | 0.314 | 0.513 | 0.332 | 0.043 | 0.234 | 0.424 |
BDD100k generated foggy images | DENSE(heavy fog) | 0.342 | 0.551 | 0.368 | 0.070 | 0.266 | 0.438 |
None | Transweather(fog) | 0.268 | 0.431 | 0.282 | 0.013 | 0.194 | 0.373 |
BDD100k clear images | Transweather(fog) | 0.314 | 0.500 | 0.341 | 0.074 | 0.238 | 0.414 |
BDD100k all images | Transweather(fog) | 0.324 | 0.527 | 0.346 | 0.051 | 0.246 | 0.432 |
BDD100k generated foggy images | Transweather(fog) | 0.343 | 0.554 | 0.364 | 0.074 | 0.264 | 0.447 |
Finetune Faster R-CNN and test on DENSE dataset’s heavy foggy images
Fine-tuning can significantly improve the performance of Faster R-CNN!
11
11
11
11
Finetune Faster R-CNN to Improve Detection
Ground Truth Finetune on clear images Generated foggy images
12
12
12
12
Application: Object Filtering Based on Tracking
Detection
Output
(with FP)
FP removal
13
13
13
13
Visualization of Trajectory Length
Current Frame
bboxes Output
Next Frame’s bboxes as
Region Proposals
Around frame 500, it becomes evident that our fine-tuned detection model enhances the length of trajectories.
14
14
14
14
Utilizing NeRF to Defog the Images
15
15
15
15
Pipeline of Using NeRF to Defog
L2-Loss
L2-Loss
Depth
Ground Truth
Atmospheric
Scattering
Model
Ray Origin o
Ray Direction d
Neural Radiance
Field
Color c
Density σ
Volume
Rendering
Pixel Color C
Pixel Depth D
Foggy Image
β: Learnable Parameter
16
16
16
16
Result
Physics-based Fog, with Depth Supervision
Defogged Image
Reconstructed Foggy Image
CARLA Foggy Image
17
17
17
17
Result
Physics-based Fog, with Depth Supervision
CARLA
Foggy
Image
Defogged
Image
18
18
18
18
Result
Controlling the amount of fog
CARLA foggy image
beta=0
0.5*beta
beta
2*beta
19
19
19
19
Reconstruction of Dense 3D Point Cloud from Images
Motivation:
20
20
20
20
Reconstruction
Image Sequences
Metric3D
Model
Metric Depth
Pipeline
Example
Input Image Sequence
Output Point Cloud
21
21
21
21
Filtering
To get the positions and points of obstacles in road work areas, we process the 3D point cloud and validates points based on projections across multiple 2D image masks. The goal is to filter points in the 3D space that are visible and valid in a specific number of consecutive frames.
SLAM Result
Filtering Based on 2D Segmentation Labels
BEV Visualization
Input Sequence
Segmentation Mask
Before Filtering
After Filtering
22
22
22
22
Result
Geo-alignment
23
(a)
(b)
The camera trajectories (a) before geo-alignment. (b) after geo-alignment.
The red trajectories represent the SLAM poses, while the blue trajectories depict the geo-aligned COLMAP poses.
23
23
23
23
Future Work
Conclusion
24
24
24
24
Thanks for Your Listening!
25
25
25
25