1 of 32

Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving

W-CODA @ ECCV 2024

2 of 32

Opening Remarks and Welcome�Songcen Xu

W-CODA @ ECCV 2024

3 of 32

W-CODA @ ECCV 2024

4 of 32

Vision-based End-to-end Driving by Imitation Learning�Antonio M. López

W-CODA @ ECCV 2024

5 of 32

Reasoning Multi-Agent Behavioral Topology for Interactive Autonomous Driving�Hongyang Li

W-CODA @ ECCV 2024

6 of 32

Simulating and Benchmarking �Self-Driving Cars�Andreas Geiger

W-CODA @ ECCV 2024

7 of 32

Poster Session and Coffee Break

W-CODA @ ECCV 2024

8 of 32

W-CODA @ ECCV 2024

9 of 32

Long-tail Scenario Generation for Autonomous Driving with World Models��Lorenzo Bertoni

W-CODA @ ECCV 2024

10 of 32

Industrial Talk on Autonomous Driving��Chufeng Tang

W-CODA @ ECCV 2024

11 of 32

W-CODA @ ECCV 2024

12 of 32

Workshop Accepted Papers (Full Track)

W-CODA @ ECCV 2024

13 of 32

Workshop Accepted Papers (Abstract Track)

W-CODA @ ECCV 2024

14 of 32

Challenge Summary & Awards

W-CODA @ ECCV 2024

15 of 32

Track 1: Corner Case Scene Understanding

  • Track 1 focuses on enhancing multimodal perception and comprehension capabilities of MLLMs for autonomous driving, emphasizing global scene understanding, local area reasoning, and actionable navigation.

  • The goal is to promote the development of more reliable and interpretable autonomous driving.

  • It utilizes the CODA-LM dataset, which includes around 10,000 images and textual descriptions covering global driving scenarios, corner case analyses, and future driving recommendations.

W-CODA @ ECCV 2024

16 of 32

Track 1: Statistics

Submission:

  • 🔥 Participants: 50+ teams, 100+ individuals, 30+ institutions
  • 📊 Submissions: 100+ entries

 

W-CODA @ ECCV 2024

17 of 32

Track 1: Winner & Award

W-CODA @ ECCV 2024

18 of 32

Track 1: Winner & Award

W-CODA @ ECCV 2024

19 of 32

Track 1: Winner & Award

W-CODA @ ECCV 2024

20 of 32

Post-Credits !!!

W-CODA @ ECCV 2024

21 of 32

EMOVA: Empowering LMs to See, Hear and Speak

  • End-to-end omni-modal foundational model
    • Capable of dealing with visual, text and speech data simultaneously
    • Semantic-acoustic disentangled speech tokenizer

W-CODA @ ECCV 2024

22 of 32

EMOVA: Empowering LMs to See, Hear and Speak

  • Omni-modal alignment benefits
    • Observation 1: image-text and speech-unit-text data benefit each other.
    • Observation 2: semantic-acoustic disentanglement benefits omni-modal alignment.
    • Observation 3: sequential alignment is not optimal.

W-CODA @ ECCV 2024

23 of 32

EMOVA: Empowering LMs to See, Hear and Speak

  • Q:

  • A:

Not trained with any road samples!

W-CODA @ ECCV 2024

24 of 32

Track 2: Corner Case Scene Generation

  • This track aims to enhance diffusion models for generating multi-view street scene videos that align with 3D geometric descriptors, such as Bird's Eye View (BEV) maps and 3D LiDAR bounding boxes.

  • Building on the MagicDrive framework, this track strives for improved scene generation capabilities in autonomous driving, emphasizing consistency, higher resolution, and extended video duration.

W-CODA @ ECCV 2024

25 of 32

Track 2: Statistics

Submission:

    • 🔥 Participants: 10+ teams, 40+ individuals, 20+ institutions
    • 📊 Submissions: 10+ entries

Performance Improvements:

    • 🚀 FVD: 218.12 → 94.60 (Stunning Improveent: 56.63%)
    • 🌟 mAP: 11.86 → 24.55 (Remarkable Boost: 106.97%)
    • 💡 mIoU: 18.34 → 35.96 (Significant Increase: 96.02%)

Impressive Results In Submissions!

W-CODA @ ECCV 2024

26 of 32

Track 2: Winner & Award

W-CODA @ ECCV 2024

27 of 32

Track 2: Winner & Award

W-CODA @ ECCV 2024

28 of 32

Track 2: Winner & Award

W-CODA @ ECCV 2024

29 of 32

Post-Credits !!!

W-CODA @ ECCV 2024

30 of 32

MagicDrive3D: Controllable Street Scene Generation

MagicDrive can only generate fixed poses

Different vehicles, different camera

Unbounded controllable driving scene generation

?

  • Solution: Geometry-free Generation + Geometry-focused Reconstruction

Step 1: Controls -> Reconstruction-friendly video

Step 2: Generated view-friendly reconstruction

Dynamic Scene Video

Static Scene Video

  • Keep ego movement
  • Remove other object movement

a. A better initial PCD

b. A better deformable representation

b. A better algorithm with appearance embedding

  • Motivation: From one static camera to any novel poses

https://gaoruiyuan.com/magicdrive3d/

31 of 32

Make perception model robust to camera pose change!

MagicDrive3D: Controllable Street Scene Generation

  • Results: From one static camera to any novel poses

https://gaoruiyuan.com/magicdrive3d/

32 of 32

Summary & Future Plans

W-CODA @ ECCV 2024