1 of 19

DIFFERENTIABLE END-TO-END AUTONOMOUS DRIVING

Darius Kianersi

2 of 19

Modular vs. End-to-End

3 of 19

Modular vs. End-to-End

MODULAR

  • Allows for parallel development
  • Interpretability
  • May not be optimized for driving
  • Compounded loss
  • Lots of human labor

END-TO-END

  • Optimizes directly for driving
  • Data-driven + annotations are cheap
  • Deep learning backbones
  • Hard to interpret

4 of 19

Imitation Learning

  • Learn from expert demonstrations (trajectories)
  • possible actions (steering angle, speed, brakes, etc.)
  • possible states (road image, sensor data)
  • Learn from
  • State distribution

5 of 19

Behavior Cloning

  • Use a state distribution provided by the expert

  • Supervised learning
  • Issue: State IID assumption

6 of 19

DAgger

  • Data Aggregation (DAgger) (Ross and Bagnell)
  • Solution: Collect on-policy data during rollout
    • Query data against expert policy to aggregate dataset
  • Improvements: critical state and replay buffer

7 of 19

ALVINN

  • Autonomous Land Vehicle in a Neural Network (ALVINN)
  • Proposed in 1988 by CMU
  • Video input + Laser range finder
  • 3 layer NN to map road images to steering angle
  • Trained on simulated road images, tested on real urban roads

8 of 19

NavLab

9 of 19

PilotNet

  • Proposed by NVIDIA (2016)
  • 3 camera angles
  • Adjust for shift and rotation
  • CNN architecture, inspired by AlexNet
    • 250k parameters
  • Trained on 72 hours of driving

10 of 19

Addressing Interpretability

  • VisualBackProp: efficient visualization of CNNs (Bojarski et al.)
  • Create visualization ”masks” using activations from inference
  • Debugged self-driving cars by shifting objects in original images

11 of 19

Conditional Imitation Learning

  • Condition state on signal (e.g. turn left) obtained via GPS
  • New objective:

  • Two proposed architectures by Codevilla et al.:

12 of 19

Inverse (RL) Optimal Control

  • Learn an unknown reward function from expert demonstrations
  • Rollout using the reward and compare to expert behavior
  • Limitations: many possible rewards; difficult to optimize
  • Generative Adversarial Imitation Learning (GAIL)
    • Reward function as adversarial objective

13 of 19

  • Markov Decision Process
    • States, actions, reward function, state distribution, discount factor
  • Markov property: future is independent of previous states given present state

Reinforcement Learning

14 of 19

Reinforcement Learning

  • Gradients from RL are generally inefficient for deep perception architectures
  • Affordances
    • Traffic light state
    • Semantic segmentation
  • Implicit affordances (Toromanoff et al.)
    • Train encoder backbone (Resnet) to predict affordances
    • Activations of encoder are used as state for RL
    • SoTA performance

15 of 19

CARLA Benchmark

CARLA 0.9.15 release

  • CARLA: open-source simulator for autonomous driving research
  • Customizability
    • Assets: Urban layouts, buildings, vehicles
    • Sensor suites, env. conditions, dynamic actors, etc.
  • Benchmark: four tasks
    • Straight, one turn, navigation, nav. dynamic
  • Benchmark: infractions
    • Opposite lane, sidewalk, collision

16 of 19

Current Directions: World Model-based RL

  • Build internal representation of the world
    • Baseball player “predicting the future”
  • 3-part architecture
    • VAE
    • Memory RNN
    • Controller
  • Trajectory forecasting

17 of 19

Current Directions: World-model based RL

  • Hu et al. proposed MILE
  • Encoder: lift and pool features to BEV
  • KL divergence to match prior and posterior dist.
  • Decoders output the reconstructed observation
  • Driving policy gives vehicle control
  • RNN computes a deterministic transition

18 of 19

Current Directions: Policy Distillation

  • Learning by Cheating (Chen et al.)
  • Two-part imitation learning
  • Teacher and student model
    • Teacher has privileged information

19 of 19

Sources

Bojarski, M., Chen, C., Daw, J., Değirmenci, A., Deri, J., Firner, B., ... & Yang, Z. (2020). The NVIDIA pilotnet experiments. arXiv preprint arXiv:2010.08776.

Bojarski, M., Choromanska, A., Choromanski, K., Firner, B., Jackel, L., Muller, U., & Zieba, K. (2016). Visualbackprop: efficient visualization of cnns. arXiv preprint arXiv:1611.05418.

Chen, D., Zhou, B., Koltun, V., & Krähenbühl, P. (2020, May). Learning by cheating. In Conference on Robot Learning (pp. 66-75). PMLR.

Codevilla, F., Müller, M., López, A., Koltun, V., & Dosovitskiy, A. (2018, May). End-to-end driving via conditional imitation learning. In 2018 IEEE international conference on robotics and automation (ICRA) (pp. 4693-4700). IEEE.

Dauner, D., Hallgarten, M., Geiger, A., & Chitta, K. (2023). Parting with Misconceptions about Learning-based Vehicle Motion Planning. arXiv preprint arXiv:2306.07962.

Ha, D., & Schmidhuber, J. (2018). World models. arXiv preprint arXiv:1803.10122.

Hu, A., Corrado, G., Griffiths, N., Murez, Z., Gurau, C., Yeo, H., ... & Shotton, J. (2022). Model-based imitation learning for urban driving. Advances in Neural Information Processing Systems35, 20703-20716.

Jin, W., Kulić, D., Mou, S., & Hirche, S. (2021). Inverse optimal control from incomplete trajectory observations. The International Journal of Robotics Research, 40(6-7), 848-865.

Toromanoff, M., Wirbel, E., & Moutarde, F. (2020). End-to-end model-free reinforcement learning for urban driving using implicit affordances. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition (pp. 7153-7162).

Yuan, Y., Cheng, H., Yang, M. Y., & Sester, M. (2023). Generating Evidential BEV Maps in Continuous Driving Space. arXiv preprint arXiv:2302.02928.