1 of 42

STATS / DATA SCI 315

Lecture 01

Course logistics and introduction to deep learning

2 of 42

Course logistics

3 of 42

What is the course about?

  • Deep learning (DL): A branch of machine learning that uses multilayer neural networks to solve problems such as object recognition and playing board games (Chess, Go)
    • Machine learning (ML): A branch of artificial intelligence (AI) that seeks to endow machines with the ability to learn from experience
    • Statistics (Stats): The science of learning from data. ML and Stats have been coming increasingly close over the past two decades
    • Data science (DS): An emerging discipline that seeks to marry statistical thinking with computational thinking to solve difficult real world problems.
  • Tensorflow (TF)/Keras: We will use popular software libraries that make it very easy to build deep learning models (which has its pros and cons)

4 of 42

Prerequisites

  • Calculus: derivatives, gradients, chain rule
    • We will try to get by without multivariable calculus
  • Linear algebra: vectors, matrices, norms
    • Covered in boot camp
  • Prob/Stats: random variables, expectations, linear regression
  • Programming: variables, data structures, loops, functions, classes, objects
  • We will NOT assume prior exposure to:
    • Machine learning
    • Python
  • The first few weeks can feel intense as we cover the necessary Stats/ML and Python material to get you started with deep learning using Tensorflow

5 of 42

Books

  • Deep Learning with Python (2nd edition) by Chollet
    • Emphasizes the programming side
  • Dive into Deep Learning by Zhang, Lipton, Li and Smola
    • Comprehensive, covers math background as well
  • Deep Learning by Goodfellow, Bengio and Courville
    • Graduate student/researcher level, just for reference

You don’t have to buy anything! Materials not on the web will be provided via Canvas

6 of 42

Course communication tools

7 of 42

Instructional team

Faculty Instructor: Ambuj Tewari, tewaria@umich.edu

GSI: Sahana Rayan, srayan@umich.edu

GSI: Jacob (Jake) Trauger, jtrauger@umich.edu

GSI: Mallory Wang, wmallory@umich.edu

8 of 42

Grading will be on a curve

  • Canvas quizzes (20%): Will drop two lowest scores
  • Homeworks (30%): Assigned roughly every other week. Will drop one lowest score
  • Midterm Exam (20%): In class, timed, multiple choice, open book
  • Final Exam (30%): In class, timed, multiple choice, open book

  • Median grade will be A-

9 of 42

Academic Integrity

  • You can discuss homeworks (but not quizzes/exams) with your classmates
  • But all submitted work, including code, must be your own
  • Misconduct will be reported to the Dean’s office
  • When in doubt, ask!

10 of 42

Accommodation for Students with Disabilities

  • Submit your VISA form as soon as possible
    • Electronic submission strongly encouraged
  • Talk to me privately if you need any other accommodations

11 of 42

Mental Health and Well-Being

  • Be aware of available resources
  • Seeking help when needed is courageous!
  • If this course is adding to your stress, talk to me privately

12 of 42

Rough Course Outline

  • Intro to Python, Numpy, TF2, Keras, Jupyter, Colab
  • Simple models that are precursors to DL: linear regression
  • Fully connected, multilayer Neural Networks (NNs)
  • Convolutional neural networks and vision
  • Sequence models and language

13 of 42

Reading assignments

  • Course schedule is at: https://ambujtewari.github.io/stats315-winter2023/
  • Every lecture has associated reading assignments
  • Material not available on the web will be under “Files” in Canvas
  • It is your responsibility to read the required materials
  • Recommendation: read it at least twice, preferably thrice. At least once before lecture and once afterwards

14 of 42

Introduction to deep learning

15 of 42

16 of 42

Beginnings of AI

Alan Turing’s seminal papers

  • Intelligent Machinery, report written for the Nat. Physical Lab, 1948 (published only in 1970)
  • Computing Machinery and Intelligence, MIND, Vol. 59, 1950

Dartmouth summer workshop proposal, 1956

The study is to proceed on the basis of the conjecture that every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it. An attempt will be made to find how to make machines use language, form abstractions and concepts, solve kinds of problems now reserved for humans, and improve themselves. We think that a significant advance can be made in one or more of these problems if a carefully selected group of scientists work on it together for a summer.

17 of 42

Symbolic AI

  • Relied on hand-crafted rules for manipulating knowledge stored in databases
    • Cyc had over 24 million rules about over 1 million objects in its ontology in 2017
  • Dominant approach in the 1950s to 1980s
  • Reached its peak with the expert systems boom of the 1980s
  • Tasks that are easy for people to do but hard to formalize proved challenging
    • Recognizing spoken words or recognizing faces in images

18 of 42

19 of 42

Machine learning

  • An ML system is trained instead of being explicitly programmed
  • Started in the 1990s
  • Has had explosive growth driven by faster hardware and larger datasets
  • Example: wake word detection in voice assistants (Alexa, Siri, Google Assistant)

20 of 42

Wake word detection

  • You do not know how to program a computer to recognize the word “Alexa”
  • You yourself are able to recognize it!
  • We can collect a huge dataset containing examples of:
    • Audio, and
    • Label, i.e. whether or not the audio contains the wake word

21 of 42

Representations

  • An ML model transforms the input data (audio) into meaningful output (is the wake word in it?)
  • Different representations are useful for different tasks: consider images in RGB (red-green-blue) vs HSV (hue-saturation-value) formats
    • “Select all red pixels” is easier in RGB, “Make the image less saturated” is easier in HSV
  • We can make tasks easier by choosing better representations

22 of 42

23 of 42

Deep learning

  • Learn successive layers of increasingly meaningful representations
  • Enables the computer to build complex concepts out of simpler concepts
  • Modern DL involves tens and sometimes hundreds of layers
  • All of the parameters in these layers are learned automatically from data

24 of 42

25 of 42

26 of 42

27 of 42

28 of 42

Deep learning: why now?

29 of 42

30 of 42

Historical roots of deep learning

  • The apparent novelty of deep learning is deceptive
  • “Deep learning” is a recent rebranding of neural networks
  • NN research has had three waves:
    • 1940s-60s: cybernetics
    • 1980s-90s: connectionism
    • Since 2006: deep learning
  • (Artificial) neural networks are engineered systems inspired by the brain
  • However, the models and algorithms we will use are NOT realistic biological models of brain functions

31 of 42

What happened around 2010?

  • Prior to 2010, algorithms were too computationally costly for available hardware
  • In 2006, Geoff Hinton showed that a kind of neural network called a “deep belief network” could be efficiently trained
  • This wave of neural networks research popularized the use of the term “deep learning”
  • Researchers were now able to train deeper neural networks than had been possible before
  • They focused their attention on the theoretical importance of depth

32 of 42

33 of 42

34 of 42

35 of 42

36 of 42

Improvements in the median accuracy of predictions in the free modelling category for the best team in each CASP. A score of around 90 GDT is informally considered to be competitive with results obtained from experimental methods.

37 of 42

Success stories

Near-human-level image classification�Near-human-level speech transcription�Near-human-level handwriting transcription�Dramatically improved machine translation�Dramatically improved text-to-speech conversion�Digital assistants such as Google Assistant and Amazon Alexa�Near-human-level autonomous driving�Improved ad targeting, as used by Google, Baidu, or Bing�Improved search results on the web�Ability to answer natural language questions�Superhuman Go playing�Dramatically improved protein structure prediction

38 of 42

Real world impact of computing and AI

  • In 2009 the top ten companies by market cap included only one big tech company, Microsoft
  • As of Jan 2023, there were four: Apple, Microsoft, Alphabet (Google), Amazon
    • 3 more – Tesla, Tencent, NVIDIA – in top twenty

39 of 42

Source: state of AI report 2022 at https://www.stateof.ai/

40 of 42

Short term hype vs long term vision

  • Expectations for what the field will be able to achieve in the next decade tend to run much higher than what will likely be possible
  • Believable dialogue systems, human-level machine translation across arbitrary languages, and human-level natural language understanding may remain elusive for a long time
  • We may be currently witnessing the third cycle of AI hype and disappointment, and we’re still in the phase of intense optimism. Might another AI Winter be lurking around?
  • Long term vision: AI will help humanity as a whole move forward, by assisting human scientists in new breakthrough discoveries across all scientific fields, from genomics to mathematics
  • Don’t believe the short-term hype, but do believe in the long-term vision!

41 of 42

42 of 42

Ethical issues in AI

  • Energy consumption, sustainability, climate change
  • Social, economic and political inequality
  • Unemployment
  • Autonomous weapons
  • Surveillance and loss of privacy