1 of 7

Machine Learning in

applied economic analysis

Introduction to the Module 6800 in the Promotionskolleg, August 23-27, 2021

Kathy Baylis,

Thomas Heckelei,

Hugo Storm

2 of 7

Intro to the course

Learning Objectives

  • understand when ML is appropriate and what it can provide complementary to your econometric toolbox
  • focus on concepts, tuning and applications, not coding from scratch
  • enable students to learn specific approaches or applications more deeply on their own

What do we do?

  • We will cover some basic concepts in ML and relate them to standard econometric practice
  • We will cover methods like LASSO, Ridge Regressions, Random forests, boosted trees and neural networks (for prediction and also to connect them to causal analysis)
  • You will get hands-on experience running ML routines using python in jupyter notebooks (You do not need to be familiar with Python beforehand, although experience with some programming is never a bad thing…)

3 of 7

Technical format of the course

General schedule (specific one will be provided before the course)

  • 8:30-17:00 with lunch break from 12:00 to 13:30, Friday modified and only until 15:00
  • Videos and slides provided in the morning for self-study
  • Live zoom sessions for Q&A after lunch
  • Live zoom sessions in afternoon with breakout rooms for labs and some interaction on your research (days 3-5)

Material

Assignments

  • are given each day and you should get a head-start on them in the lab session
  • Final versions to be uploaded latest by September 30 using this link (password provided in email)

4 of 7

Things to do before starting the class

  1. Create a google account
  2. Become familiar with Jupyter notebooks (https://jupyter.org/)
    1. Jupyter notebook guides

http://opentechschool.github.io/python-data-intro/core/notebook.html

https://www.datacamp.com/community/tutorials/tutorial-jupyter-notebook

​

5 of 7

Intro to the data (used on several days)

  • Gridded dataset: 0.01 x 0.01 degrees grid cells (~1.011 km x 1.011 km)
  • Outcome variable: Hansen’s deforestation data (Hansen et al., 2013*)
  • Each pixel takes the value of 1, if it goes from being forest to no forest
  • Forests are defined as grid cells where at least 30% of the area is covered by trees
  • We estimate the percentage of each grid cell that is deforested every year

​

​

* Hansen, M. C., Potapov, P. V., Moore, R., Hancher, M., Turubanova, S. A., Tyukavina, A., … Townshend, J. R. G. (2013). High-Resolution Global Maps of 21st-Century Forest Cover Change. Science, 342(6160), 850–853. https://doi.org/10.1126/science.1244693

6 of 7

Intro to the data ((used on day 1, 2, 4)

  • Explanatory variables: Protected areas, national and subnational borders, distance to roads, travel time to major towns, rainfall, crop suitability, population, elevation and land-use (% crop and pasture).
  • Variable description and sources
  • Spatial lags of all variables, up to the third order.
  • Download of the data

Protected Areas in 2018

7 of 7

Intro Lab (to be done before course starts)

​

    • Introduction to notebooks; illustration with OLS
    • To make sure you can run a jupyter notebook with google colab as this is how we do it during the ML course

​

Assignment, due Day 1:

  1. Open the introductory jupyter notebook
  2. Follow the instructions in the notebook
  3. Save your version of the file (add your last name at the end) when you are done under the upload link (password provided in email)

​

No worries….most of what needs to be done in the jupyter notebook is already prepared