1 of 45

Where’s the Wood?�Quantifying Aggregate Demand for Roundwood at Primary Processing Mills in the Northeast/Midwest – A Machine Learning Approach

IAN KENNEDY

1

2 of 45

Agenda

1. What is the TPO Survey?

2. Survey Biases & Limitations

3. The Modeling Process

4. The Mapping Process

2

3 of 45

What is the TPO Survey?

  • TPO = Timber Products Output
  • USDA FS implemented survey attempting to track timber removals and subsequent lumber production at mills.
  • Implementation is split across three regions; the North, South, & West. FFRC works in coordination with NRS-FIA to implement the North’s surveys.

4 of 45

What type of information is collected?

  • Basic Information (Name, Address, Contact Info, etc.)
  • Timber Procurement Volume (unprocessed)
    • ‘Species Matrix’ used to quantify species’ proportions
    • ‘Species-Origin Matrix’ used to quantify procurement-source proportions for each species…Ideally at county level, though state level if county information unknown

*More to come on this*

  • Volume of Primary Product(s) Produced
    • Proportions of Primary Product Type(s)
  • Byproduct/Residue Disposal Avenue(s)
    • Separate sections for Milling and Logging residues

4

5 of 45

The Annual Sample Draw

  1. Goals of the ‘Draw’
        • Aim for ~40% sample-size within each state
        • Aim for ~80% capture of aggregate state volume
  2. How is this done?
    • The Annual Sample is drawn from a list of ~3,000 mills, using a few ‘weighting’ schemes at the state level:
        • Volume - Higher volume enhances likelihood of being drawn.
        • Mill Count – Each state must have >= 20 mills. All mills in states with <20 mills are sampled with certainty.
        • Mill Type – Each state must have >= 5 mills for each mill ‘type’. Mills ‘types’ with <5 mills are sampled with certainty.

5

6 of 45

Survey Limitations & Biases – Sample Statistics

6

7 of 45

Procurement Volume Histogram

7

8 of 45

Mill Type Cont. (Products by Mill Type)

8

9 of 45

The Modeling Process

All modeling analyses conducted using R

10 of 45

Why use Procurement Radius when Procurement-Source Information is collected?

  • Quality of responses for the matrix aren’t great:
    • Mill owners/operators often do not have the time and/or relevant information to provide county-level procurement information.
    • Many mill owners/operators do not engage in logging, rather they receive logs from adjacent logging companies.

10

Issues with using the Procurement Radius

  • Radius question has only been asked for 2019 (Phone), 2020, & 2021 iterations:
    • Each of these iterations features different formats for the unique identifier

11 of 45

Models Assessed

 

11

All assessed models attempt to predict SqRt(Procurement Radius). Assessed models and their associated (hyper)parameters are listed below:

4. Random Forest Regression (5x CV)

        • SqRt(Pro. Radius) ~ x
          • x = log10(MCF), State, Region, & Mill Type (Sawmill/Other)
          • Tuned Hyperparameters (3x):

- # of trees

- # of pred. vars used for each train/test split

5. Machine Learning (h2o) Regression (3x CV)

        • SqRt(Procurement Radius) ~ x
          • x = log10(MCF), State, Mill Type (Sawmill/Other), Lat/Lon, # of employees, Equipment, Portable?, Exported?

12 of 45

13 of 45

14 of 45

15 of 45

Linear Regression Visualization

15

Model Formula:

SqRt(Pro. Radius)i ~ log10(MCF)i

i = Region

16 of 45

Additive Regression Visualization

16

Model Formula:

SqRt(Pro. Radius)i ~ s(log10(MCF)i)

i = Region

s = tuned smoothing (by region)

17 of 45

Decision Tree Regression Visualization

17

Model Formula:

SqRt(Pro. Radius) ~ x

x = log10(MCF), State, Region,

& Mill Type (Sawmill/Other)

Tuned Hyperparameters (3x):

  1. Tree depth (# of nodes)
  2. Minimum node Size

18 of 45

Random Forest Regression Visualization

18

Model Formula:

SqRt(Pro. Radius) ~ x

x = log10(MCF), State, Region,

& Mill Type (Sawmill/Other)

Tuned Hyperparameters (3x):

  1. # of trees
  2. # of pred. vars used for each train/test split

19 of 45

ML (h2o) Regression Visualization

19

20 of 45

21 of 45

ML (h2o) Regression VIF Scores

21

22 of 45

ML (h2o) Regression Aggregate Residuals

22

23 of 45

The Mapping Process

Mapping script completed through ArcGIS Pro (Python)

24 of 45

GIS Scripting Workflow

24

  1. Geocode Addresses:
    1. First use Mailing Address
    2. If Mailing Address is unmatched, use Physical Address
    3. If Physical Address is unmatched, use Lat/Lon
    4. If still unmatched, store in ‘Unmatched’ datasheet
  1. Merge geocoded locations to a single shapefile
  2. Buffer locations using actual/predicted procurement radii
  3. Clip buffers to outline of USA
  4. Split buffers by ID
  5. Clip each buffer by the respective 1-hr Tucking Time limit
  6. Create independent raster for each buffer
  7. Merge rasters to create final ‘Mosaic’/Aggregate raster
  8. Alter symbology

25 of 45

GIS Scripting Workflow – Working in Travel Times

25

26 of 45

26

27 of 45

Preliminary Species-Composition Modeling Results

All modeling analyses conducted using R

28 of 45

Automated ML (h2o) Species-Proportion Modeling

  • Utilizing ‘complete’ species-composition responses from the 2020/2021 TPO Survey (Northern Region)
    • ‘complete’ = respondent provided annual procurement volume and a species-composition breakdown that accounted for the entire procurement volume.

  • The ‘h2o’ Model:
    • Train/Test split (70/30) utilized…All model statistics calculated on Test data.
    • Predictor Variables:
      • log10(MCF), State, Mill Type (Sawmill/Other), Lat/Lon, # of employees, Equipment, Portable?, Urban logs?, Exported?

  • Model Parameters (subject to change)
    • Stopping Metric: RMSE
    • Stopping Tolerance: .05
    • CV Folds: 5

29 of 45

What does the Data look like prior to Modeling?

30 of 45

White Oak Group – Residuals

30

31 of 45

White Oak Group – Imputed/Observed Proportions/Volumes

31

32 of 45

White Oak Group – Imputed State Volumes & VIF Scores

32

33 of 45

Red Oak Group – Residuals

33

34 of 45

Red Oak Group – Imputed/Observed Proportions/Volumes

34

35 of 45

Red Oak Group – Imputed State Volumes & VIF Scores

35

36 of 45

White Pine – Residuals

36

37 of 45

White Pine – Imputed/Observed Proportions/Volumes

37

38 of 45

White Pine – Imputed State Volumes & VIF Scores

38

39 of 45

Red Pine – Residuals

39

40 of 45

Red Pine – Imputed/Observed Proportions/Volumes

40

41 of 45

Red Pine – Imputed State Volumes & VIF Scores

41

42 of 45

Balsam Fir – Residuals

42

43 of 45

Balsam Fir – Imputed/Observed Proportions/Volumes

43

44 of 45

Balsam Fir – Imputed/Observed Proportions/Volumes

44

45 of 45

Questions? �and/or �Suggestions?

Thanks to:

  • Brett Butler, Marla Lindsay, & Sarah Butler (FFRC)
  • Ron Piva, Dave Haugen, & Mitch Slater (USDA FS, NRS-FIA)

45