1 of 18

Upscaling Wetland Fluxes using AI

Ashley Brereton, Zelalem Mekonnen, Bhavna Arora, Bill Riley, Kunxiaojia Yuan, Yi Xu, Yu Zhang, Qing Zhu, Tyler Anthony, Adina Paytan

2 of 18

Introduction

  • Wetlands are low oxygen environments

  • This slows down the decomposition of organic matter

  • Good for keeping carbon locked in the soil

  • Bad for Methane release

  • Goal: Complicated and nonlinear system – get the computer to work it all out.

3 of 18

Typical wetland site in the Delta

4 of 18

Region of interest: �Sacramento-San Joaquim Delta

Site Code

Site Name

Water Type

Salinity

Years of Data (Full)

Start Date

US-Myb

Mayberry Wetland

Non-Tidal

Fresh

13

2010

US-Tw1

Twitchell Wetland West Pond

Non-Tidal

Fresh

12

2011

US-Tw4

Twitchell Island East End Wetland

Non-Tidal

Fresh

10

2013

  • The Delta has a disproportionately large volume of CO2 and methane flux data (35 site years)

  • Home to an expansive area of wetlands (1000km2)

  • Perfect test case for model upscaling application

5 of 18

Model Training: Schematic

Target

AI model in training

Features (predictors)

6 of 18

Model Upscaling: Schematic

Model prediction

Features

Trained AI model

7 of 18

Model targets

CO2 flux

Methane flux

Target 1:

Target 2:

8 of 18

Model features (predictors)

  • Application: Regional upscaling

  • Important consideration: Data availability at the regional scale

  • Gold standard in-situ measurements are not available at the regional scale

WLDAS (1km, daily)

  • Weather
    • Air temp
    • Solar
    • Wind

  • Subsurface
    • Soil temp
    • Soil moisture
    • Water table depth

LANDSAT (30m, weekly)

  • Vegetation
    • Cover
    • Color

9 of 18

Model suite

  • Many studies use fancy models, but leave out simpler models, such as linear regression

  • The motivation for using fancy models is that the extra complexity is justified due to superior performance over simpler models – so this must be demonstrated!

  • So we establish a model framework:

    • Training a group of models

    • From linear regression (simple) to neural networks (complex)

    • … and in between (decision trees, support vector machines, gradient boosting)

Model Name

Category

Description

Key Strengths

Linear Regression

Regression

Fits a linear relationship between predictors and fluxes

Simple baseline, easily interpretable

Random Forest

Ensemble of Decision Trees

Aggregates multiple decision trees to enhance prediction stability

Robust to nonlinearity, reduces overfitting

Support Vector Machine (SVM)

Kernel-Based Method

Uses flexible kernels to find optimal separating hyperplanes

Effective in high dimensions, adaptable kernels

LightGBM

Gradient Boosting

Employs iterative boosting with efficient tree growth

Fast, memory-efficient, handles large datasets

XGBoost

Gradient Boosting

Improves boosting with regularization and efficient computations

Manages outliers, handles sparse data well

LSTM Neural Network

Recurrent Neural Network

Captures temporal dependencies in sequential data inputs

Ideal for time-series, learns long-term patterns

GRU Neural Network

Recurrent Neural Network

Similar to LSTM but streamlined with fewer parameters

Efficient temporal modeling, lower complexity

10 of 18

Model Training: Leave-One-Site-Out (LOSO)

  • Upscaling: learn from monitored sites and apply model to unmonitored sites

  • We should train our model appropriately in the same way

Training Recipe:

  1. Take all site locations
  2. Remove one site
  3. Train model on remaining sites
  4. Predict removed site
  5. Cycle through all sites

Result:

‘Blind’ predictions of CO2 and methane flux for all sites.

Used to compare to observations to evaluate model performance (R2 , r, RMSE)

11 of 18

Model Training: Leave-One-Site-Out (LOSO)

  • We run the full model suite in LOSO training mode and compare performance

  • ‘Good’ performance all round (R2 > 0.5)

  • Complex models perform best

  • But the difference between simple and complex is small

12 of 18

Trained Models

CO2 flux

Methane flux

Tules

Rice

R2 = 0.73

R2 = 0.53

R2 = 0.51

R2 = 0.57

13 of 18

Aim

  • Now that we have a framework to upscale C fluxes, we test it on 2 types of vegetation

    • Tules (Environmentally ‘friendly’ crops)
    • Rice (Agriculture)

  • Purpose is to

    • Directly analyze the effectiveness of land use strategies
    • Implications for the environment

14 of 18

Tules Upscaling

NECB = Carbon Sequestration

  • Mostly negative

  • Carbon goes into the ground – Good!

RF = Radiative Forcing

  • Mostly negative, but variability exists

  • Greenhouse gas goes into the ground – Good!

15 of 18

Rice Upscaling

NECB = Carbon Sequestration

  • Mostly negative

  • Carbon goes into the ground – Good!

RF = Radiative Forcing

  • Mostly positive, but variability exists

  • Greenhouse gas goes into the atmosphere – Bad.

  • NOTE – Even the bad is better than the peat oxidation baseline

16 of 18

Comparing Regional Flux Predictions

TULES

RICE

NECB = -0.35 kg GHG = -0.2 kg

NECB = -0.25 kg GHG = +0.1 kg

17 of 18

Summary

  • Developed a robust Framework for upscaling C Fluxes using AI methods

  • Assessed two types of vegetation

    • Tules
    • Rice
  • Model supports wetland restoration

  • Model implies climate impact for land reclamation for agriculture

  • But model also indicates rice as a carbon sink despite it being a GHG source.

Future work – other agriculture practices in the Delta (corn, alfalfa) but more important tidal wetlands which are expected to be more complex.

18 of 18

Feed-forward feature selection

Target Variable

Step

Chosen Feature

Target Variable

Step

Chosen Feature

FCO2

TULES

1

Soil Adjusted Vegetation Index (SAVI)

0.59

FCO2

TULES

1

GNDVI

(Greenness) NDVI

0.52

2

Upward Sensible Heat Flux

0.73

2

Shortwave radiation

0.56

FCH4

TULES

1

Canopy Temperature

0.48

FCH4

RICE

1

Soil Adjusted Vegetation Index (SAVI)

0.34

2

Soil Temperature

0.52

2

Emmisivity std

0.52

3

GNDVI

(Greenness) NDVI

0.53