1 of 38

1

Agenda

Duration (CST)

Introduction + AutoGluon Tabular

2:00PM – 2:55PM

Break

2:55PM – 3:05PM

AutoGluon Multimodal

3:05PM – 4:00PM

Break

4:00PM – 4:10PM

AutoGluon Timeseries

4:10PM – 4:50PM

Additional Q&A + Feedback

4:50PM – 5:00PM

Workshop Website

*Note on Hands-on Notebooks

  • pip install autogluon

  • Restart the kernel

© 2022, Amazon Web Services, Inc. or its Affiliates. All rights reserved.

auto.gluon.ai

2 of 38

Xingjian Shi, Senior Applied Scientist, AWS

Yi Zhu, Senior Applied Scientist, AWS

AutoGluon Multimodal

2

© 2022, Amazon Web Services, Inc. or its Affiliates. All rights reserved.

auto.gluon.ai

3 of 38

Real-life Multimodal Problems: Funding or not?

3

Fund raising platform aims to bring creative projects to life.

Predict whether the project will get funded?

auto.gluon.ai

4 of 38

Real-life Multimodal Problems: Funding or not?

4

Text

Categorical

Numerical

Multimodal features!

auto.gluon.ai

5 of 38

Real-life Multimodal Problems: Pet Adoption

5

Online pet adoption platform

Predict the adoption speed.

auto.gluon.ai

6 of 38

Real-life Multimodal Problems: Pet Adoption

6

Subject Focus

0

Eyes

1

Face

1

Near

1

Action

0

Accessory

0

Group

1

Image

Tabular meta-data

Pawpularity=63.

(Range: 0 – 100)

Multimodal features!

auto.gluon.ai

7 of 38

Real-life Multimodal Problems: Product Matching

7

Product Image

Title (Text)

Score (Numerical)

Price (Numerical)

Description (Text)

auto.gluon.ai

8 of 38

Real-life Multimodal Problems: Product Matching

8

Product Image

Title (Text)

Price (Numerical)

Color (Categorical)

Same!

Uploaded image

auto.gluon.ai

9 of 38

Real-life Multimodal Problems: Product Matching

9

Product Image

Title (Text)

Price (Numerical)

Color (Categorical)

Different!

Uploaded Image

auto.gluon.ai

10 of 38

AutoGluon – Multimodal AutoML Toolkit

10

Image

Text

Tabular

auto.gluon.ai/

State-of-the-art multimodal AutoML Toolkit on Image, Text, Time-series and Tabular data.

1. Data Processing

2. Training

3. Deployment

Huge science/engineering cost!

Train models in three lines of codes

Iterative + Expensive

auto.gluon.ai

11 of 38

Input Format of AutoGluon Multimodal – DataFrame

11

Text

Image

Categorical

Numerical

Support other types like bounding boxes, named entities

auto.gluon.ai

12 of 38

Train Model with Three Lines of Code

12

OpenAI/CLIP

Automated Model Ensemble (discussed in part 1)

Support pretrained “Foundation Models”

Focus of this section

auto.gluon.ai

13 of 38

Problems Supported in AutoGluon Multimodal

  • Classification & Regression
    • Image + Text + Tabular
  • Named Entity-Recognition
  • Object Detection
  • Multimodal Matching

13

auto.gluon.ai

14 of 38

Technical Details Underlying These Three Lines

  • Fusion foundation models of image / text
  • Object detection
  • Multimodal matching
  • Advanced topics
    • Parameter-efficient finetuning
    • Model distillation
    • Hyper-parameter optimization

14

auto.gluon.ai

15 of 38

Foundation Models of Image / Text

15

Text foundation models

Image foundation models

Text+Image foundation models

Foundation Models: Trained on large-scale datasets, able to support lots of applications [1].

[1] Bommasani, Rishi, et al. "On the opportunities and risks of foundation models." arXiv preprint arXiv:2108.07258 (2021).

auto.gluon.ai

16 of 38

Foundation Models of Image / Text

16

Source: https://huggingface.co/models. Taken on 2022/08/10

Large number of public foundation models (>60000 in huggingface/transformers, >600 in TIMM)!

The idea of model zoo is not widely adopted in AutoML toolkits. AutoGluon Multimodal focus on integrating/fusion of common model zoos.

auto.gluon.ai

17 of 38

Fusion Foundation Models of Image/Text

17

Image

Text

Tabular data

FT transformer [1]

Task prediction

OpenAI/CLIP

Fusion + Task-specific heads

[1] Gorishniy, Yury, et al. "Revisiting deep learning models for tabular data.” NeurIPS 2021

auto.gluon.ai

18 of 38

Fusion Foundation Model is a Powerful Solution

18

Fusion text pretrained model + tabular feature: Score equivalent to 2nd/2380 Teams

$100,000�Total Prize

auto.gluon.ai

19 of 38

Fusion Foundation Model is a Powerful Solution

19

The model ranks 20th place (out of 3537 teams), top 1%

Ensemble with different foundation models

  • vit_large_patch16_384 (fusion)
  • swin_large_patch4_window12_384 (fusion)
  • convnext_large_384_in22ft2k
  • swin_large_patch4_window7_224

Kaggle: Petfinder Pawpularity

auto.gluon.ai

20 of 38

Model Ensemble

20

Foundation Models

+

Classic Tabular Models

How to combine them and achieve the best of both worlds?

auto.gluon.ai

21 of 38

Model Ensemble

21

a) Feature extraction

b) Weighted ensemble

c) Stack ensemble

Stack ensemble > Weighted ensemble >> Feature extraction

[1] Shi, Xingjian, et al. "Benchmarking multimodal automl for tabular data with text fields." NeurIPS 2021 Track on Datasets and Benchmarks.

Conducted experiments on a benchmark with 18 multimodal datasets.

auto.gluon.ai

22 of 38

Ensemble Foundation Model and Tree Models: Top-1 in 3 Lines

22

Winning formula

  • Traditional tabular models (xgboost, etc.)
  • Fusion foundation model
  • Stacking

auto.gluon.ai

23 of 38

Ensemble Foundation Model and Tree Models: Top-1 in 3 Lines

23

auto.gluon.ai

24 of 38

Object Detection – Task-specific Model Zoo

  • 30+ architectures, 500+ pretrained models in MMDetection
  • Train a detector with 600+ images within 5-min.
  • See hands-on notebook in the later section

24

Example: Detect potholes for better road safety. Image from Kaggle Pothole Dataset

auto.gluon.ai

25 of 38

Multimodal Matching

Twin-tower Architecture

25

Image

Image

Text

Image

Text

Text

Matching?

Matching?

CLIP

Matching?

auto.gluon.ai

26 of 38

Multimodal Matching

26

Source: CLIP paper

Directly running inference

Finetune on relevance data

- Negative Sampling

- Metric Learning

auto.gluon.ai

27 of 38

Foundation Models are Bigger

27

  • Train: Parameter-efficient finetuning
  • Deployment: Model distillation

auto.gluon.ai

28 of 38

Parameter-efficient finetuning

28

 

 

 

 

 

 

 

 

 

Image credit from LoRA [3]

 

Freezing backbone + Adapter

Effective for training large language models [2, 3]

Able to train FLAN-T5-XL [1] with single GPU-Instance (AWS g4dn.2x).

Parameter-Efficient Finetuning Tutorial

[1] Chung, Hyung Won, et al. "Scaling Instruction-Finetuned Language Models." arXiv preprint arXiv:2210.11416 (2022).

[2] Chen, Jiao, et al. “Parameter-Efficient Finetuning Design Spaces” In submission (2022).

[3] Hu, Edward J., et al. "Lora: Low-rank adaptation of large language models." ICLR 2021.

auto.gluon.ai

29 of 38

Top-1 in RAFT Leaderboard with AutoGluon Multimodal

Top-1 solution

  • FLAN-T5-XXL (11B)
  • Combining LoRA [1] and IA^3 [2] (“ia3_lora”)
  • Learning rate 1E-3

29

[1] Hu, Edward J., et al. "Lora: Low-rank adaptation of large language models." ICLR 2021.

[2] Liu, Haokun, et al. "Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning." NeurIPS 2022.

More details: link

Human: 5th

GPT-3 (175B): 16th

Ours (11B): 1st

auto.gluon.ai

30 of 38

Model Distillation

30

Student Model

Teacher Model (Large-scale)

[1] Fakoor, Rasool, et al. "Fast, accurate, and simple models for tabular data via augmented distillation." NeurIPS 2020

[2] He, Haoyu, et al. "Towards Automated Distillation: A Systematic Study of Knowledge Distillation in Natural Language Processing.” AutoML-Conf 2022 Late-Breaking Workshop

Augmented distillation [1]

Intermediate feature-level distillation

Augmenter

Data augmentation and intermediate distillation are essential to the performance [2].

auto.gluon.ai

31 of 38

Model Distillation

31

auto.gluon.ai

32 of 38

Hyper-parameter Optimization – Random Search

32

A huge choice space: foundation models, training strategies, etc. Use hyper-parameter optimization

Image Source: (Feurer and Hutter, 2018)

auto.gluon.ai

33 of 38

Hyper-parameter Optimization – Bayesian Optimization

33

Image Source: (Feurer and Hutter, 2018)

  • Probabilistic surrogate: p(loss | hyperparameters), and sequentially pick the next hyperparameters according to the best expected improvement.

auto.gluon.ai

34 of 38

Hyper-parameter Optimization – Multi-fidelity Method

34

Source: Hyperband [1]

  • Successive halving: Train with limited budget. Throw away the worst performing half and double the budget for the remaining ones.
  • Hyperband: Portfolio of successive halving

[1] Li, Lisha, et al. "Hyperband: A novel bandit-based approach to hyperparameter optimization." The Journal of Machine Learning Research 18.1 (2017): 6765-6816.

auto.gluon.ai

35 of 38

Try our hands-on notebooks!

  • Five notebooks
    • Classification
    • Object Detection
    • Named Entity Recognition
    • Text-Text Matching
    • Text-Image Matching

35

NeurIPS workshop Website

auto.gluon.ai

36 of 38

Thanks! Q&A

Get started with `pip install autogluon`

Github: github.com/autogluon/autogluon

Open-source project. Welcome to contribute!

36

AutoGluon Website

NeurIPS workshop Website

auto.gluon.ai

37 of 38

Backup Slides

37

auto.gluon.ai

38 of 38

Data Augmentation (Pick one randomly)

38

auto.gluon.ai