1
Agenda | Duration (CST) |
Introduction + AutoGluon Tabular | 2:00PM – 2:55PM |
Break | 2:55PM – 3:05PM |
AutoGluon Multimodal | 3:05PM – 4:00PM |
Break | 4:00PM – 4:10PM |
AutoGluon Timeseries | 4:10PM – 4:50PM |
Additional Q&A + Feedback | 4:50PM – 5:00PM |
Workshop Website
*Note on Hands-on Notebooks
© 2022, Amazon Web Services, Inc. or its Affiliates. All rights reserved.
auto.gluon.ai
Xingjian Shi, Senior Applied Scientist, AWS
Yi Zhu, Senior Applied Scientist, AWS
AutoGluon Multimodal
2
© 2022, Amazon Web Services, Inc. or its Affiliates. All rights reserved.
auto.gluon.ai
Real-life Multimodal Problems: Funding or not?
3
Source: https://www.kickstarter.com/
Fund raising platform aims to bring creative projects to life.
Predict whether the project will get funded?
Source: Kaggle: Kickstarter Funding
auto.gluon.ai
Real-life Multimodal Problems: Funding or not?
4
Source: Kaggle: Kickstarter Funding
Text
Categorical
Numerical
Multimodal features!
auto.gluon.ai
Real-life Multimodal Problems: Pet Adoption
5
Source: https://www.petfinder.com/
Online pet adoption platform
Predict the adoption speed.
auto.gluon.ai
Real-life Multimodal Problems: Pet Adoption
6
Subject Focus | 0 |
Eyes | 1 |
Face | 1 |
Near | 1 |
Action | 0 |
Accessory | 0 |
Group | 1 |
… | |
Image
Tabular meta-data
Pawpularity=63.
(Range: 0 – 100)
Source: Kaggle: Pet Pawpularity Score
Multimodal features!
auto.gluon.ai
Real-life Multimodal Problems: Product Matching
7
Source: https://www.amazon.com/
Product Image
Title (Text)
Score (Numerical)
Price (Numerical)
Description (Text)
auto.gluon.ai
Real-life Multimodal Problems: Product Matching
8
Product Image
Title (Text)
Price (Numerical)
Color (Categorical)
Same!
Uploaded image
auto.gluon.ai
Real-life Multimodal Problems: Product Matching
9
Product Image
Title (Text)
Price (Numerical)
Color (Categorical)
Different!
Uploaded Image
auto.gluon.ai
AutoGluon – Multimodal AutoML Toolkit
10
Image
Text
Tabular
State-of-the-art multimodal AutoML Toolkit on Image, Text, Time-series and Tabular data.
1. Data Processing
2. Training
3. Deployment
Huge science/engineering cost!
Train models in three lines of codes
Iterative + Expensive
auto.gluon.ai
Input Format of AutoGluon Multimodal – DataFrame
11
Text
Image
Categorical
Numerical
Support other types like bounding boxes, named entities
auto.gluon.ai
Train Model with Three Lines of Code
12
OpenAI/CLIP
Automated Model Ensemble (discussed in part 1)
Support pretrained “Foundation Models”
Focus of this section
auto.gluon.ai
Problems Supported in AutoGluon Multimodal
13
auto.gluon.ai
Technical Details Underlying These Three Lines
14
auto.gluon.ai
Foundation Models of Image / Text
15
Text foundation models
Image foundation models
Foundation Models: Trained on large-scale datasets, able to support lots of applications [1].
[1] Bommasani, Rishi, et al. "On the opportunities and risks of foundation models." arXiv preprint arXiv:2108.07258 (2021).
auto.gluon.ai
Foundation Models of Image / Text
16
Source: https://huggingface.co/models. Taken on 2022/08/10
Source: https://github.com/rwightman/pytorch-image-models/blob/master/results/results-imagenet.csv. Taken on 2022/08/10
Large number of public foundation models (>60000 in huggingface/transformers, >600 in TIMM)!
The idea of model zoo is not widely adopted in AutoML toolkits. AutoGluon Multimodal focus on integrating/fusion of common model zoos.
auto.gluon.ai
Fusion Foundation Models of Image/Text
17
Image
Text
Tabular data
FT transformer [1]
Task prediction
OpenAI/CLIP
Fusion + Task-specific heads
[1] Gorishniy, Yury, et al. "Revisiting deep learning models for tabular data.” NeurIPS 2021
auto.gluon.ai
Fusion Foundation Model is a Powerful Solution
18
Fusion text pretrained model + tabular feature: Score equivalent to 2nd/2380 Teams
$100,000�Total Prize
auto.gluon.ai
Fusion Foundation Model is a Powerful Solution
19
The model ranks 20th place (out of 3537 teams), top 1%
Ensemble with different foundation models
Kaggle: Petfinder Pawpularity
auto.gluon.ai
Model Ensemble
20
Foundation Models
+
Classic Tabular Models
How to combine them and achieve the best of both worlds?
auto.gluon.ai
Model Ensemble
21
a) Feature extraction
b) Weighted ensemble
c) Stack ensemble
Stack ensemble > Weighted ensemble >> Feature extraction
[1] Shi, Xingjian, et al. "Benchmarking multimodal automl for tabular data with text fields." NeurIPS 2021 Track on Datasets and Benchmarks.
Conducted experiments on a benchmark with 18 multimodal datasets.
auto.gluon.ai
Ensemble Foundation Model and Tree Models: Top-1 in 3 Lines
22
Winning formula
Source: https://machinehack.com/hackathons/product_sentiment_classification_weekend_hackathon_19/overview
auto.gluon.ai
Ensemble Foundation Model and Tree Models: Top-1 in 3 Lines
23
Source: https://machinehack.com/hackathons/predict_the_data_scientists_salary_in_india_hackathon/leaderboard
auto.gluon.ai
Object Detection – Task-specific Model Zoo
24
Example: Detect potholes for better road safety. Image from Kaggle Pothole Dataset
auto.gluon.ai
Multimodal Matching
Twin-tower Architecture
25
Image
Image
Text
Image
Text
Text
Matching?
Matching?
CLIP
Matching?
auto.gluon.ai
Multimodal Matching
26
Source: CLIP paper
Directly running inference
Finetune on relevance data
- Negative Sampling
- Metric Learning
auto.gluon.ai
Foundation Models are Bigger
27
auto.gluon.ai
Parameter-efficient finetuning
28
Image credit from LoRA [3]
Freezing backbone + Adapter
Effective for training large language models [2, 3]
Able to train FLAN-T5-XL [1] with single GPU-Instance (AWS g4dn.2x).
[1] Chung, Hyung Won, et al. "Scaling Instruction-Finetuned Language Models." arXiv preprint arXiv:2210.11416 (2022).
[2] Chen, Jiao, et al. “Parameter-Efficient Finetuning Design Spaces” In submission (2022).
[3] Hu, Edward J., et al. "Lora: Low-rank adaptation of large language models." ICLR 2021.
auto.gluon.ai
Top-1 in RAFT Leaderboard with AutoGluon Multimodal
Top-1 solution
29
Source: https://huggingface.co/spaces/ought/raft-leaderboard. Taken on 2022/11/26
[1] Hu, Edward J., et al. "Lora: Low-rank adaptation of large language models." ICLR 2021.
[2] Liu, Haokun, et al. "Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning." NeurIPS 2022.
More details: link
Human: 5th
GPT-3 (175B): 16th
Ours (11B): 1st
auto.gluon.ai
Model Distillation
30
Student Model
Teacher Model (Large-scale)
[1] Fakoor, Rasool, et al. "Fast, accurate, and simple models for tabular data via augmented distillation." NeurIPS 2020
[2] He, Haoyu, et al. "Towards Automated Distillation: A Systematic Study of Knowledge Distillation in Natural Language Processing.” AutoML-Conf 2022 Late-Breaking Workshop
Augmented distillation [1]
Intermediate feature-level distillation
Augmenter
Data augmentation and intermediate distillation are essential to the performance [2].
auto.gluon.ai
Model Distillation
31
auto.gluon.ai
Hyper-parameter Optimization – Random Search
32
A huge choice space: foundation models, training strategies, etc. Use hyper-parameter optimization
Image Source: (Feurer and Hutter, 2018)
HPO Tutorial. Based on ray[tune].
auto.gluon.ai
Hyper-parameter Optimization – Bayesian Optimization
33
HPO Tutorial. Based on ray[tune].
Image Source: (Feurer and Hutter, 2018)
auto.gluon.ai
Hyper-parameter Optimization – Multi-fidelity Method
34
Source: Hyperband [1]
[1] Li, Lisha, et al. "Hyperband: A novel bandit-based approach to hyperparameter optimization." The Journal of Machine Learning Research 18.1 (2017): 6765-6816.
HPO Tutorial. Based on ray[tune].
auto.gluon.ai
Try our hands-on notebooks!
35
NeurIPS workshop Website
auto.gluon.ai
Thanks! Q&A
Get started with `pip install autogluon`
Github: github.com/autogluon/autogluon
Open-source project. Welcome to contribute!
36
AutoGluon Website
NeurIPS workshop Website
auto.gluon.ai
Backup Slides
37
auto.gluon.ai
Data Augmentation (Pick one randomly)
38
auto.gluon.ai