Learning Transferable Visual Models From Natural Language Supervision
By Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger and Ilya Sutskever
Presented by Balaji Balasubramanian and Eshwanth Baskaran
Introduction
L. Jing and Y. Tian, ‘‘Self-supervised visual feature learning with deep neural networks: A survey,’’
CLIP - What is it?
CLIP - Motivation
CLIP - Dataset curation
Only around 100,000 images
Noisy and sparse metadata - filtering reduces data by 6x
MS-COCO
YFCC100M
CLIP - Dataset curation
CLIP - Problem formulation
Generative model
Predict
Input
Features
Features
Text encoder
Image encoder
Contrastive loss
Input
Input
Predict exact words…
HARD TASK
MORE COMPUTE
Predict which text pairs with which image
EASIER TASK
LESS COMPUTE
EFFECTIVE
VS
Tian, Y., Krishnan, D., and Isola, P. Contrastive
multiview coding. arXiv preprint arXiv:1906.05849, 2019.
CLIP - Architecture
CLIP - Models
Visual Transformer (ViT) - Architecture
CLIP vision transformers are about 3x more compute efficient than CLIP ResNets
Text Encoder - GPT-2
CLIP - Training
1
2
3
4
Training - Temperature parameter
https://openaccess.thecvf.com/content_cvpr_2018/CameraReady/0801.pdf
CLIP - Training configuration
Want the minibatch size to be larger, but hardware constraints
CLIP - Test phase (zero-shot)
Test image
Test categories
Prediction
CLIP - Comparison with classification models
CLIP
Classifier
Input
CLIP as a linear classifier
Related works- Virtex
Desai et al.- VirTex: Learning Visual Representations from Textual Annotations
Related works- ConVIRT
Yuhao Zhang et al- Contrastive Learning of Medical Visual Representations from Paired Images and Text
Related works- ICMLM
Mert Bulent Sariyildiz et al- Learning Visual Representations with Caption Annotations
Experiments
Prompt Engineering and Ensembling
Experiments
Zero Shot CLIP vs Linear Probe on Resnet 50
Experiments
Zero Shot CLIP vs Linear Probe on Resnet 50
Experiments
Zero Shot CLIP vs Linear Probe on Resnet 50
Question
Question
Yuan et al.- Florence: A New Foundation Model for Computer Vision
Liu et al.- Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
Experiments
CLIP with Linear Probe
Experiments
CLIP with Linear Probe
Experiments
CLIP with Linear Probe
Xie et al. - Self-training with Noisy Student improves ImageNet classification
Experiments
Compute NSL for CLIP
Experiments
Compute NSL for CLIP
Experiments
Experiments
Broader Impacts
Bias
K. Kärkkäinen and J. Joo, “FairFace: Face attribute dataset for
balanced race, gender, and age,” Aug. 2019.
Bias
K. Kärkkäinen and J. Joo, “FairFace: Face attribute dataset for
balanced race, gender, and age,” Aug. 2019.
Mahajan et al. : Exploring the Limits of Weakly Supervised Pretraining
Race, Age, Gender Classification for Non-White Categories
Applications- Image Search
https://mikhalevi.ch/rclip-an-ai-powered-command-line-photo-search-tool/
Applications- Image Similarity
https://blog.roboflow.com/openai-clip/
Applications- Deciphering Corrupted Images
Sriram et al. - Inverse Problems Leveraging Pre-trained Contrastive Representations
CLIP - How is it better?
CLIP - Limitations
CLIP - Limitations
3. Poor generalization to images not covered in its pre-training dataset.
4. Zero-shot classification - sensitive to wording or phrasing
CLIP - Limitations
5. Typographic attacks
https://arxiv.org/pdf/2103.10480.pdf
Conclusion