Multimodal and Contrastive Learning
AIDA Technical workshop
Erik Ylipää
Outline
Has Deep Learning solved Computer Vision?
3
Betteridge's law of headlines: "Any headline that ends in a question mark can be answered by the word no."
Computer vision using Deep Learning
4
Goodfellow, Ian J., Jonathon Shlens, and Christian Szegedy. "Explaining and harnessing adversarial examples." arXiv preprint arXiv:1412.6572 (2014).
https://www.eff.org/files/AI-progress-metrics.html
Note, “human performance” has N=1: Andrej Karpathy
http://karpathy.github.io/2014/09/02/what-i-learned-from-competing-against-a-convnet-on-imagenet/
Excellent generalization to a test set
Brittle in surprising ways (adversarial examples)
What features are used?
The current incarnation of deep neural networks exhibit a tendency to learn surface statistical regularities as opposed to higher level abstractions in the dataset. For tasks such as object recognition, due to the strong statistical properties of natural images, these superficial cues that the deep neural network have learned are sufficient for high performance generalization, but in a narrow distributional sense.
Jo, Jason, and Yoshua Bengio. "Measuring the tendency of cnns to learn surface statistical regularities." arXiv preprint arXiv:1711.11561 (2017).
What features are used?
Wang, Haohan, et al. "High-frequency component helps explain the generalization of convolutional neural networks." Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.
Learning Representations using Supervised Learning
Representation space
Representation Learning
8
Artificial Intelligence (AI) and Deep Learning
9
Deep Learning
Example: Recurrent Neural Networks
Representation Learning
Exemple: Word vectors (Word2Vec)
Machine Learning
Exemple: Random Forest
Artificial Intelligence
Example: Expert system
Bengio, Yoshua, Ian Goodfellow, and Aaron Courville. Deep learning. Vol. 1. MIT press, 2017. https://www.deeplearningbook.org/
Deep Learning - Complex data can be modeled as a hierarchy of simpler factors through learning
10
Salakhutdinov, Ruslan, and Hugo Larochelle. "Efficient learning of deep Boltzmann machines." Proceedings of the thirteenth international conference on artificial intelligence and statistics. JMLR Workshop and Conference Proceedings, 2010.
Representation is fundamental for AI
11
Representation Learning
Transfer learning
Representation Learning can be thought of as an “inverse” problem - example inverse rendering
14
Yildirim, Ilker, et al. "Efficient inverse graphics in biological face processing." Science advances 6.10 (2020): eaax5979. https://www.science.org/doi/10.1126/sciadv.aax5979
Dynamically computed “layers”
15
Focus on the head of the model
16
Typical focus for CNNs for computer vision
Our focus for now
Intuiting the dot-product
17
Highlighting the dot product
18
A softmax “layer”
19
MNIST Example
20
Dot product and angles
21
Slight detour - Capacity of high dimensional space
22
For a deeper dive, see the first chapter of Blum, A., Hopcroft, J., & Kannan, R. (2020). Foundations of Data Science. Cambridge: Cambridge University Press. doi:10.1017/9781108755528. Preprint online version here.
Neural Networks solves problems by aligning vectors
23
Softmax regression and the dot product
24
Notebook 1
Contrastive Learning
Let’s do away with the softmax vectors
27
What can we do with this “non-parametric” softmax?
28
Wu, Zhirong, et al. "Unsupervised feature learning via non-parametric instance discrimination." Proceedings of the IEEE conference on computer vision and pattern recognition. 2018.
Triplet Loss
29
Contrastive Learning
30
Tian, Yonglong, Dilip Krishnan, and Phillip Isola. "Contrastive multiview coding." arXiv preprint arXiv:1906.05849 (2019).
Multimodal Learning
31
The way in which something happens or is experienced.
https://sites.google.com/site/multiml2016cvpr/MMML-Tutorial-Part1.pdf?attredirects=0
Multimodal Learning
32
Arandjelovic, Relja, and Andrew Zisserman. "Look, listen and learn." Proceedings of the IEEE International Conference on Computer Vision. 2017.
Multimodal learning - CLIP
33
https://openai.com/blog/clip/
Notes on training: Full dataset is 400 million (text, image) pairs scraped from the internet. Each training batch contains ~32000 images.
Multimodal learning - CLIP
34
Radford, Alec, et al. "Learning transferable visual models from natural language supervision." International Conference on Machine Learning. PMLR, 2021.
The problem of annotated data
35
Notebook 2
36
Self-supervised Learning
37
Supervised vs Unsupervised learning
38
Self-supervised learning through prediction
39
Yann LeCun, https://www.youtube.com/watch?v=7I0Qt7GALVk
Self-supervised learning through prediction
40
The Autoencoder
41
Issues with autoencoders�reconstructing pixels might learn the wrong thing
42
Bengio, Yoshua, Ian Goodfellow, and Aaron Courville. Deep learning. Vol. 1, page 542. MIT press, 2017. https://www.deeplearningbook.org/
Gato
43
Reed, Scott, et al. "A Generalist Agent." arXiv preprint arXiv:2205.06175 (2022).
Self-supervision and Contrastive Learning
44
https://ai.googleblog.com/2020/04/advancing-self-supervised-and-semi.html
Gordon, Daniel, et al. "Watching the world go by: Representation learning from unlabeled videos." arXiv preprint arXiv:2003.07990 (2020).
Thank you!
erik.ylipaa@scilifelab.se