Object Detection
with Deep Learning
CVPR 2014 Tutorial
Pierre Sermanet, Google Research
Tutorial on Deep Learning for Vision, CVPR 2014 June 23, 2014
What is object detection?
difficulty
Why is object detection important?
Is it deployed?
What datasets for detection?
The pascal visual object classes (voc) challenge. Everingham, Mark, Luc Van Gool, Christopher KI Williams, John Winn, and Andrew Zisserman. International journal of computer vision 88, no. 2 (2010): 303-338.
Imagenet: A large-scale hierarchical image database. Deng, Jia, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. In Computer Vision and Pattern Recognition, 2009. CVPR 2009. IEEE Conference on, pp. 248-255. IEEE, 2009.
SUN Database: Large-scale Scene Recognition from Abbey to Zoo, Jianxiong Xiao, James Hays, Krista Ehinger, Aude Oliva, and Antonio Torralba.IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2010.
Microsoft COCO: Common Objects in Context, Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, C. Lawrence Zitnick, http://arxiv.org/abs/1405.0312, May 2014
| # classes | average # categories per image | average # instances per image | average object scale | average resolution | # images | # objects | ||||||
total | train | val | test | total | train | val | test | ||||||
PASCAL | 20 | 1.521 | 2.711 | 0.207 | 469x387 | 22k | 6k | 6k | 10k | 42k? | 14k | 14k | - |
ImageNet13 | 200 | 1.534 | 2.758 | 0.170 | 482x415 | 516k | 456k | 20k | 40k | 648k? | 480k | 56k | - |
Sun | 4919 | 9.8 | 16.9 | 0.1040 | 732x547 | 16873 | - | 285k | - | ||||
COCO | 91 | 3.5 | 7.6 | 0.117 | 578x483 | 328k | 164k | 82k | 82k | 2500k | ~1250k | ~625k | ~625k |
What datasets for detection?
Microsoft COCO: Common Objects in Context, Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, C. Lawrence Zitnick, http://arxiv.org/abs/1405.0312, May 2014
What datasets for detection?
Microsoft COCO: Common Objects in Context, Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, C. Lawrence Zitnick, http://arxiv.org/abs/1405.0312, May 2014
Recent history of object detection
VOC’10
VOC’12
VOC’07
VOC’09
VOC’08
VOC’11
Rich feature hierarchies for accurate object detection and semantic segmentation. Girshick, Ross, Jeff Donahue, Trevor Darrell, and Jitendra Malik. arXiv preprint arXiv:1311.2524 (2013).
The PASCAL Visual Object Classes Challenge - a Retrospective, Everingham, M., Eslami, S. M. A., Van Gool, L., Williams, C. K. I., Winn, J. and Zisserman, A. Accepted for International Journal of Computer Vision, 2014
DPM
DPM,
HOG+BOW
DPM,
MKL
DPM++
DPM++,
MKL,
Selective
Search
Selective
Search,
DPM++,
MKL
41%
41%
37%
28%
23%
17%
SegDPM (2013)
Regionlets (2013)
Regionlets
(2013)
R-CNN
58.5%
R-CNN
53.7%
R-CNN
53.3%
ConvNets (2014)
Recent history of object detection
Rich feature hierarchies for accurate object detection and semantic segmentation. Girshick, Ross, Jeff Donahue, Trevor Darrell, and Jitendra Malik. arXiv preprint arXiv:1311.2524 (2013).
Overfeat: Integrated recognition, localization and detection using convolutional networks. Sermanet, Pierre, David Eigen, Xiang Zhang, Michaël Mathieu, Rob Fergus, and Yann LeCun. arXiv preprint arXiv:1312.6229 (2013), International Conference on Learning Representations (ICLR)` 2014.
No ConvNets
Deformable Parts Models (DPM) ->
ConvNets breakthroughs for visual tasks
| Dataset | Performance | Score |
[Sermanet et al 2014]: OverFeat (fine-tuned features for each task) (tasks are ordered by increasing difficulty) | |||
| ImageNet LSVRC 2013 Dogs vs Cats Kaggle challenge 2014 ImageNet LSVRC 2013 ImageNet LSVRC 2013 | competitive state of the art state of the art competitive | 13.6 % error 98.9% 29.9% error 24.3% mAP |
[Razavian et al, 2014]: public OverFeat library (no retraining) + SVM (simplest approach possible on purpose, no attempt at more complex classifiers) (tasks are ordered by “distance” from classification task on which OverFeat was trained) | |||
(search by image similarity) | Pascal VOC 2007 MIT-67 Caltech-UCSD Birds 200-2011 Oxford 102 Flowers UIUC 64 object attributes H3D Human Attributes Oxford 5k buildings Paris 6k buildings Sculp6k Holidays UKBench | competitive state of the art competitive state of the art state of the art competitive state of the art state of the art competitive state of the art state of the art | 77.2% mAP 69% mAP 61.8% mAP 86.8% mAP 91.4% mAUC 73% mAP 68% mAP? 79.5% mAP? 42.3% mAP? 84.3% mAP? 91.1% mAP? |
Pierre Sermanet, David Eigen, Xiang Zhang, Michael Mathieu, Rob Fergus, Yann LeCun, OverFeat: Integrated Recognition, Localization and Detection using Convolutional Networks, http://arxiv.org/abs/1312.6229, ICLR 2014 Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, Stefan Carlsson, CNN Features off-the-shelf: an Astounding Baseline for Recognition, http://arxiv.org/abs/1403.6382, DeepVision CVPR 2014 workshop | |||
ConvNets breakthroughs for visual tasks
| Dataset | Performance | Score |
[Zeiler et al 2013]
| ImageNet LSVRC 2013 Caltech-101 (15, 30 samples per class) Caltech-256 (15, 60 samples per class) Pascal VOC 2012 | state of the art competitive state of the art competitive | 11.2% error 83.8%, 86.5% 65.7%, 74.2% 79% mAP |
[Donahue et al, 2014]: DeCAF+SVM
| Caltech-101 (30 classes) Amazon -> Webcam, DSLR -> Webcam Caltech-UCSD Birds 200-2011 SUN-397 | state of the art state of the art state of the art competitive | 86.91% 82.1%, 94.8% 65.0% 40.9% |
[Girshick et al, 2013]
| Pascal VOC 2007 Pascal VOC 2010 (comp4) ImageNet LSVRC 2013 Pascal VOC 2011 (comp6) | state of the art state of the art state of the art state of the art | 48.0% mAP 43.5% mAP 31.4% mAP 47.9% mAP |
[Oquab et al, 2013]
| Pascal VOC 2007 Pascal VOC 2012 Pascal VOC 2012 (action classification) | state of the art state of the art state of the art | 77.7% mAP 82.8% mAP 70.2% mAP |
M.D. Zeiler, R. Fergus, Visualizing and Understanding Convolutional Networks, Arxiv 1311.2901 http://arxiv.org/abs/1311.2901 J. Donahue, Y. Jia, O. Vinyals, J. Hoffman, N. Zhang, E. Tzeng, and T. Darrell. Decaf: A deep convolutional activation feature for generic visual recognition. In ICML, 2014, http://arxiv.org/abs/1310.1531 R. B. Girshick, J. Donahue, T. Darrell, and J. Malik. Rich feature hierarchies for accurate object detection and semantic segmentation. arxiv:1311.2524 [cs.CV], 2013, http://arxiv.org/abs/1311.2524 M. Oquab, L. Bottou, I. Laptev, and J. Sivic. Learning and transferring mid-level image representations using convolutional neural networks. Technical Report HAL-00911179, INRIA, 2013. http://hal.inria.fr/hal-00911179 | |||
ConvNets breakthroughs for visual tasks
| Dataset | Performance | Score |
[Khan et al 2014]
| UCF CMU UIUC | state of the art state of the art state of the art | 90.56% 88.79% 93.16% |
[Sander Dieleman, 2014]
| Kaggle Galaxy Zoo challenge | state of the art | 0.07492 |
S. H. Khan, M. Bennamoun, F. Sohel, R. Togneri. Automatic Feature Learning for Robust Shadow Detection, CVPR 2014 Sander Dieleman, Kaggle Galaxy Zoo challenge 2014 http://benanne.github.io/2014/04/05/galaxy-zoo.html | |||
ConvNets breakthroughs for visual tasks
[Razavian et al, 2014]:
"It can be concluded that from now on, deep learning with CNN has to be considered as the primary candidate in essentially any visual recognition task."
Ali Sharif Razavian, Hossein Azizpour, Josephine Sullivan, Stefan Carlsson, CNN Features off-the-shelf: an Astounding Baseline for Recognition, http://arxiv.org/abs/1403.6382, DeepVision CVPR 2014 workshop
Perceptron
Neural Net
Boosting
SVM
GMM
SP
BayesNP
Convolutional
Neural Net
Recurrent Neural Net
Autoencoder Neural Net
Sparse Coding
Restricted BM
Deep Belief Net
Deep (sparse/denoising) Autoencoder
UNSUPERVISED
SUPERVISED
DEEP
SHALLOW
Slide: M. Ranzato
History of detection with ConvNets
LeCun, Huang, Bottou 2004
NORB dataset
Cireşan et al. 2013
Mitosis detection
Sermanet et al. 2013
Pedestrian detection
Vaillant, Monrocq, LeCun 1994
Osadchy, LeCun, Miller 2004
Face detection with pose estimation!
PASCAL detection
Girshick et al. 2013
Szegedy, Toshev, Erhan 2013
ImageNet detection
Girshick et al. 2014 (R-CNN)
Sermanet et al. 2014 (OverFeat)
slide: Girshick/Sermanet
History of ConvNets
slide: Girshick
Fukushima 1980
Neocognitron
LeCun et al. 1989-1998
Hand-written digit reading
Rumelhart, Hinton, Williams 1986
“T” versus “C” problem
...
Krizhevksy, Sutskever, Hinton 2012
ImageNet classification breakthrough
“SuperVision” CNN
Pedestrian detection with ConvNets (video)
What are the weakest links in detection?
Finding the weakest link in person detectors. Parikh, Devi, and C. Lawrence Zitnick. Computer Vision and Pattern Recognition (CVPR), 2011 IEEE Conference on. IEEE, 2011.
ConvNets vs Primates
Deep Neural Networks Rival the Representation of Primate IT Cortex for Core Visual Object Recognition. Cadieu, Charles F., Ha Hong, Daniel LK Yamins, Nicolas Pinto, Diego Ardila, Ethan A. Solomon, Najib J. Majaj, and James J. DiCarlo. arXiv preprint arXiv:1406.3284 (2014).
Learned convolutional filters: Stage 1
Visualizing and understanding convolutional neural networks. Zeiler, Matthew D., and Rob Fergus. arXiv preprint arXiv:1311.2901 (2013).
9 patches with strongest
activation
learned filters
(7x7x96)
Strongest activations: Stage 2
Visualizing and understanding convolutional neural networks. Zeiler, Matthew D., and Rob Fergus. arXiv preprint arXiv:1311.2901 (2013).
Strongest activations: Stage 3
Visualizing and understanding convolutional neural networks. Zeiler, Matthew D., and Rob Fergus. arXiv preprint arXiv:1311.2901 (2013).
Strongest activations: Stage 4
Visualizing and understanding convolutional neural networks. Zeiler, Matthew D., and Rob Fergus. arXiv preprint arXiv:1311.2901 (2013).
Strongest activations: Stage 5
Visualizing and understanding convolutional neural networks. Zeiler, Matthew D., and Rob Fergus. arXiv preprint arXiv:1311.2901 (2013).
What is a ConvNet?
ConvNets 2.0
Imagenet classification with deep convolutional neural networks. Krizhevsky, Alex, Ilya Sutskever, and Geoffrey E. Hinton. Advances in neural information processing systems. 2012.
Why are ConvNets good for detection?
ImageNet pre-training
CNN Features off-the-shelf: an Astounding Baseline for Recognition. Razavian, Ali Sharif, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson. arXiv preprint arXiv:1403.6382 (2014).
ImageNet pre-training
Bird Species Categorization Using Pose Normalized Deep Convolutional Nets. Branson, Steve, Grant Van Horn, Serge Belongie, and Pietro Perona. arXiv preprint arXiv:1406.2952 (2014).
How much does fine-tuning matter?
| VOC 2007 | VOC 2010 |
Regionlets (Wang et al. 2013) | 41.7% | 39.7% |
SegDPM (Fidler et al. 2013) | | 40.4% |
R-CNN pool5 | 44.2% | |
R-CNN fc6 | 46.2% | |
R-CNN fc7 | 44.7% | |
R-CNN FT pool5 | 47.3% | |
R-CNN FT fc6 | 53.1% | |
R-CNN FT fc7 | 54.2% | 50.2% |
metric: mean average precision (higher is better)
fine-tuned
CNN Features off-the-shelf: an Astounding Baseline for Recognition. Razavian, Ali Sharif, Hossein Azizpour, Josephine Sullivan, and Stefan Carlsson. arXiv preprint arXiv:1403.6382 (2014).
Rich feature hierarchies for accurate object detection and semantic segmentation. Girshick, Ross, Jeff Donahue, Trevor Darrell, and Jitendra Malik. arXiv preprint arXiv:1311.2524 (2013).
R-CNN: Regions with CNN features
Input
image
Extract region
proposals (~2k / image)
Rich feature hierarchies for accurate object detection and semantic segmentation. Girshick, Ross, Jeff Donahue, Trevor Darrell, and Jitendra Malik. arXiv preprint arXiv:1311.2524 (2013).
R-CNN: bounding-box regression
Rich feature hierarchies for accurate object detection and semantic segmentation. Girshick, Ross, Jeff Donahue, Trevor Darrell, and Jitendra Malik. arXiv preprint arXiv:1311.2524 (2013).
Linear regression
on CNN features
Original
proposal
Predicted
object bounding box
Objectness / Selective Search
(old idea used for detection of faces, traffic signs, etc)
van de Sande, K. E., Uijlings, J. R., Gevers, T., & Smeulders, A. W. (2011, November). Segmentation as selective search for object recognition. ICCV’11
Dumitru Erhan, Christian Szegedy, Alexander Toshev, Dragomir Anguelov, Scalable Object Detection using Deep Neural Networks, CVPR’14
How to debug object detection?
Diagnosing error in object detectors. Hoiem, Derek, Yodsawalai Chodpathumwan, and Qieyun Dai. Computer Vision–ECCV 2012. Springer Berlin Heidelberg, 2012. 340-353.
Rich feature hierarchies for accurate object detection and semantic segmentation. Girshick, Ross, Jeff Donahue, Trevor Darrell, and Jitendra Malik. arXiv preprint arXiv:1311.2524 (2013).
no fine-tuning fine-tuning b-box regression
OverFeat: dense detection
OverFeat: dense detection
OverFeat: dense detection
Localization regression
Pose regression with Convnet
[Osadchy’04]:
Detection: localization voting for multiple objects
Augmenting sliding density of a ConvNet
Augmenting sliding density of a ConvNet
A. Giusti, D. C. Ciresan, J. Masci, L. M. Gambardella, and J. Schmidhuber. Fast image scanning with deep max-pooling convolutional neural networks. In International Conference on Image Processing (ICIP), 2013.
Online bootstrapping
How do R-CNN and OverFeat differ?
Rich feature hierarchies for accurate object detection and semantic segmentation. Girshick, Ross, Jeff Donahue, Trevor Darrell, and Jitendra Malik. arXiv preprint arXiv:1311.2524 (2013).
Overfeat: Integrated recognition, localization and detection using convolutional networks. Sermanet, Pierre, David Eigen, Xiang Zhang, Michaël Mathieu, Rob Fergus, and Yann LeCun. arXiv preprint arXiv:1312.6229 (2013), International Conference on Learning Representations (ICLR)` 2014.
Warping
Rich feature hierarchies for accurate object detection and semantic segmentation. Girshick, Ross, Jeff Donahue, Trevor Darrell, and Jitendra Malik. arXiv preprint arXiv:1311.2524 (2013).
Warping
Bird Species Categorization Using Pose Normalized Deep Convolutional Nets. Branson, Steve, Grant Van Horn, Serge Belongie, and Pietro Perona. arXiv preprint arXiv:1406.2952 (2014).
ImageNet detection examples
ImageNet detection examples: occlusions
OverFeat • Pierre Sermanet • New York University
ImageNet detection examples: occlusions
ImageNet detection examples
ImageNet detection failures that make sense
ImageNet detection failures that make sense
ImageNet detection: some hard ones
Conclusions
Questions
Additional Information
Pedestrian detection with ConvNets
`
Pedestrian detection with unsupervised multi-stage feature learning. Sermanet, Pierre, Koray Kavukcuoglu, Soumith Chintala, and Yann LeCun. In Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on, pp. 3626-3633. IEEE, 2013.
Unsupervised learning for detection
Sparse Coding
Convolutional Sparse Coding (CPSD)
Pedestrian detection with unsupervised multi-stage feature learning. Sermanet, Pierre, Koray Kavukcuoglu, Soumith Chintala, and Yann LeCun. In Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on, pp. 3626-3633. IEEE, 2013.
Multi-stage features (skip layer)
Pedestrian detection with unsupervised multi-stage feature learning. Sermanet, P., Kavukcuoglu, K., Chintala, S., & LeCun, Y. (2013, June). InComputer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on(pp. 3626-3633). IEEE.
Traffic sign recognition with multi-scale convolutional networks. Sermanet, Pierre, and Yann LeCun. Neural Networks (IJCNN), The 2011 International Joint Conference on. IEEE, 2011.
Convolutional neural networks applied to house numbers digit classification. Sermanet, Pierre, Soumith Chintala, and Yann LeCun. Pattern Recognition (ICPR), 2012 21st International Conference on. IEEE, 2012.
Deep Learning Face Representation from Predicting 10,000 Classes. Sun, Yi, Xiaogang Wang, and Xiaoou Tang.
Dogs vs Cats: results and approaches
Team | Error % | Software | Approach | ImageNet pre-training | deep learning |
Pierre Sermanet | 1.09 | OverFeat | 7 models average + multi-scale + drop fully connected layers | yes | yes |
anton (post competition) | 1.68 | OverFeat | 2 models average “about 15min of coding + 20 bucks to Amazon to get the images through the nets” | yes | yes |
orchid | 1.70 | OverFeat + Decaf | OverFeat & Decaf models + hand-crafted features (haralick/zernikemoments/lbp/pftas/tas/surf) | yes | yes |
Owen | 1.83 | | ? | | |
Paul Covington | 1.83 | | ? | | |
Maxim Milakov | 1.86 | nnForge | “rather deep” ConvNet with dropout and enriching training data by various distortions. | no | yes |
we've been in KAIST | 1.90 | Caffe (Decaf) | L2-SVM hinge loss instead of softmax loss + fine tuning entire caffe model except for 1st convolution | yes | yes |
Doug Koch | 1.94 | | ? | | |
fastml.com/ cats-and-dogs | 2.00 | OverFeat+ Decaf | 9 models ensemble | yes | yes |
13/28