Mobile Vision Learning�How cross channel pooling causes information loss in CNN ?�: application to Mobilenet v2�
Jaewook Kang, Ph.D.
jwkang10@gmail.com
June. 2018
1
© 2018
MoT Lab
All Rights Reserved
누구나 TensorFlow!
J. Kang Ph.D.
소 개
2
Jaewook Kang, et al., "Bayesian Hypothesis Test using Nonparametric Belief Propagation for Noisy Sparse Recovery," IEEE Trans. on Signal process., Feb. 2015
Jaewook Kang et al., "Fast Signal Separation of 2D Sparse Mixture via Approximate Message-Passing," IEEE Signal Processing Letters, Nov. 2015
Jaewook Kang (강재욱)
누구나 TensorFlow!
J. Kang Ph.D.
1. Inverted Residual and Linear bottleneck:�MobileNet v2 � �� Mark Sandler et al. “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, CoRR, 2017. ��Maniford embedding, linear bottleneck, inverted residual 으악!��
3
누구나 TensorFlow!
J. Kang Ph.D.
Some References
4
누구나 TensorFlow!
J. Kang Ph.D.
Further Reducing model size& complexity
5
Improving the state of the art performance of mobile models!!!
누구나 TensorFlow!
J. Kang Ph.D.
Further Reducing model size& complexity
6
누구나 TensorFlow!
J. Kang Ph.D.
What is Manifold ?
7
누구나 TensorFlow!
J. Kang Ph.D.
What is Manifold ?
8
누구나 TensorFlow!
J. Kang Ph.D.
What is Manifold ?
9
누구나 TensorFlow!
J. Kang Ph.D.
What is Manifold ?
10
Representation= 3차원
Manifold = 2차원
누구나 TensorFlow!
J. Kang Ph.D.
Cross-channel and Spatial Correlation
11
누구나 TensorFlow!
J. Kang Ph.D.
Cross-channel and Spatial Correlation
12
High cross-channel correlation!
누구나 TensorFlow!
J. Kang Ph.D.
Cross-channel and Spatial Correlation
13
Very? Low cross-channel correlation!
누구나 TensorFlow!
J. Kang Ph.D.
Cross-channel and Spatial Correlation
14
누구나 TensorFlow!
J. Kang Ph.D.
Cross-channel and Spatial Correlation
15
누구나 TensorFlow!
J. Kang Ph.D.
Cross-channel and Spatial Correlation
16
Conv filter
Low correlation
High correlation
누구나 TensorFlow!
J. Kang Ph.D.
Cross-channel and Spatial Correlation
17
Conv filter
누구나 TensorFlow!
J. Kang Ph.D.
Depthwise Separable Convolution
18
누구나 TensorFlow!
J. Kang Ph.D.
Depthwise Separable Conv
19
Depthwise Conv
Pointwise Conv
Dwise Filter Size : K x K x 1x(L)
(K=3)
Pwise Filter Size : 1 x 1 x L (x M)
(L=3)
+
누구나 TensorFlow!
J. Kang Ph.D.
Depthwise Separable Conv
a NxNx1 2D input channel
K x K x 1 x(L) 2D filter
from NxNxM output channel (M < L)
2D convolution with 1x1xLx(M)
1D conv filters
20
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
21
X: 3x3xL
Input features
W: 1x1xL Filter
→ Single 1x1xL conv filters (L=3,M=1)
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
22
Z1
X: 3x3xL
Input features
+
w11
w12
w13
W: 1x1xL Filter
→ Single 1x1xL conv filters (L=3,M=1)
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
23
Z1
Z1
X: 3x3xL
Input features
Z: 3x3x1
Logit features
+
W: 1x1xL Filter
→ Single 1x1xL conv filters (L=3,M=1)
w11
w12
w13
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
24
Z1
Z1
Z2
X: 3x3xL
Input features
Z: 3x3x2 (M=2)
Logit features
+
+
W: 1x1xL Filter
→ Two 1x1xL conv filters (L=3,M=2)
w11
Z2
w12
w13
w21
w22
w23
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
25
Z1
Z1
Z2
X: 3x3xL
Input features
Y1
Z: 3x3xM
Logit features
Y: 3x3xM
Output features
+
+
W: 1x1xL Filter
→ Two 1x1xL conv filters (L=3,M=2)
w11
Z2
Y2
w12
w13
w21
w22
w23
Relu activation
Relu activation
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
26
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
27
Lx1x1 local
patch vector
Two different
logits scalars
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
28
1x1xL conv1
1x1xL conv2
Lx1x1 local
patch vector
Two different
logits scalars
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
29
1x1xL conv1
1x1xL conv2
.
Input channels
1x1xNxK conv filters
Output logit
Before activation
Lx1x1 vector X
1x1xLxM
filter matrix, W
�
MX1 output logit Z
�
=
=
output
채널방향
input�채널방향
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
30
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
31
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
32
Where “dim” indicates the dimension of activation space span by W.
Note: Activation space- (선형변환 후 feature space)
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
33
누구나 TensorFlow!
J. Kang Ph.D.
1x1 Convolution Revisit!
34
누구나 TensorFlow!
J. Kang Ph.D.
Compressed Sensing Nutshell
35
- D. Donoho, “Compressed sensing,” IEEE TIT, 2006.
- E. J. Candes, J. Romberg, T. Tao, "Robust uncertainty principles: exact signal reconstruction from highly incomplete frequency information," IEEE TIT, 2006
누구나 TensorFlow!
J. Kang Ph.D.
Compressed Sensing Nutshell
36
- D. Donoho, “Compressed sensing,” IEEE TIT, 2006.
- E. J. Candes, J. Romberg, T. Tao, "Robust uncertainty principles: exact signal reconstruction from highly incomplete frequency information," IEEE TIT, 2006
누구나 TensorFlow!
J. Kang Ph.D.
Compressed Sensing Nutshell
37
누구나 TensorFlow!
J. Kang Ph.D.
Compressed Sensing Nutshell
38
누구나 TensorFlow!
J. Kang Ph.D.
Compressed Sensing Nutshell
39
누구나 TensorFlow!
J. Kang Ph.D.
Cross Channel Pooling Revisit!
40
누구나 TensorFlow!
J. Kang Ph.D.
Compressed Sensing Nutshell
41
누구나 TensorFlow!
J. Kang Ph.D.
Compressed Sensing Nutshell
42
누구나 TensorFlow!
J. Kang Ph.D.
How ReLu Restricts Feature Space
�
43
1x1xL conv1
1x1xL conv2
1x1xL conv3
1x1xL conv4
=
.
Features
After Dwise conv
Set of 1x1xL conv filters
Output logit
Before activation
1X1XLXM
filter matrix, W
�
M X 1 output logit Z
(M=4)�
누구나 TensorFlow!
J. Kang Ph.D.
How ReLu Restricts Feature Space
�
44
1x1xL conv1
1x1xL conv2
1x1xL conv3
1x1xL conv4
=
.
Features
After Dwise conv
Set of 1x1xL conv filters
Output logit
Before activation
1X1XLXM
filter matrix, W
�
M X 1 output logit Z
(M=4)�
ReLu !!!!
누구나 TensorFlow!
J. Kang Ph.D.
How ReLu Restricts Feature Space
�
45
1x1xL conv1
1x1xL conv2
1x1xL conv3
1x1xL conv4
=
.
Features
After Dwise conv
Set of 1x1xL conv filters
Output
After activation
1X1XLXM
filter matrix, W
�
M X 1 output Y
(M=4)�
J. Kang Ph.D. presents
How ReLu Restricts Feature Space
46
X manifold 차원수 << Relu 하기 전 activation space (WX) 차원수
누구나 TensorFlow!
J. Kang Ph.D.
Mobilenet v2!
- Mark Sandler et al. “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, CoRR, 2017.
47
누구나 TensorFlow!
J. Kang Ph.D.
Mobilenet v2!
- Mark Sandler et al. “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, CoRR, 2017.
48
누구나 TensorFlow!
J. Kang Ph.D.
Linear Bottleneck Block
49
Dwise Conv
3x3x1xL
Pwise Conv
1x1xLxM
BN →ReLu6
BN →ReLu6
Ch in
X
NxNxL
Ch out
Y
NxNxM
Feature maps
NxNxL
Feature map
NxNxM
Spatial
Feature extraction
Channel
Projection
누구나 TensorFlow!
J. Kang Ph.D.
Linear Bottleneck Block
50
Dwise Conv
3x3x1xL
Linear Bottleneck
1x1xTxM
(T>M)
Ch in
X
NxNxL
Ch out
Y
NxNxM
Feature maps
NxNxL
Feature map
NxNxT
Spatial
Feature extraction
Channel
Projection
BN →ReLu6
Pwise Conv
1x1xLxT
(L < T)
BN
BN →ReLu6
Channel
expansion
누구나 TensorFlow!
J. Kang Ph.D.
Inverted Residual Block
Dwise Conv
3x3x1xL
Linear Bottleneck
1x1xTxM
(T>M)
Ch in
X
NxNxL
Ch out
Y
NxNxM
Feature maps
NxNxL
Feature map
NxNxT
Spatial
Feature extraction
Channel
Projection
BN →ReLu6
Pwise Conv
1x1xLxT
(L < T)
BN
BN →ReLu6
Channel
expansion
누구나 TensorFlow!
J. Kang Ph.D.
Inverted Residual Block
52
Dwise Conv
3x3x1xT
Ch in
X
NxNxL
Ch out
Y
NxNxM
Expanded
Feature maps
NxNxT
Feature map
NxNxM
Spatial
Feature extraction
Channel
projection
BN
Pwise Conv
1x1xTxM
(T>M)
BN→ReLu6
Linear Bottleneck
1x1xLxT
(L<T)
Expand
Input channels
BN→ Relu6
누구나 TensorFlow!
J. Kang Ph.D.
Inverted Residual Block
53
기존 | 개선 |
Dwise conv : N^2*L + K^2 Pwise conv : N^2*L + LT Linear bottleneck : N^2*T + TM | Dwise conv : N^2*T+ K^2 Pwise conv : N^2*T + TM Linear bottleneck : N^2*L + LT |
Dwise Conv
3x3x1xT
Ch in
X
NxNxL
Ch out
Y
NxNxM
Expanded
Feature maps
NxNxT
Feature map
NxNxM
Spatial
Feature extraction
Channel
projection
BN
Pwise Conv
1x1xTxM
(T>M)
BN→ReLu6
Linear Bottleneck
1x1xLxT
(L<T)
Expand
Input channels
BN→ Relu6
누구나 TensorFlow!
J. Kang Ph.D.
Inverted Residual Block
54
+
Shortcut connection for residual learning
Dwise Conv
3x3x1xT
Ch in
X
NxNxL
Ch out
Y
NxNxM
Spatial
Feature extraction
Channel
Projection
BN
Pwise Conv
1x1xTxM
(T>M)
BN→ReLu6
Linear Bottleneck
1x1xLxT
(L<T)
Expand
Input channels
BN →
Relu6
누구나 TensorFlow!
J. Kang Ph.D.
Inverted Residual Block
55
누구나 TensorFlow!
J. Kang Ph.D.
MobileNetv2 Architecture
56
t: expansion layer scale
c: # of input channels
n : # of layer repetitions
- Each seq. has n layers.
s: stride parameter of
the first layer of the squ.
where 1st layer of each
Seq. has a stride ”s”,
all other use stride1.
Spatial conv filters are
All 3x3
누구나 TensorFlow!
J. Kang Ph.D.
MobileNetv2 Architecture
57
Where utilize dropout and
Batch normalization for
all conv layers
Both blocks are used.
(b) Linear Bottleneck Block
(a) Bottleneck Residual Block
누구나 TensorFlow!
J. Kang Ph.D.
ImageNet Benchmark Comparison
58
Remarks:
누구나 TensorFlow!
J. Kang Ph.D.
MSCOCO Benchmark Comparison
59
- Over Google Pixel 1 phone using TFLite
Remarks:
- 26% reduction in CPU time from v1
누구나 TensorFlow!
J. Kang Ph.D.
Keys for success of MobileNetv2
60
누구나 TensorFlow!
J. Kang Ph.D.
모두연 MoT랩 소개
61
누구나 TensorFlow!
J. Kang Ph.D.