1 of 61

Mobile Vision Learning�How cross channel pooling causes information loss in CNN ?�: application to Mobilenet v2

Jaewook Kang, Ph.D.

jwkang10@gmail.com

June. 2018

1

© 2018

MoT Lab

All Rights Reserved

누구나 TensorFlow!

J. Kang Ph.D.

2 of 61

소 개

    • GIST EEC Ph.D. (2015)
    • 신호처리 과학자, 삽질러
    • 모두의 연구소 MoT 연구실 리더
    • https://www.facebook.com/jwkkang
    • 좋아하는 것:
      • 통계적 신호처리 / 무선통신 신호처리
      • C++ Native 라이브러리 구현
      • Mobile Machine learning
      • 수영 덕력 6년

2

  • 대표논문:

Jaewook Kang, et al., "Bayesian Hypothesis Test using Nonparametric Belief Propagation for Noisy Sparse Recovery,"  IEEE Trans. on Signal process., Feb. 2015

Jaewook Kang et al., "Fast Signal Separation of 2D Sparse Mixture via Approximate Message-Passing," IEEE Signal Processing Letters, Nov. 2015

Jaewook Kang (강재욱)

누구나 TensorFlow!

J. Kang Ph.D.

3 of 61

1. Inverted Residual and Linear bottleneck:MobileNet v2 Mark Sandler et al. “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, CoRR, 2017. ��Maniford embedding, linear bottleneck, inverted residual 으악!

3

누구나 TensorFlow!

J. Kang Ph.D.

4 of 61

Some References

4

누구나 TensorFlow!

J. Kang Ph.D.

5 of 61

Further Reducing model size& complexity

  • 목적: 모바일에도 올릴 수 있는 효율적인 모델을 만들고 싶다!

5

Improving the state of the art performance of mobile models!!!

누구나 TensorFlow!

J. Kang Ph.D.

6 of 61

Further Reducing model size& complexity

  • 목적: 모바일에도 올릴 수 있는 효율적인 모델을 만들고 싶다!
  • Room: ReLu non-linearity에 의한 정보 손실 (Representation bottleneck)이 발생한다!
    • 입력 채널 정보 (the input manifold of interest)를 activation space(또는 feature space)에 온전히 담지 못하는 현상에서 발생

    • 원인1: 과도한 cross-channel pooling
      • 연산량을 줄이기 위해서 1x1 conv로 cross-channel pooling 수행
      • Activation space의 차원수 (dimensionality)를 제한 → 정보 손실

    • 원인2: ReLu non-linearity에 의한 정보손실
      • Conv filter이후의 Relu activation는 non-zero만을 가져옴
      • Activation space의 차원수 (dimensionality)를 제한 → 정보 손실

6

누구나 TensorFlow!

J. Kang Ph.D.

7 of 61

What is Manifold ?

  • 여기서 잠깐! Manifold가 몬가요!
    • 두가지 개념 (서로 반대 개념)

7

누구나 TensorFlow!

J. Kang Ph.D.

8 of 61

What is Manifold ?

  • 여기서 잠깐! Manifold가 몬가요!
    • 두가지 개념 (서로 반대 개념)
    • 수학적 정의 말고 쉬운 정의
      • Manifold: 어떤 data/feature가 고유로 가지고 있는 위상 공간

      • Representation: 어떤 data/feature의 주어진 공간 구조에서의 대수적 표현

8

누구나 TensorFlow!

J. Kang Ph.D.

9 of 61

What is Manifold ?

  • 여기서 잠깐! Manifold가 몬가요!
    • An Example: Representation차원과 manifold 차원는?

9

누구나 TensorFlow!

J. Kang Ph.D.

10 of 61

What is Manifold ?

  • 여기서 잠깐! Manifold가 몬가요!
    • Representation차원과 manifold 차원는?

10

Representation= 3차원

Manifold = 2차원

누구나 TensorFlow!

J. Kang Ph.D.

11 of 61

Cross-channel and Spatial Correlation

  • 개념 복습
    • Cross-channel correlation

    • Spatial Correlation

11

누구나 TensorFlow!

J. Kang Ph.D.

12 of 61

Cross-channel and Spatial Correlation

  • Cross-channel correlation:
    • conv layer에 입력되는 채널 간의 비슷한 정도

12

High cross-channel correlation!

누구나 TensorFlow!

J. Kang Ph.D.

13 of 61

Cross-channel and Spatial Correlation

  • Cross-channel correlation:
    • conv layer에 입력되는 채널 간의 비슷한 정도

13

Very? Low cross-channel correlation!

누구나 TensorFlow!

J. Kang Ph.D.

14 of 61

Cross-channel and Spatial Correlation

  • Cross-channel correlation: conv layer의 입력채널 간의 비슷한 정도
    • High cross-channel correlation:
      • 입력 채널 간의 상관도 높다.
      • → 입력 채널간 비슷한 특징을 가진다.
      • → feature space의 차원수가 작다.
      • → conv layer의 표현력이 떨어진다.
      • → 다양한 입력 채널 정보을 보존하기 어렵다.

14

누구나 TensorFlow!

J. Kang Ph.D.

15 of 61

Cross-channel and Spatial Correlation

  • Cross-channel correlation: conv layer의 입력채널 간의 비슷한 정도
    • Low cross-channel correlation:
      • 입력 채널 간의 상관도 낮다.
      • → 입력 채널간 구별되는 특징을 가진다.
      • → feature space의 차원수가 높다.
      • → conv layer의 표현력이 우수하다.
      • → 다양한 특징을 가지는 채널 정보를 보존할 수 있다.

15

누구나 TensorFlow!

J. Kang Ph.D.

16 of 61

Cross-channel and Spatial Correlation

  • Spatial correlation:
    • conv filter와 입력 채널 사이의 상관도

16

Conv filter

Low correlation

High correlation

누구나 TensorFlow!

J. Kang Ph.D.

17 of 61

Cross-channel and Spatial Correlation

  • Spatial correlation:
    • conv filter와 입력 채널 사이의 상관도

    • High spatial correlation: 특정 conv filter로 feature extraction이 잘된다.

    • Low spatial correlation: 특정 conv filter로 feature extraction이 잘 안된다.

17

Conv filter

누구나 TensorFlow!

J. Kang Ph.D.

18 of 61

Depthwise Separable Convolution

  • Depthwise separable conv의 핵심가설:
    • The mapping of “cross-channels correlation” and “spatial correlation” can be entirely decoupled!

    • 그러니깐 “cross-channels correlation”과 “spatial correlation” 분리해서 다루겠다는 것!

18

누구나 TensorFlow!

J. Kang Ph.D.

19 of 61

Depthwise Separable Conv

  • Depthwise separable conv
    • Dwise conv + Pwise conv로 분할 계산 (L개 채널입력)
      • Input channel : N x N x L
      • Dwise Filter size (weight) : K x K x 1 x(L)
      • Pwise 1x1 filter: 1 x 1 x L X (M)

19

Depthwise Conv

Pointwise Conv

Dwise Filter Size : K x K x 1x(L)

(K=3)

Pwise Filter Size : 1 x 1 x L (x M)

(L=3)

+

누구나 TensorFlow!

J. Kang Ph.D.

20 of 61

Depthwise Separable Conv

  • Depthwise separable conv
    • 1) Depthwise Convolution:
      • Extract spatial correlation from

a NxNx1 2D input channel

      • 2D convolution with a single

K x K x 1 x(L) 2D filter

    • 2) Pointwise Convolution:
      • Compress NxNxL input channel to

from NxNxM output channel (M < L)

      • Weighted linear mixing using

2D convolution with 1x1xLx(M)

1D conv filters

20

누구나 TensorFlow!

J. Kang Ph.D.

21 of 61

1x1 Convolution Revisit!

  • 1x1 conv filter는 어떻케 동작하는가?
    • X : input (L=3개채널), Z: logit with M=2, 1x1xL conv filters,
    • Y: activation,

21

X: 3x3xL

Input features

W: 1x1xL Filter

→ Single 1x1xL conv filters (L=3,M=1)

누구나 TensorFlow!

J. Kang Ph.D.

22 of 61

1x1 Convolution Revisit!

  • 1x1 conv filter는 어떻케 동작하는가?
    • X : input (L=3개채널), Z: logit with M=2, 1x1xL conv filters,
    • Y: activation,

22

Z1

X: 3x3xL

Input features

+

w11

w12

w13

W: 1x1xL Filter

→ Single 1x1xL conv filters (L=3,M=1)

누구나 TensorFlow!

J. Kang Ph.D.

23 of 61

1x1 Convolution Revisit!

  • 1x1 conv filter는 어떻케 동작하는가?
    • X : input (L=3개채널), Z: logit with M=2, 1x1xL conv filters,
    • Y: activation,

23

Z1

Z1

X: 3x3xL

Input features

Z: 3x3x1

Logit features

+

W: 1x1xL Filter

→ Single 1x1xL conv filters (L=3,M=1)

w11

w12

w13

누구나 TensorFlow!

J. Kang Ph.D.

24 of 61

1x1 Convolution Revisit!

  • 1x1 conv filter는 어떻케 동작하는가?
    • X : input (L=3개채널), Z: logit with M=2, 1x1xL conv filters,
    • Y: activation,

24

Z1

Z1

Z2

X: 3x3xL

Input features

Z: 3x3x2 (M=2)

Logit features

+

+

W: 1x1xL Filter

Two 1x1xL conv filters (L=3,M=2)

w11

Z2

w12

w13

w21

w22

w23

누구나 TensorFlow!

J. Kang Ph.D.

25 of 61

1x1 Convolution Revisit!

  • 1x1 conv filter는 어떻케 동작하는가?
    • X : input (L=3개채널), Z: logit with M=2, 1x1xL conv filters,
    • Y: activation,

25

Z1

Z1

Z2

X: 3x3xL

Input features

Y1

Z: 3x3xM

Logit features

Y: 3x3xM

Output features

+

+

W: 1x1xL Filter

→ Two 1x1xL conv filters (L=3,M=2)

w11

Z2

Y2

w12

w13

w21

w22

w23

Relu activation

Relu activation

누구나 TensorFlow!

J. Kang Ph.D.

26 of 61

1x1 Convolution Revisit!

  • 1x1 conv filter는 어떻케 동작하는가?
    • Is 1x1 conv filtering equivalent to matrix multiplication ?
    • We can look the convolution at a 1x1 pixel point (i=i*,j=j*).
    • Consider a M=2, L=3 case
    • Given
      • Lx1x1 local patch vector, [X1,X2,X3]^T from X (3x3xL, L=3)
      • two 1x1 conv filters (1x1xLxM, M=2,L=3)
      • we have two different logits scalars: Z1, Z2

26

누구나 TensorFlow!

J. Kang Ph.D.

27 of 61

1x1 Convolution Revisit!

  • 1x1 conv filter는 어떻케 동작하는가?
    • Is 1x1 conv filtering equivalent to matrix multiplication ?
    • We can look the convolution at a 1x1 pixel point (i=i*,j=j*).
    • Consider a M=2, L=3 case
    • Given
      • Lx1x1 local patch vector, [X1,X2,X3]^T from X (3x3xL, L=3)
      • two 1x1 conv filters (1x1xLxM, M=2,L=3)
      • we have two different logits scalars: Z1, Z2
    • These conv operations recast into a matrix-vector form:

27

Lx1x1 local

patch vector

Two different

logits scalars

누구나 TensorFlow!

J. Kang Ph.D.

28 of 61

1x1 Convolution Revisit!

  • 1x1 conv filter는 어떻케 동작하는가?
    • Is 1x1 conv filtering equivalent to matrix multiplication ?
    • We can look the convolution at a 1x1 pixel point (i=i*,j=j*).
    • Consider a M=2, L=3 case
    • Given
      • Lx1x1 local patch vector, [X1,X2,X3]^T from X (3x3xL, L=3)
      • two 1x1 conv filters (1x1xLxM, M=2,L=3)
      • we have two different logits scalars: Z1, Z2
    • These conv operations recast into a matrix-vector form:

28

1x1xL conv1

1x1xL conv2

Lx1x1 local

patch vector

Two different

logits scalars

누구나 TensorFlow!

J. Kang Ph.D.

29 of 61

1x1 Convolution Revisit!

  • 1x1 conv filter는 어떻케 동작하는가?
    • Is 1x1 conv filtering equivalent to matrix multiplication ?
    • Consider a M=2 case

29

1x1xL conv1

1x1xL conv2

.

Input channels

1x1xNxK conv filters

Output logit

Before activation

Lx1x1 vector X

1x1xLxM

filter matrix, W

MX1 output logit Z

=

=

output

채널방향

input�채널방향

누구나 TensorFlow!

J. Kang Ph.D.

30 of 61

1x1 Convolution Revisit!

  • 원인1: input manifold를 온전히 담지 못하는 activation space
    • Given 1x1xL(xM) conv filters
    • 1x1xL conv filter 개수, M → linear transform W의 row 개수

      • M < L channel pooling → 정보 손실
      • M>= L → channel expansion → 정보 보존

    • M<L 경우 linear transform W이 충분한 개수의 independent basis을 가질 수 없음
      • W가 span하는 feature space의 dimensionality가 X의 정보를 보존하기에 충분하지 않을 수도 있음

30

누구나 TensorFlow!

J. Kang Ph.D.

31 of 61

1x1 Convolution Revisit!

  • 원인1: input manifold를 온전히 담지 못하는 activation space
    • 연산량을 줄이기 위해서 1x1 conv filter의 개수 M을 과도하게 작게 한 경우에 해당 (M<<L)
      • L : 입력 채널 개수
      • M: 출력 채널 개수
    • Depth multiplier를 과도하게 작게 작게 하는 경우

31

누구나 TensorFlow!

J. Kang Ph.D.

32 of 61

1x1 Convolution Revisit!

  • 원인1: input manifold를 온전히 담지 못하는 activation space

32

Where “dim” indicates the dimension of activation space span by W.

Note: Activation space- (선형변환 후 feature space)

누구나 TensorFlow!

J. Kang Ph.D.

33 of 61

1x1 Convolution Revisit!

  • 원인1: input manifold를 온전히 담지 못하는 activation space
    • 조금 어렵고 잼난 얘기를 쉽게 해볼께요 ☺

33

누구나 TensorFlow!

J. Kang Ph.D.

34 of 61

1x1 Convolution Revisit!

  • 원인1: input manifold를 온전히 담지 못하는 activation space
    • “M < L” 인 경우에도 X의 manifold의 차원수 (dimensionality)가 M보다 작다면 충분히 X의 정보를 activation space에서 보존 할 수 있다.

34

누구나 TensorFlow!

J. Kang Ph.D.

35 of 61

Compressed Sensing Nutshell

  • 원인1: input manifold를 온전히 담지 못하는 activation space
    • “M < L” 인 경우에도 X의 manifold의 차원수 (dimensionality)가 M보다 작다면 충분히 X의 정보를 activation space에서 보존 할 수 있다.
    • An example: compressed sensing !

35

- D. Donoho, “Compressed sensing,” IEEE TIT, 2006.

- E. J. Candes, J. Romberg, T. Tao, "Robust uncertainty principles: exact signal reconstruction from highly incomplete frequency information," IEEE TIT, 2006

누구나 TensorFlow!

J. Kang Ph.D.

36 of 61

Compressed Sensing Nutshell

  • Compressed sensing
    • 목적: y로 부터 x를 복원하는 것

36

- D. Donoho, “Compressed sensing,” IEEE TIT, 2006.

- E. J. Candes, J. Romberg, T. Tao, "Robust uncertainty principles: exact signal reconstruction from highly incomplete frequency information," IEEE TIT, 2006

누구나 TensorFlow!

J. Kang Ph.D.

37 of 61

Compressed Sensing Nutshell

  • Compressed sensing
    • 근데 M < N 면 \PHI matrix 가 singular임 (역행렬 존재X)
      • 통상적인 방법으로 Y로부터 X복원 안됨

37

누구나 TensorFlow!

J. Kang Ph.D.

38 of 61

Compressed Sensing Nutshell

  • Compressed sensing
    • 근데 가만히 보니 X가 sparse!
      • Nonzero개수가 3이다
      • X의 manifold의 차원수가 3
    • X의 nonzero만 보면 복원이 가능할 것도 같기도 한데!
    • X의 정보를 어떻케 하면 linear transform이후에 보존 할수 있을까?

38

누구나 TensorFlow!

J. Kang Ph.D.

39 of 61

Compressed Sensing Nutshell

  • Compressed sensing
    • 근데 가만히 보니 X가 sparse!
      • Nonzero개수가 3이다
      • X의 manifold의 차원수가 3
    • X의 nonzero만 보면 복원이 가능할 것도 같기도 한데!
    • X의 정보를 어떻케 하면 linear transform이후에 보존 할수 있을까?

39

누구나 TensorFlow!

J. Kang Ph.D.

40 of 61

Cross Channel Pooling Revisit!

  • 원인1: 과도한 cross-channel pooling
    • “M < L” 인 경우에도 X의 manifold의 차원수 (dimensionality)가 M보다 작다면 충분히 X의 정보를 activation space에서 보존 할 수 있다.

40

누구나 TensorFlow!

J. Kang Ph.D.

41 of 61

Compressed Sensing Nutshell

  • Compressed sensing
    • 정보가 보존된다는 것의 의미
      • 1) 임의의 입력 채널 X1, X2가 있을때 linear transform이후에서 두 채널 사이의 거리가 유지 되어야 한다.

      • 2) 그렇케 되도록 linear transform matrix \Phi를 설계하면 Y로 부터 X를 L1 minimization을 통해서 복원할 수 있다.

41

누구나 TensorFlow!

J. Kang Ph.D.

42 of 61

Compressed Sensing Nutshell

  • 원인1: input manifold를 온전히 담지 못하는 activation space
    • “M < L” 인 경우에도 X의 manifold의 차원수 (dimensionality)가 M보다 작다면 충분히 X의 정보를 activation space에서 보존 할 수 있다.

    • 다시 돌아와서 적용하면
      • X의 manifold 차원수가 충분히 작으면 M <L 의 1x1 conv filter( linear transform M) 을 가지고 X의 정보를 보할 수 있다.

      • 결론적으로 activation space (WX) 의 차원수가 X manifold 차원수 보다 충분히 커야 정보 보존이 된다.

42

누구나 TensorFlow!

J. Kang Ph.D.

43 of 61

How ReLu Restricts Feature Space

  • 원인2: ReLu non-linearity에 의한 정보손실
    • ReLu는 0보다 작은 값을 출력하는 1x1 conv filtering 결과를 중요하지 않은 정보라고 판단하고 zero-mapping한다.
    • An 1x1 conv example:

43

1x1xL conv1

1x1xL conv2

1x1xL conv3

1x1xL conv4

=

.

Features

After Dwise conv

Set of 1x1xL conv filters

Output logit

Before activation

1X1XLXM

filter matrix, W

M X 1 output logit Z

(M=4)�

누구나 TensorFlow!

J. Kang Ph.D.

44 of 61

How ReLu Restricts Feature Space

  • 원인2: ReLu non-linearity에 의한 정보손실
    • ReLu는 0보다 작은 값을 출력하는 1x1 conv filtering 결과를 중요하지 않은 정보라고 판단하고 zero-mapping한다.
    • An 1x1 conv example:

44

1x1xL conv1

1x1xL conv2

1x1xL conv3

1x1xL conv4

=

.

Features

After Dwise conv

Set of 1x1xL conv filters

Output logit

Before activation

1X1XLXM

filter matrix, W

M X 1 output logit Z

(M=4)�

ReLu !!!!

누구나 TensorFlow!

J. Kang Ph.D.

45 of 61

How ReLu Restricts Feature Space

  • 원인2: ReLu non-linearity에 의한 정보손실
    • ReLu는 0보다 작은 값을 출력하는 1x1 conv filtering 결과를 중요하지 않은 정보라고 판단하고 zero-mapping한다.
    • An 1x1 conv example:

45

1x1xL conv1

1x1xL conv2

1x1xL conv3

1x1xL conv4

=

.

Features

After Dwise conv

Set of 1x1xL conv filters

Output

After activation

1X1XLXM

filter matrix, W

M X 1 output Y

(M=4)�

J. Kang Ph.D. presents

46 of 61

How ReLu Restricts Feature Space

  • 원인2: ReLu non-linearity에 의한 정보손실
    • ReLu는 0보다 작은 값을 출력하는 1x1 conv filtering 결과를 중요하지 않은 정보라고 판단하고 zero-mapping한다.

    • Input manifold를 충분히 담지 못하는 space에서 Relu를 수행하면 정보손실이 발생한다.

    • 차원수가 충분히 큰 space에서 Relu를 하면 정보가 손실될 가능성이 더 작다

46

X manifold 차원수 << Relu 하기 전 activation space (WX) 차원수

누구나 TensorFlow!

J. Kang Ph.D.

47 of 61

Mobilenet v2!

  • 핵심질문: 입력채널의 manifold를 효과적으로 보존하기 위해서는 conv layer의 구조를 어떻케 구성해야 하는가?

- Mark Sandler et al. “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, CoRR, 2017.

  • 방향1: Linear bottleneck 삽입을 통한 ReLu에 의한 정보손실 해결
    • 마지막 ReLu전에 channel expansion하여 input manifold를 충분히 큰 space에 담어 놓고 relu하면 정보손실을 최소화 할 수 있다.
    • channel expansion: 1x1 pwise conv에서 channel을 확장 L → T
    • channel projection : ReLu 통과시킨 후에 추가적인 1x1 conv 삽입하여 projection T → M

  • 방향2: Inverted Residuals with shortcut connection
    • spatial feature extraction (Dwise conv)하기 전에 channel expansion
    • Pwise conv에서 channel projection하여 출력 채널수 결정
    • Shortcut connection사용!

47

누구나 TensorFlow!

J. Kang Ph.D.

48 of 61

Mobilenet v2!

  • 핵심질문: 입력채널의 manifold를 효과적으로 보존하기 위해서는 conv layer의 구조를 어떻케 구성해야 하는가?

- Mark Sandler et al. “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, CoRR, 2017.

  • Mobilenet 계열의 논문은 매우 불친절하다. 잘 안가르쳐준다. ㅠㅠ

  • 그래서 제가 한번 파헤쳐 봤습니다.!

48

누구나 TensorFlow!

J. Kang Ph.D.

49 of 61

Linear Bottleneck Block

  • 핵심질문: Linear bottleneck의 삽입은 어떻케 ReLu non-linearity에 의한 정보손실을 해결하는가?
  • 기존 Depthwise-separable conv (Mobilenet v1):

    • 마지막 Relu를 수행 전에 channel expansion하여 Relu non-linearity로 인한 정보손실을 를 최소화 하자!

49

Dwise Conv

3x3x1xL

Pwise Conv

1x1xLxM

BN →ReLu6

BN →ReLu6

Ch in

X

NxNxL

Ch out

Y

NxNxM

Feature maps

NxNxL

Feature map

NxNxM

Spatial

Feature extraction

Channel

Projection

누구나 TensorFlow!

J. Kang Ph.D.

50 of 61

Linear Bottleneck Block

  • 핵심질문: Linear bottleneck의 삽입은 어떻케 ReLu non-linearity에 의한 정보손실을 해결하는가?
  • Linear Bottleneck (Mobilenet v2):

    • Linear Bottleneck: 충분히 channel expansion하고 relu하자!
      • Depthwise-separable conv 내부 Pwise conv에서 channel expansion
      • ReLu non-linearity를 linear bottleneck에서 보완할 수 있도록 학습됨 (soft pooling)

50

Dwise Conv

3x3x1xL

Linear Bottleneck

1x1xTxM

(T>M)

Ch in

X

NxNxL

Ch out

Y

NxNxM

Feature maps

NxNxL

Feature map

NxNxT

Spatial

Feature extraction

Channel

Projection

BN →ReLu6

Pwise Conv

1x1xLxT

(L < T)

BN

BN →ReLu6

Channel

expansion

누구나 TensorFlow!

J. Kang Ph.D.

51 of 61

Inverted Residual Block

  • 핵심질문: 더 효율적인 channel expansion + projection 구조는?

  • 방향1: Linear bottleneck을 마지막에 두지 말고 앞으로 가져와서 squeezeNet에서 하듯이 expansion → projection 구조를 사용

  • Space expansion (Linear bottleneck) → spatial feature extraction (Dwise conv) → Space projection (Pwise conv)

Dwise Conv

3x3x1xL

Linear Bottleneck

1x1xTxM

(T>M)

Ch in

X

NxNxL

Ch out

Y

NxNxM

Feature maps

NxNxL

Feature map

NxNxT

Spatial

Feature extraction

Channel

Projection

BN →ReLu6

Pwise Conv

1x1xLxT

(L < T)

BN

BN →ReLu6

Channel

expansion

누구나 TensorFlow!

J. Kang Ph.D.

52 of 61

Inverted Residual Block

  • 핵심질문: 더 효율적인 channel expansion + projection 구조는?

52

Dwise Conv

3x3x1xT

Ch in

X

NxNxL

Ch out

Y

NxNxM

Expanded

Feature maps

NxNxT

Feature map

NxNxM

Spatial

Feature extraction

Channel

projection

BN

Pwise Conv

1x1xTxM

(T>M)

BN→ReLu6

Linear Bottleneck

1x1xLxT

(L<T)

Expand

Input channels

BN→ Relu6

누구나 TensorFlow!

J. Kang Ph.D.

53 of 61

Inverted Residual Block

  • 핵심질문: 더 효율적인 channel expansion + projection 구조는?

    • 메모리 사용량 계산 비교: 개선 기법이 더 사용 (T > L) ;;;;

53

기존

개선

Dwise conv : N^2*L + K^2

Pwise conv : N^2*L + LT

Linear bottleneck : N^2*T + TM

Dwise conv : N^2*T+ K^2

Pwise conv : N^2*T + TM

Linear bottleneck : N^2*L + LT

Dwise Conv

3x3x1xT

Ch in

X

NxNxL

Ch out

Y

NxNxM

Expanded

Feature maps

NxNxT

Feature map

NxNxM

Spatial

Feature extraction

Channel

projection

BN

Pwise Conv

1x1xTxM

(T>M)

BN→ReLu6

Linear Bottleneck

1x1xLxT

(L<T)

Expand

Input channels

BN→ Relu6

누구나 TensorFlow!

J. Kang Ph.D.

54 of 61

Inverted Residual Block

  • 핵심질문: 더 효율적인 channel expansion + projection 구조는?

  • 방향2: shortcut connection의 사용
    • Degradation 문제 해결을 위해서 사용

54

+

Shortcut connection for residual learning

Dwise Conv

3x3x1xT

Ch in

X

NxNxL

Ch out

Y

NxNxM

Spatial

Feature extraction

Channel

Projection

BN

Pwise Conv

1x1xTxM

(T>M)

BN→ReLu6

Linear Bottleneck

1x1xLxT

(L<T)

Expand

Input channels

BN →

Relu6

누구나 TensorFlow!

J. Kang Ph.D.

55 of 61

Inverted Residual Block

  • 핵심질문: 어떻케 하면 효율적으로 shortcut connection을 linear bottleneck구조에 적용할 수 있을까?
  • Where
    • T = tk, M = k’,

55

누구나 TensorFlow!

J. Kang Ph.D.

56 of 61

MobileNetv2 Architecture

56

t: expansion layer scale

c: # of input channels

n : # of layer repetitions

- Each seq. has n layers.

s: stride parameter of

the first layer of the squ.

where 1st layer of each

Seq. has a stride ”s”,

all other use stride1.

Spatial conv filters are

All 3x3

누구나 TensorFlow!

J. Kang Ph.D.

57 of 61

MobileNetv2 Architecture

57

Where utilize dropout and

Batch normalization for

all conv layers

Both blocks are used.

  • Linear bottleneck block
  • Bottleneck residual block

(b) Linear Bottleneck Block

(a) Bottleneck Residual Block

누구나 TensorFlow!

J. Kang Ph.D.

58 of 61

ImageNet Benchmark Comparison

58

  • Over Google Pixel 1 phone using TFLite
  • Use RMSprop optimizer
  • Use weight decay with 4E-5
  • Batch size 96

Remarks:

  • Mobilebnet v2 achieves below:
    • 20% reduction in param from v1
    • 48% reduction in complexity from v1
    • 2% improvement in Top1 accuracy
  • Comparable to Shufflenet in all aspects.

누구나 TensorFlow!

J. Kang Ph.D.

59 of 61

MSCOCO Benchmark Comparison

59

- Over Google Pixel 1 phone using TFLite

Remarks:

  • Mnet v2 + SSDLite achieves below:
    • 16% reduction in param from v1
    • 39% reduction in complexity from v1

- 26% reduction in CPU time from v1

  • Advantage of v2 is less compared to the ImageNet comparison. (why?)

누구나 TensorFlow!

J. Kang Ph.D.

60 of 61

Keys for success of MobileNetv2

  • Channel pooling →ReLu non-linearity 구조가 activation space를 제한하여 정보손실을 야기하는 것에 주목

  • Linear bottleneck을 추가하여 relu non-linearity를 보완하면서 channel pooling을 할 수 있게 함
    • Linear bottleneck block:
    • Depthwise separable conv → Relu → linear bottleneck (channel pooling)

  • 대세인 shortcut connection + grid reduction with stride conv2 사용!
    • Bottleneck residual block!
    • Expansion → squeezing 구조 사용

  • Per-layer parameters는 증가하나 channel 수 ( 1x1 conv filter개수)를 linear bottleneck구조를 가지고 현저하게 줄일 수 있음!
    • 연산량과 사이즈를 절감 가능!

60

누구나 TensorFlow!

J. Kang Ph.D.

61 of 61

모두연 MoT랩 소개

  • Keywords:
    • Thin CNN Model
    • Model optimization
    • Tensorflow + lite
    • Embedded Sys. (IoT)
    • Android Mobile/Things

61

누구나 TensorFlow!

J. Kang Ph.D.