1 of 143

ML Visuals

​

​

2 of 143

0~1s

Rest

Rest

ERP

PSD

T0

T1/T2

Motor Imagery

Time

T0

T1/T2

1~3s

Visual

Stimulus

Visual

Stimulus

3 of 143

Rest

Motor Imagery

Rest

Time

4-sec

Visual

Stimulus

Visual

Stimulus

An EEG Trial

4 of 143

5 of 143

L

L

L

L

L

L

L

L

L

L

C

C

C

C

C

C

C

C

C

C

L

L

L

L

L

L

L

L

L

L

A

A

A

A

A

A

A

A

A

A

6 of 143

L

L

L

L

L

L

L

L

L

L

C

C

C

C

C

C

C

C

C

C

L

L

L

L

L

L

L

L

L

L

A

A

A

A

A

A

A

A

A

A

Spectral-Stream

Spatial-Stream

7 of 143

L

L

L

L

L

L

L

L

C

C

C

C

C

C

C

C

L

L

L

L

L

L

L

L

A

A

A

A

A

A

A

A

Spectral-Stream

Spatial-Stream

EEG of all channels in time domain

(0-1s, downsampling to 100Hz)

EEG of all channels in frequency domain

(0-50Hz, average of 0-4s (one trial))

Spatial Topographic Maps

(25 frames)

Spectral Topographic Maps

(25 frames)

(a) EEG Topographic Maps Generation

8 of 143

L

L

L

L

L

L

L

L

C

C

C

C

C

C

C

C

L

L

L

L

L

L

L

L

A

A

A

A

A

A

A

A

Spectral-Stream

Temporal-Stream

EEG of all channels in time domain

(0-1s, downsampling to 100Hz)

EEG of all channels in frequency domain

(0-50Hz, average of 0-4s (one trial))

Spatial Topographic Maps

(25 frames)

Spectral Topographic Maps

(25 frames)

  1. EEG Topographic Maps Generation

(b) Spatial-Spectral-Temporal Feature Learning

BiLSTM

CNN

ANN

FC

Layer

Softmax

Layer

9 of 143

L

L

L

L

L

L

L

L

C

C

C

C

C

C

C

C

L

L

L

L

L

L

L

L

A

A

A

A

A

A

A

A

Spectral-Stream

EEG of all channels in time domain

(0-1s, downsampling to 100Hz)

EEG of all channels in frequency domain

(0-50Hz, average of 0-4s (one trial))

Spatial Topographic Maps

(25 frames)

Spectral Topographic Maps

(25 frames)

(a) EEG Topographic Maps Generation

(b) Spatial-Spectral-Temporal Feature Learning

Bi-LSTM

CNN

ANN

FC

Layer

Softmax

Layer

Temporal-Stream

0ms

40ms

80ms

960ms

0Hz

2Hz

4Hz

48Hz

10 of 143

L

L

L

L

L

L

L

L

C

C

C

C

C

C

C

C

L

L

L

L

L

L

L

L

A

A

A

A

A

A

A

A

Spectral-Stream

Amplitude of all channels in time domain

(0-1s, downsampling to 100Hz)

PSD of all channels in frequency domain

(0-50Hz, average of 0-4s (one trial))

Spatial Topographic Maps

(25 frames)

Spectral Topographic Maps

(25 frames)

(a) EEG Topographic Maps Generation

(b) Spatial-Spectral-Temporal Feature Learning

Bi-LSTM

CNN

ANN

FC

Layer

Softmax

Layer

Temporal-Stream

11 of 143

Bicubic

Amplitude Distribution

(Interval: 40ms)

20ms

60ms

100ms

940ms

980ms

Power Distribution

(Interval: 2Hz)

1Hz

3Hz

5Hz

47Hz

49Hz

Spectral EEG Topographic Maps

(25 frames)

Temporal EEG Topographic Maps

(25 frames)

EEG of all channels in time domain (0-1s, downsampling to 100Hz)

EEG of all channels in frequency domain (0-50Hz, average of 0-4s (one trial))

Bicubic

12 of 143

13 of 143

L

L

L

L

L

L

L

L

L

L

C

C

C

C

C

C

C

C

C

C

L

L

L

L

L

L

L

L

L

L

A

A

A

A

A

A

A

A

A

A

14 of 143

head-like

topographic map

sensors position

box-like

topographic map

15 of 143

16 of 143

17 of 143

18 of 143

19 of 143

20 of 143

Basic components

21 of 143

22 of 143

23 of 143

24 of 143

25 of 143

Softmax

Convolve

Sharpen

26 of 143

Softmax

Convolve

Sharpen

27 of 143

Architectures

28 of 143

L

L

L

L

L

L

L

C

C

C

C

C

C

C

FC

SM

L

L

L

C

C

C

29 of 143

C

C

C

C

C

C

C

FC

SM

C

C

C

Conv

30 of 143

Level 1

Sample1

Sample2

Level 2

Sample1

Sample2

Level 3

Sample1

Sample2

Level 4

Sample1

Sample2

Frame1

Frame2

Frame3

Frame4

Frame5

Frame6

Frame7

Frame8

Frame9

Frame10

31 of 143

Frame1

Frame2

Frame3

Frame4

Frame5

Frame6

Frame7

Frame8

Frame9

Frame10

S7

S8

32 of 143

Frame1

Frame2

Frame3

Frame4

Frame5

Frame6

Frame7

Frame8

Frame9

Frame10

S7

S8

Level 1

Sample1

Sample2

Level 2

Sample1

Sample2

Level 3

Sample1

Sample2

Level 4

Sample1

Sample2

Frame1

Frame2

Frame3

Frame4

Frame5

Frame6

Frame7

Frame8

Frame9

Frame10

33 of 143

34 of 143

L

L

L

L

L

L

L

C

C

C

C

C

C

C

FC

SM

L

L

L

C

C

C

Conv

AM

35 of 143

L

L

L

L

L

L

L

C

C

C

C

C

C

C

FC

SM

L

L

L

C

C

C

Attention

36 of 143

A) CNN-LSTM

B) CNN-LSTM/1D-Conv

D) CNN-ANN-BiLSTM (ARCNN)

C) CNN-BiLSTM

37 of 143

A) LSTM

B) LSTM / 1D-Conv

D) Att-BiLSTM

C) BiLSTM

38 of 143

L

L

L

L

L

L

L

C

FC

SM

L

L

L

C

C

C

C

C

C

C

C

C

39 of 143

L

L

L

L

L

L

L

C

C

C

C

C

C

C

FC

SM

L

L

L

C

C

C

Conv

40 of 143

L

L

L

L

L

L

L

C

C

C

C

C

C

C

FC

SM

L

L

L

C

C

C

Conv

AM

L

L

L

L

L

L

L

L

L

L

41 of 143

L

L

L

L

L

L

L

C

C

C

C

C

C

C

FC

SM

L

L

L

C

C

C

Conv

AM

L

L

L

L

L

L

L

L

L

L

42 of 143

43 of 143

L

L

L

L

L

L

L

C

C

C

C

C

C

C

FC

SM

L

L

L

C

C

C

Conv

44 of 143

C

C

C

C

C

C

C

LSTM layer

Attention layer

CNN

layer

C

C

C

45 of 143

C

LSTM layer

Attention layer

CNN

layer

C

C

C

C

C

C

C

C

C

46 of 143

C

LSTM layer

Attention layer

CNN

layer

C

C

C

C

C

C

C

C

C

Output

47 of 143

L

L

L

L

L

L

L

C

FC

SM

L

L

L

C

C

C

C

C

C

C

C

C

48 of 143

L

L

L

L

L

L

L

AM

FC

SM

Conv

C

C

C

C

C

C

C

49 of 143

50 of 143

Input Layer

Hidden Layers

Output Layer

X = A[0]

a[4]

A[1]

A[3]

X

Ŷ

a[1]1

a[1]2

a[1]3

a[1]n

a[2]1

a[2]2

a[2]3

a[2]n

a[3]1

a[3]2

a[3]3

a[3]n

A[2]

A[4]

51 of 143

Input Layer

Hidden Layers

Output Layer

X = A[0]

a[4]

A[1]

A[3]

X

Ŷ

[1a]1

a[1]2

a[1]3

a[1]n

a[2]1

a[2]2

a[2]3

a[2]n

a[3]1

a[3]2

a[3]3

a[3]n

A[2]

A[4]

52 of 143

Input Layer

Hidden Layers

Output Layer

X = A[0]

a[4]

A[1]

A[3]

X

Ŷ

a[1]1

a[1]2

a[1]3

a[1]n

a[2]1

a[2]2

a[2]3

a[2]n

a[3]1

a[3]2

a[3]3

a[3]n

A[2]

A[4]

53 of 143

NxNx3

+b1

+b2

MxM

MxM

+b1

+b2

ReLU

ReLU

a[l]

MxMX2

a[l-1]

CONV operation

54 of 143

NxNx3

+b1

+b2

MxM

MxM

+b1

+b2

ReLU

ReLU

MxMX2

CONV operation

55 of 143

NxNx3

+b1

+b2

MxM

MxM

+b1

+b2

ReLU

ReLU

MxMX2

CONV operation

56 of 143

S=1

S=2

Striding in CONV

57 of 143

58 of 143

59 of 143

NxNx192

NxNx64

NxNx32

NxNx128

NxNx192

1x1 Same

3x3 Same

5x5 Same

MaxPool Same s=1

Inception Module

60 of 143

61 of 143

  1. Retraining w/o expansion

t-1

t

​

62 of 143

  1. No-Retraining w/ expansion
  1. Partial Retraining w/ expansion

63 of 143

  1. No-Retraining w/ expansion
  1. Partial Retraining w/ expansion

t-1

t

t

t-1

64 of 143

  1. No-Retraining expansion

t

  1. Partial Retraining expansion

t

t-1

  1. Retraining expansion

t-1

t

​

t-1

65 of 143

Positional

Encoding

Masked

Multi-Head

Attention

Add & Norm

Output

Embedding

Multi-Head

Attention

Add & Norm

Outputs(shifted right)

Positional

Encoding

Multi-Head

Attention

Add & Norm

Input

Embedding

Feed

Forward

Add & Norm

Inputs

Feed

Forward

Add & Norm

Linear

Softmax

66 of 143

Multi-Head

Attention

Add & Norm

Input

Embedding

Output

Embedding

Feed

Forward

Add & Norm

Masked

Multi-Head

Attention

Add & Norm

Multi-Head

Attention

Add & Norm

Feed

Forward

Add & Norm

Linear

Softmax

Inputs

Outputs (shifted right)

Positional

Encoding

Positional

Encoding

67 of 143

Tokenize

I

love

coding

and

writing

“I love coding and writing”

68 of 143

ML Concepts

69 of 143

Size

#bed

ZIP

Wealth

Family?

Walk?

School

PRICE ŷ

X

Y

X

Ŷ = 0

Ŷ = 1

How does NN work (Insprired from Coursera)

Logistic Regression

Basic Neuron Model

70 of 143

Size

$

Size

$

Linear regression

ReLU(x)

71 of 143

NxN

NxN

NxN

NxN

256

225

56

.

.

.

214

210

211

R-G-B

Unrolling Feature vectors

72 of 143

Large NN

Med NN

Small NN

SVM,LR etc

η

Amount of Data

Why does Deep learning work?

73 of 143

a[1]1

a[1]2

a[1]3

Input

Hidden

Output

X = A[0]

a[1]4

a[2]

A[1]

A[2]

X

Ŷ

One hidden layer neural network

74 of 143

a[1]1

a[1]2

x[1]

a[2]

x[2]

x[2]

x[3]

x[1]

Neural network templates

75 of 143

Train

Valid

Test

x1

x2

x1

x2

x1

x2

Train-Dev-Test vs. Model fitting

Underfitting

Good fit

Overfitting

76 of 143

x[2]

x[3]

x[1]

a[L]

x1

x2

r=1

x1

x2

DropOut

Normalization

77 of 143

w1

w1

w2

J

w1

w2

J

w1

w2

w2

Before Normalization

After Normalization

Early stopping

Dev

​

Train

Err

it.

78 of 143

x1

x2

w[1]

w[2]

w[L-2]

w[L-1]

w[L]

FN

TN

TP

FP

Deep neural networks

Understanding

Precision & Recall

​

79 of 143

w1

w2

SGD

BGD

w1

w2

SGD

Batch vs. Mini-batch �Gradient Descent

Batch �Gradient Descent vs. SGD

80 of 143

x[2]

x[3]

x[1]

p[1]

p[2]

Softmax Prediction with 2 outputs

81 of 143

Abstract backgrounds

82 of 143

83 of 143

dair.ai

84 of 143

85 of 143

86 of 143

Gradient Backgrounds

87 of 143

88 of 143

89 of 143

90 of 143

ML and Health

91 of 143

ICA

Conv

Conv

Conv

Conv

Conv

Conv

Time slice

Conv

Spatial Feature Learning

EEG Images

Alpha

RNN+ANN

CNN

CNN

CNN

CNN

CNN

CNN

CNN

CNN

CNN

CNN

Temporal Feature Aggregation

Pain Intensity Assessment

Delta

Alpha

Beta

Spectral Topography Maps

EEG Time Series

92 of 143

ICA

Conv

Conv

Conv

Conv

Conv

Conv

Time slice

Conv

Spatial Feature Learning

EEG Images

Alpha

RNN+ANN

CNN

CNN

CNN

CNN

CNN

CNN

CNN

CNN

CNN

CNN

Temporal Feature Aggregation

Pain Intensity Assessment

Delta

Alpha

Beta

Spectral Topography Maps

EEG Time Series

Pain

Localization

No pain

93 of 143

94 of 143

Conv

Conv

Conv

Conv

Conv

Conv

Conv

Conv

Conv

Conv

95 of 143

96 of 143

97 of 143

98 of 143

Level 1

Level 2

Level 3

Level 4

No Pain

Low Pain

Moderate Pain

High Pain

Signal Segmentation

Time

Withdraw

hand

Immerse

hand

99 of 143

Level-1

Level-2

Level-3

Level-4

No Pain

Time

Immerse

hand

Withdraw

hand

Low Pain

Medium Pain

High Pain

…

Sliding Window

Event-Marker

5s

100 of 143

Level 1

Level 2

Level 3

Level 4

No Pain

Low Pain

Medium Pain

High Pain

Time

Immerse

hand

Withdraw

hand

101 of 143

102 of 143

(a)

(b)

103 of 143

AEP

104 of 143

FFT

AEP

Delta

Alpha

Beta

Bicubic

EEG Image

(32x32x3)

105 of 143

FFT

PSD

Theta(4~8Hz)

Alpha(8~13Hz)

Beta(13~30Hz)

Image Generation

AEP

Bicubic

Spectral Topography Map

106 of 143

Conv3-32

x4

Maxpool

(2x2)

Conv3-64

x2

Maxpool

(2x2)

Conv3-

128

Maxpool

(2x2)

FC-512

Feature

Vector

Input

32x32x3

Output

ConvNet Configuration

107 of 143

Conv3-32

Conv3-32

Conv3-32

Max-Pool

Conv3-32

Conv3-128

Max-Pool

Conv3-64

Conv3-64

Max-Pool

Input

Conv

Conv

Max-Pool

Max-Pool

FC

Layer1

Softmax

FC

Layer2

Layer3

Layer4

Stack1

Stack2

Stack3

Input

Conv3-32

Conv3-32

Max-Pool

Conv3-128

Conv3-64

Conv3-64

Max-Pool

Layer1

Layer2

Layer3

FC-512

Output

Max-Pool

ConvNet Configuration

Stack4

Conv3-32

Conv3-32

Conv3-32

Conv3-32

Conv3-32

Conv3-32

Output

108 of 143

Input

Conv3-32

Conv3-32

Max-Pool

Conv3-128

Conv3-64

Conv3-64

Max-Pool

Layer1

Layer2

Layer3

Max-Pool

ConvNet Configuration

Conv3-32

Conv3-32

Conv3-32

Conv3-32

Conv3-32

Conv3-32

Output

Conv3-32

Conv3-32

Max-Pool

Conv3-128

Conv3-64

Conv3-64

Max-Pool

Layer1

Layer2

Layer3

Max-Pool

ConvNet Configuration

Conv3-32

Conv3-32

Conv3-32

Conv3-32

Conv3-32

Conv3-32

Output

109 of 143

Conv3-32

Conv3-32

Conv3-32

Max-Pool

Conv3-32

Conv3-128

Max-Pool

Conv3-64

Conv3-64

Max-Pool

Stack1

Stack2

Stack3

FC-512

Output

Stack4

110 of 143

Level 1

Level 2

Level 3

Level 4

Time

Level 5

No Pain

LowPain

MediumPain

High

Pain

Unbearable

Pain

(a)

(b)

111 of 143

Level 1

Level 2

Level 3

Level 4

Time

Level 5

No Pain

LowPain

MediumPain

High

Pain

Unbearable

Pain

(a)

(b)

112 of 143

113 of 143

114 of 143

Miscellaneous

115 of 143

3

64

16

16

32

32

64

128

128

256

256

128+256

128

1

64+128

64

32+64

32

16+32

16

16

Convolution 3x3

Max Pooling 2x2

Convolution 1x1

Skip connection

Up Sampling 2x2

Block copied

Dropout 0.1

Dropout 0.2

Dropout 0.3

116 of 143

Conv3-32

Conv3-32

Conv3-32

Max-Pool

Conv3-32

Conv3-128

Max-Pool

Conv3-64

Conv3-64

Max-Pool

Input

Conv

Conv

Max-Pool

Max-Pool

FC

Layer1

Softmax

FC

Layer2

Layer3

Layer4

Layer1

Layer2

Layer3

Layer4

Input

Conv3-32

Conv3-32

Conv3-32

Max-Pool

Conv3-32

Conv3-128

Conv3-64

Conv3-64

Max-Pool

Layer1

Layer2

Layer3

Feature

Vector

FC-512

Output

Max-Pool

FC-512

Output

117 of 143

Previous layer

1x1 convolutions

1x1 convolutions

3x3 convolutions

1x1 convolutions

5x5 convolutions

3x3 max pooling

1x1 convolutions

Filter concatenation

118 of 143

Previous layer

1x1 convolutions

1x1 convolutions

3x3 convolutions

1x1 convolutions

5x5 convolutions

3x3 max pooling

1x1 convolutions

Filter concatenation

119 of 143

Input

1x11 conv

1x11 conv

Inception 1

Inception 2

1x7 conv

1x7 conv

FC

FC

Output

Inception 2

Inception 2

120 of 143

Previous layer

1x3 conv,

1 padding

1x5 conv,

2 padding

1x3 conv,

1 padding

1x7 conv,

3 padding

Filter concatenation

1x3 conv,

1 padding

1x3 conv,

1 padding

121 of 143

Input

Conv

Max-Pool

Max-Pool

Max-Pool

Inception

Inception

Max-Pool

Conv

Max-Pool

Conv

Inception

Inception

Inception

Inception

Inception

Inception

Inception

Avg-Pool

Conv

FC

FC

Softmax

Avg-Pool

Conv

FC

Softmax

Avg-Pool

Conv

FC

FC

Softmax

Auxiliary Classifier

Auxiliary Classifier

122 of 143

Previous layer

1x1 conv.

1x1 conv.

3x3 conv.

1x1 conv.

3x3 conv.

Pool

1x1 conv.

Filter concatenation

3x3 conv.

Previous layer

1x1 conv.

1x1 conv.

1x1 conv.

3x3 conv.

Pool

1x1 conv.

Filter concatenation

1x3 conv.

3x1 conv.

1x3 conv.

3x1 conv.

(a)

(b)

123 of 143

R1

R2

R3

R1

R2

R3

R1

R1

R1

R2

Stacked layers

Previous input

x

F(x)

y=F(x)

Stacked layers

Previous input

x

F(x)

y=F(x)+x

x

identity

+

124 of 143

Input

Conv

Avg-Pool

Dense Block 2

​

​

​

Dense Block 3

​

​

​

Conv

Avg-Pool

Conv

Dense Block 1

​

​

​

Avg-Pool

FC

Softmax

Transition layers

125 of 143

3x3 conv

(a)

add

identity

3x3 conv

5x5 conv

3x3 avg

identity

3x3 avg

3x3 avg

3x3 conv

5x5 conv

add

add

add

add

Filter concatenation

hi

hi-1

...

hi+1

hi

hi-1

...

7x7 conv

5x5 conv

7x7 conv

3x3 max

5x5 conv

3x3 avg

add

add

add

identity

3x3 avg

3x3 max

3x3 conv

add

add

Filter concatenation

hi+1

(b)

126 of 143

127 of 143

128 of 143

1

1

2

4

5

6

7

8

3

2

1

0

1

2

3

4

6

8

3

4

Max(1,1,5,6) = 6

Image Representation

Y

X

Pooling performed with a 2x2 kernel and a stride of 2

129 of 143

ML System Design / Infrastructure

130 of 143

EEG

疼痛识别

疼痛等级

疼痛位置

APP

治疗力度

治疗方案

疼痛治疗仪

使用者的治疗时长、治疗方案、治疗反馈记录

131 of 143

Pain Recognition

EEG

Deep

Learning

Pain Localization

疼痛强度

APP

Page Design

Data

Mining

Reinforce

Learning

Pain Treatment Apparatus

Appearance

Design

Circuit

Design

治疗力度

治疗方案

Pain Management

治疗反馈

脑电检测

Closed-loop Control

132 of 143

应用场景

  • 普通患者:能够自己感觉到疼痛并报告出来。

​

    • 对于该类患者可以脱离脑电波头盔而单独存在,用户根据自身感觉,使用该疼痛治疗仪和app,将tens片放置疼痛源部位,可自行选择不同的治疗模式或使用我们的智能推荐模式进行治疗,有效地减缓疼痛。在该应用场景下,我们的产品依赖于人体的主观感觉和反馈,在强化学习和数据挖掘的基础上进行智能治疗。

​

    • 优点:使用范围广,成本低,无创伤,智能治疗

​

  • 特殊患者:无法感觉疼痛或无法报告出来,如老年痴呆症患者,婴幼儿,手术前/后,或处于昏迷状态的患者。

​

    • 对于该类患者,我们引入了脑机接口技术,通过检测患者的脑电活动,识别出疼痛强度和疼痛位置,将信息传至app端,便可根据该结果进行准确治疗。治疗过程中,患者脑电波能够做出一定的反馈,根据该反馈进行强化学习不断调整治疗参数,以达到有效且精确的治疗。

​

    • 优点:更为客观的基于脑电波的疼痛识别,较高的临床价值,无创伤,智能治疗
    • 缺点:脑电头盔成本高,使用范围较为局限

​

  • 脑电监测:仅需监控,无需治疗,可使用我们的疼痛识别算法系统(软件),可应用于手术中患者的疼痛监测等。

133 of 143

APP

页面设计

数据挖掘

强化学习

疼痛治疗仪

外观设计

电路设计

治疗力度

治疗方案

治疗反馈

主观疼痛感受

疼痛治疗

普通用户

134 of 143

脑电信号

采集/反馈

疼痛识别

EEG

深度学习

疼痛定位

疼痛强度

APP

页面设计

数据挖掘

强化学习

疼痛治疗仪

外观设计

电路设计

治疗力度

治疗方案

疼痛治疗

特殊患者

脑电监测

135 of 143

脑电信号

采集/反馈

疼痛识别

EEG

深度学习

疼痛定位

疼痛强度

APP

页面设计

数据挖掘

强化学习

疼痛治疗仪

外观设计

电路设计

治疗力度

治疗方案

疼痛治疗

主动反馈

136 of 143

EEG

机器挖掘

​

强化学习

​

治疗参数自动调整

治疗参数

治疗反馈

信号处理

​

特征提取

​

疼痛分类

可视化

治疗参数人为调整

服务器

用户端

疼痛缓解

治疗系统

Tens片

治疗方案

疼痛源

识别结果

​

治疗记录

​

个人信息

疼痛检测系统

精确治疗系统

137 of 143

不同场景—— 痛经治疗仪

138 of 143

139 of 143

140 of 143

脑电信号

采集/反馈

疼痛识别

EEG

深度学习

疼痛定位

疼痛强度

APP

页面

设计

数据挖掘

强化学习

疼痛治疗仪

外观设计

电路设计

治疗力度

治疗方案

疼痛治疗

治疗反馈

主观疼痛感受

特殊患者

141 of 143

142 of 143

143 of 143