1 of 24

Data Mining_Anoop Chaturvedi

1

Swayam Prabha

Course Title

Multivariate Data Mining- Methods and Applications

Lecture 26

Recurrent neural network and Projection Pursuit Regression

By

Anoop Chaturvedi

Department of Statistics, University of Allahabad

Prayagraj (India)

Slides can be downloaded from https://sites.google.com/view/anoopchaturvedi/swayam-prabha

2 of 24

Advantage of weight sharing in CNNs⇒

For a 4x4 image, if a 2x2 filter is passed with no stride, then the filter with four weights one per pixel is applied nine times, requiring 36 weights in all.

If same weights are used for all nine filters, it requires four weights.

Thus, weight sharing (i) reduces the number of weights to be learned (reduces model training time and cost).

(ii) Makes feature search insensitive to feature location in the image.

Data Mining_Anoop Chaturvedi

2

3 of 24

  •  

Data Mining_Anoop Chaturvedi

3

4 of 24

Local connectivity is the concept of each neural connected only to a subset of the input image.

Parameter sharing is the sharing of weights by all neurons in a particular feature map. It helps to reduce the number of parameters in the whole system and makes the computation more efficient.

Based on the assumption that if one feature is useful to compute at some spatial position, then it should also be useful to compute at a different position.

Data Mining_Anoop Chaturvedi

4

5 of 24

Receptive field: In a convolutional layer, each neuron receives input from a restricted area of the previous layer, called the neuron's receptive field.

In a fully connected layer, the receptive field is the entire previous layer.

Pooling layer

  • Reduce dimensions by combining the outputs of neuron clusters at one layer into a single neuron in the next layer. Local pooling combines small clusters, whereas global pooling acts on all the neurons of the feature map.

Data Mining_Anoop Chaturvedi

5

6 of 24

  •  

Data Mining_Anoop Chaturvedi

6

7 of 24

  •  

Data Mining_Anoop Chaturvedi

7

 

Cycas Flower

8 of 24

Sobel Kernel Matrix ⇒ Emphasize regions of high spatial intensity change in both horizontal and vertical directions�Combining 8x8 pixel input matrix with a 3x3 kernel. Output ⇒ 6x6 matrix data

Horizontal Vertical

Data Mining_Anoop Chaturvedi

8

9 of 24

Data Mining_Anoop Chaturvedi

9

10 of 24

Data Mining_Anoop Chaturvedi

10

Emboss kernel matrix ⇒Employed for generating a 3D effect on an image.

Central element corresponds to the current pixel being processed.

Surrounding elements represent the differences in intensity values between the current pixel and its neighboring pixels.

Negative coefficients indicate regions of lower intensity. Positive coefficients indicate regions of higher intensity.

11 of 24

Data Mining_Anoop Chaturvedi

11

Sharpening Kernel Matrix ⇒ Emphasize regions of high spatial intensity change in horizontal and vertical directions. Enhances the edges by increasing the difference in intensity between adjacent pixels.

Central element (6) represents the weight of the current pixel.

Surrounding elements (-1, -1) compute the difference in intensity between the central pixel and its neighbors.

The example uses Laplacian Kernel

12 of 24

Recurrent neural network and Convolutional neural network exhibit temporal dynamic behavior.

Convolutional neural network ⇒ Class of networks with a finite impulse response.

Impulse response ⇒ Output of a dynamic system when presented with a brief input signal

Recurrent neural network

  • Class of networks with an infinite impulse response
  • A bi-directional artificial neural network. �

Econometrics_Anoop Chaturvedi

12

13 of 24

  • Allows the output from some nodes to affect subsequent input to the same nodes.
  • Uses internal state (memory) to process arbitrary sequences of inputs.
  • Applicable to tasks such as unsegmented, connected handwriting recognition or speech recognition.

Variants of RNNs�Fully recurrent neural networks (FRNN) ⇒ Connect the outputs of all neurons to the inputs of all neurons. All other topologies can be represented by setting some connection weights to zero.

Econometrics_Anoop Chaturvedi

13

14 of 24

 

Econometrics_Anoop Chaturvedi

14

15 of 24

Econometrics_Anoop Chaturvedi

15

Input

Hidden

Units

Output

Delay Units

Output

Hidden

Units

Input

Delay Units

 

 

 

 

 

 

 

 

 

Elman

Jordan

16 of 24

  •  

Data Mining_Anoop Chaturvedi

16

17 of 24

  •  

Data Mining_Anoop Chaturvedi

17

18 of 24

  •  

Data Mining_Anoop Chaturvedi

18

19 of 24

  •  

Data Mining_Anoop Chaturvedi

19

20 of 24

  •  

Data Mining_Anoop Chaturvedi

20

21 of 24

  •  

Data Mining_Anoop Chaturvedi

21

22 of 24

  •  

Data Mining_Anoop Chaturvedi

22

23 of 24

  • PPR model first projects the data matrix of input variables in the optimal direction. Then applies smoothing functions to these input variables.
  • Uses univariate regression functions, which allows for simple and efficient estimation and effectively deals with the curse of dimensionality
  • PPR effectively ignores variables with low explanatory power.
  • Transformations of variables in PPR are data driven whereas in a single-layer neural network transformations are fixed.

Data Mining_Anoop Chaturvedi

23

24 of 24

Projection pursuit regression and neural networks

  • Both approaches take nonlinear functions of linear combinations (derived features) of the inputs.
  • Work well for regression and classification and compete with the best learning methods on many problems.
  • Effective in problems with a high signal-to-noise ratio.
  • Not effective, if the goal is to describe the physical process that generated the data and the roles of individual inputs.

Econometrics_Anoop Chaturvedi

24