1 of 1

higher frequencies

past time points

  • Normalization in auditory systems is thought to be local, whereas deep neural networks commonly normalize using global statistics
  • In place of standard normalization, we learn normalization kernels, representing a local neighborhood of the receptive field, to simulate divisive normalization in CNNs
  • Divisive normalization is a canonical neural computation known to enhance local contrast and reduce statistical dependencies in input stimuli, both critical to many auditory tasks[1,2]
  • Questions we'd like to ask:

Do structured normalization kernels emerge from task optimization?

Do they produce suppressive effects across time and frequency?

Do they support human-level task performance and generalization ability?

Introduction

Modeling Divisive Normalization in the Central Auditory

Pathway With Convolutional Neural Networks (CNNs)

Chanel Cheng, Annesya Banerjee, Josh McDermott

McGovern Institute for Brain Research, CBMM, MIT BCS

Methodology

  • Normal hearing (NH) cochlea model produces nervegram representations of sounds across time and frequency
  • CNN model with divisive normalization receives nervegram input and computes corresponding task label
  • ๐›พ, ฮฒ, ๐€, ษ‘ are learnable parameters that shape the local divisive neighborhood

Learned Network Results

This work was supported by MSRP Bio 2023, with funding from MIT's School of Science, the National Science Foundation, and the Simons Center for the Social Brain.

  • CNN trained to identify word spoken at the midpoint of sentence excerpts with background noise
  • 794 word labels and ~250k training samples
  • Signal to Noise Ratio from -10 to +10 dB

2 sec. sentence excerpt

background noise

Suppression effect observed from neuro- physiological recordings

conv 3

conv 0

conv 1

conv 2

conv 4

conv 5

conv 6

dense

div norm input

div norm 0

div norm 2

div norm 1

div norm 3

div norm 4

div norm 5

div norm 6

central stochasticity (additive noise)

Speech Recognition Task

Conclusion

  • Learned local normalization kernels reproduce effects of adaptation
  • Suppressive effects found in biology were not fully replicated, for reasons we do not currently understand
  • Human-level task performance is exhibited with improved generalizability to new inputs

xi

((๐€/ษ‘)ฮฃj ๐›พj xj + ๐ˆ)แต

yi =

  • 1D kernel weights were learned independently for each dimension
  • Causal kernel only normalized across past time points
  • Frequency-time points contributing more to normalization were assigned higher weights

Weights relevant to speech recognition task learned across time and frequency

Learned Normalization Kernels

Auditory nerve response is the CNN input

References

References can be found by scanning the QR code to the right or by visiting: https://qrco.de/beCngc

CNN has 54,943,911 optimizable parameters

Divisive norm:

behavioral judgement โ€”โ€” task label

  • Response to excitatory tone is reduced by suppressor tone[6]

Two-tone suppression effect

Simulated central auditory pathway

Simulated normal cochlea

  • Response to ongoing stimuli decays to a steady- state level[7]

Adaptation to ongoing stimuli

Two-tone suppression effect is mostly absent after training for divisive norm

abnormal response pattern

limited suppression effect

standard norm

divisive norm

excitatory tone (2kHz),

suppressor tone (3kHz)

Learned divisive norm produces greater adaptation in response to continuous stimuli

firing rate reaches different steady-state levels

human data [-9 to +3 dB]

Both models perform similarly to humans on task

Performance for out-of-distribution speech is improved for model with divisive norm

whispered speech

time-reversed speech

noise-vocoded speech

Time (ms)

Frequency (kHz)

๐€

target neuron xi

๐›พ