Normalization in auditory systems is thought to be local, whereas deep neural networks commonly normalize using global statistics
In place of standard normalization, we learn normalization kernels, representing a local neighborhood of the receptive field, to simulate divisive normalizationin CNNs
Divisive normalization is a canonical neural computation known to enhance local contrast and reduce statistical dependencies in input stimuli, both critical to many auditory tasks[1,2]
Questions we'd like to ask:
Do structured normalization kernels emerge from task optimization?
Do they produce suppressive effects across time and frequency?
Do they support human-level task performance and generalization ability?
Introduction
Modeling Divisive Normalization in the Central Auditory
Pathway With Convolutional Neural Networks (CNNs)
Chanel Cheng, Annesya Banerjee, Josh McDermott
McGovern Institute for Brain Research, CBMM, MIT BCS
Methodology
Normal hearing (NH) cochlea model produces nervegram representations of sounds across time and frequency
CNN model with divisive normalization receives nervegram input and computes corresponding task label
๐พ, ฮฒ, ๐, ษ are learnable parameters that shape the local divisive neighborhood
Learned Network Results
This work was supported by MSRP Bio 2023, with funding from MIT's School of Science, the National Science Foundation, and the Simons Center for the Social Brain.
CNN trained to identify word spoken at the midpoint of sentence excerptswith background noise
794 word labels and ~250k training samples
Signal to Noise Ratio from -10 to +10 dB
2 sec. sentence excerpt
background noise
Suppression effect observed from neuro- physiological recordings
conv 3
conv 0
conv 1
conv 2
conv 4
conv 5
conv 6
dense
div norm input
div norm 0
div norm 2
div norm 1
div norm 3
div norm 4
div norm 5
div norm 6
central stochasticity (additive noise)
Speech Recognition Task
Conclusion
Learned local normalization kernels reproduce effects of adaptation
Suppressive effects found in biology were not fully replicated, for reasons we do not currently understand
Human-level task performance is exhibited with improved generalizability to new inputs
xi
((๐/ษ)ฮฃj๐พjxj + ๐)แต
yi =
1D kernel weights were learned independently for each dimension
Causal kernel only normalized across past time points
Frequency-time points contributing more to normalization were assigned higher weights
Weights relevant to speech recognition task learned across time and frequency
Learned Normalization Kernels
Auditory nerve response is the CNN input
References
References can be found by scanning the QR code to the right or by visiting: https://qrco.de/beCngc
CNN has 54,943,911 optimizable parameters
Divisive norm:
behavioral judgement โโ task label
Response to excitatory tone is reduced by suppressor tone[6]
Two-tone suppression effect
Simulated central auditory pathway
Simulated normal cochlea
Response to ongoing stimuli decays to a steady- state level[7]
Adaptation to ongoing stimuli
Two-tone suppression effect is mostly absent after training for divisive norm
abnormal response pattern
limited suppression effect
standard norm
divisive norm
excitatory tone (2kHz),
suppressor tone (3kHz)
Learned divisive norm produces greater adaptation in response to continuous stimuli
firing rate reaches different steady-state levels
human data [-9 to +3 dB]
Both models perform similarly to humans on task
Performance for out-of-distribution speech is improved for model with divisive norm