Entropy and Private Language Models
Nandan Kumar Jha, Brandon Reagen
04/09/2025
Private Inference (PI)
Client's input privacy is preserved, and the server’s model is protected
Client
Raw Input
Encrypted Input
Encryption
Decryption
Encrypted Input
#2sf$6x98z
#2sf$6x98z
“Diabetes 0.82”
Encrypted prediction
Encrypted prediction
Trained model
Server
Encrypted Input
A Brief History of Cryptographically Secure Private Inference
3
2009
FHE using ideal lattices [STOC]
2016
CryptoNets [ICML]
2018
GAZELLE
[USENIX Security]
2020
Delphi
[USENIX Security]
2021
DeepReDuce
[ICML]
2022
SNL
[ICML]
2023
SENets
[ICLR]
2024
DeepReShape
[TMLR]
Private Inference on Language Models
4
Motivation: Private-Inference on LLMs
5
High Latency & Bandwidth :
1 Token = 8.2 min, 25.3 GB (GPT-2, 125M)
Nonlinearities are the Key Bottleneck for PI1
Talk Agenda
6
Entropy Dynamics of Language Models
7
Shannon’s Entropy for Attention-Score Distribution
8
Attention Matrix
Attention Score
Entropic measure of Attention Spread
Higher Entropy
Exploratory Attention
Lower Entropy
Focused Attention
Entropy offers an information-theoretic lens to analyze fine-grained information flow, and identify critical failure modes (e.g, entropy collapse) in Transformer architecture
Zhai et al., Stabilizing Transformer Training by Preventing Attention Entropy Collapse, ICML 2023.
Well-behaved (Attention-Weights) Entropy Distribution
9
SM+LN+G
SM+LN+R
GPT-2 model, trained on 2.1B tokens
Normalization-Free Language Models
10
ReLU Outperforms GELU in Normalization-Free LLMs
11
NeurIPS 2024, ATTRIB Workshop
Normalization-free models exhibit the opposite trend
Entropic Overload in Normalization-free LLMs
12
SM+G
SM+R
Near-Zero Negative Slopes Emerge Without Normalization
13
Layerwise Learnable Negative Slope
Global Learnable Negative Slope
Leaky ReLU w/ 5e-2 Negative Slope
LayerNorm and GELU-free Models
(Softmax-Only Models)
14
Entropy Dynamics and Training Instability in Softmax-only LLM
15
Early layers: Entropic Overload; Deeper layers: Entropy Collapses 1
Degenerated Attention Patterns
Training Instability
Addressing Training Instability in Softmax-Only Model
16
Weight Normalization
Spectral Normalization
Scaled FFN
Eval PPL:
3.640
3.624
3.478
Salimans et al., Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks, NeurIPS 2016.
Miyato et al., Spectral Normalization for Generative Adversarial Networks, ICLR 2018.
Addressing Entropic Overload in Softmax-only LLMs
17
Softmax learnable temperature
Attention-head Entropy is a function of learnable temp. value
t > 1
t < 1
More Uniform Softmax score distribution
More Peaked Softmax score distribution
[AAAI 2025, PPAI Workshop]
Key Design Principles for Entropy Regularization
18
19
Learnable Threshold for each head
Tolerance margin to prevent over-regaulization
Dynamic Thresholds with Head-Specific Adaptation
20
Learnable-thresholds foster Attention-head diversity in the absence of crucial nonlinearities in LLMs
Effect of Tolerance Margins in Entropy-Regularization (½)
21
Effect of Tolerance Margins in Entropy-Regularization (2/2)
22
Summary: Entropy-Guided Attention for Private LLMs
23
Entropy regularization is needed for preventing Entropic overload
Softmax-only Private LLM
Experimental Results
24
Experimental Results (½)
25
GPT-2 (L=12, H=12, d=768) Trained on 2.1B Tokens (CodeParrot dataset)
SOTA: He et al., Simplifying Transformer Blocks, ICLR 2024
Experimental Results (2/2)
26
GPT-2 (L=12, H=12, d=768) Trained on Languini-book dataset
Breakdown of Latency and Communication Savings
27
GPT-2 (L=12, H=12, d=768)
System Setup: AMD EPYC 7502 server (2.5 GHz, 32 cores, & 256 GB RAM)
WAN setting: 100Mbps, 80ms (#Threads = 32)
Lu et al., Bumblebee: Secure two-party inference framework for large transformers, NDSS 2025
Conclusion and Key-Takeaways
28
29
Future Work: Feed-Forward Latent Space Dynamics
30
Spectral Entropy
Participation Ratio
Jensen–Shannon divergence
Eigenvalue Early Enrichment
How FFNs Adapt in the Absence of Normalization?
31
EEE (Post activations)
Thanks for Your Attention !��Q & A
32