1 of 32

Entropy and Private Language Models

Nandan Kumar Jha, Brandon Reagen

04/09/2025

2 of 32

Private Inference (PI)

Client's input privacy is preserved, and the server’s model is protected

Client

Raw Input

Encrypted Input

Encryption

Decryption

Encrypted Input

#2sf$6x98z

#2sf$6x98z

“Diabetes 0.82”

Encrypted prediction

Encrypted prediction

Trained model

Server

Encrypted Input

3 of 32

A Brief History of Cryptographically Secure Private Inference

3

2009

FHE using ideal lattices [STOC]

2016

CryptoNets [ICML]

2018

GAZELLE

[USENIX Security]

2020

Delphi

[USENIX Security]

2021

DeepReDuce

[ICML]

2022

SNL

[ICML]

2023

SENets

[ICLR]

2024

DeepReShape

[TMLR]

4 of 32

Private Inference on Language Models

4

5 of 32

Motivation: Private-Inference on LLMs

5

High Latency & Bandwidth :

1 Token = 8.2 min, 25.3 GB (GPT-2, 125M)

Nonlinearities are the Key Bottleneck for PI1

  1. Hou et al., CipherGPT: Secure two-party GPT Inference.

6 of 32

Talk Agenda

  • Entropy Dynamics for Understanding the Role of Nonlinearities in LLMs
  • Normalization-Free LLM Architecture
  • Softmax-only LLM Architecture
  • Inference-efficient Alternatives for Normalization in LLMs
  • Entropy-guided Attention for Private LLMs
  • Key Design Principles for Entropy Regularization in Softmax-only Private LLMs
  • Conclusion and Key Takeaways
  • Future work: Latent Space Dynamics in Feed-Forward Networks

6

7 of 32

Entropy Dynamics of Language Models

7

8 of 32

Shannon’s Entropy for Attention-Score Distribution

8

Attention Matrix

Attention Score

Entropic measure of Attention Spread

Higher Entropy

Exploratory Attention

Lower Entropy

Focused Attention

Entropy offers an information-theoretic lens to analyze fine-grained information flow, and identify critical failure modes (e.g, entropy collapse) in Transformer architecture

Zhai et al., Stabilizing Transformer Training by Preventing Attention Entropy Collapse, ICML 2023.

9 of 32

Well-behaved (Attention-Weights) Entropy Distribution

9

SM+LN+G

SM+LN+R

GPT-2 model, trained on 2.1B tokens

10 of 32

Normalization-Free Language Models

10

11 of 32

ReLU Outperforms GELU in Normalization-Free LLMs

11

NeurIPS 2024, ATTRIB Workshop

Normalization-free models exhibit the opposite trend

12 of 32

Entropic Overload in Normalization-free LLMs

12

SM+G

SM+R

13 of 32

Near-Zero Negative Slopes Emerge Without Normalization

13

Layerwise Learnable Negative Slope

Global Learnable Negative Slope

Leaky ReLU w/ 5e-2 Negative Slope

14 of 32

LayerNorm and GELU-free Models

(Softmax-Only Models)

14

15 of 32

Entropy Dynamics and Training Instability in Softmax-only LLM

15

Early layers: Entropic Overload; Deeper layers: Entropy Collapses 1

  1. Zhai et al., Stabilizing Transformer Training by Preventing Attention Entropy Collapse, ICML 2023.

Degenerated Attention Patterns

Training Instability

16 of 32

Addressing Training Instability in Softmax-Only Model

16

Weight Normalization

Spectral Normalization

Scaled FFN

Eval PPL:

3.640

3.624

3.478

Salimans et al., Weight Normalization: A Simple Reparameterization to Accelerate Training of Deep Neural Networks, NeurIPS 2016.

Miyato et al., Spectral Normalization for Generative Adversarial Networks, ICLR 2018.

17 of 32

Addressing Entropic Overload in Softmax-only LLMs

17

Softmax learnable temperature

Attention-head Entropy is a function of learnable temp. value

t > 1

t < 1

More Uniform Softmax score distribution

More Peaked Softmax score distribution

[AAAI 2025, PPAI Workshop]

18 of 32

Key Design Principles for Entropy Regularization

18

19 of 32

19

Learnable Threshold for each head

Tolerance margin to prevent over-regaulization

20 of 32

Dynamic Thresholds with Head-Specific Adaptation

20

Learnable-thresholds foster Attention-head diversity in the absence of crucial nonlinearities in LLMs

21 of 32

Effect of Tolerance Margins in Entropy-Regularization (½)

21

22 of 32

Effect of Tolerance Margins in Entropy-Regularization (2/2)

22

23 of 32

Summary: Entropy-Guided Attention for Private LLMs

23

Entropy regularization is needed for preventing Entropic overload

Softmax-only Private LLM

24 of 32

Experimental Results

24

25 of 32

Experimental Results (½)

25

GPT-2 (L=12, H=12, d=768) Trained on 2.1B Tokens (CodeParrot dataset)

SOTA: He et al., Simplifying Transformer Blocks, ICLR 2024

26 of 32

Experimental Results (2/2)

26

GPT-2 (L=12, H=12, d=768) Trained on Languini-book dataset

27 of 32

Breakdown of Latency and Communication Savings

27

GPT-2 (L=12, H=12, d=768)

System Setup: AMD EPYC 7502 server (2.5 GHz, 32 cores, & 256 GB RAM)

WAN setting: 100Mbps, 80ms (#Threads = 32)

Lu et al., Bumblebee: Secure two-party inference framework for large transformers, NDSS 2025

28 of 32

Conclusion and Key-Takeaways

28

29 of 32

  1. Nonlinearities are the lifeblood of LLMs: their removal results in
  2. Entropy collapse in deeper layers (training instability)
  3. Entropic overload in early layers (hampers attention diversity )

  • Parametric normalization in FFNs is an effective technique to prevents entropy collapses in Softmax-only models

  • Entropy-regularization, when applied strategically, can prevent entropic overload in Softmax-only private LLMs.

29

30 of 32

Future Work: Feed-Forward Latent Space Dynamics

30

Spectral Entropy

Participation Ratio

Jensen–Shannon divergence

Eigenvalue Early Enrichment

31 of 32

How FFNs Adapt in the Absence of Normalization?

31

EEE (Post activations)

32 of 32

Thanks for Your Attention !��Q & A

32