In Praise of Zeros
Sparsity across the reinforcement learning pipeline.
1. Signal
2. Rollout
3. Rewards
4. Updates
5. Sync
Ø
Ø
Ø
θt
xt
Abhishek Harshvardhan Mishra · @tokenbender
SHEET 01
Personal Context
NAME
Abhishek Harshvardhan Mishra
BIO
Research for my lab Deus Experiments as @tokenbender
FOCUS
RL post-training
Sparsity & efficiency
Capability extraction
BUILDS
avataRL (Reinforcement learning from random weights)
PRISM (Extracting function calling with less than 30% of the network)
CodeCherryPop (First smart Llama2 coder)
PIC (One of the earliest general purpose and tool calling models in open source)
@tokenbender
SHEET 02
Sparsity Across the RL Pipeline
→
Compute is spent across the whole loop
└→
Sparse signal
└→
Sparse rollout
└→
Sparse rewards
└→
Sparse updates
→
Sparsity is where we decide how to spend the compute
SHEET 03
Sparse Signal to Selective Prompts
→
Not every prompt is worth RL compute
└→
Too easy, got it already
no gains
└→
Too hard, no reward
no reps
└→
Almost there, train here
growth zone
SHEET 04
Sparse Rollout to Cheaper Generation
→
Expert sparsity: only a few experts needed per token
└→
Higher throughput
└→
Expert parallelism
└→
Router replay
→
Sparse attention / KV cache: each token looks at fewer tokens
└→
Cheaper long-context rollouts
prefill ×1
Prompt prefix
rollout 1
rollout 2
rollout N
Run prefill once on the shared prefix,
reuse the KV cache across rollouts
SHEET 05
Sparse Rewards to Cheaper Judging
→
Selective reward evaluation
└→
Judge only promising rollouts
└→
Use cheaper heuristics first
→
Verifier pipeline
└→
Check process or outcome
└→
Skip redundant traces
rollouts
Cheap checks
Ø
Ø
Ø
Expensive
judge
Most rollouts end in a zero;
only useful evidence gets judged
SHEET 06
Sparse Updates to Cheaper Learning
→
Delta or adapter sync
└→
Don't resend the whole model if only a small update changed
→
Precision-aware rollouts
└→
Rollouts might run in BF16 or quantized, but you need
to watch the train vs inference match
Trainer
BF16
Rollout workers
FP8 / quantized
full weights
Ø
sync deltas / adapters
Ship the delta, not the model
SHEET 07
Takeaway
“
”
Sparsity is not just LoRA or FP8 compression.
It is a loop-wide design principle
and the key to scalable RL infrastructure.
Ø
SHEET 08