1 of 8

In Praise of Zeros

Sparsity across the reinforcement learning pipeline.

1. Signal

2. Rollout

3. Rewards

4. Updates

5. Sync

Ø

Ø

Ø

θt

xt

Abhishek Harshvardhan Mishra · @tokenbender

SHEET 01

2 of 8

Personal Context

NAME

Abhishek Harshvardhan Mishra

BIO

Research for my lab Deus Experiments as @tokenbender

FOCUS

RL post-training

Sparsity & efficiency

Capability extraction

BUILDS

avataRL (Reinforcement learning from random weights)

PRISM (Extracting function calling with less than 30% of the network)

CodeCherryPop (First smart Llama2 coder)

PIC (One of the earliest general purpose and tool calling models in open source)

@tokenbender

SHEET 02

3 of 8

Sparsity Across the RL Pipeline

Compute is spent across the whole loop

└→

Sparse signal

└→

Sparse rollout

└→

Sparse rewards

└→

Sparse updates

Sparsity is where we decide how to spend the compute

SHEET 03

4 of 8

Sparse Signal to Selective Prompts

Not every prompt is worth RL compute

└→

Too easy, got it already

no gains

└→

Too hard, no reward

no reps

└→

Almost there, train here

growth zone

SHEET 04

5 of 8

Sparse Rollout to Cheaper Generation

Expert sparsity: only a few experts needed per token

└→

Higher throughput

└→

Expert parallelism

└→

Router replay

Sparse attention / KV cache: each token looks at fewer tokens

└→

Cheaper long-context rollouts

prefill ×1

Prompt prefix

rollout 1

rollout 2

rollout N

Run prefill once on the shared prefix,

reuse the KV cache across rollouts

SHEET 05

6 of 8

Sparse Rewards to Cheaper Judging

Selective reward evaluation

└→

Judge only promising rollouts

└→

Use cheaper heuristics first

Verifier pipeline

└→

Check process or outcome

└→

Skip redundant traces

rollouts

Cheap checks

Ø

Ø

Ø

Expensive

judge

Most rollouts end in a zero;

only useful evidence gets judged

SHEET 06

7 of 8

Sparse Updates to Cheaper Learning

Delta or adapter sync

└→

Don't resend the whole model if only a small update changed

Precision-aware rollouts

└→

Rollouts might run in BF16 or quantized, but you need

to watch the train vs inference match

Trainer

BF16

Rollout workers

FP8 / quantized

full weights

Ø

sync deltas / adapters

Ship the delta, not the model

SHEET 07

8 of 8

Takeaway

Sparsity is not just LoRA or FP8 compression.

It is a loop-wide design principle

and the key to scalable RL infrastructure.

Ø

SHEET 08