Stream2LLM: Overlap Context Streaming and Prefill for Reduced Time-to-First-Token
Rajveer Bachkaniwala, Chengqi Luo, Richard So, Divya Mahajan, Kexin Rong
Georgia Institute of Technology
1
LLMs benefit from context retrieval
2
Retrieval
External
Knowledge Source
I have a question…
parametric
knowledge
External
Context
Context retrieval hurts request latency
Retrieval latency
3
Doc
Retrieval
Generation
t=0
Prefill (Doc)
TTFT
+ Prefill latency
Time to First Token =
Decode
Time
Stream2LLM: Streaming LLM Inference
Goal: Overlap retrieval and prefill to reduce TTFT
Deployment setting: prefill instances in P/D disaggregation
* Implemented on top of vLLM
4
D1
D2
Prefill (D1)
Prefill (D2)
Decode
Retrieval
Generation
t=0
Prefill (D1)
Prefill (D2)
Decode
Streaming
Baseline
TTFT
TTFT
Challenges with Context Streaming
Each iteration, the scheduler allocates compute (token budget) and memory (KV cache) across requests without knowing:
Prior work:
5
Gap: concurrent, streaming inference.
Characterize Retrieval Workloads
6
Two distinct behavior modes:
Require different scheduling and cache management mechanisms
Two-Phase Scheduling
7
Phase 1: What to run Select requests based on priority; check feasibility �
Phase 1: Priority Ordering & Feasibility
Scheduling Priority
�LCAS | FCFS | MCPS | …
Feasibility check:
1. Token Budget
2. GPU Block Budget
Scheduled Requests
Unfinished��Requests
Separate scheduling decisions from resource allocation mechanisms
Two-Phase Scheduling
8
Phase 1: What to run Select requests based on priority; check feasibility �
Phase 2: How to allocate If allocation fails: preempt lowest-priority running request;
Phase 1: Priority Ordering & Feasibility
Scheduling Priority
�LCAS | FCFS | MCPS | …
Feasibility check:
1. Token Budget
2. GPU Block Budget
Phase 2: Resource Acquisition & Preemption
Scheduled Requests
GPU Block �Allocator
Unfinished��Requests
Success
Allocated
Free
GPU Block Pool
CPU Block Pool
Separate scheduling decisions from resource allocation mechanisms
Two-Phase Scheduling
9
Phase 1: What to run Select requests based on priority; check feasibility�
Phase 2: How to allocate If allocation fails: preempt lowest-priority running request;
Phase 1: Priority Ordering & Feasibility
Scheduling Priority
�LCAS | FCFS | MCPS | …
Feasibility check:
1. Token Budget
2. GPU Block Budget
Phase 2: Resource Acquisition & Preemption
Scheduled Requests
GPU Block �Allocator
Unfinished��Requests
Allocated
Free
GPU Block Pool
CPU Block Pool
Adaptive Preemption
Eviction Strategy
Failure
(OOM)
Separate scheduling decisions from resource allocation mechanisms
Two Phase Scheduling
10
Phase 1- What to run: Select requests based on priority; check feasibility�
Phase 2 - How to allocate: If allocation fails: preempt lowest-priority running request; choose recompute or swap (based on cost model)
Phase 1: Priority Ordering & Feasibility
Scheduling Priority
�LCAS | FCFS | MCPS | …
Feasibility check:
1. Token Budget
2. GPU Block Budget
Phase 2: Resource Acquisition & Preemption
Scheduled Requests
GPU Block �Allocator
Unfinished��Requests
Adaptive Preemption
Eviction Strategy
Failure
(OOM)
CPU Block Pool
Allocated
Recompute
Free
Recompute
GPU Block Pool
Separate scheduling decisions from resource allocation mechanisms
Two Phase Scheduling
11
Phase 1- What to run: Select requests based on priority; compute scheduling feasibility�
Phase 2 - How to allocate: If allocation fails: preempt lowest-priority running request; choose recompute or swap (based on cost model)
Phase 1: Priority Ordering & Feasibility
Scheduling Priority
�LCAS | FCFS | MCPS | …
Feasibility check:
1. Token Budget
2. GPU Block Budget
Phase 2: Resource Acquisition & Preemption
Scheduled Requests
GPU Block �Allocator
Unfinished��Requests
Adaptive Preemption
Eviction Strategy
Failure
(OOM)
Swap
CPU Block Pool
Allocated
Recompute
Swap
Free
Recompute
Swap
GPU Block Pool
Separate scheduling decisions from resource allocation mechanisms
Scheduling Policies�
Default (FIFO)
Most processed tokens (MCPS)
First-Come-First-Served (FCFS)
Last chunk arrival time (LCAS)
12
Two-tier approach:
�*Eviction in reverse scheduling priority
Evaluation Setup
Model: Llama-3.1-8B-Instruct
Hardware: H200 (TP=2)
Baselines:
�Retrieval Workloads:
13
Evaluation (Append): TTFT vs QPS
14
Methods of Comparison
better
Evaluation (Update): TTFT vs QPS
15
Methods of Comparison
Stream2LLM Summary
Overlapping retrieval with prefill to reduce TTFT
16
D1
D2
Prefill (D1)
Prefill (D2)
Decode
Retrieval
Generation
t=0
Prefill (D1)
Prefill (D2)
Decode
Streaming
Baseline
TTFT
TTFT
Scan for Code!
Evaluation: Memory Pressure
Stress test: chunk inter-arrival delays to saturate GPU block pool
17
Naïve streaming algorithms could collapse under memory pressure!