LLM Continuous Batching: Refilling Slots
Batching requests can reduce the overhead of processing them separately, but requests have different lengths. Whether a finished short request’s slot can be assigned promptly to another request affects utilization. Continuous batching reorganizes active requests between generation iterations.
A two-slot schedule
Suppose A, B and C are already waiting and require 1, 3 and 2 abstract steps. At most two requests can run together. A static batch processes A and B, waits for B, then starts C, taking 3+2=5 steps. Replacing A with C after A finishes completes all three in three steps.
Total useful work is 1+3+2=6 slot-steps. The static schedule provides 2×5=10 slot-steps; ideal refilling provides 2×3=6. This explains idle capacity but does not establish a 5/3 speedup for a real engine.
Reproduce the schedule, not a benchmark
The executed code checks five static steps, three refill steps and six units of useful work. It omits prefill, sequence-dependent iteration costs, memory allocation, arrival times and scheduling overhead. Its simple list queue is for explanation, not a production scheduler.
lengths = [1,3,2]
slots = 2
static_steps = sum(max(lengths[i:i+slots]) for i in range(0,len(lengths),slots))
queue = list(lengths)
active = []
steps = 0
while queue or active:
while queue and len(active)<slots:
active.append(queue.pop(0))
active = [n-1 for n in active if n>1]
steps += 1
assert static_steps == 5 and steps == 3
assert sum(lengths)==6
print('static / refill steps:',static_steps,steps)
print('useful slot steps:',sum(lengths))
The official vLLM project lists continuous batching as a feature. That does not mean equal gains for every workload. Similar request lengths, insufficient queued work, memory limits or increased per-step cost make the ideal refill model inadequate for performance prediction.
Measure throughput and request latency together
Higher throughput can coexist with a longer wait for some requests. Record arrival, first token, later gaps, completion and failures. The client definitions in TTFT and token intervals can be compared with server scheduling metrics.
When varying concurrency, keep input and output length distributions comparable. Report completed requests and the measurement window; requests merely started are not completed throughput. Memory planning must include KV state for active sequences rather than relying only on slot count.
Continuous batching creates a more flexible scheduling opportunity. Actual throughput and tail-latency effects still require validation on the model, hardware and workload being served.
References
Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the planned 2026-10-10 19:00 article date. The code is an original teaching check, not a real-model performance test.
Support
If this article helped you, you can support this site.
Click a QR code to enlarge. More options: support page。


