LLM Continuous Batching: Refilling Slots

黎 浩然/ 10 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

Batching requests can reduce the overhead of processing them separately, but requests have different lengths. Whether a finished short request’s slot can be assigned promptly to another request affects utilization. Continuous batching reorganizes active requests between generation iterations.

A two-slot schedule

Suppose A, B and C are already waiting and require 1, 3 and 2 abstract steps. At most two requests can run together. A static batch processes A and B, waits for B, then starts C, taking 3+2=5 steps. Replacing A with C after A finishes completes all three in three steps.

Step 1: A, B
Step 2: C, B
Step 3: C, B
Original ideal schedule: C replaces completed A. Each request advances one unit per step; this is not GPU timing.

Total useful work is 1+3+2=6 slot-steps. The static schedule provides 2×5=10 slot-steps; ideal refilling provides 2×3=6. This explains idle capacity but does not establish a 5/3 speedup for a real engine.

Reproduce the schedule, not a benchmark

The executed code checks five static steps, three refill steps and six units of useful work. It omits prefill, sequence-dependent iteration costs, memory allocation, arrival times and scheduling overhead. Its simple list queue is for explanation, not a production scheduler.

lengths = [1,3,2]
slots = 2
static_steps = sum(max(lengths[i:i+slots]) for i in range(0,len(lengths),slots))
queue = list(lengths)
active = []
steps = 0
while queue or active:
    while queue and len(active)<slots:
        active.append(queue.pop(0))
    active = [n-1 for n in active if n>1]
    steps += 1
assert static_steps == 5 and steps == 3
assert sum(lengths)==6
print('static / refill steps:',static_steps,steps)
print('useful slot steps:',sum(lengths))

The official vLLM project lists continuous batching as a feature. That does not mean equal gains for every workload. Similar request lengths, insufficient queued work, memory limits or increased per-step cost make the ideal refill model inadequate for performance prediction.

Measure throughput and request latency together

Higher throughput can coexist with a longer wait for some requests. Record arrival, first token, later gaps, completion and failures. The client definitions in TTFT and token intervals can be compared with server scheduling metrics.

When varying concurrency, keep input and output length distributions comparable. Report completed requests and the measurement window; requests merely started are not completed throughput. Memory planning must include KV state for active sequences rather than relying only on slot count.

Continuous batching creates a more flexible scheduling opportunity. Actual throughput and tail-latency effects still require validation on the model, hardware and workload being served.

中文版

References

Official documentation

Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the planned 2026-10-10 19:00 article date. The code is an original teaching check, not a real-model performance test.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*