LLM Latency: TTFT and Token Intervals

黎 浩然/ 10 10 月, 2026/ 大语言模型/LARGELANGUAGEMODEL/LLM, 机器学习/MACHINELEARNING, 研究生/POSTGRADUATE, 计算机/COMPUTER/ 0 comments

A streaming answer may begin quickly and then stall. Another service may start slowly but deliver smoothly afterward. Total duration alone cannot distinguish these experiences. Record first-token latency and later token intervals separately.

Define the measurement boundary

On the client, let t0 be request dispatch, t1 the first actual content token’s arrival, and tn the last. Here TTFT=t1−t0, last-token latency=tn−t0, and mean subsequent interval=(tn−t1)/(n−1), requiring at least two tokens. With one token, the subsequent interval is undefined rather than zero.

Dispatch: 0 s
First token: 0.6 s
Next: 0.8, 1.0 s
Last token: 1.5 s
Original arrival timeline with invented timestamps; each event represents one token in this example.

These client definitions do not automatically equal internal inference-server timing. The official vLLM metrics distinguish first-token, inter-token and end-to-end latency. Verify version and boundaries before integrating monitoring, including whether queues, networking and final-response events are counted.

How averages hide pauses

The example has TTFT 0.6 seconds and last-token latency 1.5 seconds. Later gaps are approximately 0.2, 0.2 and 0.5 seconds, averaging 0.3. The mean hides the longer final wait; inspect within-request gaps and the overall distribution too.

from math import isclose

def measure(start, arrivals):
    if not arrivals or arrivals[0] < start or any(b<a for a,b in zip(arrivals,arrivals[1:])):
        raise ValueError('invalid timestamps')
    gaps = [b-a for a,b in zip(arrivals,arrivals[1:])]
    return arrivals[0]-start, arrivals[-1]-start, gaps, (sum(gaps)/len(gaps) if gaps else None)

ttft, end, gaps, mean_gap = measure(0, [.6,.8,1.,1.5])
assert isclose(ttft,.6) and isclose(end,1.5)
assert isclose(mean_gap,.3) and isclose(max(gaps),.5)
assert measure(0,[.6])[3] is None
print('TTFT:', ttft, 'last-token latency:',end)
print('gaps:',gaps, 'mean:',mean_gap)

The executed code checks first- and last-token latency, mean and maximum gaps, and the one-token boundary. A network chunk may contain multiple tokens. A streaming callback is not inherently one token; if only chunk timestamps are available, label the metric as chunk interval.

Match optimization to the waiting stage

For slow first tokens, investigate queuing, input processing and transport. For long later gaps, investigate generation, scheduling and client buffering. These numbers alone do not prove a component fault; server-side evidence is needed.

Compare deployments with controlled input length, output length, concurrency and cache state. Include tail latency and failures instead of selecting only successful short answers. End-to-end duration can include cleanup after the last token, so distinguish it from this example’s last-token latency.

Prefix caching can help the input stage without removing every output pause. Identify the metric that changed before claiming a better user experience.

中文版

References

Official documentation

Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the planned 2026-10-10 14:00 article date. The code is an original teaching check, not a real-model performance test.

Support

If this article helped you, you can support this site.

WeChat support QR code; click to enlarge
WeChat
Alipay support QR code; click to enlarge
Alipay
Buy Me a Coffee; support this site
Buy Me a Coffee

Click a QR code to enlarge. More options: support page。

Share this Post

Leave a Comment

您的邮箱地址不会被公开。 必填项已用 * 标注

*
*