LLM Latency: TTFT and Token Intervals
A streaming answer may begin quickly and then stall. Another service may start slowly but deliver smoothly afterward. Total duration alone cannot distinguish these experiences. Record first-token latency and later token intervals separately.
Define the measurement boundary
On the client, let t0 be request dispatch, t1 the first actual content token’s arrival, and tn the last. Here TTFT=t1−t0, last-token latency=tn−t0, and mean subsequent interval=(tn−t1)/(n−1), requiring at least two tokens. With one token, the subsequent interval is undefined rather than zero.
These client definitions do not automatically equal internal inference-server timing. The official vLLM metrics distinguish first-token, inter-token and end-to-end latency. Verify version and boundaries before integrating monitoring, including whether queues, networking and final-response events are counted.
How averages hide pauses
The example has TTFT 0.6 seconds and last-token latency 1.5 seconds. Later gaps are approximately 0.2, 0.2 and 0.5 seconds, averaging 0.3. The mean hides the longer final wait; inspect within-request gaps and the overall distribution too.
from math import isclose
def measure(start, arrivals):
if not arrivals or arrivals[0] < start or any(b<a for a,b in zip(arrivals,arrivals[1:])):
raise ValueError('invalid timestamps')
gaps = [b-a for a,b in zip(arrivals,arrivals[1:])]
return arrivals[0]-start, arrivals[-1]-start, gaps, (sum(gaps)/len(gaps) if gaps else None)
ttft, end, gaps, mean_gap = measure(0, [.6,.8,1.,1.5])
assert isclose(ttft,.6) and isclose(end,1.5)
assert isclose(mean_gap,.3) and isclose(max(gaps),.5)
assert measure(0,[.6])[3] is None
print('TTFT:', ttft, 'last-token latency:',end)
print('gaps:',gaps, 'mean:',mean_gap)
The executed code checks first- and last-token latency, mean and maximum gaps, and the one-token boundary. A network chunk may contain multiple tokens. A streaming callback is not inherently one token; if only chunk timestamps are available, label the metric as chunk interval.
Match optimization to the waiting stage
For slow first tokens, investigate queuing, input processing and transport. For long later gaps, investigate generation, scheduling and client buffering. These numbers alone do not prove a component fault; server-side evidence is needed.
Compare deployments with controlled input length, output length, concurrency and cache state. Include tail latency and failures instead of selecting only successful short answers. End-to-end duration can include cleanup after the last token, so distinguish it from this example’s last-token latency.
Prefix caching can help the input stage without removing every output pause. Identify the metric that changed before claiming a better user experience.
References
Publication note: actually published as a catch-up on October 10, 2026 (Beijing time), retaining the planned 2026-10-10 14:00 article date. The code is an original teaching check, not a real-model performance test.
Support
If this article helped you, you can support this site.
Click a QR code to enlarge. More options: support page。


