LLM latency has two distinct components, and conflating them is a common mistake.
Time-to-first-token (TTFT) is the wait before output begins. It includes prompt processing, network time, queueing, retrieval, and checks that run before generation. TTFT grows with input length.
Total latency (or end-to-end time) is TTFT plus the decode phase, where tokens are generated one at a time. Decode is memory-bandwidth-bound and roughly linear in the number of output tokens.
Why both matter: for interactive UX, TTFT is what determines whether the app feels responsive. With streaming, the user starts reading as soon as the first tokens arrive.
Design implication: to cut TTFT, shrink the prompt (trim retrieved context, use prompt caching), and stream. To cut total time, cap output length, use a faster model, or split work.
This answer doesn't lend itself to a diagram - it reads best . No credits were charged.
Why there's no diagram: “”
The interactive diagram is below the answer - jump to diagram ↓ · Below it, the related concept . Jump to it ↓
The diagram below the answer is the concept . Jump to it ↓