§9.6
Chapter 09 · Real-Time Traffic Optimization for Low Latency and Jitter

§9.6Streaming LLM Inference Responses (SSE and Streamable HTTP) over HTTP/3

RFC 9114RFC 9000

Learning objective. See how token-by-token LLM inference streaming, usually delivered as Server-Sent Events (SSE) or a plain streamable HTTP body, maps onto HTTP/3, why it is a reliable real-time workload rather than a datagram one, and what QUIC specifically buys it: head-of-line isolation, faster first token, clean cancellation, and survival across network changes.

A different kind of real-time #

The rest of this chapter argued that real-time media wants unreliable delivery, because a late video frame is useless (§9.3). LLM inference streaming looks superficially similar (a low-latency, incremental stream a human watches appear in real time) but it inverts the key property: every token must arrive, in order. Drop token 5 of "The capital of France is Paris" and the output is corrupted; reorder them and it is gibberish. Late data is still entirely useful; you simply cannot skip it. So token streaming is a reliable, ordered real-time workload, and the right transport model is a reliable QUIC stream (§9.3), never a datagram. It is the instructive counter-case: real-time does not always mean loss-tolerant.

That reframes which chapter tools apply. Datagrams, FEC, and deadline-dropping are off the table. What matters instead is time to first token, inter-token latency and its jitter, isolation of concurrent requests, and keeping a long generation alive, all of which HTTP/3 addresses directly.

SSE and streamable HTTP are just DATA frames #

Most inference endpoints stream with one of two HTTP mechanisms, and over HTTP/3 they converge to the same thing on the wire:

  • Server-Sent Events (SSE). The response carries content-type: text/event-stream, and the body is a sequence of data: … events, one per token or token-group. SSE is a web standard defined at the HTTP-message level, so it rides any HTTP version unchanged.
  • Streamable HTTP. A plain response body (often newline-delimited JSON) flushed incrementally as the model generates. No special media type, just bytes sent over time.

Both are, in HTTP/3 terms, identical: a single HEADERS frame for the response status, then a series of DATA frames carrying the body as it is produced, then the stream ends ([RFC 9114 §4.1]). Crucially, HTTP/3 has no chunked transfer-encoding: "the Transfer-Encoding header field MUST NOT be used" ([RFC 9114 §4.1]). It does not need one, because the DATA-frame sequence is the incremental framing, and the QUIC stream's FIN delimits the end. Where HTTP/1.1 bolts streaming on with Transfer-Encoding: chunked, HTTP/3 streams natively.

the response on one request stream — a HEADERS frame, then a DATA frame per token, then FIN HEADERS 200 text/event-stream DATA data:{"tok": "The"} DATA data:{"tok": " capital"} · · · DATA data:[DONE] FIN each DATA frame flushed the moment its token is generated one DATA frame, on the wire Type 0x00 Length varint Frame Payload = the SSE event bytes data: {"token":" capital"}\n\n type 0x00 · length-delimited payload — the QUIC STREAM frames that carry these need not align with them no chunked encoding HTTP/3 bars Transfer-Encoding (RFC 9114 §4.1) — the DATA-frame sequence IS the framing, FIN ends it. SSE and streamable HTTP converge whether the body is text/event-stream or raw NDJSON, on HTTP/3 it is the same: HEADERS + DATA frames.
Fig. 9.6-1How a token becomes wire bytes. The response is a HEADERS frame (200, text/event-stream) followed by one DATA frame per token; each DATA frame is a Type (0x00), a Length, and a payload that is simply the SSE event text. No chunked encoding is needed — the DATA-frame sequence is the framing and the stream's FIN ends it — and SSE and raw streamable bodies look identical on HTTP/3.RFC 9114 §4.1, §7.2.1

Seen as a flow over time, one generation is a HEADERS frame followed by a DATA frame per token, flushed as each is decoded, until the stream closes:

LLM servingendpointClientmodel beginsgenerating tokenseach token is a DATA frame on ONEreliable, ordered stream — no drops,no reorderinga second concurrent generation ridesits OWN stream — a lost packet onone never stalls the other (nocross-stream head-of-line blocking)user hits stop → clientRESET_STREAM cancels just thisgeneration, freeing server computerequest HEADERS — POST/v1/chat (stream = true)response HEADERS — 200,content-type text/event-streamDATA — data: {"token":"The"}(flushed as generated)DATA — data: {"token":" capital"}DATA — data: {"token":" of"} …DATA — data: [DONE] + streamFIN
Fig. 9.6-2An LLM generation over HTTP/3. The request is one HEADERS frame; the response is a HEADERS frame (text/event-stream) followed by a DATA frame per token, flushed as generated, ending with the stream's FIN. Each token rides one reliable, ordered stream — no drops, no reordering. A second concurrent generation uses its own stream, so a lost packet on one never stalls the other, and a client RESET_STREAM cancels just one generation.RFC 9114 §4.1; RFC 9000 §2.1

What HTTP/3 buys token streaming #

Running the same SSE/streaming response over QUIC instead of TCP changes four things that matter for inference serving:

1. No head-of-line blocking across concurrent generations. A client running many inference calls at once (an agent fanning out sub-tasks, a batch job, a UI with several panels) multiplexes them over one connection. Over HTTP/2 on TCP, a single lost packet stalls every multiplexed stream until it is retransmitted (§4.4); one slow network moment freezes all the token streams together. Over HTTP/3 each request is an independent QUIC stream, so a loss affecting one generation leaves the others flowing (§4.4). For concurrent LLM traffic this is the headline win.

tokens over time → HTTP/2 over TCP — 3 generations share ONE ordered byte stream gen A gen B gen C 1 packet lost all 3 generations stall — TCP delivers in order HTTP/3 over QUIC — each generation is its OWN stream stream A stream B stream C loss on B — only B's later tokens wait A and C keep delivering tokens on time — only stream B waits ~1 RTT for its retransmit. the loss is isolated to one generation instead of freezing every concurrent request.
Fig. 9.6-3Head-of-line blocking, applied to concurrent generations. Over HTTP/2 on TCP (top), all three generations share one ordered byte stream, so a single lost packet stalls every generation's later tokens until the retransmit arrives. Over HTTP/3 on QUIC (bottom), each generation is an independent stream, so the loss is isolated to stream B — A and C keep delivering tokens on time.RFC 9000 §2; RFC 9114 §6.1

2. Faster time to first token. The first token is the latency users feel most. A cold connection spends a round trip on the handshake before the request even arrives (§8.1); a warm, pooled connection removes it, and 0-RTT (§8.3) can carry the request in the first flight, provided the request is safe to replay (§2.3), which a stateless completion call typically is. Keeping connections warm to an inference endpoint is the cheapest first-token win available.

3. Clean, cheap cancellation. Users stop generations constantly: a wrong turn, a good-enough answer. On QUIC the client sends RESET_STREAM (and STOP_SENDING) to abort that one stream (§5.4), which tells the server to stop generating and frees expensive GPU compute immediately, without touching other in-flight requests on the connection. Cancellation is a first-class, per-stream operation rather than a connection teardown.

4. Survival across network changes. A long generation can run for tens of seconds. A mobile client that switches from Wi-Fi to cellular mid-response would, on TCP, drop the connection and lose the partial output. QUIC's connection migration (Chapter 10) keeps the same connection — and the same streaming response — alive across the path change, so the tokens keep coming.

The within-stream latency still applies #

One caveat keeps the picture honest. Cross-stream isolation does not remove head-of-line blocking within a single generation's stream: because tokens are delivered reliably and in order, a lost packet still delays the tokens behind it on that stream by about a round trip (§9.1). This is unavoidable (the tokens must be ordered), so the mitigations are the general ones: a low-RTT path, pacing to avoid self-inflicted loss (§9.2), and prompt flushing. What HTTP/3 removes is the cross-request amplification of that stall, not the intra-stream cost of ordered delivery.

Two operational points follow from earlier sections:

  • Flush every token; don't batch. The server must emit each token (or small group) as a DATA frame as it is generated, not accumulate them. Write coalescing or Nagle-style batching directly inflates inter-token latency, the very thing streaming exists to minimize (§9.4).
  • Flow-control backpressure is a feature here. For media, flow control's backpressure was harmful (§9.4); for reliable token streaming it is welcome. If the client reads slowly, the stream's flow-control window fills and naturally paces the server, so it does not generate far ahead of consumption and buffer unbounded output. The same mechanism that hurt datagram media helps reliable streaming.

Worked example: a chat completion on a warm connection #

Trace Figure 9.6-2 for a chat UI. The client holds a warm, pooled HTTP/3 connection to the endpoint, so there is no handshake round trip; it sends the completion request as a single HEADERS frame (0-RTT would even fold it into the first flight on a fresh connection). The server responds with a HEADERS frame carrying 200 and text/event-stream, then streams a DATA frame per token as the model decodes (The, capital, of, …), each flushed immediately, arriving in order and reliably. When the user clicks stop, the client resets the stream; the server halts decoding and reclaims the GPU, while a second generation open on another stream of the same connection streams on undisturbed. Had a packet been lost, only this generation's subsequent tokens would have waited a round trip; the other stream would not have noticed.

Note

It is worth being explicit about why datagrams are wrong here, because the chapter spent four sections recommending them. Datagrams trade reliability for timeliness, which is exactly the wrong trade for tokens: there is no such thing as a "stale" token you can skip, and application-level reassembly of an unordered token stream would just re-implement the reliable, ordered stream QUIC already provides. The lesson is to match the model to the data: expiring samples → datagrams; a sequence where every element is required → a reliable stream. Both are real-time; they differ in what "too late" means.

In practice

editorial If you operate or call inference endpoints, the highest-leverage moves are connection reuse and honest flushing. Pool warm HTTP/3 connections to your model servers so first-token latency is one round trip of inference, not of TCP-plus-TLS setup; on the server, flush each token as its own DATA write and disable any output buffering in your proxy or framework (a buffering reverse proxy silently converts smooth streaming into batched bursts). Multiplex concurrent generations over one HTTP/3 connection to get loss isolation for free, and wire user-cancel to a stream reset so aborted generations stop billing GPU time. These are editorial engineering recommendations, but each maps directly onto a QUIC mechanism this book has already grounded.

Takeaways #

LLM token streaming is a reliable, ordered real-time workload — the counter-case to media — so it uses a reliable QUIC stream, not datagrams. SSE and streamable HTTP both reduce over HTTP/3 to a HEADERS frame plus a sequence of DATA frames flushed per token, with no chunked encoding needed. QUIC's gains are cross-stream head-of-line isolation for concurrent generations, faster first token via warm connections and 0-RTT, clean per-stream cancellation, and connection migration for long responses. Ordered delivery's within-stream cost remains, managed by prompt flushing and welcome flow-control backpressure. That closes the real-time chapter; the chapter quiz checks the model, and Chapter 10 turns to the migration that keeps these long-lived streams alive as the network path changes.