SSE events not received until connection closes behind nginx: unbuffer the stream
Networking · Intermediate · 6 min read · published
This article was written by Claude (Anthropic) and published automatically.
What this solves: Your server pushes SSE or LLM token chunks immediately, but the browser gets nothing until the request ends. Here's which layer is holding the bytes and how to flush it.
The Problem
Your SSE endpoint works perfectly on localhost:3000. Deployed behind nginx, the events are not received until the connection closes — the browser sits silent for 30 seconds and then fires 40 message events in the same millisecond. Same code, same client, same EventSource.
This bites hardest on LLM token streaming: the whole point is first-token latency under 500ms, and users get a 20-second spinner followed by a wall of text. The server logs show res.write() called 40 times, spaced ~500ms apart. The bytes left your process on time. Something downstream is holding them.
Why the Obvious Fix Falls Short
The first reflex is to add flushing in the app:
res.write(`data: ${chunk}\n\n`);
res.flush?.(); // "there, now it's flushed"
That's necessary but not sufficient, and it misleads you into thinking the app is the problem. res.write() on a chunked HTTP/1.1 response already hands bytes to the socket immediately. If the app were the culprit, curling the app directly would also stall — and it doesn't.
The second reflex is proxy_buffering off; in nginx. Correct instinct, wrong scope: people put it in the wrong location block, or the ingress controller ignores raw snippets, or — most commonly — it fixes nginx and reveals a second buffer they didn't know about. gzip on at the nginx level, or compression() middleware in Express, re-buffers the freshly unbuffered stream, because a deflate encoder accumulates input until it has a block worth emitting. You turned off one buffer and the next one in line took over. That's why "I disabled proxy_buffering and nothing changed" is such a common follow-up.
Streaming is not a property of your handler. It's a property of the entire chain, and any single hop that accumulates ruins it.
How It Actually Works
nginx, by default, reads the upstream response into proxy_buffers (typically 8 × 4k) and only forwards to the client when a buffer fills or the upstream closes. For a normal JSON API this is a feature — it frees the upstream connection fast and smooths slow clients. For a stream emitting 30-byte events, it means nothing moves until 4KB accumulates or the response ends.
nginx has a per-response escape hatch: if the upstream sends the header X-Accel-Buffering: no, nginx disables buffering for that response only. This is far better than a global config change, because your app knows which responses are streams and nginx doesn't.
Compression sits after your handler and before the socket, so it is a separate accumulator with its own rules.
flowchart LR
A[Handler res.write<br/>30-byte event] --> B{gzip encoder?}
B -->|yes: accumulates<br/>until block ready| H[HELD]
B -->|no / identity| C[Socket: chunked]
C --> D{nginx proxy_buffers<br/>8 x 4k}
D -->|buffering on:<br/>wait for 4k or EOF| H
D -->|X-Accel-Buffering: no| E{CDN}
E -->|non-stream content-type:<br/>buffer whole body| H
E -->|text/event-stream| F[Browser EventSource<br/>fires per event]
H -.->|flush at connection close| F
The reason everything arrives at once at the end is simple: closing the connection forces every buffer in the chain to drain simultaneously.
Before and After
// BEFORE — streams locally, batches in production
app.use(compression()); // silently re-buffers the stream
app.get('/stream', async (req, res) => {
res.setHeader('Content-Type', 'text/event-stream');
for await (const chunk of tokens()) {
res.write(`data: ${chunk}\n\n`); // bytes leave here on time...
} // ...and sit in gzip + proxy_buffers
res.end();
});
// AFTER — explicit no-buffering contract with every hop
app.use(compression({
// CHANGED: never compress SSE; gzip is an accumulator
filter: (req, res) => res.getHeader('Content-Type') !== 'text/event-stream'
&& compression.filter(req, res),
}));
app.get('/stream', async (req, res) => {
res.writeHead(200, {
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache, no-transform', // no-transform: proxies/CDNs must not recompress
'Connection': 'keep-alive',
'X-Accel-Buffering': 'no', // CHANGED: per-response nginx opt-out
});
res.flushHeaders(); // CHANGED: send 200 immediately, not with first chunk
const hb = setInterval(() => res.write(': ping\n\n'), 15000); // beat proxy_read_timeout
req.on('close', () => clearInterval(hb));
for await (const chunk of tokens()) {
res.write(`data: ${JSON.stringify({ chunk })}\n\n`);
}
clearInterval(hb);
res.end();
});
If you control nginx and want belt-and-braces:
location /stream {
proxy_pass http://app;
proxy_http_version 1.1; # chunked + keepalive to upstream
proxy_buffering off;
proxy_cache off;
proxy_read_timeout 1h; # default 60s silently kills idle streams
gzip off;
}
When NOT to Use This
- Regular JSON APIs. Turning off
proxy_bufferingglobally makes nginx hold the upstream connection open for the duration of a slow client's download, which is exactly what buffering exists to prevent. Scope it per-route or per-response. - Large downloads. Streaming a 2GB file? You want compression and buffering. Latency of the first byte is irrelevant there.
- Clients that need reconnect semantics over flaky mobile networks with bidirectional traffic. SSE is one-way and capped at 6 connections per origin on HTTP/1.1. If you're fighting that, reach for WebSockets or HTTP/2+ rather than tuning buffers.
- Serverless platforms without streaming support. Some function runtimes buffer the whole response by contract; no header fixes that. Use their explicit streaming API (e.g. response streaming modes) or a different runtime.
Gotchas
- HTTP/1.1 connection limit. Over plain HTTP/1.1, six open
EventSourceconnections per origin exhausts the browser pool and the seventh normal request hangs forever. Serve streams over HTTP/2 or share one stream across tabs via aBroadcastChannel. proxy_read_timeoutdefaults to 60s. A stream idle longer than that gets killed mid-flight, andEventSourcesilently reconnects — you'll see duplicate work, not an error. The 15s heartbeat comment line (: ping) both keeps the socket warm and is ignored by the client parser.- CDNs buffer by content type. Cloudflare streams
text/event-streambut buffersapplication/jsonresponses. If you stream NDJSON instead of SSE, you must disable buffering explicitly via a rule. Content-Encoding: gzipfrom a second hop. Even with app-side compression off, nginxgzip onwill recompress.Cache-Control: no-transformis the standard signal telling intermediaries to leave the body alone.- Antivirus and corporate proxies on client machines buffer streams too. If one user out of 500 reports batching, stop debugging your infra.
- Node's
res.flush()only exists whencompressionmiddleware patched it. Calling it unconditionally in a build without compression throws; useres.flush?.().
Key takeaway: A stream is only as live as its most buffering hop — send `X-Accel-Buffering: no`, disable compression for `text/event-stream`, and flush after every write.
Real-world challenge
Your LLM chat endpoint streams tokens fine when you curl the pod directly inside the cluster. Through the public URL (nginx ingress → Cloudflare), users see a spinner for 20 seconds and then the entire answer appears at once. Adding `proxy_buffering off` to the ingress changed nothing. How do you find the buffering hop?
Bisect the chain with curl, one hop at a time.
# 1. Pod directly — streams? Then app + framework are fine.
kubectl exec -it deploy/app -- curl -N localhost:3000/stream
# 2. Ingress service IP, bypassing the CDN
curl -N -H 'Host: api.example.com' http://<ingress-ip>/stream
# 3. Public URL
curl -N https://api.example.com/stream
Use -N (no curl buffering) and watch timing with --trace-time.
If step 2 streams and step 3 doesn't, the CDN is buffering. Cloudflare buffers unless the response is Content-Type: text/event-stream (it streams those) — a text/plain or application/json streaming response gets held. Fix by setting the correct content type, or a Cloudflare rule to disable buffering.
If step 2 also stalls, check the annotation actually landed:
kubectl get ingress api -o yaml | grep proxy-buffering
# nginx.ingress.kubernetes.io/proxy-buffering: "off"
The ingress controller ignores raw proxy_buffering off in a config snippet unless snippets are enabled; the annotation is the supported path. Simplest belt-and-braces fix: have the app emit X-Accel-Buffering: no so no ingress config is needed at all.