Diagnosing Buffered SSE Output in App Servers Permalink to this section
Part of Buffer Management & Chunked Transfer Encoding, under Backend Stream Generation & Connection Management.
“Events arrive in bursts” is the most common Server-Sent Events bug report, and the least specific. Somewhere between the line of code that writes an event and the browser that renders it, a buffer is holding bytes until it fills or the response ends. There are usually five or six candidate layers, and guessing wastes days. This guide gives a bisection procedure that finds the buffering layer in minutes, and the fix for each layer you are likely to find.
Symptom & Developer Intent Permalink to this section
- Events appear in batches every few seconds, or only when the stream closes.
- Everything works with the development server and fails in production.
- A heartbeat seems to “release” several queued events at once.
curl -Nagainst the app port works, but through the load balancer it does not.- The first event is delayed until roughly 4 KB, 8 KB or 16 KB of output has been produced.
The intent is to identify the exact layer that buffers and apply a targeted fix, rather than changing settings everywhere and hoping.
Root Cause Analysis Permalink to this section
Every layer in the output path may buffer for efficiency. For request/response traffic, buffering is invisible because the response ends and everything is flushed. For a stream, a buffer that waits to fill waits indefinitely.
The telltale sizes help: a first event delayed until about 4 KB points at a runtime or framework write buffer; 8 KB or 16 KB often matches a gzip or proxy buffer; “only when the response ends” means something is reading the whole body, such as a response-capturing middleware or a CDN that does not stream.
Step-by-Step Resolution Permalink to this section
Step 1 — Build a timing probe Permalink to this section
A stream that emits one small, timestamped event per second makes buffering visible. Add a debug route (disabled in production) or use an existing low-rate stream:
app.get('/debug/tick', (req, res) => {
res.writeHead(200, { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache' });
const t = setInterval(() => res.write(`data: ${Date.now()}\n\n`), 1000);
req.on('close', () => clearInterval(t));
});
# Print local arrival time next to each event's send time; the difference is the delay.
probe() { curl -sN "$1" | while IFS= read -r l; do
[ "${l#data: }" != "$l" ] && echo "$(date +%s%3N) sent=${l#data: } lag=$(( $(date +%s%3N) - ${l#data: } ))ms"; done; }
Step 2 — Bisect from the inside out Permalink to this section
Run the probe against each boundary, starting closest to the code:
probe http://127.0.0.1:3000/debug/tick # 1. app server port, on the host
probe http://10.0.1.11:3000/debug/tick # 2. from another host (network path, no proxy)
probe http://10.0.1.5/debug/tick # 3. through the reverse proxy
probe https://app.example.com/debug/tick # 4. through the CDN / public edge
The first step where lag jumps from single-digit milliseconds to seconds, or where events arrive in groups, has the buffer. If step 1 already buffers, the problem is inside the process: repeat with middleware disabled to separate framework from server.
Step 3 — Fix the layer you found Permalink to this section
Application code. Flush after every event. In Go, call Flush() on the http.Flusher; in ASP.NET Core, await Response.Body.FlushAsync(); in Express, res.write sends immediately unless compression middleware is present; in Java servlets, writer.flush() or SseEmitter, which flushes for you.
Framework middleware. Exclude the route from gzip, response logging and anything that reads or wraps the body. Serving SSE from Express with compression enabled shows the pattern for Express; the others follow the same idea.
Application server. Python WSGI servers with sync workers stream, but only one response per worker; PHP-FPM behind FastCGI buffers unless output_buffering is off and the web server’s FastCGI buffering is disabled; servlet containers need async mode for long responses. Uvicorn, Hypercorn, Kestrel, Netty and Go’s net/http stream flushed writes immediately.
Runtime writer. Nagle’s algorithm can delay small writes by up to 40 ms waiting for an ACK. Most HTTP servers set TCP_NODELAY; if yours does not, set it on accepted sockets.
Reverse proxy. nginx: proxy_buffering off; for the location, or send X-Accel-Buffering: no from the app. Also ensure gzip off for text/event-stream in the proxy. Other proxies are covered in proxy and CDN configuration for SSE.
CDN. Enable streaming or bypass the cache for the route, and send Cache-Control: no-cache, no-transform so the edge does not transform the body. See disabling CDN buffering for event streams.
Step 4 — Rule out the client Permalink to this section
If every server-side boundary is clean but a particular user still sees bursts, suspect their environment: corporate TLS-inspecting proxies and some antivirus products buffer HTTP responses on the client machine. Test from that network with curl; if curl bursts too, the buffering is on their network path. The only robust mitigation is a fallback — for instance, a fetch-based client with a watchdog that detects bursts and switches to polling.
Validation & Monitoring Permalink to this section
Keep the probe as a synthetic monitor: an external check that opens the stream through the public URL, measures the lag of a heartbeat or timestamped event, and alerts if it exceeds a second. Configuration changes to proxies and CDNs are the most common way buffering returns, and a synthetic check catches it within minutes.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Why does a heartbeat release several events at once?
The heartbeat adds enough bytes to fill a buffer that was waiting for more data, so everything queued behind it is flushed together. It is a strong sign of size-based buffering in the path.
Is Transfer-Encoding: chunked enough to prevent buffering?
No. Chunked encoding describes how the body is framed on the wire, not when intermediaries forward it. A proxy can receive chunks and still buffer them before forwarding.
Can padding the first event work around buffering?
Sending a few kilobytes of comment padding can push a fixed-size buffer to flush once, and was a common workaround for old proxies. It does not fix continuous buffering and wastes bandwidth; find and fix the layer instead.
Does HTTP/2 avoid proxy buffering?
No. HTTP/2 proxies can buffer DATA frames just as HTTP/1.1 proxies buffer chunks. The same configuration switches apply.