Configuring HAProxy for SSE Permalink to this section

Part of Proxy & CDN Configuration for SSE, under SSE Protocol Fundamentals & Architecture.

HAProxy forwards response data as it arrives, so unlike nginx it does not need buffering turned off for Server-Sent Events. What it does need is timeouts that fit a response that lasts hours, a compression filter that leaves event streams alone, a balancing algorithm suited to long-lived connections, and a drain procedure for reloads. This guide builds a configuration for an HAProxy 2.x frontend in front of an SSE backend pool.

Symptom & Developer Intent Permalink to this section

  • Streams close after exactly 50 seconds (or whatever timeout server is set to) when no events are sent.
  • One backend server holds most of the connections while others are idle.
  • Enabling compression on the frontend made events arrive in batches.
  • A configuration reload or backend restart drops every stream at once.
  • The stats page shows session counts but nothing about stream health.

The intent is an HAProxy configuration where streams survive quiet periods, spread evenly, reconnect gracefully on reloads and are visible in metrics.

Root Cause Analysis Permalink to this section

HAProxy applies inactivity timeouts to each side of a connection: timeout client for the browser side and timeout server for the backend side. An SSE stream with no events for longer than either is closed.

HAProxy timeouts that apply to an SSE stream Layers listing HAProxy timeouts in the order a stream encounters them: connect, client inactivity, server inactivity and the HTTP keep-alive timeout between requests. HAProxy timeouts that apply to an SSE stream timeout connect backend TCP connect, a few seconds timeout client browser side inactivity timeout server backend side inactivity timeout http-keep-alive idle between requests only
With 15-second heartbeats, 60-second inactivity timeouts are enough. Without heartbeats, no finite timeout is safe for a quiet stream.

Balancing is the second issue. roundrobin distributes connections, and SSE connections persist, so whatever imbalance exists after a restart persists too. leastconn sends each new connection to the server with the fewest active ones, which continuously corrects imbalance. The compression filter buffers output to compress it efficiently, which batches small events.

Step-by-Step Resolution Permalink to this section

Step 1 — Route streams to their own backend Permalink to this section

frontend fe_https
    bind :443 ssl crt /etc/haproxy/certs/site.pem alpn h2,http/1.1
    mode http
    timeout client 65s                       # > 4 × a 15 s heartbeat
    acl is_sse path_beg /api/stream /api/events
    acl is_sse req.hdr(accept) -m sub text/event-stream
    use_backend be_sse if is_sse
    default_backend be_api

backend be_sse
    mode http
    balance leastconn                        # long-lived connections: least active first
    timeout connect 5s
    timeout server 65s                       # backend-side inactivity, beaten by heartbeats
    option httpchk GET /healthz
    http-response set-header X-Accel-Buffering no
    server sse1 10.0.1.11:8080 check maxconn 20000
    server sse2 10.0.1.12:8080 check maxconn 20000
    server sse3 10.0.1.13:8080 check maxconn 20000

With heartbeats every 15 seconds, 65-second inactivity timeouts are safe on both sides. Keeping them finite matters: they are how HAProxy frees connections to clients and servers that silently disappeared.

If you cannot guarantee heartbeats, raise the backend timeout for stream requests only, leaving the rest of the site’s limits intact:

backend be_sse
    http-request set-timeout server 1h if { path_beg /api/stream }

Step 2 — Keep compression away from event streams Permalink to this section

backend be_api
    compression algo gzip
    compression type application/json text/html text/css application/javascript
# be_sse has no compression directives at all.

List MIME types explicitly rather than using a broad text/ match, so text/event-stream is never included.

Step 3 — Enable HTTP/2 to browsers Permalink to this section

The alpn h2,http/1.1 on the bind line lets browsers use HTTP/2, which removes the six-connections-per-origin limit for SSE. HAProxy speaks HTTP/1.1 to the backend unless configured otherwise; that is fine. Make sure the backend does not send Connection: keep-alive — HAProxy removes connection-specific headers when converting to HTTP/2, but see debugging HTTP/2 stream resets if resets appear.

One stream through HAProxy Flow from the browser over HTTP/2 to the HAProxy frontend, the SSE routing rule, the leastconn backend, and an SSE server over HTTP/1.1. One stream through HAProxy Browser HTTP/2 TLS, ALPN h2 fe_https acl is_sse use_backend be_sse leastconn, 65 s forward SSE server HTTP/1.1 stream
HAProxy forwards bytes as they arrive. The configuration's job is to keep the connection alive, route it well and leave the body untouched.

Step 4 — Drain gracefully on reloads and deploys Permalink to this section

HAProxy reloads start a new process and let the old one finish existing connections, up to hard-stop-after. Set it long enough for streams to reconnect naturally, and let the application recycle connections:

global
    hard-stop-after 5m          # old process keeps streams up to 5 minutes after a reload

To take a backend server out for a deploy, put it in drain mode through the runtime API; it receives no new connections while existing streams continue until the application closes them with a retry hint:

echo "set server be_sse/sse2 state drain" | socat stdio /run/haproxy/admin.sock

Step 5 — Protect the pool with connection limits and queues Permalink to this section

Each open stream holds a backend connection for its whole life, so maxconn on a server is a hard ceiling on the streams that server will carry. When every server is at its limit, HAProxy queues new connections for up to timeout queue and then returns a 503. For SSE, a 503 makes EventSource stop reconnecting, so a queue timeout during an overload silently turns off live updates for the affected users. Keep timeout queue short, alert when queues form, and prefer letting the application shed load softly — a 200 with a long retry hint, as in using retry hints for load shedding — by setting maxconn slightly above the application’s own admission limit.

backend be_sse
    timeout queue 5s
    server sse1 10.0.1.11:8080 check maxconn 22000   # app sheds softly at 20,000

Step 6 — Watch the right metrics Permalink to this section

frontend stats
    bind :8404
    http-request use-service prometheus-exporter if { path /metrics }

The useful series for SSE are current sessions per backend server (haproxy_server_current_sessions), sessions ended by timeouts versus normally, and connection errors. A sawtooth in current sessions with a period equal to a timeout indicates heartbeats are not reaching HAProxy.

Validation & Monitoring Permalink to this section

# Idle survival: hold a stream for 5 minutes with no application events.
timeout 300 curl -sN https://app.example.com/api/stream | grep -c '^:'       # ~20 heartbeats

# Distribution: sessions per server should converge with leastconn.
echo "show stat" | socat stdio /run/haproxy/admin.sock | awk -F, '$1=="be_sse"{print $2, $5}'
Sessions per server one hour after a rolling restart Bar chart comparing the spread of SSE sessions across three servers with roundrobin and leastconn balancing one hour after a rolling restart. Sessions per server one hour after a rolling restart roundrobin 51 % leastconn 35 % largest server's share of sessions (even = 33 %)
Round-robin preserves the post-restart skew; leastconn pulls new connections toward the emptier servers.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Does HAProxy buffer SSE responses?

No. It forwards response data as it receives it. Buffering appears only if a compression filter or another content-processing filter is applied to the stream.

Should I use timeout tunnel for SSE?

timeout tunnel applies to connections in tunnel mode, such as WebSockets after an upgrade. SSE stays in HTTP mode, so timeout client and timeout server are the relevant settings.

How do I keep a user on the same backend?

You rarely need to if streams are resumable from shared storage. If you do, use a stick table on a user cookie, accepting that it works against even distribution.

What maxconn should each server have?

Slightly below what the server can hold with healthy latency, as measured by a load test. HAProxy queues excess connections rather than overloading the server.