Rebalancing SSE Connections After a Deploy Permalink to this section

Part of Connection Pooling for SSE Servers, under Backend Stream Generation & Connection Management.

Load balancers balance new connections. Server-Sent Events connections are not new for long: they last hours. After a rolling deploy, the first node to restart collects the reconnects of every node restarted after it, and the last node to restart starts nearly empty. After a scale-out, new nodes receive only the trickle of fresh connections while the old ones stay full. Nothing rebalances until clients happen to reconnect. This guide makes rebalancing deliberate.

Symptom & Developer Intent Permalink to this section

  • After each deploy, one node holds three or four times the connections of the others.
  • Autoscaling adds nodes during a spike, but CPU on the original nodes does not fall.
  • The busiest node hits its file descriptor or memory limit while others are idle.
  • Connection counts even out only slowly over the following day.
  • Scaling in (removing nodes) causes a thundering herd of reconnects onto the survivors.

The intent is connection counts that converge to an even spread within minutes of any deploy or scaling event, without visible disruption to users.

Root Cause Analysis Permalink to this section

A rolling deploy replaces nodes one at a time. When node 1 restarts, its clients reconnect and are spread across nodes 2 to N, which are all still running. When node 2 restarts, its clients — now including some of node 1’s — spread over the remaining nodes, and so on. The node restarted last receives only the connections of the node restarted just before it, while earlier survivors have accumulated reconnects from every earlier step.

Connections per node after a rolling deploy of five nodes Bar chart of connections held by each of five nodes after a rolling deploy with no rebalancing, showing a strong skew towards the nodes that stayed up longest. Connections per node after a rolling deploy of five nodes Node 1 (restarted first) 4,500 Node 2 7,000 Node 3 10,500 Node 4 14,000 Node 5 (restarted last) 14,000 connections held per node immediately after the deploy
With 50,000 clients and no rebalancing, the spread after the deploy is uneven by a factor of several. Round-robin at the balancer cannot fix connections it never sees again.

The exact shape depends on deploy order and balancer algorithm, but the lesson is general: long-lived connections preserve whatever imbalance existed when they were made. The fix is to make connections not permanent — give each one a bounded lifetime so that, continuously, a fraction of clients reconnect and are placed afresh.

Step-by-Step Resolution Permalink to this section

Step 1 — Give every stream a maximum age with jitter Permalink to this section

End each stream after a randomised lifetime, sending a short retry: hint first so the client reconnects promptly and resumes from Last-Event-ID:

const MIN_AGE_MS = 20 * 60 * 1000;
const MAX_AGE_MS = 40 * 60 * 1000;

function scheduleRecycle(res) {
  const age = MIN_AGE_MS + Math.random() * (MAX_AGE_MS - MIN_AGE_MS);   // jitter spreads reconnects
  return setTimeout(() => {
    res.write('retry: 500\n: recycling connection\n\n');
    res.end();
  }, age);
}

With a 20–40 minute lifetime, about 3–5 % of connections recycle every minute, and each recycled connection is placed by the balancer on current information. Imbalance decays steadily instead of persisting for hours. Because the stream is resumable, the recycle is invisible to users.

Step 2 — Route new connections by least connections Permalink to this section

Round-robin places connections without regard to how many each node already holds. For long-lived streams, least-connections routing makes each reconnect correct the imbalance:

upstream sse_nodes {
    least_conn;                   # send each new stream to the node with fewest active
    server 10.0.1.11:8080;
    server 10.0.1.12:8080;
    server 10.0.1.13:8080;
    keepalive 64;
}

Cloud balancers expose the same idea under different names — “least outstanding requests” on AWS ALB target groups, for example. With least-connections routing and recycling together, the spread converges quickly after any disturbance.

Busiest node's share of connections after a deploy Line chart over two hours of the busiest node's connection share after a rolling deploy, comparing no recycling, 30-minute jittered recycling with round-robin, and recycling with least-connections routing. Busiest node's share of connections after a deploy no recycling recycle + round-robin recycle + least-conn 0 10 20 30 40 0 24 48 72 96 120 even spread minutes after the deploy busiest node share (%)
With five nodes, an even spread is 20 %. Recycling alone converges within an hour; recycling plus least-connections converges in minutes.

Step 3 — Drain deliberately when a node leaves Permalink to this section

When a node is removed — deploy or scale-in — do not let its connections drop all at once. Stop accepting new streams, then close existing ones gradually over a drain window:

async function drain(streams, windowMs = 30_000) {
  server.close();                                       // stop accepting new connections
  const list = [...streams];
  const gap = windowMs / Math.max(1, list.length);
  for (const res of list) {
    res.write(`retry: ${1000 + Math.floor(Math.random() * 4000)}\n\n`);   // spread the returns
    res.end();
    await new Promise((r) => setTimeout(r, gap));
  }
}
process.on('SIGTERM', () => drain(openStreams));

Set the orchestrator’s termination grace period longer than the drain window. In Kubernetes, that is terminationGracePeriodSeconds, typically paired with a preStop hook that starts the drain.

Step 4 — Actively shed load from hot nodes Permalink to this section

For fast convergence after scale-out, have overloaded nodes shed a fraction of connections. A node that holds more than, say, 120 % of the fleet average can end a small random share of its streams each minute; the balancer sends them to the new, emptier nodes.

setInterval(async () => {
  const avg = await fleetAverageConnections();            // from a shared metric or registry
  const mine = openStreams.size;
  if (mine > avg * 1.2) {
    const excess = Math.min(mine - avg, Math.ceil(mine * 0.02));   // at most 2 % per minute
    for (const res of pickRandom(openStreams, excess)) { res.write('retry: 1000\n\n'); res.end(); }
  }
}, 60_000);

Cap the shed rate so the process never produces a reconnect storm of its own.

Where the fleet average comes from matters less than it seems. A gauge scraped by Prometheus and read back through its API, a small Redis hash where each node writes its count every ten seconds, or the load balancer’s own per-target connection metric all work; the rebalancer only needs a number that is roughly right and roughly current. Avoid making nodes query each other directly, which couples every node to every other and turns a slow node into a slow fleet.

Validation & Monitoring Permalink to this section

# Per-node connection counts during and after a staged rolling deploy.
watch -n 10 'for n in 10.0.1.11 10.0.1.12 10.0.1.13; do echo -n "$n "; curl -s $n:9100/metrics | grep ^sse_open_streams; done'

Graph the coefficient of variation (standard deviation divided by mean) of connections per node. It should spike at each deploy and return below 0.1 within the recycle window. Also graph reconnect rate: recycling adds a steady, predictable baseline, and drains add short, bounded bumps rather than cliffs.

A rolling deploy with drain and recycling Timeline of thirty minutes showing three nodes each draining over thirty seconds during a rolling deploy, followed by recycling that evens out the spread. A rolling deploy with drain and recycling Node A Node B Recycling serving (after drain) serving (after drain) ~4 % per min 0 6 12 18 24 30 minutes deploy done
Each node leaves gradually, and recycling continues afterwards so the survivors do not keep the skew the deploy created.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Won't recycling connections hurt the user experience?

Not if streams are resumable. The client reconnects within the retry interval, sends Last-Event-ID and receives anything it missed, which typically takes well under a second.

What maximum age should I use?

Long enough that reconnect overhead is negligible and short enough that imbalance decays within your tolerance. Twenty to sixty minutes with jitter suits most services.

Is least-connections enough on its own?

No. It only affects new connections, and without recycling there are very few of them after a deploy. The two work together: recycling creates reconnects, least-connections places them well.

Does HTTP/2 change this?

Yes, in one way: many streams share one connection from each browser, so the balancer sees fewer, heavier connections. Recycling at the stream level still works, but balancing happens per connection, so imbalance can be coarser.