Resuming Progress Streams after Reconnect Permalink to this section

Part of Progress Streaming for Long-Running Jobs, under Real-Time Application Patterns.

A job that runs for twenty minutes will outlive at least one connection in a meaningful share of sessions: a phone switching from Wi-Fi to mobile data, a laptop lid closed and reopened, a deploy that restarts the node holding the stream. EventSource reconnects automatically. Whether the reconnect shows the right progress, or a bar reset to zero, or a 404, depends on how the server answers it. This guide makes every reconnect land on the job’s true current state, from any node, and makes the stream stop cleanly once there is nothing left to watch.

Symptom & Developer Intent Permalink to this section

  • After a network blip the progress bar drops back to 0 % and then climbs again.
  • Reconnecting to a different node returns 404 because only the original node knew the job.
  • After a deploy, every watching client shows “job failed” even though the jobs kept running on the workers.
  • Once the job finishes, the browser keeps reconnecting every few seconds for as long as the tab stays open.
  • Reloading the page mid-job starts a new job instead of reattaching to the running one.

The intent is that any disconnection — network, node, or page — is invisible except for a short pause, and that finished jobs are never reconnected to.

Root Cause Analysis Permalink to this section

The failures share one cause: progress state that lives only in the process holding the stream. When the stream handler keeps the job’s progress in memory, or when the stream handler is the process doing the work, a reconnect to anywhere else has nothing to report.

Where job state lives decides what a reconnect sees Two panels comparing in-process job state, which is lost on reconnect or restart, with shared-store job state, which any node can read. Where job state lives decides what a reconnect sees State in the stream handler reconnect to other node: 404 deploy: job looks failed reload: progress starts at 0 finished: nothing to say State in a shared store any node reads current state deploy: workers unaffected reload: bar at true position finished: terminal replayed
Moving state out of the stream handler is the whole fix. Everything else in this guide follows from it.

The second cause is the reconnect loop after completion. The specification tells EventSource to reconnect whenever a response ends, unless the response was not a 200 with the right content type, or the client called close(). A server that ends the response after the final event, and answers the next request with the same final event and another end, creates a loop that lasts as long as the tab. The way out, per the spec, is HTTP 204 No Content: it tells the browser to stop and not reconnect.

The third cause is the reload. A page that starts jobs in its load handler, rather than on an explicit user action whose result is remembered, starts a second job on every reload.

Step-by-Step Resolution Permalink to this section

Step 1 — Keep the latest frame of every job in a shared store Permalink to this section

Every progress report from the worker writes the complete latest frame — not a delta — to a key any node can read, and publishes the same frame.

// worker side
async function report(jobId, frame) {
  const key = `job:${jobId}`;
  await redis.multi()
    .set(`${key}:latest`, JSON.stringify(frame), 'EX', 86400)
    .publish(key, JSON.stringify(frame))
    .exec();
}

Progress frames must be self-contained ({"done":24100,"total":48210,"pct":50,"stage":"rows"}), so that the latest one alone is a complete picture.

Step 2 — Answer every connection with the latest frame first Permalink to this section

app.get('/api/jobs/:id/events', requireJobAccess, async (req, res) => {
  const key = `job:${req.params.id}`;
  const sub = redis.duplicate();
  await sub.subscribe(key);                                   // subscribe first
  const latest = JSON.parse(await redis.get(`${key}:latest`) ?? 'null');
  if (!latest) { await sub.quit(); return res.status(404).end(); }

  const lastSeen = req.get('Last-Event-ID');
  if (latest.terminal && lastSeen === String(latest.id)) {
    await sub.quit();
    return res.status(204).end();                             // client already has the ending
  }

  openStream(res, { retryMs: 2000 });
  write(res, latest);                                         // true current state, from any node
  if (latest.terminal) return close();

  sub.on('message', (_c, raw) => {
    const f = JSON.parse(raw);
    if (f.id <= latest.id) return;
    write(res, f);
    if (f.terminal) close();
  });
  req.on('close', close);
  function close() { sub.quit().catch(() => {}); res.end(); }
});
A reconnect to a different node mid-job Sequence diagram in which a browser loses its stream on node A, reconnects to node B, receives the latest frame from the shared store, and continues receiving live progress. A reconnect to a different node mid-job Browser Node A Node B Store progress 48 % id 31 node restarts (deploy) reconnect Last-Event-ID: 31 GET latest, SUBSCRIBE progress 52 % id 34 progress 52 % id 34
Node B never saw this job before. It does not need to — the store has the latest frame and the channel has everything after it.

Step 3 — Close on the client and make terminal events distinguishable Permalink to this section

es.addEventListener('done', (e) => { es.close(); showResult(JSON.parse(e.data)); });
es.addEventListener('failed', (e) => { es.close(); showError(JSON.parse(e.data)); });

The client close() is the primary stop. The server’s 204 for acknowledged terminal ids is the safety net for clients that miss the event — for example, a reconnect that happens in the instant between the terminal frame being written and the listener running.

Step 4 — Drain nodes gracefully during deploys Permalink to this section

A restarting node should not make watching clients believe their jobs failed. Before shutdown, send a short retry: hint and end streams, so clients reconnect promptly to healthy nodes rather than waiting for a TCP timeout:

process.on('SIGTERM', () => {
  for (const res of openStreams) {
    res.write('retry: 500\n: node draining\n\n');          // reconnect quickly, elsewhere
    res.end();
  }
  server.close();
});

Because the worker, not the node, owns the job, the job itself is unaffected, and the reconnect lands on the latest frame from step 2.

Step 5 — Remember the job, not the action, across reloads Permalink to this section

// Starting a job is an explicit action; its id goes in the URL so reloads reattach.
async function startExport() {
  const { jobId } = await (await fetch('/api/exports', { method: 'POST' })).json();
  history.replaceState(null, '', `?job=${jobId}`);
  watchJob(jobId);
}

const existing = new URLSearchParams(location.search).get('job');
if (existing) watchJob(existing);                           // reload → same job, current state

Putting the id in the URL also makes the page shareable: a colleague opening the link sees the same progress.

What users saw after a mid-job deploy Bar chart comparing the share of watching clients that showed a wrong state after a deploy, before and after moving job state to a shared store and adding graceful drain. What users saw after a mid-job deploy In-process state 86 % Shared store 9 % (slow reconnect) Shared store + drain < 1 % share of watching clients showing a wrong state after the deploy
With state in the stream handler, most clients showed a failure or a reset bar. With a shared store and drain, the deploy showed up only as a pause of under a second.

Validation & Monitoring Permalink to this section

# Reconnect with an old id to a second node: expect the latest frame, not a 404.
curl -sN -H 'Last-Event-ID: 3' --resolve app.example.com:443:10.0.1.12 \
  https://app.example.com/api/jobs/exp_7Qk/events | head -2

# Acknowledged terminal id: expect 204 and no body.
curl -s -o /dev/null -w '%{http_code}\n' -H 'Last-Event-ID: 99' \
  https://app.example.com/api/jobs/exp_7Qk/events

Run a deploy during a load test with a hundred watched jobs and count client-visible failures. In production, track reconnects per stream and the share of streams that end with a terminal event versus a client disconnect; watch for requests that return 204, which should appear at the rate jobs finish, and never as a steady stream from the same client.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Do progress streams need a replay buffer?

No. Progress is state, so the latest frame is enough. Store it, send it first on every connection, and let the live channel carry what follows.

Why return 204 rather than 200 with the final event?

A 200 response that ends is treated by EventSource as a dropped connection, so it reconnects. A 204 is defined by the specification as the signal to stop reconnecting, which ends the loop without relying on client code.

Is sticky routing still useful?

It reduces reconnect cost slightly when the original node is healthy, but it is not needed for correctness once state is in a shared store, and it must never be relied on across deploys.

How long should finished jobs stay in the store?

Long enough for users to come back to them — a day is common. After that, the job's result page can be served from the database rather than the stream.