Resuming Progress Streams after Reconnect Permalink to this section
Part of Progress Streaming for Long-Running Jobs, under Real-Time Application Patterns.
A job that runs for twenty minutes will outlive at least one connection in a meaningful share of sessions: a phone switching from Wi-Fi to mobile data, a laptop lid closed and reopened, a deploy that restarts the node holding the stream. EventSource reconnects automatically. Whether the reconnect shows the right progress, or a bar reset to zero, or a 404, depends on how the server answers it. This guide makes every reconnect land on the job’s true current state, from any node, and makes the stream stop cleanly once there is nothing left to watch.
Symptom & Developer Intent Permalink to this section
- After a network blip the progress bar drops back to 0 % and then climbs again.
- Reconnecting to a different node returns 404 because only the original node knew the job.
- After a deploy, every watching client shows “job failed” even though the jobs kept running on the workers.
- Once the job finishes, the browser keeps reconnecting every few seconds for as long as the tab stays open.
- Reloading the page mid-job starts a new job instead of reattaching to the running one.
The intent is that any disconnection — network, node, or page — is invisible except for a short pause, and that finished jobs are never reconnected to.
Root Cause Analysis Permalink to this section
The failures share one cause: progress state that lives only in the process holding the stream. When the stream handler keeps the job’s progress in memory, or when the stream handler is the process doing the work, a reconnect to anywhere else has nothing to report.
The second cause is the reconnect loop after completion. The specification tells EventSource to reconnect whenever a response ends, unless the response was not a 200 with the right content type, or the client called close(). A server that ends the response after the final event, and answers the next request with the same final event and another end, creates a loop that lasts as long as the tab. The way out, per the spec, is HTTP 204 No Content: it tells the browser to stop and not reconnect.
The third cause is the reload. A page that starts jobs in its load handler, rather than on an explicit user action whose result is remembered, starts a second job on every reload.
Step-by-Step Resolution Permalink to this section
Step 1 — Keep the latest frame of every job in a shared store Permalink to this section
Every progress report from the worker writes the complete latest frame — not a delta — to a key any node can read, and publishes the same frame.
// worker side
async function report(jobId, frame) {
const key = `job:${jobId}`;
await redis.multi()
.set(`${key}:latest`, JSON.stringify(frame), 'EX', 86400)
.publish(key, JSON.stringify(frame))
.exec();
}
Progress frames must be self-contained ({"done":24100,"total":48210,"pct":50,"stage":"rows"}), so that the latest one alone is a complete picture.
Step 2 — Answer every connection with the latest frame first Permalink to this section
app.get('/api/jobs/:id/events', requireJobAccess, async (req, res) => {
const key = `job:${req.params.id}`;
const sub = redis.duplicate();
await sub.subscribe(key); // subscribe first
const latest = JSON.parse(await redis.get(`${key}:latest`) ?? 'null');
if (!latest) { await sub.quit(); return res.status(404).end(); }
const lastSeen = req.get('Last-Event-ID');
if (latest.terminal && lastSeen === String(latest.id)) {
await sub.quit();
return res.status(204).end(); // client already has the ending
}
openStream(res, { retryMs: 2000 });
write(res, latest); // true current state, from any node
if (latest.terminal) return close();
sub.on('message', (_c, raw) => {
const f = JSON.parse(raw);
if (f.id <= latest.id) return;
write(res, f);
if (f.terminal) close();
});
req.on('close', close);
function close() { sub.quit().catch(() => {}); res.end(); }
});
Step 3 — Close on the client and make terminal events distinguishable Permalink to this section
es.addEventListener('done', (e) => { es.close(); showResult(JSON.parse(e.data)); });
es.addEventListener('failed', (e) => { es.close(); showError(JSON.parse(e.data)); });
The client close() is the primary stop. The server’s 204 for acknowledged terminal ids is the safety net for clients that miss the event — for example, a reconnect that happens in the instant between the terminal frame being written and the listener running.
Step 4 — Drain nodes gracefully during deploys Permalink to this section
A restarting node should not make watching clients believe their jobs failed. Before shutdown, send a short retry: hint and end streams, so clients reconnect promptly to healthy nodes rather than waiting for a TCP timeout:
process.on('SIGTERM', () => {
for (const res of openStreams) {
res.write('retry: 500\n: node draining\n\n'); // reconnect quickly, elsewhere
res.end();
}
server.close();
});
Because the worker, not the node, owns the job, the job itself is unaffected, and the reconnect lands on the latest frame from step 2.
Step 5 — Remember the job, not the action, across reloads Permalink to this section
// Starting a job is an explicit action; its id goes in the URL so reloads reattach.
async function startExport() {
const { jobId } = await (await fetch('/api/exports', { method: 'POST' })).json();
history.replaceState(null, '', `?job=${jobId}`);
watchJob(jobId);
}
const existing = new URLSearchParams(location.search).get('job');
if (existing) watchJob(existing); // reload → same job, current state
Putting the id in the URL also makes the page shareable: a colleague opening the link sees the same progress.
Validation & Monitoring Permalink to this section
# Reconnect with an old id to a second node: expect the latest frame, not a 404.
curl -sN -H 'Last-Event-ID: 3' --resolve app.example.com:443:10.0.1.12 \
https://app.example.com/api/jobs/exp_7Qk/events | head -2
# Acknowledged terminal id: expect 204 and no body.
curl -s -o /dev/null -w '%{http_code}\n' -H 'Last-Event-ID: 99' \
https://app.example.com/api/jobs/exp_7Qk/events
Run a deploy during a load test with a hundred watched jobs and count client-visible failures. In production, track reconnects per stream and the share of streams that end with a terminal event versus a client disconnect; watch for requests that return 204, which should appear at the rate jobs finish, and never as a steady stream from the same client.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Do progress streams need a replay buffer?
No. Progress is state, so the latest frame is enough. Store it, send it first on every connection, and let the live channel carry what follows.
Why return 204 rather than 200 with the final event?
A 200 response that ends is treated by EventSource as a dropped connection, so it reconnects. A 204 is defined by the specification as the signal to stop reconnecting, which ends the loop without relying on client code.
Is sticky routing still useful?
It reduces reconnect cost slightly when the original node is healthy, but it is not needed for correctness once state is in a shared store, and it must never be relied on across deploys.
How long should finished jobs stay in the store?
Long enough for users to come back to them — a day is common. After that, the job's result page can be served from the database rather than the stream.