Distinguishing Fatal from Transient SSE Errors Permalink to this section

Part of Error Handling & Reconnection UX, under Frontend Consumption & Client Patterns.

EventSource reports every problem the same way: an error event with no details. A Wi-Fi blip that the browser will fix in three seconds, an expired session that needs a login, a deleted resource that will never come back and an overloaded server that needs a minute to recover all look identical to the onerror handler. Treating them the same is how applications end up either retrying forever against a 401 or giving up after a harmless network drop. This guide classifies SSE errors, shows how to get the information EventSource hides, and maps each class to the right reaction.

Symptom & Developer Intent Permalink to this section

  • After a session expires, the live view silently stops updating and never recovers.
  • A “Connection lost” banner flashes on every brief network hiccup.
  • A deleted document keeps reconnecting in a tight loop in the background.
  • During an outage, every client retries at once when the service returns.
  • Error telemetry shows thousands of error events with no way to tell what they mean.

The intent is a client that recovers silently from transient problems, stops or re-authenticates on permanent ones, backs off on overload, and reports each class distinctly.

Root Cause Analysis Permalink to this section

The browser’s behaviour depends on what went wrong, and that behaviour is the only signal available:

Classifying an EventSource error Decision tree using readyState after an error event: CONNECTING means the browser is retrying a transient failure, CLOSED means a permanent failure whose cause must be probed. Classifying an EventSource error readyState is CONNECTING? Transient: browser retries yes no Probe returns 401? Re-authenticate, reopen yes no Probe returns 403 / 404 / 410? Fatal: stop, tell user yes no Server trouble: back off, reopen
readyState after the error is the first and cheapest classifier. For CLOSED, a probe request reveals the status the browser hid.
  • Transient. A network error or a stream that ended normally. The browser sets readyState to CONNECTING and reconnects after the retry interval. The application need do nothing.
  • Permanent, as far as the browser knows. A non-200 status or a wrong content type. The browser sets readyState to CLOSED and never retries. The status is not exposed.

So the first question is always the readyState inside the error handler. For CLOSED, the application must find out why, because the right reaction depends on the status: 401 means re-authenticate, 403/404/410 mean stop, 429/5xx or a proxy’s 502 mean back off and reopen.

Step-by-Step Resolution Permalink to this section

Step 1 — Classify with readyState and a debounce Permalink to this section

es.onerror = () => {
  if (es.readyState === EventSource.CONNECTING) {
    scheduleStatus('reconnecting', 2000);   // show only if it lasts more than 2 s
    return;                                 // the browser is handling it
  }
  onPermanentFailure();                     // CLOSED: find out why
};
es.onopen = () => { cancelScheduledStatus(); setStatus('live'); };

Debouncing the visible status avoids flashing warnings for the many sub-second reconnects that users never notice. The status UX itself is covered in showing connection status in the UI.

Step 2 — Probe the status the browser hid Permalink to this section

async function onPermanentFailure() {
  let status = 0;
  try {
    // A cheap HEAD or a dedicated endpoint with the same auth and authorisation checks as the stream.
    const r = await fetch(streamUrl, { method: 'HEAD', credentials: 'include', cache: 'no-store' });
    status = r.status;
  } catch { status = 0; }                   // network down: treat as transient
  react(status);
}

Make sure the stream route answers HEAD (or add a small /stream/status endpoint) and applies exactly the same authentication and authorisation, so the probe’s status is the status the stream would return.

Step 3 — React per class Permalink to this section

function react(status) {
  switch (true) {
    case status === 401:
      telemetry.count('sse_error', { cls: 'auth' });
      return auth.refresh().then(reopen, () => setStatus('signed-out'));
    case status === 403 || status === 404 || status === 410:
      telemetry.count('sse_error', { cls: 'fatal', status });
      return setStatus('unavailable');                    // stop: retrying cannot help
    case status === 429 || status >= 500 || status === 0:
      telemetry.count('sse_error', { cls: 'overload', status });
      return setTimeout(reopen, backoff.next());          // jittered exponential backoff
    default:
      telemetry.count('sse_error', { cls: 'unknown', status });
      return setTimeout(reopen, backoff.next());
  }
}

Reset the backoff after a stream has been open and healthy for a while, not merely on open, so a server that accepts connections and immediately drops them does not reset the delay each time. Exponential backoff with jitter for SSE reconnects covers the backoff itself.

Error classes and the reaction to each Matrix of five error classes with how they appear to EventSource, who retries, and what the user should see. Error classes and the reaction to each Class EventSource sees Who retries User sees Network blip CONNECTING browser nothing (debounced) Session expired (401) CLOSED app, after refresh sign-in if refresh fails Gone (403/404/410) CLOSED nobody clear message Overload (429/5xx) CLOSED app, backoff reconnecting after delay Proxy error (502/504) CLOSED app, backoff reconnecting after delay
The user sees something only when their action is required or the outage is real. Everything else is handled quietly.

Step 4 — Let the server help the classification Permalink to this section

Servers can make classification easier: return accurate statuses before streaming; shed load with a 200 plus a long retry: rather than a 503, so overload becomes a transient case the browser handles itself (see using retry hints for load shedding); and send an explicit event before ending a stream for a reason the client should know, such as event: access-revoked.

The words matter as much as the logic. For each class, decide in advance what the interface says: nothing at all for transient blips under a couple of seconds; a quiet “Reconnecting…” for longer ones and for server overload; a specific, actionable message for authentication (“Your session expired — sign in to resume live updates”); and a plain statement for permanent loss (“This document was deleted”). Generic “Connection error” text for every class trains users to ignore it, which defeats the purpose on the rare occasions it matters.

Step 5 — Or use a fetch-based client that sees statuses directly Permalink to this section

If probing feels indirect, a fetch-based SSE client receives the real status of every connection attempt and can classify without a second request, at the cost of implementing reconnection itself.

Validation & Monitoring Permalink to this section

Simulate each class against a test server: drop the connection (transient), return 401, 404 and 503, and put a proxy in front that returns 502. Assert the client’s final state and the number of reconnect attempts for each. In production, graph sse_error by class; a spike in auth at a round interval points at token lifetimes, a spike in fatal after a deploy points at a routing change, and overload should correlate with server incidents.

Reconnect attempts in 10 minutes against a deleted resource Bar chart comparing reconnect attempts against a stream that returns 404, for a client that always reopens on error, a client with backoff only, and a client that classifies 404 as fatal. Reconnect attempts in 10 minutes against a deleted resource Always reopen immediately ~600 Backoff, no classification 14 Classified as fatal 2 (stream + probe) requests to the stream endpoint from one client
Without classification, a permanent error becomes a permanent background load. Classification stops it at the first probe.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Why does EventSource not expose the HTTP status?

The API was designed to be simple and to avoid leaking cross-origin information. The only signal it gives is readyState, which distinguishes retrying from given up.

Is a 204 response an error?

No. 204 No Content tells EventSource to stop reconnecting cleanly. Servers use it to end finished streams; the client should treat it as completion, not failure.

Should the client ever give up on transient errors?

Not automatically while the page is open; networks come back. Cap the backoff delay, and pause attempts while the page is hidden or offline, resuming on visibility and online events.

How do I test proxy errors locally?

Put nginx or a fault-injection proxy in front of the test server and return 502 or 504 from it. EventSource treats them like any other non-200 status.