Keeping SSE Alive on Flaky Mobile Networks Permalink to this section

Part of Mobile Background Tab Handling, under Frontend Consumption & Client Patterns.

On a desktop broadband connection, a Server-Sent Events stream can stay open for hours without incident. On a phone moving through a city, the network changes underneath it constantly: Wi-Fi to cellular, one cell tower to the next, a tunnel, a lift, a carrier NAT that silently forgets idle connections. The browser’s automatic reconnection handles clean failures, but mobile failures are often silent — the connection looks open while nothing can flow. This guide adds the detection, recovery and payload discipline that make a stream usable on a phone.

Symptom & Developer Intent Permalink to this section

  • On the train, the live view freezes for minutes while the page still shows “connected”.
  • After walking out of a building onto cellular, updates only resume when the user pulls to refresh.
  • The stream reconnects repeatedly in weak coverage, each attempt taking many seconds.
  • Mobile data usage from the live feature is higher than expected.
  • Users on mobile see stale prices or statuses without knowing they are stale.

The intent is a stream that detects silent failure within seconds, reconnects promptly when the network returns, resumes without data loss, uses little bandwidth, and tells the user honestly when data is not current.

Root Cause Analysis Permalink to this section

Clean failures — a reset, a closed connection — trigger EventSource’s reconnection immediately. Mobile networks frequently produce unclean ones instead:

A silent failure on a mobile network Timeline of three minutes in which a phone switches from Wi-Fi to cellular, the old connection stops delivering without closing, and a silence watchdog triggers a reconnect on the new network. A silent failure on a mobile network Wi-Fi stream Dead socket Cellular stream events + heartbeats looks open, nothing arrives resumed from Last-Event-ID 0 36 72 108 144 180 seconds network switch watchdog fires
Without the watchdog, EventSource would wait on the dead connection until TCP gives up, which can take many minutes.
  • Network switches. When the device moves from Wi-Fi to cellular, connections bound to the old interface stop working. The browser may not notice until TCP retransmission times out, which can take minutes.
  • NAT and carrier timeouts. Carrier-grade NATs drop mappings for idle connections, sometimes after 30 seconds or less. The server’s next write goes nowhere and nobody is told.
  • Weak coverage. Connections succeed slowly and fail often; tight retry loops waste battery and data.

The only reliable detector for silent failure is the absence of expected traffic: if heartbeats are sent every 15 seconds and nothing has arrived for 40, the connection is dead regardless of what readyState says.

Step-by-Step Resolution Permalink to this section

Step 1 — Send a visible heartbeat and run a silence watchdog Permalink to this section

Comment heartbeats are invisible to EventSource listeners, so send a tiny named event the client can observe:

event: hb
data: 
const HEARTBEAT_MS = 15000;
let lastTraffic = Date.now();
const touch = () => { lastTraffic = Date.now(); };

function watch(es) {
  es.addEventListener('hb', touch);
  es.addEventListener('open', touch);
  EVENT_TYPES.forEach((t) => es.addEventListener(t, touch));
}

setInterval(() => {
  if (document.visibilityState !== 'visible') return;
  if (Date.now() - lastTraffic > HEARTBEAT_MS * 2.5) reconnectNow('silence');
}, 5000);

reconnectNow closes the current EventSource and opens a new one, passing the last event id as a query parameter, since a new instance does not carry Last-Event-ID over.

Step 2 — Reconnect immediately on network change Permalink to this section

window.addEventListener('online', () => reconnectNow('online'));
navigator.connection?.addEventListener?.('change', () => reconnectNow('network-change'));   // where supported
document.addEventListener('visibilitychange', () => {
  if (document.visibilityState === 'visible' && Date.now() - lastTraffic > HEARTBEAT_MS) reconnectNow('visible');
});

The online event fires when the browser regains connectivity; the Network Information API’s change event (not available in every browser) fires when the connection type changes. Reconnecting on these signals avoids waiting for the watchdog. Pause attempts while navigator.onLine is false — they cannot succeed and only drain the battery.

Every path back to a live stream Flow from four triggers — silence watchdog, online event, network change and tab becoming visible — into a single reconnect function that resumes from the stored last event id. Every path back to a live stream Triggers watchdog, online, change once reconnectNow() close old source open New EventSource ?after=lastId server Replay missed events caught up Live status: live
Several independent detectors, one reconnect path. Whichever notices first wins; the resume makes the result identical.

Step 3 — Resume fast, and keep the replay cheap Permalink to this section

Every mobile reconnect replays what was missed, so replay must be quick: an indexed range query or a stream-log read, not a scan. For state-shaped data, send a compact snapshot on reconnect instead of every intermediate event, as in snapshot plus delta streaming for dashboards. A reconnect should cost one small response, not a flood.

Step 4 — Spend bandwidth deliberately Permalink to this section

Mobile data costs users money. Keep event payloads compact (short keys, only changed fields), conflate high-rate data on the server, and consider a lower update rate for clients that report a slow connection:

const slow = navigator.connection && ['slow-2g', '2g', '3g'].includes(navigator.connection.effectiveType);
const es = new EventSource(`/api/stream?rate=${slow ? 'low' : 'normal'}&after=${encodeURIComponent(lastId)}`);

The server then conflates to, for example, one update per second instead of ten for rate=low. Heartbeats add a little traffic, but at one small event every 15 seconds they are negligible next to real data.

Step 5 — Be honest about freshness Permalink to this section

Show when data was last updated, and switch the indicator to “offline” or “reconnecting” when the watchdog fires, not only when EventSource reports an error. Users on the move accept stale data far more readily when they can see it is stale.

Validation & Monitoring Permalink to this section

Test on a real phone on real networks: walk out of Wi-Fi range with the page open, ride a train, toggle airplane mode. In DevTools, simulate offline periods and throttled profiles, and use remote debugging to watch reconnects on the device.

Median time frozen after a Wi-Fi to cellular switch Bar chart comparing how long the live view stays frozen after a network switch with plain EventSource, with a silence watchdog, and with watchdog plus online and network-change triggers. Median time frozen after a Wi-Fi to cellular switch Plain EventSource ~140 s + silence watchdog ~38 s + online / change triggers ~3 s seconds until live updates resume after the switch
Illustrative field measurements. The watchdog bounds the worst case; network-change triggers make the common case nearly instant.

Report reconnect causes (silence, online, network-change, visible, browser-initiated) and time-to-live after reconnect from real mobile sessions. A high silence share on a particular carrier is a sign of an aggressive NAT timeout; shortening the heartbeat for mobile clients may help.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Why does EventSource not detect dead mobile connections itself?

It relies on TCP to report failures, and a connection bound to a vanished network interface can take minutes to fail at the TCP level. Only application-level heartbeats reveal it quickly.

Should I shorten the heartbeat for mobile clients?

Only if you observe carrier NAT timeouts shorter than your interval. Shorter heartbeats cost battery and data; 15 seconds is a reasonable default.

Is WebSocket better on mobile networks?

It faces the same silent-failure problem and needs the same ping, watchdog and resume logic. SSE's built-in Last-Event-ID resume is an advantage once reconnection is handled.

What about battery usage?

An open stream keeps the radio more active. Close the stream when the page is hidden and reconnect on return, which the visibility handling above already supports.