Measuring End-to-End Event Latency Permalink to this section

Part of Observability & Metrics for SSE, under Backend Stream Generation & Connection Management.

For a real-time feature, latency is the product. A dashboard that is connected, error-free and forty seconds behind is broken, yet it passes every health check. Request latency metrics do not help, because a stream has no request per event. What needs measuring is the time from the moment something changes at the source to the moment the user can see it — across the producer, the broker, the SSE node, the network and the browser. This guide stamps events at each hop, corrects for clock differences, reports from real users, and turns the result into alerts that fire when users would notice.

Symptom & Developer Intent Permalink to this section

  • Users say the live view “lags” and nothing in the monitoring shows a problem.
  • The team cannot tell whether a delay is in the broker, the SSE tier, the network or the browser.
  • After a deploy, latency is suspected to have regressed, with no numbers to confirm it.
  • Latency figures computed in the browser are sometimes negative or absurdly large.
  • Averages look fine while a subset of users is consistently far behind.

The intent is a latency measurement per hop and end to end, corrected for clock skew, reported as percentiles, and alertable.

Root Cause Analysis Permalink to this section

Latency hides because no single component sees the whole path. Each hop measures its own processing time at best, and the sum of those is not the end-to-end delay: queueing between hops — in the broker, in socket buffers, in a proxy — is invisible to all of them.

Timestamps along the path of one event Flow of an event from the source through the broker, the SSE node and the network to the browser, with a timestamp recorded at each hop. Timestamps along the path of one event Source t0 change produce Broker t1 publish deliver SSE node t2 write network Browser t3 receive frame Render t4 painted
Stamp once at the source and once at every hop. Differences between consecutive stamps attribute the delay; the first-to-last difference is what users experience.

Browser-side measurements involve two clocks: the server’s and the user’s. User clocks are often wrong by seconds, occasionally by minutes, which produces the negative and absurd values. Correcting requires estimating the offset between the two clocks, which the stream itself can do.

Step-by-Step Resolution Permalink to this section

Step 1 — Stamp the source time into every event Permalink to this section

// At the source: record when the change happened, not when it was published.
await bus.publish('orders', JSON.stringify({ ...change, ts: { src: change.committedAt } }));

Use the business timestamp (commit time, trade time) when available. That makes the measurement include any delay before publication, such as a change-data-capture lag.

Step 2 — Add a stamp at each server hop Permalink to this section

// SSE node: record when the event is written to each socket.
function writeEvent(res, evt) {
  const now = Date.now();
  evt.ts.write = now;
  metrics.srcToWrite.observe(now - evt.ts.src);       // server-side latency, one clock domain
  res.write(`id: ${evt.id}\nevent: ${evt.type}\ndata: ${JSON.stringify(evt)}\n\n`);
}

Server clocks should be synchronised with NTP or chrony, which keeps them within a millisecond or two on a well-run network, so srcToWrite is reliable server-side.

Step 3 — Estimate the client’s clock offset from the stream Permalink to this section

Send the server time in the stream’s opening event and in heartbeats, and compute a smoothed offset on the client:

let offset = 0;                          // serverTime ≈ Date.now() + offset
let samples = 0;
es.addEventListener('time', (e) => {
  const serverNow = JSON.parse(e.data).now;
  const est = serverNow - Date.now();    // ignores one-way network delay: a small, known bias
  offset = samples === 0 ? est : offset * 0.8 + est * 0.2;
  samples++;
});

The estimate includes the one-way network delay, so it biases measured latency down by roughly half a round trip. For user-facing latency that bias is acceptable; for precision, measure round-trip time with a timed fetch and subtract half.

Step 4 — Measure receive and render in the browser Permalink to this section

es.addEventListener('order', (e) => {
  const evt = JSON.parse(e.data);
  const receivedAt = Date.now() + offset;
  applyToStore(evt);
  requestAnimationFrame(() => requestAnimationFrame(() => {   // after the frame containing the change
    const renderedAt = Date.now() + offset;
    latencySamples.push({
      srcToWrite: evt.ts.write - evt.ts.src,
      network: receivedAt - evt.ts.write,
      render: renderedAt - receivedAt,
      total: renderedAt - evt.ts.src,
    });
  }));
});

The double requestAnimationFrame waits until the frame that includes the state change has been produced, a common approximation of “painted”.

Step 5 — Sample and report Permalink to this section

Sending a beacon per event would double traffic. Sample — every event for a small share of sessions, or one event in N for all sessions — and batch:

const SAMPLE = Math.random() < 0.05;                 // 5 % of sessions report
setInterval(() => {
  if (!SAMPLE || !latencySamples.length) return;
  navigator.sendBeacon('/rum/sse-latency', JSON.stringify(latencySamples.splice(0)));
}, 30_000);
Where the time went, p95 by hop Bar chart of p95 latency contributed by each hop for a live order feed: source to broker, broker to SSE write, network, and render. Where the time went, p95 by hop Source → broker 35 ms Broker → SSE write 20 ms Network 180 ms Receive → render 420 ms p95 milliseconds per hop, live order feed
The breakdown turns "it feels slow" into a specific hop. In this example the render step, not the backend, dominated the tail.

Step 6 — Add a synthetic probe for continuous coverage Permalink to this section

Real-user data depends on traffic: at night or in a small region there may be too few sessions to compute percentiles. A synthetic probe fills the gap. Publish a canary event every few seconds from the source side, keep a probe client connected from several regions, and record the canary’s end-to-end delay. Because both ends are under your control and clock-synchronised, the probe gives a clean, always-available baseline, and a divergence between probe latency and real-user latency points at client-side causes — slow devices, heavy rendering — rather than the pipeline.

# Minimal probe: measure canary delay through the public URL every run.
curl -sN https://app.example.com/api/stream?topic=canary \
  | grep --line-buffered '^data:' \
  | while read -r l; do ts=$(echo "${l#data: }" | jq .ts.src); echo "lag=$(( $(date +%s%3N) - ts ))ms"; done

Validation & Monitoring Permalink to this section

Alert on percentiles, per hop and end to end, with thresholds that reflect what users notice:

# Prometheus rules (server-side hop) and RUM-derived series (client hops)
- alert: SSEServerLagHigh
  expr: histogram_quantile(0.99, sum by (le) (rate(sse_src_to_write_ms_bucket[5m]))) > 1000
  for: 5m
- alert: SSEEndToEndLagHigh
  expr: histogram_quantile(0.95, sum by (le) (rate(rum_sse_total_ms_bucket[10m]))) > 2000
  for: 10m
End-to-end p95 around a deploy Line chart of end-to-end p95 latency over six hours, with a jump after a deploy that added per-subscriber serialisation and a recovery after the fix. End-to-end p95 around a deploy end-to-end p95 server hop p95 0 400 800 1200 1600 0 1.2 2.4 3.6 4.8 6 alert threshold hours p95 latency (ms)
Without a per-hop breakdown, the jump would read as "the network got slower". The server hop showed the regression immediately.

Slice the percentiles by region, network type (from the Network Information API where available) and client version. A tail that belongs to one carrier is a network condition; a tail across everyone after a release is a regression.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Why not just measure server processing time?

Because most delay in a streaming system is queueing between components and in the network, which no single component's processing time includes. Only timestamps carried with the event capture it.

Can I trust Date.now() in the browser?

Not in absolute terms — user clocks are frequently off by seconds. Use it only after correcting with an offset estimated from server time sent in the stream.

Does adding timestamps bloat events?

A few numeric fields add tens of bytes. If that matters, send them only on sampled events or strip them for sessions that do not report.

What latency should a live feature target?

For human-facing updates, under a second end to end at p95 feels live. Dashboards tolerate a few seconds; collaborative cursors and trading screens want a few hundred milliseconds.