Measuring End-to-End Event Latency Permalink to this section
Part of Observability & Metrics for SSE, under Backend Stream Generation & Connection Management.
For a real-time feature, latency is the product. A dashboard that is connected, error-free and forty seconds behind is broken, yet it passes every health check. Request latency metrics do not help, because a stream has no request per event. What needs measuring is the time from the moment something changes at the source to the moment the user can see it — across the producer, the broker, the SSE node, the network and the browser. This guide stamps events at each hop, corrects for clock differences, reports from real users, and turns the result into alerts that fire when users would notice.
Symptom & Developer Intent Permalink to this section
- Users say the live view “lags” and nothing in the monitoring shows a problem.
- The team cannot tell whether a delay is in the broker, the SSE tier, the network or the browser.
- After a deploy, latency is suspected to have regressed, with no numbers to confirm it.
- Latency figures computed in the browser are sometimes negative or absurdly large.
- Averages look fine while a subset of users is consistently far behind.
The intent is a latency measurement per hop and end to end, corrected for clock skew, reported as percentiles, and alertable.
Root Cause Analysis Permalink to this section
Latency hides because no single component sees the whole path. Each hop measures its own processing time at best, and the sum of those is not the end-to-end delay: queueing between hops — in the broker, in socket buffers, in a proxy — is invisible to all of them.
Browser-side measurements involve two clocks: the server’s and the user’s. User clocks are often wrong by seconds, occasionally by minutes, which produces the negative and absurd values. Correcting requires estimating the offset between the two clocks, which the stream itself can do.
Step-by-Step Resolution Permalink to this section
Step 1 — Stamp the source time into every event Permalink to this section
// At the source: record when the change happened, not when it was published.
await bus.publish('orders', JSON.stringify({ ...change, ts: { src: change.committedAt } }));
Use the business timestamp (commit time, trade time) when available. That makes the measurement include any delay before publication, such as a change-data-capture lag.
Step 2 — Add a stamp at each server hop Permalink to this section
// SSE node: record when the event is written to each socket.
function writeEvent(res, evt) {
const now = Date.now();
evt.ts.write = now;
metrics.srcToWrite.observe(now - evt.ts.src); // server-side latency, one clock domain
res.write(`id: ${evt.id}\nevent: ${evt.type}\ndata: ${JSON.stringify(evt)}\n\n`);
}
Server clocks should be synchronised with NTP or chrony, which keeps them within a millisecond or two on a well-run network, so srcToWrite is reliable server-side.
Step 3 — Estimate the client’s clock offset from the stream Permalink to this section
Send the server time in the stream’s opening event and in heartbeats, and compute a smoothed offset on the client:
let offset = 0; // serverTime ≈ Date.now() + offset
let samples = 0;
es.addEventListener('time', (e) => {
const serverNow = JSON.parse(e.data).now;
const est = serverNow - Date.now(); // ignores one-way network delay: a small, known bias
offset = samples === 0 ? est : offset * 0.8 + est * 0.2;
samples++;
});
The estimate includes the one-way network delay, so it biases measured latency down by roughly half a round trip. For user-facing latency that bias is acceptable; for precision, measure round-trip time with a timed fetch and subtract half.
Step 4 — Measure receive and render in the browser Permalink to this section
es.addEventListener('order', (e) => {
const evt = JSON.parse(e.data);
const receivedAt = Date.now() + offset;
applyToStore(evt);
requestAnimationFrame(() => requestAnimationFrame(() => { // after the frame containing the change
const renderedAt = Date.now() + offset;
latencySamples.push({
srcToWrite: evt.ts.write - evt.ts.src,
network: receivedAt - evt.ts.write,
render: renderedAt - receivedAt,
total: renderedAt - evt.ts.src,
});
}));
});
The double requestAnimationFrame waits until the frame that includes the state change has been produced, a common approximation of “painted”.
Step 5 — Sample and report Permalink to this section
Sending a beacon per event would double traffic. Sample — every event for a small share of sessions, or one event in N for all sessions — and batch:
const SAMPLE = Math.random() < 0.05; // 5 % of sessions report
setInterval(() => {
if (!SAMPLE || !latencySamples.length) return;
navigator.sendBeacon('/rum/sse-latency', JSON.stringify(latencySamples.splice(0)));
}, 30_000);
Step 6 — Add a synthetic probe for continuous coverage Permalink to this section
Real-user data depends on traffic: at night or in a small region there may be too few sessions to compute percentiles. A synthetic probe fills the gap. Publish a canary event every few seconds from the source side, keep a probe client connected from several regions, and record the canary’s end-to-end delay. Because both ends are under your control and clock-synchronised, the probe gives a clean, always-available baseline, and a divergence between probe latency and real-user latency points at client-side causes — slow devices, heavy rendering — rather than the pipeline.
# Minimal probe: measure canary delay through the public URL every run.
curl -sN https://app.example.com/api/stream?topic=canary \
| grep --line-buffered '^data:' \
| while read -r l; do ts=$(echo "${l#data: }" | jq .ts.src); echo "lag=$(( $(date +%s%3N) - ts ))ms"; done
Validation & Monitoring Permalink to this section
Alert on percentiles, per hop and end to end, with thresholds that reflect what users notice:
# Prometheus rules (server-side hop) and RUM-derived series (client hops)
- alert: SSEServerLagHigh
expr: histogram_quantile(0.99, sum by (le) (rate(sse_src_to_write_ms_bucket[5m]))) > 1000
for: 5m
- alert: SSEEndToEndLagHigh
expr: histogram_quantile(0.95, sum by (le) (rate(rum_sse_total_ms_bucket[10m]))) > 2000
for: 10m
Slice the percentiles by region, network type (from the Network Information API where available) and client version. A tail that belongs to one carrier is a network condition; a tail across everyone after a release is a regression.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Why not just measure server processing time?
Because most delay in a streaming system is queueing between components and in the network, which no single component's processing time includes. Only timestamps carried with the event capture it.
Can I trust Date.now() in the browser?
Not in absolute terms — user clocks are frequently off by seconds. Use it only after correcting with an offset estimated from server time sent in the stream.
Does adding timestamps bloat events?
A few numeric fields add tens of bytes. If that matters, send them only on sampled events or strip them for sessions that do not report.
What latency should a live feature target?
For human-facing updates, under a second end to end at p95 feels live. Dashboards tolerate a few seconds; collaborative cursors and trading screens want a few hundred milliseconds.