Live Dashboards & Metrics Feeds Permalink to this section
Part of Real-Time Application Patterns.
A live dashboard is the canonical Server-Sent Events use case: a page of numbers, gauges and charts that update without the user pressing refresh. It is also the pattern where a working demo and a production-grade feature diverge the most. The demo pushes every metric change straight to the browser and re-renders on each message. In production the same design shows stale values after every network blip, freezes the tab when an incident makes the metrics busy — exactly when people are staring at the dashboard — and costs one database query per viewer per update. This guide is for engineers building operations consoles, business KPI boards, IoT fleet views or admin panels, and it covers the stream design that avoids all three failures: a snapshot on connect followed by deltas, sampling and coalescing on the server, and rendering locked to the display’s frame rate on the client.
How It Works Permalink to this section
A dashboard models current state. The user does not care about the fourteen intermediate values the CPU gauge passed through while their laptop was asleep; they care about the value now. That single observation shapes the protocol usage.
The stream carries two event types. snapshot holds the whole view — every value the dashboard displays — and is always the first frame of a connection. delta holds only the fields that changed since the previous frame. The id is a version counter for the view as a whole. The server does not need to replay anything: a reconnecting client receives a new snapshot and is immediately correct.
retry: 3000
event: snapshot
id: v900
data: {"cpu":38.2,"mem":61.0,"rps":1180,"err_rate":0.002,"p95_ms":204,"hosts_up":42}
event: delta
id: v901
data: {"rps":1193,"p95_ms":211}
: heartbeat 1726650000
event: delta
id: v902
data: {"cpu":39.0,"err_rate":0.003}
The Last-Event-ID the browser sends on reconnect is still useful, just not for replay. If the version it names equals the current one, the server may skip the snapshot; if it is older, a snapshot is due. That optimisation matters little for a small dashboard and a lot for a large one with hundreds of tiles.
Server-Side Implementation Permalink to this section
The server’s job splits into two independent loops. A sampler reads the metrics source on a fixed cadence and computes what changed. A broadcaster writes the resulting delta to every connected viewer. Crucially, the sampler runs once per node regardless of how many viewers there are.
// dashboard-stream.js — Node.js, one sampler shared by every viewer on this process.
import { EventEmitter } from 'node:events';
const hub = new EventEmitter();
hub.setMaxListeners(0); // thousands of viewers is normal
let view = {}; // latest full snapshot
let version = 0;
async function sample() {
const next = await metrics.readDashboard(); // ONE read per tick, not per viewer
const delta = {};
for (const [k, v] of Object.entries(next)) {
// Ignore changes below display precision — they would be invisible anyway.
if (view[k] === undefined || Math.abs(v - view[k]) >= precision(k)) delta[k] = v;
}
view = next;
if (Object.keys(delta).length) {
version += 1;
hub.emit('delta', { id: `v${version}`, data: delta });
}
}
setInterval(() => sample().catch((e) => log.warn({ e }, 'sample failed')), 1000);
app.get('/dash/stream', (req, res) => {
res.writeHead(200, {
'Content-Type': 'text/event-stream',
'Cache-Control': 'no-cache',
'X-Accel-Buffering': 'no',
});
res.write('retry: 3000\n\n');
// Every connection starts from the truth.
res.write(`event: snapshot\nid: v${version}\ndata: ${JSON.stringify(view)}\n\n`);
const onDelta = (f) => {
// Skip viewers whose socket buffer is full; they will get the next delta.
if (res.writableNeedDrain) return;
res.write(`event: delta\nid: ${f.id}\ndata: ${JSON.stringify(f.data)}\n\n`);
};
hub.on('delta', onDelta);
const beat = setInterval(() => res.write(`: hb ${Date.now()}\n\n`), 15000);
req.on('close', () => { hub.off('delta', onDelta); clearInterval(beat); });
});
Three details in that handler are deliberate. The precision filter drops changes too small to display, which on a noisy metric can remove most of the traffic. The writableNeedDrain check means a viewer on a slow link skips deltas instead of accumulating an unbounded write queue; because deltas only ever overwrite values, a skipped delta costs nothing once the next one arrives — but it does mean a skipped field can stay stale until it changes again, so production versions send a fresh snapshot to a viewer once it drains. And the heartbeat keeps idle proxies from closing a quiet dashboard at night.
The same structure translates directly to Python. With FastAPI and an asyncio sampler task, each viewer gets its own bounded queue; when the queue is full the viewer is marked for a resnapshot instead of blocking the sampler.
# dashboard.py — FastAPI; one sampler task, one small queue per viewer.
import asyncio, json
from fastapi import FastAPI, Request
from fastapi.responses import StreamingResponse
app = FastAPI()
view: dict = {}
version = 0
viewers: set[asyncio.Queue] = set()
async def sampler():
global view, version
while True:
nxt = await read_dashboard() # one read per tick
delta = {k: v for k, v in nxt.items() if view.get(k) != v}
view = nxt
if delta:
version += 1
frame = f"event: delta\nid: v{version}\ndata: {json.dumps(delta)}\n\n"
for q in list(viewers):
if q.full():
q.resnapshot = True # fall back to a snapshot later
else:
q.put_nowait(frame)
await asyncio.sleep(1)
@app.on_event("startup")
async def start():
asyncio.create_task(sampler())
@app.get("/dash/stream")
async def stream(request: Request):
q: asyncio.Queue = asyncio.Queue(maxsize=32)
q.resnapshot = False
viewers.add(q)
async def gen():
try:
yield "retry: 3000\n\n"
yield f"event: snapshot\nid: v{version}\ndata: {json.dumps(view)}\n\n"
while not await request.is_disconnected():
try:
frame = await asyncio.wait_for(q.get(), timeout=15)
except asyncio.TimeoutError:
yield ": hb\n\n" # keep idle proxies honest
continue
if q.resnapshot:
q.resnapshot = False
yield f"event: snapshot\nid: v{version}\ndata: {json.dumps(view)}\n\n"
else:
yield frame
finally:
viewers.discard(q)
return StreamingResponse(gen(), media_type="text/event-stream",
headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})
In Go the natural shape is a hub goroutine that owns the viewer set and a buffered channel per viewer, with a non-blocking send so a slow viewer never stalls the hub. That pattern is developed in full in fanning out events with a Go hub goroutine; the dashboard only changes what the hub broadcasts.
In a multi-node deployment the sampler still runs per node. If the metrics source cannot tolerate one reader per node, move the sampler behind a broker: a single leader samples and publishes deltas, and every node subscribes. The Redis pub/sub fan-out topic covers that relay; the only dashboard-specific addition is that each node must also cache the latest snapshot so it can greet new connections without asking the leader.
Client-Side Consumption Permalink to this section
The client keeps one state object and applies frames to it. Rendering is decoupled from arrival: frames update the state as they arrive, and a requestAnimationFrame loop paints at most once per display frame.
// useDashboard.js — framework-free core; wrap it in a React hook or Vue composable.
export function connectDashboard(url, paint) {
let state = null;
let dirty = false;
const es = new EventSource(url);
es.addEventListener('snapshot', (e) => { state = JSON.parse(e.data); dirty = true; });
es.addEventListener('delta', (e) => {
if (!state) return; // a delta before any snapshot is meaningless
Object.assign(state, JSON.parse(e.data));
dirty = true;
});
let raf;
const loop = () => {
if (dirty) { dirty = false; paint(state); } // at most one paint per frame
raf = requestAnimationFrame(loop);
};
raf = requestAnimationFrame(loop);
return () => { es.close(); cancelAnimationFrame(raf); };
}
Wrapped in a framework, the same idea becomes a hook that returns the state and a connection status. The React EventSource hooks topic builds the generic hook; the dashboard version only adds the snapshot/delta reducer and the frame-rate gate. Throttling dashboard updates to the frame rate measures what that gate saves.
Charts need one more rule: a bounded history. A line chart that appends every point forever eventually holds millions of points and takes seconds to redraw. Keep a fixed-size ring buffer per series, sized to the chart’s visible window, and discard the rest — the server remains the source of history.
Edge Cases & Network Interference Permalink to this section
Dashboards are often left open for days on wall-mounted screens and in background tabs, which exposes every intermediary’s idle behaviour.
- Idle proxies closing quiet streams. A dashboard that only changes during business hours goes silent at night, and a load balancer with a 60-second idle timeout closes it. Heartbeat comments every 15 seconds keep it open; see choosing heartbeat intervals.
- Buffering proxies batching deltas. A proxy with buffering enabled delivers ten seconds of deltas at once, which looks like a frozen dashboard followed by a jump. Disable buffering on the route and send
X-Accel-Buffering: no; the proxy and CDN configuration topic has per-proxy settings. - Background tabs. Browsers throttle timers and animation frames in hidden tabs, so the paint loop stops — which is exactly right. But the
EventSourcekeeps receiving, and state keeps being updated. When the tab becomes visible, the first frame paints the latest state. If a hidden tab should not hold a connection at all, close it onvisibilitychangeand reconnect on return; the fresh snapshot makes that safe. - Clock-sensitive values. “Last updated 3 seconds ago” computed from server timestamps drifts if the client clock is wrong. Send server time in the heartbeat and compute an offset.
- Wall screens and sleep. A kiosk that sleeps and wakes may keep a dead connection in
OPENstate. Run a silence watchdog: if no byte, not even a heartbeat, arrives in twice the heartbeat interval, close and reopen.
Mitigation checklist:
Performance & Scale Considerations Permalink to this section
The cost of a dashboard stream is dominated by write fan-out: frames per second multiplied by viewers. There are three levers, and they multiply.
- Coalesce per tick. Emit one delta per sampling interval containing every changed field, not one frame per metric change. This is the difference between frame count scaling with metrics and scaling with time.
- Filter below display precision. A gauge showing whole percentages does not need a frame when the value moves from 38.21 to 38.24.
- Tier the cadence. Error rates and request rates earn a one-second tick; disk usage and host counts can tick every ten seconds. Two samplers with different intervals feeding the same stream is simple and effective.
Memory per viewer is small because the stream holds no per-viewer state beyond the socket and one listener. File descriptors are the first limit a busy dashboard node hits — which is a connection pooling problem, not a dashboard one. Snapshot plus delta streaming goes further into versioning and large views.
Validation & Debugging Permalink to this section
Validate the stream contract with curl before touching the interface. The first frame must be a snapshot, and deltas must arrive at the sampling cadence, not in bursts.
# The first event must be a snapshot; deltas should follow roughly once per second.
curl -sN -H 'Accept: text/event-stream' https://app.example.com/dash/stream \
| awk '/^event:/ { printf "%s %s\n", strftime("%H:%M:%S"), $2; fflush() }' | head -20
# Reconnect with a stale cursor: expect a fresh snapshot, not a replay.
curl -sN -H 'Last-Event-ID: v1' https://app.example.com/dash/stream | head -4
Bursty timestamps in the first command mean something between the server and curl is buffering. In Chrome DevTools, the Network panel’s EventStream tab on the stream request shows each event with its type and arrival time; a column of identical times is the same symptom.
Log one line per connection lifecycle rather than per frame:
{"evt":"dash_stream_open","viewer":"u_1842","cursor":"v900","snapshot_bytes":1840}
{"evt":"dash_stream_close","viewer":"u_1842","duration_s":8412,"frames":8190,"skipped_slow":14,"reason":"client_gone"}
Because the client logic is a reducer, the most valuable automated test needs no network at all. Record a real stream once with curl, commit it as a fixture, and assert on the state it produces:
// dashboard.test.js — replay a recorded stream through the same parsing and reducer code.
import { readFileSync } from 'node:fs';
import { parseEventStream } from './parse.js';
import { dashboardReducer } from './reducer.js';
test('a delta after a reconnect snapshot never resurrects old values', () => {
const frames = parseEventStream(readFileSync('fixtures/dash-reconnect.txt', 'utf8'));
const state = frames.reduce(dashboardReducer, { version: null, values: {} });
expect(state.version).toBe('v917');
expect(state.values.rps).toBe(1302); // from the post-reconnect snapshot
});
A non-zero skipped_slow is expected on mobile networks; a large one on office networks points at a proxy or a server that cannot keep up. The Performance panel’s frame chart confirms the client side: with the frame-rate gate in place there should be at most one paint per frame even when the stream is busy.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Why not send the whole dashboard on every update?
For a small view it is fine and simpler. Once the view exceeds a few kilobytes and updates every second, full frames multiply bandwidth by the number of viewers for no benefit. Snapshot on connect plus deltas afterwards keeps the correctness of full frames at the cost of deltas.
Do dashboards need Last-Event-ID replay?
No. A dashboard models current state, so a reconnecting client needs a fresh snapshot rather than the deltas it missed. The id is still useful as a version, letting the server skip the snapshot when the client is already current.
How often should a dashboard update?
No faster than people can read and no faster than the metric changes meaningfully. One second is a good default for operational metrics; business KPIs are usually fine at five to thirty seconds. Anything faster than the display frame rate is wasted work.
Can one stream serve dashboards with different tiles for different users?
Yes, if the projector filters per viewer. Keep one sampler producing the union of all metrics, and have each viewer's handler pick out the fields that viewer is allowed and has chosen to see before writing. Filtering in the browser instead would leak metrics to anyone who opens DevTools.
Should dashboard values be rounded on the server or the client?
On the server, to the precision the interface displays. Rounding first lets the precision filter suppress invisible changes, which is often the largest bandwidth saving available, and it keeps every viewer showing identical numbers.
What happens to a dashboard in a background tab?
The stream keeps delivering and state keeps updating, but animation frames stop, so nothing is painted. When the tab returns, the next frame paints the latest state. If background connections are too expensive, close on visibility change and rely on the snapshot when reopening.