Live Dashboards & Metrics Feeds Permalink to this section

Part of Real-Time Application Patterns.

A live dashboard is the canonical Server-Sent Events use case: a page of numbers, gauges and charts that update without the user pressing refresh. It is also the pattern where a working demo and a production-grade feature diverge the most. The demo pushes every metric change straight to the browser and re-renders on each message. In production the same design shows stale values after every network blip, freezes the tab when an incident makes the metrics busy — exactly when people are staring at the dashboard — and costs one database query per viewer per update. This guide is for engineers building operations consoles, business KPI boards, IoT fleet views or admin panels, and it covers the stream design that avoids all three failures: a snapshot on connect followed by deltas, sampling and coalescing on the server, and rendering locked to the display’s frame rate on the client.

How It Works Permalink to this section

A dashboard models current state. The user does not care about the fourteen intermediate values the CPU gauge passed through while their laptop was asleep; they care about the value now. That single observation shapes the protocol usage.

Snapshot first, deltas after, snapshot again on reconnect Sequence diagram of a browser opening a dashboard stream, receiving a snapshot then deltas, losing the connection and receiving a fresh snapshot on reconnect. Snapshot first, deltas after, snapshot again on reconnect Browser SSE node Metrics store GET /dash/stream read current view event: snapshot id: v900 event: delta id: v901 connection lost EventSource waits the retry interval, then reconnects GET Last-Event-ID: v901 event: snapshot id: v917
The reconnect needs no replay buffer. The server answers every new connection with the current state, so a client is never more than one frame away from the truth.

The stream carries two event types. snapshot holds the whole view — every value the dashboard displays — and is always the first frame of a connection. delta holds only the fields that changed since the previous frame. The id is a version counter for the view as a whole. The server does not need to replay anything: a reconnecting client receives a new snapshot and is immediately correct.

retry: 3000

event: snapshot
id: v900
data: {"cpu":38.2,"mem":61.0,"rps":1180,"err_rate":0.002,"p95_ms":204,"hosts_up":42}

event: delta
id: v901
data: {"rps":1193,"p95_ms":211}

: heartbeat 1726650000

event: delta
id: v902
data: {"cpu":39.0,"err_rate":0.003}

The Last-Event-ID the browser sends on reconnect is still useful, just not for replay. If the version it names equals the current one, the server may skip the snapshot; if it is older, a snapshot is due. That optimisation matters little for a small dashboard and a lot for a large one with hundreds of tiles.

Server-Side Implementation Permalink to this section

The server’s job splits into two independent loops. A sampler reads the metrics source on a fixed cadence and computes what changed. A broadcaster writes the resulting delta to every connected viewer. Crucially, the sampler runs once per node regardless of how many viewers there are.

One sampler per node, not one per viewer A metrics store is read by a single sampler loop on the node, which computes a delta and writes it to every open viewer connection. One sampler per node, not one per viewer Metrics store Prometheus, SQL, cache Sampler loop every 1 s, diff 1 query / s Viewer 1 delta frame Viewer 2 delta frame Viewer 5,000 delta frame write
Viewer count multiplies socket writes, never database reads. At 5,000 viewers the store still sees one query per second from each node.
// dashboard-stream.js — Node.js, one sampler shared by every viewer on this process.
import { EventEmitter } from 'node:events';

const hub = new EventEmitter();
hub.setMaxListeners(0);                 // thousands of viewers is normal

let view = {};                          // latest full snapshot
let version = 0;

async function sample() {
  const next = await metrics.readDashboard();     // ONE read per tick, not per viewer
  const delta = {};
  for (const [k, v] of Object.entries(next)) {
    // Ignore changes below display precision — they would be invisible anyway.
    if (view[k] === undefined || Math.abs(v - view[k]) >= precision(k)) delta[k] = v;
  }
  view = next;
  if (Object.keys(delta).length) {
    version += 1;
    hub.emit('delta', { id: `v${version}`, data: delta });
  }
}
setInterval(() => sample().catch((e) => log.warn({ e }, 'sample failed')), 1000);

app.get('/dash/stream', (req, res) => {
  res.writeHead(200, {
    'Content-Type': 'text/event-stream',
    'Cache-Control': 'no-cache',
    'X-Accel-Buffering': 'no',
  });
  res.write('retry: 3000\n\n');
  // Every connection starts from the truth.
  res.write(`event: snapshot\nid: v${version}\ndata: ${JSON.stringify(view)}\n\n`);

  const onDelta = (f) => {
    // Skip viewers whose socket buffer is full; they will get the next delta.
    if (res.writableNeedDrain) return;
    res.write(`event: delta\nid: ${f.id}\ndata: ${JSON.stringify(f.data)}\n\n`);
  };
  hub.on('delta', onDelta);
  const beat = setInterval(() => res.write(`: hb ${Date.now()}\n\n`), 15000);

  req.on('close', () => { hub.off('delta', onDelta); clearInterval(beat); });
});

Three details in that handler are deliberate. The precision filter drops changes too small to display, which on a noisy metric can remove most of the traffic. The writableNeedDrain check means a viewer on a slow link skips deltas instead of accumulating an unbounded write queue; because deltas only ever overwrite values, a skipped delta costs nothing once the next one arrives — but it does mean a skipped field can stay stale until it changes again, so production versions send a fresh snapshot to a viewer once it drains. And the heartbeat keeps idle proxies from closing a quiet dashboard at night.

The same structure translates directly to Python. With FastAPI and an asyncio sampler task, each viewer gets its own bounded queue; when the queue is full the viewer is marked for a resnapshot instead of blocking the sampler.

# dashboard.py — FastAPI; one sampler task, one small queue per viewer.
import asyncio, json
from fastapi import FastAPI, Request
from fastapi.responses import StreamingResponse

app = FastAPI()
view: dict = {}
version = 0
viewers: set[asyncio.Queue] = set()

async def sampler():
    global view, version
    while True:
        nxt = await read_dashboard()                     # one read per tick
        delta = {k: v for k, v in nxt.items() if view.get(k) != v}
        view = nxt
        if delta:
            version += 1
            frame = f"event: delta\nid: v{version}\ndata: {json.dumps(delta)}\n\n"
            for q in list(viewers):
                if q.full():
                    q.resnapshot = True                  # fall back to a snapshot later
                else:
                    q.put_nowait(frame)
        await asyncio.sleep(1)

@app.on_event("startup")
async def start():
    asyncio.create_task(sampler())

@app.get("/dash/stream")
async def stream(request: Request):
    q: asyncio.Queue = asyncio.Queue(maxsize=32)
    q.resnapshot = False
    viewers.add(q)

    async def gen():
        try:
            yield "retry: 3000\n\n"
            yield f"event: snapshot\nid: v{version}\ndata: {json.dumps(view)}\n\n"
            while not await request.is_disconnected():
                try:
                    frame = await asyncio.wait_for(q.get(), timeout=15)
                except asyncio.TimeoutError:
                    yield ": hb\n\n"                     # keep idle proxies honest
                    continue
                if q.resnapshot:
                    q.resnapshot = False
                    yield f"event: snapshot\nid: v{version}\ndata: {json.dumps(view)}\n\n"
                else:
                    yield frame
        finally:
            viewers.discard(q)

    return StreamingResponse(gen(), media_type="text/event-stream",
                             headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})

In Go the natural shape is a hub goroutine that owns the viewer set and a buffered channel per viewer, with a non-blocking send so a slow viewer never stalls the hub. That pattern is developed in full in fanning out events with a Go hub goroutine; the dashboard only changes what the hub broadcasts.

In a multi-node deployment the sampler still runs per node. If the metrics source cannot tolerate one reader per node, move the sampler behind a broker: a single leader samples and publishes deltas, and every node subscribes. The Redis pub/sub fan-out topic covers that relay; the only dashboard-specific addition is that each node must also cache the latest snapshot so it can greet new connections without asking the leader.

Client-Side Consumption Permalink to this section

The client keeps one state object and applies frames to it. Rendering is decoupled from arrival: frames update the state as they arrive, and a requestAnimationFrame loop paints at most once per display frame.

// useDashboard.js — framework-free core; wrap it in a React hook or Vue composable.
export function connectDashboard(url, paint) {
  let state = null;
  let dirty = false;

  const es = new EventSource(url);
  es.addEventListener('snapshot', (e) => { state = JSON.parse(e.data); dirty = true; });
  es.addEventListener('delta', (e) => {
    if (!state) return;                          // a delta before any snapshot is meaningless
    Object.assign(state, JSON.parse(e.data));
    dirty = true;
  });

  let raf;
  const loop = () => {
    if (dirty) { dirty = false; paint(state); }  // at most one paint per frame
    raf = requestAnimationFrame(loop);
  };
  raf = requestAnimationFrame(loop);

  return () => { es.close(); cancelAnimationFrame(raf); };
}

Wrapped in a framework, the same idea becomes a hook that returns the state and a connection status. The React EventSource hooks topic builds the generic hook; the dashboard version only adds the snapshot/delta reducer and the frame-rate gate. Throttling dashboard updates to the frame rate measures what that gate saves.

Charts need one more rule: a bounded history. A line chart that appends every point forever eventually holds millions of points and takes seconds to redraw. Keep a fixed-size ring buffer per series, sized to the chart’s visible window, and discard the rest — the server remains the source of history.

Client memory with and without a bounded series buffer Line chart over eight hours showing heap size growing linearly for an unbounded chart series and staying flat for a ring buffer sized to the visible window. Client memory with and without a bounded series buffer unbounded array ring buffer (1 h window) 0 200 400 600 800 0 1.6 3.2 4.8 6.4 8 mobile tab kill zone hours the dashboard has been open tab heap (MB)
One point per second across twenty series adds up. The ring buffer holds the visible window and nothing else; older history comes from the server on demand.

Edge Cases & Network Interference Permalink to this section

Dashboards are often left open for days on wall-mounted screens and in background tabs, which exposes every intermediary’s idle behaviour.

  • Idle proxies closing quiet streams. A dashboard that only changes during business hours goes silent at night, and a load balancer with a 60-second idle timeout closes it. Heartbeat comments every 15 seconds keep it open; see choosing heartbeat intervals.
  • Buffering proxies batching deltas. A proxy with buffering enabled delivers ten seconds of deltas at once, which looks like a frozen dashboard followed by a jump. Disable buffering on the route and send X-Accel-Buffering: no; the proxy and CDN configuration topic has per-proxy settings.
  • Background tabs. Browsers throttle timers and animation frames in hidden tabs, so the paint loop stops — which is exactly right. But the EventSource keeps receiving, and state keeps being updated. When the tab becomes visible, the first frame paints the latest state. If a hidden tab should not hold a connection at all, close it on visibilitychange and reconnect on return; the fresh snapshot makes that safe.
  • Clock-sensitive values. “Last updated 3 seconds ago” computed from server timestamps drifts if the client clock is wrong. Send server time in the heartbeat and compute an offset.
  • Wall screens and sleep. A kiosk that sleeps and wakes may keep a dead connection in OPEN state. Run a silence watchdog: if no byte, not even a heartbeat, arrives in twice the heartbeat interval, close and reopen.

Mitigation checklist:

Performance & Scale Considerations Permalink to this section

The cost of a dashboard stream is dominated by write fan-out: frames per second multiplied by viewers. There are three levers, and they multiply.

Frames written per second for 2,000 viewers of a 40-tile dashboard Bar chart showing write rate falling as three optimisations are applied: raw per-metric push, one delta per sampling tick, precision filtering, and a two-second tick for slow tiles. Frames written per second for 2,000 viewers of a 40-tile dashboard Push every metric change 80,000 / s One delta per 1 s tick 2,000 / s + precision filter 1,100 / s + 2 s tick for slow tiles 700 / s frames per second across the node
Coalescing to one delta per tick is the largest single saving. Precision filtering and per-tile cadence take it the rest of the way.
  1. Coalesce per tick. Emit one delta per sampling interval containing every changed field, not one frame per metric change. This is the difference between frame count scaling with metrics and scaling with time.
  2. Filter below display precision. A gauge showing whole percentages does not need a frame when the value moves from 38.21 to 38.24.
  3. Tier the cadence. Error rates and request rates earn a one-second tick; disk usage and host counts can tick every ten seconds. Two samplers with different intervals feeding the same stream is simple and effective.

Memory per viewer is small because the stream holds no per-viewer state beyond the socket and one listener. File descriptors are the first limit a busy dashboard node hits — which is a connection pooling problem, not a dashboard one. Snapshot plus delta streaming goes further into versioning and large views.

Validation & Debugging Permalink to this section

Validate the stream contract with curl before touching the interface. The first frame must be a snapshot, and deltas must arrive at the sampling cadence, not in bursts.

# The first event must be a snapshot; deltas should follow roughly once per second.
curl -sN -H 'Accept: text/event-stream' https://app.example.com/dash/stream \
  | awk '/^event:/ { printf "%s %s\n", strftime("%H:%M:%S"), $2; fflush() }' | head -20

# Reconnect with a stale cursor: expect a fresh snapshot, not a replay.
curl -sN -H 'Last-Event-ID: v1' https://app.example.com/dash/stream | head -4

Bursty timestamps in the first command mean something between the server and curl is buffering. In Chrome DevTools, the Network panel’s EventStream tab on the stream request shows each event with its type and arrival time; a column of identical times is the same symptom.

Log one line per connection lifecycle rather than per frame:

{"evt":"dash_stream_open","viewer":"u_1842","cursor":"v900","snapshot_bytes":1840}
{"evt":"dash_stream_close","viewer":"u_1842","duration_s":8412,"frames":8190,"skipped_slow":14,"reason":"client_gone"}

Because the client logic is a reducer, the most valuable automated test needs no network at all. Record a real stream once with curl, commit it as a fixture, and assert on the state it produces:

// dashboard.test.js — replay a recorded stream through the same parsing and reducer code.
import { readFileSync } from 'node:fs';
import { parseEventStream } from './parse.js';
import { dashboardReducer } from './reducer.js';

test('a delta after a reconnect snapshot never resurrects old values', () => {
  const frames = parseEventStream(readFileSync('fixtures/dash-reconnect.txt', 'utf8'));
  const state = frames.reduce(dashboardReducer, { version: null, values: {} });
  expect(state.version).toBe('v917');
  expect(state.values.rps).toBe(1302);          // from the post-reconnect snapshot
});

A non-zero skipped_slow is expected on mobile networks; a large one on office networks points at a proxy or a server that cannot keep up. The Performance panel’s frame chart confirms the client side: with the frame-rate gate in place there should be at most one paint per frame even when the stream is busy.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Why not send the whole dashboard on every update?

For a small view it is fine and simpler. Once the view exceeds a few kilobytes and updates every second, full frames multiply bandwidth by the number of viewers for no benefit. Snapshot on connect plus deltas afterwards keeps the correctness of full frames at the cost of deltas.

Do dashboards need Last-Event-ID replay?

No. A dashboard models current state, so a reconnecting client needs a fresh snapshot rather than the deltas it missed. The id is still useful as a version, letting the server skip the snapshot when the client is already current.

How often should a dashboard update?

No faster than people can read and no faster than the metric changes meaningfully. One second is a good default for operational metrics; business KPIs are usually fine at five to thirty seconds. Anything faster than the display frame rate is wasted work.

Can one stream serve dashboards with different tiles for different users?

Yes, if the projector filters per viewer. Keep one sampler producing the union of all metrics, and have each viewer's handler pick out the fields that viewer is allowed and has chosen to see before writing. Filtering in the browser instead would leak metrics to anyone who opens DevTools.

Should dashboard values be rounded on the server or the client?

On the server, to the precision the interface displays. Rounding first lets the precision filter suppress invisible changes, which is often the largest bandwidth saving available, and it keeps every viewer showing identical numbers.

What happens to a dashboard in a background tab?

The stream keeps delivering and state keeps updating, but animation frames stop, so nothing is painted. When the tab returns, the next frame paints the latest state. If background connections are too expensive, close on visibility change and rely on the snapshot when reopening.

Deep Dives