Collaborative Presence & Live Updates Permalink to this section

Part of Real-Time Application Patterns.

Collaboration features — avatars of who else is viewing a document, live cursors, “Ana is typing…”, comments and edits appearing as colleagues make them — are usually assumed to need WebSockets. Most of them do not. The traffic is lopsided: each user produces a small number of changes and consumes everyone else’s, so the natural split is ordinary HTTP requests for the upstream and a Server-Sent Events stream for the downstream. That split keeps the upstream on the same authenticated, rate-limited, observable request path as the rest of the API, and keeps the downstream on plain HTTP that every proxy understands. This guide covers the pieces: presence modelled as expiring leases rather than join and leave events, a per-document channel that carries edits, presence and ephemeral signals, ordering and conflict handling for concurrent edits, and the client that merges it all. It is written for teams adding collaboration to documents, boards, tickets or dashboards, not for teams writing a real-time game engine.

How It Works Permalink to this section

Each open document has a channel. Every client viewing the document holds one SSE stream subscribed to it and sends its own actions — edits, cursor moves, heartbeats — as POST requests. The server validates each action, applies it, and publishes the result to the channel, where every stream (including the sender’s) receives it.

HTTP up, SSE down Sequence diagram in which two clients hold SSE streams on a document channel; one posts an edit, the server applies and publishes it, and both streams receive the confirmed edit with a sequence number. HTTP up, SSE down Ana API Channel Ben POST /docs/42/ops {op, baseSeq 118} validate, apply, seq 119 publish op 119 event: op id: 119 event: op id: 119
The sender learns its edit was accepted from the same stream everyone else reads. That single ordered stream is what keeps every client's view consistent.

The channel carries three kinds of messages, and they have different durability requirements:

Event Example Durable Has id On reconnect
op text inserted, card moved, comment added yes, stored sequence replay after cursor
presence roster of viewers, their colours and status no, derived none send current roster
signal cursor position, selection, typing no, ephemeral none drop, next one follows

Only op events carry an id, so only they advance the browser’s Last-Event-ID. A reconnect replays missed operations from the document’s operation log, then sends the current roster; cursors and typing signals from the gap are simply gone, which is correct — a cursor position from forty seconds ago is noise. Pairing SSE with POST for collaborative edits covers the operation path in detail.

Presence is a set of leases. A viewer is present while their lease is fresh; the stream renews it. A viewer who closes the laptop stops renewing and disappears when the lease expires — there is no reliance on a leave message that a dead connection can never send. Showing who is online with SSE presence builds the lease store.

Server-Side Implementation Permalink to this section

The stream handler joins the document, replays operations, sends the roster, then relays live messages, renewing the viewer’s presence lease on each heartbeat.

// doc-stream.js
app.get('/api/docs/:id/stream', requireDocAccess, async (req, res) => {
  const docId = req.params.id;
  const viewer = { user: req.user.id, name: req.user.name, conn: crypto.randomUUID() };
  openStream(res, { retryMs: 2000 });

  const sub = await bus.subscribe(`doc:${docId}`, (raw) => {
    const m = JSON.parse(raw);
    if (m.type === 'op') {
      if (m.seq <= cursor) return;
      cursor = m.seq;
      res.write(`event: op\nid: ${m.seq}\ndata: ${raw}\n\n`);
    } else if (m.type === 'signal') {
      if (m.conn === viewer.conn) return;                 // never echo your own cursor
      if (!res.writableNeedDrain) res.write(`event: signal\ndata: ${raw}\n\n`);   // droppable
    } else {
      res.write(`event: ${m.type}\ndata: ${raw}\n\n`);
    }
  });

  // Replay the operation log after the client's cursor.
  let cursor = Number(req.get('Last-Event-ID') ?? req.query.after ?? 0);
  for (const op of await ops.after(docId, cursor, 1000)) {
    cursor = op.seq;
    res.write(`event: op\nid: ${op.seq}\ndata: ${JSON.stringify(op)}\n\n`);
  }

  // Presence: take a lease, announce the roster, renew on every heartbeat.
  await presence.renew(docId, viewer);
  res.write(`event: presence\ndata: ${JSON.stringify(await presence.roster(docId))}\n\n`);
  await bus.publish(`doc:${docId}`, JSON.stringify({ type: 'presence', roster: await presence.roster(docId) }));

  const beat = setInterval(async () => {
    res.write(': hb\n\n');
    await presence.renew(docId, viewer);                  // lease survives only while we beat
  }, 10_000);

  req.on('close', async () => {
    clearInterval(beat);
    sub.unsubscribe();
    await presence.release(docId, viewer);                // fast path; expiry is the backstop
  });
});

Signals are the only messages the handler drops under backpressure: if a viewer’s socket is backed up, another cursor update is worthless and the next one will supersede it. Operations are never dropped; a viewer who cannot keep up with operations is disconnected and replays on reconnect.

The upstream path is ordinary request handling. Cursor updates are high-frequency, so they are rate-limited and never stored:

app.post('/api/docs/:id/signal', requireDocAccess, rateLimit({ perSecond: 20 }), async (req, res) => {
  const { kind, data } = req.body;                        // kind: 'cursor' | 'selection' | 'typing'
  if (!['cursor', 'selection', 'typing'].includes(kind)) return res.status(400).end();
  await bus.publish(`doc:${req.params.id}`, JSON.stringify({
    type: 'signal', kind, data, user: req.user.id, conn: req.get('X-Conn-Id'),
  }));
  res.status(204).end();
});

For Python services the structure is identical; with FastAPI, the heartbeat renewal becomes an asyncio task started with the generator and cancelled in its finally block, as in the FastAPI SSE implementation guide.

Ordering is assigned in one place Permalink to this section

Every operation receives its sequence number from a single authority per document — a database sequence, a row lock on the document, or a single-writer actor such as a Durable Object — before it is published. Clients never assign order. This is what makes the SSE id a valid replay cursor and what lets every client apply operations in the same order.

-- Allocate the next sequence and store the op atomically.
WITH next AS (
  UPDATE documents SET last_seq = last_seq + 1 WHERE id = $1 RETURNING last_seq
)
INSERT INTO doc_ops (doc_id, seq, user_id, base_seq, op)
SELECT $1, last_seq, $2, $3, $4 FROM next
RETURNING seq;

When the document already lives in an edge runtime, a single-writer object per document gives the same guarantee without a database round trip; see fanning out SSE with Durable Objects.

Client-Side Consumption Permalink to this section

The client keeps three stores fed by one stream: the document (from op events), the roster (from presence), and ephemeral overlays (from signal).

// collab-client.js
export function joinDocument(docId, { applyOp, setRoster, setCursor, setTyping }) {
  const connId = crypto.randomUUID();
  const es = new EventSource(`/api/docs/${docId}/stream`, { withCredentials: true });

  es.addEventListener('op', (e) => applyOp(JSON.parse(e.data)));          // ordered, durable
  es.addEventListener('presence', (e) => setRoster(JSON.parse(e.data).roster ?? JSON.parse(e.data)));
  es.addEventListener('signal', (e) => {
    const s = JSON.parse(e.data);
    if (s.kind === 'cursor') setCursor(s.user, s.data);
    if (s.kind === 'typing') setTyping(s.user, s.data.until);
  });

  // Upstream: throttle cursor updates to ~10 per second.
  let pendingCursor = null, timer = null;
  function sendCursor(pos) {
    pendingCursor = pos;
    timer ??= setTimeout(() => {
      timer = null;
      fetch(`/api/docs/${docId}/signal`, {
        method: 'POST', credentials: 'include', keepalive: true,
        headers: { 'Content-Type': 'application/json', 'X-Conn-Id': connId },
        body: JSON.stringify({ kind: 'cursor', data: pendingCursor }),
      });
    }, 100);
  }
  return { sendCursor, close: () => es.close() };
}

Remote cursors should fade after a few seconds without an update, which also cleans up cursors from users whose connection dropped without a presence change yet. Typing indicators with SSE applies the same “expire unless renewed” idea to the typing signal.

Three client stores, three lifetimes Three panels describing the document store, the roster store and the overlay store, with what feeds each and how long its data lives. Three client stores, three lifetimes Document fed by op events ordered by seq replayed on reconnect lives as long as the doc Roster fed by presence replaced wholesale resent on reconnect lives while leases do Overlays fed by signal latest per user wins never replayed fades after seconds
Mixing these lifetimes is the most common source of collaboration bugs — a cursor stored like an edit, or an edit dropped like a cursor.

Edge Cases & Network Interference Permalink to this section

  • Ghost viewers. A browser that crashes, or a laptop that sleeps, never sends a close. Without leases the avatar stays forever; with a 30-second lease renewed every 10 seconds, it disappears within half a minute.
  • Duplicate self. A user with two tabs is two connections. Show one avatar per user and count connections in the roster, or the user sees themselves twice.
  • Echo of own signals. A client that renders its own cursor from the stream shows a lagging second cursor. Tag signals with a connection id and filter them out.
  • Upstream blocked by the connection limit. On HTTP/1.1, a page with several open streams can exhaust the six-connection limit, and cursor POSTs queue behind them. Serve over HTTP/2, or share one stream across tabs; see diagnosing the six-connection limit.
  • Revoked access. Removing a user from a document must close their stream promptly, not at their next reconnect. Publish a revoke on the user’s control channel, and have the node end that user’s streams for the document with a final revoked event so the client can show why the view stopped updating.

Mitigation checklist:

Performance & Scale Considerations Permalink to this section

Collaboration traffic is dominated by signals. A document with twenty active editors, each sending ten cursor updates per second, produces 200 upstream requests per second and fans each of them out to nineteen streams — 3,800 frames per second, for one document.

Frames per second for one document with 20 active editors Bar chart comparing downstream frame rates for unthrottled cursor updates, client-throttled updates at 10 per second, and server-coalesced updates at 5 per second per user. Frames per second for one document with 20 active editors Every mousemove (~60 Hz) 22,800 / s Client throttle 10 Hz 3,800 / s + server coalesce 5 Hz 1,900 / s downstream frames per second across all viewers
Signals, not edits, set the cost of a busy document. Throttle on the client and coalesce on the server before fan-out.

The levers, in order of impact:

  1. Throttle signals on the client to ten per second at most; nobody can follow a remote cursor faster.
  2. Coalesce on the server: keep the latest signal per user per document and publish on a timer, so bursts collapse into one frame.
  3. Drop signals for backed-up viewers rather than queueing them.
  4. Scale by document: route all streams for a document to the same node when possible, so fan-out happens in memory rather than through the broker.

Operations are comparatively rare and cheap, but the operation log grows forever. Snapshot documents periodically and replay only operations after the snapshot; a reconnect with a cursor older than the latest snapshot receives the snapshot and the operations after it.

Bounding reconnect replay with periodic snapshots Timeline of a document's operation sequence with snapshots every thousand operations, and a reconnecting client whose cursor predates the latest snapshot. Bounding reconnect replay with periodic snapshots Snapshots Replay sent snapshot 1 snapshot 2 snapshot 3 +500 ops 0 700 1400 2100 2800 3500 operation sequence client cursor latest snapshot
A client more than one snapshot behind gets the snapshot plus a short tail of operations, never the document's whole history.
// On connect: decide between plain replay and snapshot + tail.
const snap = await snapshots.latest(docId);                  // { seq, state }
if (cursor < snap.seq) {
  res.write(`event: snapshot\nid: ${snap.seq}\ndata: ${JSON.stringify(snap.state)}\n\n`);
  cursor = snap.seq;
}
for (const op of await ops.after(docId, cursor, 1000)) { /* …as before… */ }

// A background job snapshots any document with more than 1,000 ops since its last snapshot.
async function compact(docId) {
  const snap = await snapshots.latest(docId);
  const tail = await ops.after(docId, snap.seq, 5000);
  if (tail.length < 1000) return;
  const state = tail.reduce(applyOp, snap.state);
  await snapshots.save(docId, { seq: tail[tail.length - 1].seq, state });
}

The snapshot event carries an id like an operation does, so the browser’s cursor jumps to the snapshot’s sequence and any later reconnect replays only operations after it. The client treats snapshot as “replace the document”, which it already has to support for first load.

Conflicts are resolved on the server, displayed on the client Permalink to this section

When two users edit the same element at once, the server sees both operations with the same baseSeq. It must either transform the later one against the earlier (operational transformation), merge them (CRDTs do this by construction), or reject it with a 409 so the client can rebase and retry. For structured data — form fields, card positions, ticket status — last-writer-wins per field with a 409 on stale baseSeq is simple and usually enough:

app.post('/api/docs/:id/ops', requireDocAccess, async (req, res) => {
  const { opId, baseSeq, field, value } = req.body;
  const current = await fields.version(req.params.id, field);     // seq of the field's last change
  if (current > baseSeq) return res.status(409).json({ field, current });   // someone got there first
  const seq = await ops.append(req.params.id, { opId, user: req.user.id, field, value, baseSeq });
  res.status(202).json({ seq });
});

The client that receives a 409 has already seen the winning operation arrive on the stream, so it can show “Ben changed this a moment ago” and let the user decide, rather than silently overwriting.

Validation & Debugging Permalink to this section

# Two viewers: open both streams, post a cursor from one, see it only on the other.
curl -sN -b ana.txt https://app.example.com/api/docs/42/stream -H 'X-Conn-Id: a1' > ana.log &
curl -sN -b ben.txt https://app.example.com/api/docs/42/stream > ben.log &
curl -s -b ana.txt -X POST -H 'Content-Type: application/json' -H 'X-Conn-Id: a1' \
  -d '{"kind":"cursor","data":{"line":12,"ch":4}}' https://app.example.com/api/docs/42/signal
grep -c signal ana.log ben.log     # ana.log: 0, ben.log: 1

# Ghost cleanup: kill a viewer without closing, wait past the lease, check the roster.
kill -9 %2; sleep 35
curl -s -b ana.txt https://app.example.com/api/docs/42/presence | jq '.[].user'

Log per document: connected viewers, operations per minute, signals in versus signals out. A large gap between signals in and out is the coalescing working; operations per minute that approach the ordering authority’s throughput is the signal to shard documents across writers.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Do collaborative features need WebSockets?

Not usually. Each client's upstream is modest, and ordinary POST requests handle it with the API's existing authentication and rate limits. WebSockets become worthwhile when upstream rates are high and latency-critical, such as real-time drawing or games.

How do I stop the sender seeing its own edit twice?

Apply edits locally as pending, then treat the stream's copy — identified by the client-supplied operation id — as the confirmation that assigns its sequence number. The pending entry is replaced rather than duplicated.

How many viewers can one document stream support?

Hundreds comfortably, as long as signals are throttled and coalesced. Beyond that — an all-hands document or a live event page — hide individual cursors, show an aggregate viewer count instead of every avatar, and keep only operations on the stream.

What happens to an edit posted while the stream is reconnecting?

It is accepted normally, because the upstream is an independent HTTP request. When the stream reconnects it replays operations after its cursor, which includes the confirmed edit, so the pending local copy is resolved as usual.

Should cursor positions be stored?

No. They are only meaningful for a few seconds. Publish them, never persist them, and never replay them after a reconnect.

Can presence work across several documents at once?

Yes. Keep leases keyed by document and publish roster changes on each document's channel. A workspace-wide "who is online" view reads a separate user-level lease set, renewed by whichever stream the user currently has open, so one heartbeat covers both.

How do CRDT libraries fit with SSE?

Well. CRDT updates are opaque binary or JSON blobs that can be posted upstream and fanned out downstream as base64 or JSON in op events. The server still assigns a sequence for replay, even though the CRDT itself does not need ordering.

Deep Dives