Testing & Load Testing SSE Endpoints Permalink to this section
Part of Backend Stream Generation & Connection Management.
Streaming endpoints are under-tested in most codebases for a simple reason: the usual tools assume a request ends. Unit test frameworks want a return value, HTTP test clients wait for the body to complete, and load generators count requests per second. A Server-Sent Events endpoint never completes, its correctness lives in sequences of frames and in what happens across reconnects, and its capacity is measured in concurrent connections and delivery latency. This guide lays out a testing strategy that fits: pure unit tests for the frame encoder, integration tests that read a stream incrementally and exercise Last-Event-ID, browser end-to-end tests that drive a real EventSource, chaos tests that put real proxies and network faults in the path, and load tests that hold tens of thousands of connections open while measuring per-event latency. Each level catches a different class of bug, and together they cover the failures this site’s other pages describe.
How It Works Permalink to this section
Think of the tests as a pyramid, with a twist at the top: the most expensive layer is not the browser test but the load test, because it needs its own infrastructure and runs for minutes.
The key technique at every level above unit tests is incremental reading: consume the response body as it arrives, parse frames with a spec-compliant parser, and assert on the sequence — then cancel the request. A test that waits for the body to end will hang forever.
// Read the first N events from a live stream, then abort. Works in Node 18+ (fetch + web streams).
export async function readEvents(url, n, { headers = {}, timeoutMs = 5000 } = {}) {
const ac = new AbortController();
const timer = setTimeout(() => ac.abort(), timeoutMs);
const res = await fetch(url, { headers: { Accept: 'text/event-stream', ...headers }, signal: ac.signal });
const events = [];
const parser = createParser((evt) => events.push(evt)); // spec-compliant parser
const reader = res.body.pipeThrough(new TextDecoderStream()).getReader();
try {
while (events.length < n) {
const { value, done } = await reader.read();
if (done) break;
parser.feed(value);
}
} finally {
clearTimeout(timer);
ac.abort(); // stop the stream
}
return { status: res.status, headers: res.headers, events };
}
The parser should implement the specification’s algorithm — the same one used by the eventsource-parser package — rather than a regex, because tests are where you want parsing that is exactly as strict as a browser’s. The event stream parsing guide walks through the algorithm.
Server-Side Implementation Permalink to this section
Testability is mostly a server design question. Three seams make SSE code easy to test:
- A pure frame encoder.
formatEvent({ id, event, data, retry })returns a string. All the tricky rules — splitting data on newlines, rejecting newlines in ids and event names, handling\r— are tested here without any I/O. See unit testing event stream serialization. - An injectable event source. The handler reads events from an interface (a hub, a broker client) that tests replace with an in-memory implementation they can publish into on demand.
- An injectable clock. Heartbeat intervals, retry hints and timeouts use a clock the tests can advance, so a 15-second heartbeat test runs in milliseconds.
// server.js — dependencies passed in, so tests can control them.
export function createApp({ hub, clock = realClock, heartbeatMs = 15000 }) {
const app = express();
app.get('/events', (req, res) => {
res.writeHead(200, { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache' });
const after = req.get('Last-Event-ID');
for (const e of hub.replay(after)) res.write(formatEvent(e));
const off = hub.subscribe((e) => res.write(formatEvent(e)));
const hb = clock.setInterval(() => res.write(': hb\n\n'), heartbeatMs);
req.on('close', () => { off(); clock.clearInterval(hb); });
});
return app;
}
An integration test then starts the app on an ephemeral port with an in-memory hub:
import { test, expect } from 'vitest';
test('replays events after Last-Event-ID, then streams live', async () => {
const hub = createMemoryHub();
hub.publish({ id: '1', event: 'order', data: '{"n":1}' });
hub.publish({ id: '2', event: 'order', data: '{"n":2}' });
const server = createApp({ hub }).listen(0);
const url = `http://127.0.0.1:${server.address().port}/events`;
const pending = readEvents(url, 2, { headers: { 'Last-Event-ID': '1' } });
setTimeout(() => hub.publish({ id: '3', event: 'order', data: '{"n":3}' }), 50);
const { status, headers, events } = await pending;
expect(status).toBe(200);
expect(headers.get('content-type')).toMatch(/^text\/event-stream/);
expect(events.map((e) => e.id)).toEqual(['2', '3']); // replay then live, no gap, no dup
server.close();
});
This single test covers the replay seam, which is where most resumable-stream bugs live.
The same test in Python uses httpx’s streaming API against the app served by uvicorn on an ephemeral port. A real server matters here: in-process test transports may collect the whole response body before returning, which never happens for an infinite stream, so the test would hang. The live_url fixture starts uvicorn in a background task and yields its address:
# test_stream.py — pytest + httpx, reading events incrementally.
import asyncio, httpx, pytest
async def read_events(client, url, n, headers=None, timeout=5.0):
events, cur = [], {}
async with client.stream("GET", url, headers=headers or {}) as res:
assert res.status_code == 200
assert res.headers["content-type"].startswith("text/event-stream")
async def consume():
async for line in res.aiter_lines():
if line == "":
if "data" in cur: events.append(dict(cur))
cur.clear()
if len(events) >= n: return
elif not line.startswith(":"):
field, _, value = line.partition(":")
cur[field] = value[1:] if value.startswith(" ") else value
await asyncio.wait_for(consume(), timeout) # fail, don't hang
return events
@pytest.mark.anyio
async def test_replay_then_live(live_url, hub):
hub.publish(id="1", event="order", data='{"n":1}')
hub.publish(id="2", event="order", data='{"n":2}')
async with httpx.AsyncClient(base_url=live_url, timeout=None) as c:
task = asyncio.create_task(read_events(c, "/events", 2, {"Last-Event-ID": "1"}))
await asyncio.sleep(0.05)
hub.publish(id="3", event="order", data='{"n":3}')
events = await task
assert [e["id"] for e in events] == ["2", "3"]
The inline parser is deliberately minimal; for anything beyond a test helper, use a library parser so tests and production agree on edge cases. Testing FastAPI SSE endpoints with httpx covers the disconnect and lifespan details specific to FastAPI.
Contract tests between producer and consumer teams Permalink to this section
When one team owns the stream and others consume it — a mobile team, a partner integration, another backend — the stream’s event names, id format and payload shapes are an API contract. Treat it like one. Record a canonical sample stream as a fixture in the producer’s repository, validate every data: payload in it against a JSON Schema per event type, and run the consumer’s parser and reducer against the same fixture in the consumer’s CI. A change that renames an event type or drops a field then fails in the producer’s pipeline before any client breaks.
# stream-contract.yaml — versioned alongside the endpoint
endpoint: /api/orders/stream
id: "monotonic integer, per user; resumable via Last-Event-ID for 72h"
retry_ms: 5000
events:
order.updated: { schema: schemas/order-updated.json }
order.deleted: { schema: schemas/order-deleted.json }
resync: { schema: schemas/resync.json, meaning: "refetch /api/orders" }
heartbeat: "comment line every 15s"
The contract is also the right place to write down the promises that no schema captures: how long ids remain resumable, what resync means, and that unknown event types must be ignored rather than treated as errors — the rule that lets producers add event types without breaking old clients.
Client-Side Consumption Permalink to this section
Client code deserves the same treatment. The reconnect logic, the reducer that folds events into state, and the UI’s connection indicator are all testable without a server:
- Reducers are pure: feed recorded frames, assert on state.
- Reconnect logic with a fake
EventSourceclass injected into the module: emitopen,errorand messages on demand, advance fake timers, assert on backoff and on theLast-Event-IDused. - Real browser behaviour — the actual
EventSourcereconnecting after a server-side close, named event dispatch,withCredentials— needs a browser. End-to-end testing SSE with Playwright shows how to drive those scenarios deterministically.
// A minimal fake for unit tests of client modules.
class FakeEventSource {
static last;
constructor(url) { this.url = url; this.readyState = 0; this.listeners = {}; FakeEventSource.last = this; }
addEventListener(t, fn) { (this.listeners[t] ??= []).push(fn); }
emit(t, data, lastEventId = '') { (this.listeners[t] ?? []).forEach((fn) => fn({ data, lastEventId })); }
close() { this.readyState = 2; }
}
Edge Cases & Network Interference Permalink to this section
The bugs that reach production are rarely in your code alone; they come from the path between the server and the browser. A test environment without that path cannot find them. Put the real reverse proxy (nginx, Envoy, the cloud load balancer’s local equivalent) in front of the test server, and add a fault-injection proxy such as Toxiproxy to simulate latency, bandwidth limits, idle timeouts and resets.
Chaos testing SSE through proxies builds that harness with Docker Compose. Checklist for the harness:
Performance & Scale Considerations Permalink to this section
Load testing an SSE service answers three questions, none of which is “requests per second”: how many concurrent connections can a node hold; how long does an event take from publish to every client at that concurrency; and what happens to healthy clients when some clients are slow.
Most HTTP load tools cannot hold streaming responses or observe individual events. k6 with the xk6-sse extension can, and so can small purpose-built clients in Go. Measure latency by having the publisher stamp a timestamp in each event and the load client compute now - stamp on receipt, with clocks synchronised (or both running on the same host). Load testing SSE with k6 has the scripts.
A realistic load profile has at least four ingredients, and leaving any out makes the result optimistic:
- A connection mix. Most real clients are fast, but a few percent are on congested mobile links. Include a cohort of clients that read slowly, so the test measures whether slow clients harm fast ones.
- Reconnect churn. Real clients disconnect and reconnect constantly — tab switches, network changes, laptop lids. A test in which connections are opened once and held forever never exercises replay, authentication at connect time, or registry cleanup.
- A deploy in the middle. Restart one node during the test and measure the reconnect wave: how long until every client is back, and what the replay load does to the remaining nodes.
- Realistic payloads. Serialisation and compression cost scale with payload size; a test with three-byte payloads measures framing overhead, not your application.
Two practical notes. The load generator itself needs raised file descriptor limits and often several source IP addresses, because one IP can open at most around 28,000 connections to one destination port with default ephemeral port ranges. And ramp slowly: opening 50,000 connections in one second tests the accept queue and TLS handshake capacity, which is a different question from steady-state capacity.
Validation & Debugging Permalink to this section
The test suite itself needs validation: a test that passes because the stream never started is worse than no test. Guard against it by asserting on status and content type before events, by bounding every stream read with a timeout that fails rather than returns empty, and by including at least one negative test (an unauthorised request must get 401, not an open stream).
# Quick manual check that mirrors the integration test.
curl -sN -H 'Last-Event-ID: 1' http://localhost:3000/events | head -6
# Confirm tests abort their streams: no lingering connections after the suite.
ss -tnp state established '( sport = :3000 )' | wc -l
In CI, give streaming tests their own job with a generous but finite timeout, run them in parallel only if each uses its own ephemeral port and in-memory hub, and fail the job if any test process leaves connections open at exit. Keep the chaos and load suites out of the per-commit pipeline: run chaos tests when proxy configuration or connection-handling code changes, and load tests nightly or before releases, recording the latency-knee number so a regression in capacity is visible as a trend rather than discovered during an incident.
When a streaming test is flaky, the cause is almost always timing: publishing before the subscription exists, or asserting before a flush. Publish only after the response headers arrive, and prefer waiting for a specific event over sleeping.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Why do my SSE tests hang?
They wait for the response body to finish, and an SSE body never finishes. Read the stream incrementally, stop after the events you need, and abort the request with a timeout that fails the test.
Can supertest or similar libraries test SSE?
Only with care: most buffer the whole response by default. Use a streaming fetch against a server on an ephemeral port, or the library's streaming mode, and parse frames as they arrive.
What is the single most valuable SSE test?
The replay seam: connect with a Last-Event-ID while a new event is published, and assert that the client receives exactly the missed events followed by the new one — no gap and no duplicate.
How many connections should a load test reach?
At least your expected peak per node plus a safety margin, and past the latency knee at least once so you know where it is. Test with a realistic mix of fast and slow clients, not only ideal ones.
How do I test heartbeats without waiting 15 seconds?
Inject the clock or the heartbeat interval. With a fake clock, advance time and assert that a comment line was written; with a configurable interval, set it to a few milliseconds in tests. Keep one slow, real-time test in the chaos suite to prove the production interval survives the proxy's idle timeout.
Do I need to test through the real proxy?
Yes. Buffering, idle timeouts and header stripping are properties of the proxy configuration, and they are the most common cause of SSE incidents. A test without the proxy cannot catch them.