Unit Testing Event Stream Serialization Permalink to this section

Part of Testing & Load Testing SSE Endpoints, under Backend Stream Generation & Connection Management.

Every SSE server has a function that turns an event into bytes, usually a template string written in five minutes: `id: ${id}\nevent: ${type}\ndata: ${data}\n\n`. It works for every payload the author thought of and silently corrupts the stream for the ones they did not — a newline in a message body, a carriage return from a Windows client, an event name containing a newline supplied by user input. Because the encoder is a pure function, it is the cheapest part of the system to test exhaustively. This guide writes the encoder properly and tests it with examples and with round-trip property tests against a spec-compliant parser.

Symptom & Developer Intent Permalink to this section

Encoder bugs look like data bugs downstream:

  • A message containing a line break arrives in the browser truncated at the break.
  • Some events never reach their listener; others arrive with the wrong event type.
  • A user-controlled field makes the client receive events that the server never sent.
  • Occasionally the whole stream stops dispatching until the next reconnect.
  • Non-ASCII text arrives garbled on some clients.

The intent is an encoder whose output, parsed by the specification’s algorithm, always yields exactly the event that was encoded — for every possible input — and a test suite that proves it.

Root Cause Analysis Permalink to this section

The event stream format is line-oriented. A line ends at LF, CR, or CRLF; a blank line dispatches the event. Any of those characters inside a field value therefore ends the field early. In data the specification provides the escape hatch: split on line breaks and emit one data: line per piece, and the parser rejoins them with LF. In id and event there is no escape — the value must not contain line breaks at all.

Anatomy of a correctly encoded multi-line event Layers of an encoded event: an id line, an event line, two data lines produced by splitting a two-line payload, and the blank dispatch line. Anatomy of a correctly encoded multi-line event id: 42 no CR or LF allowed event: note no CR or LF allowed data: first line payload line 1 data: second line payload line 2 (blank line) dispatch
Every line break in the payload becomes a new data: line. The parser joins them back with a single LF, so the client sees the original text exactly.

Injection is the security form of the same bug. If an event name or id comes from user input and contains \n\ndata: {"admin":true}\n\n, the client parses an extra event the server never meant to send. Stream-stopping bugs come from \r alone: a CR is a line terminator, so a payload ending in CR followed by the encoder’s own LF produces an unexpected blank line, and data can be dispatched early.

Step-by-Step Resolution Permalink to this section

Step 1 — Write a strict encoder Permalink to this section

// sse-format.js
const LINE_BREAK = /\r\n|\r|\n/;

export function formatEvent({ id, event, data, retry, comment } = {}) {
  let out = '';
  if (comment !== undefined) {
    for (const line of String(comment).split(LINE_BREAK)) out += `: ${line}\n`;
  }
  if (id !== undefined) {
    const v = String(id);
    if (LINE_BREAK.test(v) || v.includes('\0')) throw new TypeError('SSE id must not contain CR, LF or NUL');
    out += `id: ${v}\n`;
  }
  if (event !== undefined) {
    const v = String(event);
    if (LINE_BREAK.test(v)) throw new TypeError('SSE event name must not contain CR or LF');
    out += `event: ${v}\n`;
  }
  if (retry !== undefined) {
    if (!Number.isInteger(retry) || retry < 0) throw new TypeError('SSE retry must be a non-negative integer');
    out += `retry: ${retry}\n`;
  }
  if (data !== undefined) {
    for (const line of String(data).split(LINE_BREAK)) out += `data: ${line}\n`;
  }
  return out + '\n';
}

The NUL check on ids follows the specification, which tells parsers to ignore an id field containing NUL. Throwing on invalid ids and names turns an injection into a loud error in your code instead of a silent one in every client.

Step 2 — Test the rules with examples Permalink to this section

import { test, expect } from 'vitest';
import { formatEvent } from './sse-format.js';

test('single-line data', () => {
  expect(formatEvent({ data: 'hi' })).toBe('data: hi\n\n');
});

test('each line break in data becomes a data line', () => {
  expect(formatEvent({ data: 'a\nb\r\nc\rd' })).toBe('data: a\ndata: b\ndata: c\ndata: d\n\n');
});

test('empty data still dispatches', () => {
  expect(formatEvent({ data: '' })).toBe('data: \n\n');
});

test('ids and event names cannot smuggle lines', () => {
  expect(() => formatEvent({ id: '1\n\ndata: x', data: 'ok' })).toThrow();
  expect(() => formatEvent({ event: 'a\rb', data: 'ok' })).toThrow();
});

test('retry must be an integer', () => {
  expect(() => formatEvent({ retry: 1.5 })).toThrow();
  expect(formatEvent({ retry: 3000 })).toBe('retry: 3000\n\n');
});

test('UTF-8 passes through unchanged', () => {
  expect(formatEvent({ data: 'naïve — 日本 ✓' })).toBe('data: naïve — 日本 ✓\n\n');
});

Step 3 — Round-trip with a spec-compliant parser Permalink to this section

Examples only cover the cases you think of. A property test generates thousands of arbitrary events, encodes them, parses the bytes with a parser that implements the specification (such as eventsource-parser), and checks that the result equals the input.

import fc from 'fast-check';
import { createParser } from 'eventsource-parser';

const noBreaks = fc.string().filter((s) => !/[\r\n\0]/.test(s));

test('encode → parse is the identity', () => {
  fc.assert(fc.property(
    fc.record({
      id: fc.option(noBreaks, { nil: undefined }),
      event: fc.option(noBreaks.filter((s) => s.length > 0), { nil: undefined }),
      data: fc.string(),
    }),
    (evt) => {
      const got = [];
      const p = createParser({ onEvent: (e) => got.push(e) });
      p.feed(formatEvent(evt));
      expect(got).toHaveLength(1);
      // The parser normalises CR and CRLF inside data to LF.
      expect(got[0].data).toBe(evt.data.replace(/\r\n|\r/g, '\n'));
      expect(got[0].event ?? undefined).toBe(evt.event);
      if (evt.id !== undefined) expect(got[0].id).toBe(evt.id);
    },
  ), { numRuns: 5000 });
});

Property tests find surprising cases quickly: leading spaces in data (the parser strips exactly one space after the colon, which the encoder’s ": " accounts for), lone CRs, and empty event names.

Encoder bugs found by each kind of test Bar chart comparing the number of distinct encoder bugs found in an audit of several real encoders by example tests alone and by adding round-trip property tests. Encoder bugs found by each kind of test Example tests only 5 + round-trip property tests 14 distinct bugs found across six audited encoders
Examples catch the obvious newline bug. Property tests catch lone CR, leading-space and empty-name cases that example suites rarely include.

Step 4 — Test chunk boundaries on the parsing side Permalink to this section

Streams arrive in arbitrary chunks. If you also maintain a client-side parser, feed the same encoded stream split at every possible byte offset and assert that the parsed events are identical — including splits in the middle of a multi-byte UTF-8 character and between the CR and LF of a CRLF.

test('parsing is independent of chunk boundaries', () => {
  const stream = formatEvent({ id: '1', data: 'héllo\r\nworld' }) + formatEvent({ data: '✓' });
  const bytes = new TextEncoder().encode(stream);
  const whole = parseAll([bytes]);
  for (let i = 1; i < bytes.length; i++) {
    expect(parseAll([bytes.slice(0, i), bytes.slice(i)])).toEqual(whole);
  }
});

parseAll must decode with a streaming TextDecoder ({ stream: true }) so split characters are reassembled; the fetch-based stream parsing guide shows the pattern.

Validation & Monitoring Permalink to this section

Run the property test in CI with a fixed seed for reproducibility and a nightly job with random seeds and more runs. Make any failing input a permanent example test. In production, count encoder exceptions: a non-zero rate means some code path is trying to put user input into an id or event name, which is worth investigating as a potential injection attempt rather than just a bug.

npx vitest run sse-format        # example + property tests, seconds
FC_NUM_RUNS=100000 npx vitest run sse-format --reporter=verbose   # nightly
Encoder rules at a glance Three panels listing the rules for data fields, for id and event fields, and for retry and comment fields. Encoder rules at a glance data split on CRLF, CR, LF one data: line per piece empty string is valid id / event reject CR and LF id: reject NUL never from raw user input retry / comment retry: integer ≥ 0 comments split like data blank line ends the event
Three short lists cover the whole specification as it applies to a server. Every rule has a test in the suite above.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Is JSON.stringify output always safe as data?

Yes, as long as indentation is off: JSON escapes newlines inside strings, so the output is a single line. Pretty-printed JSON contains line breaks and must go through the multi-line data rule.

Should the encoder escape instead of throwing for bad ids?

There is no escape syntax for id or event fields. Throwing is safer than silently altering an id, because a changed id breaks resume in ways that are hard to debug.

Do I need to handle the byte order mark?

Only when parsing. A parser must ignore a leading UTF-8 BOM at the start of the stream; encoders should simply never emit one.

Which parser should property tests use?

One that implements the WHATWG algorithm faithfully, such as eventsource-parser in JavaScript. Using your own parser for both sides only proves the two agree with each other.