Character Encoding and BOM Handling in SSE Permalink to this section

Part of Understanding the Event Stream Format, under SSE Protocol Fundamentals & Architecture.

Server-Sent Events settle the encoding question once: the stream is UTF-8, always. That removes a whole class of charset negotiation bugs, and introduces a few of its own. Servers that emit Latin-1 or Windows-1252 text produce garbled characters. A byte order mark at the wrong place corrupts the first field name. Custom parsers that decode each network chunk separately mangle any multi-byte character that happens to straddle a chunk boundary. This guide explains the rules and fixes each failure.

Symptom & Developer Intent Permalink to this section

  • Accented letters appear as é or � in the browser.
  • The very first event of a stream is never dispatched, or its first field is ignored.
  • A custom fetch-based parser occasionally shows � in the middle of otherwise correct text, at random positions.
  • Emoji or CJK text displays correctly in EventSource but not in a Node.js or Python consumer.
  • Setting charset=iso-8859-1 in the Content-Type has no effect.

The intent is a stream that carries any Unicode text correctly to every client — browsers and custom parsers alike.

Root Cause Analysis Permalink to this section

The specification tells user agents to decode the stream with UTF-8 decode, which ignores any charset parameter on the Content-Type and strips a single leading byte order mark (U+FEFF, bytes EF BB BF) at the very start of the stream. Consequences:

How an event stream is decoded Layers of stream decoding from raw bytes through removal of a leading BOM, UTF-8 decoding with replacement of invalid sequences, line splitting and field parsing. How an event stream is decoded Raw bytes from the network, any chunking Leading BOM stripped once, at stream start UTF-8 decode invalid bytes → U+FFFD Line splitting CRLF, CR, LF Field parsing name:value
Only UTF-8 is ever used, whatever the Content-Type says. A BOM is tolerated only as the first bytes of the whole stream.
  • Text encoded in anything other than UTF-8 is misinterpreted. Latin-1 é (0xE9) is not valid UTF-8 on its own and becomes U+FFFD; double-encoded UTF-8 shows as é.
  • A BOM anywhere other than the start of the stream is just a character. If a server writes a BOM at the start of each event, the second and later events begin with U+FEFF, so their first field name is data rather than data, and the field is ignored.
  • A custom parser that decodes each chunk with a fresh, non-streaming decoder splits a multi-byte character when the chunk boundary falls inside it, producing replacement characters.

Step-by-Step Resolution Permalink to this section

Step 1 — Emit UTF-8, and say so Permalink to this section

res.writeHead(200, { 'Content-Type': 'text/event-stream; charset=utf-8' });
res.write(`data: ${JSON.stringify({ name: 'Zoë', city: '東京' })}\n\n`);   // Node writes strings as UTF-8

The charset=utf-8 parameter is ignored by EventSource but helps proxies, debugging tools and non-browser clients that do look at it. In languages where the default encoding is platform-dependent, encode explicitly:

// Java servlet: never rely on the platform default.
response.setContentType("text/event-stream");
response.setCharacterEncoding("UTF-8");
# Python: yield str from an ASGI generator (encoded as UTF-8 by the framework), or bytes you encoded yourself.
yield f"data: {json.dumps(obj, ensure_ascii=False)}\n\n".encode("utf-8")

ensure_ascii=False sends characters as UTF-8 rather than \uXXXX escapes. Both decode correctly; escapes are larger but immune to encoding mistakes further down the pipeline.

Step 2 — Never emit a BOM except, optionally, once Permalink to this section

Most servers never produce a BOM. If a template, file or library prepends one, make sure it occurs at most once, at byte zero. The safest rule is never:

const BOM = '';
function frame(s) {
  if (s.startsWith(BOM)) s = s.slice(1);          // strip accidental BOMs from templated fragments
  return s;
}

Step 3 — Decode with a streaming decoder in custom parsers Permalink to this section

// Correct: one TextDecoder in streaming mode for the whole response.
const decoder = new TextDecoder('utf-8');
let buffer = '';
for await (const chunk of response.body) {
  buffer += decoder.decode(chunk, { stream: true });   // holds partial characters until complete
  buffer = parseCompleteLines(buffer);
}
buffer += decoder.decode();                             // flush at the end

{ stream: true } tells the decoder to keep an incomplete trailing byte sequence and prepend it to the next chunk. response.body.pipeThrough(new TextDecoderStream()) does the same. The equivalent in Python is an incremental decoder:

import codecs
dec = codecs.getincrementaldecoder("utf-8")(errors="replace")
async for chunk in response.aiter_bytes():
    text = dec.decode(chunk)                              # partial characters carried over

Custom parsers must also strip a leading BOM once, to match browser behaviour.

A character split across two network chunks Sequence diagram showing the three bytes of a multi-byte character arriving in two chunks, a per-chunk decoder producing replacement characters, and a streaming decoder producing the correct character. A character split across two network chunks Network Per-chunk decoder Streaming decoder chunk 1 ends with E6 9D emits U+FFFD chunk 1 ends with E6 9D holds 2 bytes chunk 2 starts with B1 emits 東
The server sent valid UTF-8. Only the per-chunk decoder turned it into replacement characters.

Step 4 — Treat ids and event names as UTF-8 strings too Permalink to this section

Event ids and names may contain any Unicode characters except line breaks (and NUL, for ids). They are compared as strings by addEventListener and echoed back byte-for-byte in Last-Event-ID. Non-ASCII ids work, but an HTTP header carrying non-ASCII text can be mangled by intermediaries that assume Latin-1. Keep ids ASCII — numbers, or base64url of anything richer — and names ASCII identifiers.

Step 5 — Watch for double encoding in the pipeline Permalink to this section

Double encoding happens when text that is already UTF-8 bytes is treated as Latin-1 characters and encoded to UTF-8 again. It typically creeps in at boundaries: a database connection configured with the wrong client encoding, a message broker client that decodes bytes with a platform default, or a template engine that receives bytes where it expected a string. The telltale sign is that every non-ASCII character becomes two or more characters starting with à or Â. Fix it at the boundary where bytes became text with the wrong decoder, not by adding a compensating decode at the end, which only moves the bug. Set explicit encodings on every connection string and client (client_encoding=UTF8 for PostgreSQL, charset=utf8mb4 for MySQL — the older utf8 there cannot store emoji), and treat any payload from a broker as bytes to be decoded exactly once, as UTF-8.

Validation & Monitoring Permalink to this section

# Check the bytes on the wire: é must be C3 A9, not E9; no EF BB BF except possibly at offset 0.
curl -sN http://localhost:3000/events | head -c 400 | xxd | head -20
Encoding bugs and where they come from Matrix of four encoding symptoms, the bytes seen on the wire, and the component responsible. Encoding bugs and where they come from Symptom Bytes on the wire Culprit é shown as U+FFFD E9 server not using UTF-8 é shown as é C3 83 C2 A9 double encoding First field ignored EF BB BF mid-stream BOM per event Random U+FFFD C3 A9 correct client per-chunk decode
The wire bytes identify the culprit immediately. Look at them before changing any code.

Add a fixture to the parser’s tests that contains multi-byte characters and is split at every byte offset, asserting identical output — the chunk-boundary test from parsing SSE from a fetch ReadableStream.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Can an SSE stream use UTF-16 or Latin-1?

No. Browsers decode event streams as UTF-8 regardless of the declared charset. Anything else will be misread.

Should JSON in data fields escape non-ASCII characters?

Either works. Raw UTF-8 is smaller; \u escapes survive any encoding mistake in the pipeline. Raw UTF-8 is fine when the whole server path is known to be UTF-8.

Why does EventSource handle split characters but my parser does not?

The browser decodes the stream incrementally, carrying partial characters between chunks. A parser that decodes each chunk independently loses that state; use a streaming decoder.

Is a BOM at the start of the stream an error?

No. The specification says to ignore one leading BOM, so browsers accept it. A BOM anywhere else becomes part of the text and breaks the field it prefixes.