Filtering a Live Log Stream on the Server Permalink to this section
Part of Log Tailing & CI Output Streaming, under Real-Time Application Patterns.
A user looking for one error in a noisy service does not want the other ninety-nine lines out of every hundred. Filtering those out in the browser means they still cross the network, still get parsed, and still cost the server a write each. Filtering on the server sends only what the user asked for. This guide adds level, text and regular-expression filters to a Server-Sent Events log stream, keeps user-supplied patterns from becoming a denial-of-service vector, adds grep-style context lines, and handles the user changing the filter while the stream is running.
Symptom & Developer Intent Permalink to this section
- A filtered log view is sluggish, and the network panel shows megabytes per minute arriving for a view that displays a few lines.
- Mobile users of the admin panel cannot use the live logs at all on busy services.
- A user pastes a pathological regular expression and the log server’s CPU pins at 100 %.
- Matches are shown without the surrounding lines that explain them.
- Changing the filter drops the lines that arrived during the switch, or duplicates them.
The intent is server-side filtering that cuts traffic in proportion to selectivity, is safe against hostile patterns, shows context, and lets the filter change without losing the position in the log.
Root Cause Analysis Permalink to this section
The cost of a log stream is paid per line sent: the server serialises and writes it, the network carries it, the client parses and renders it. Client-side filtering only removes the rendering cost. For a filter that keeps 1 % of lines, 99 % of the remaining cost is waste.
Regular expressions introduce a specific risk. JavaScript’s regex engine uses backtracking, and some patterns — nested quantifiers such as (a+)+$ — take exponential time on certain inputs. A user-supplied pattern evaluated against every log line on the main event loop can stall every stream on the process. The mitigation is to restrict what users can express, run matching somewhere that can be interrupted, or use a linear-time engine.
Filter changes are a sequencing problem. If changing the filter closes the stream and opens a new one from “now”, lines in between are lost; if it reopens from the start, the user sees everything again. The stream’s cursor solves it: reopen with the new filter from the old cursor.
Step-by-Step Resolution Permalink to this section
Step 1 — Define filters as query parameters with strict validation Permalink to this section
const LEVELS = ['trace', 'debug', 'info', 'warn', 'error', 'fatal'];
function parseFilter(q) {
const f = {};
if (q.level) {
const i = LEVELS.indexOf(q.level);
if (i < 0) throw new BadRequest('level');
f.minLevel = i;
}
if (q.q) {
if (q.q.length > 200) throw new BadRequest('q too long');
f.text = q.q.toLowerCase(); // plain substring: always linear
}
if (q.re) {
if (q.re.length > 120) throw new BadRequest('re too long');
f.re = compileSafe(q.re); // see step 2
}
f.context = Math.min(Number(q.context) || 0, 10);
return f;
}
Filters live in the URL, so EventSource sends them on every reconnect automatically — no client code is needed to preserve them across a dropped connection.
Step 2 — Make regular expressions safe Permalink to this section
The most robust choice is a linear-time engine. RE2 (via the re2 package in Node.js, or google-re2 in Python) guarantees matching time proportional to input length by forbidding backreferences and lookaround.
import RE2 from 're2';
function compileSafe(pattern) {
try {
return new RE2(pattern, 'i'); // throws on unsupported constructs
} catch {
throw new BadRequest('unsupported pattern');
}
}
If adding a native dependency is not an option, reject patterns containing nested quantifiers and backreferences with a conservative check, cap line length before matching, and measure matching time per batch — closing the stream with an error event if a pattern exceeds its budget.
Step 3 — Apply filters in the tailer, before batching Permalink to this section
function makeMatcher(f) {
return (entry) => {
if (f.minLevel != null && entry.level < f.minLevel) return false;
if (f.text && !entry.lower.includes(f.text)) return false;
if (f.re && !f.re.test(entry.text.slice(0, 4096))) return false; // cap input length
return true;
};
}
Parse each line’s level once (structured logs make this trivial; plain text needs a tolerant prefix parser) and keep a lower-cased copy for substring matching.
Step 4 — Add context lines with a small ring buffer Permalink to this section
Grep’s -C option — show N lines before and after each match — is the feature that makes filtered logs readable. Keep the last N unmatched lines in a ring buffer and a countdown for lines after a match:
function withContext(match, n) {
const before = [];
let after = 0;
return (entry, emit) => {
if (match(entry)) {
before.forEach((e) => emit({ ...e, ctx: true })); // flush leading context
before.length = 0;
emit(entry);
after = n;
} else if (after > 0) {
emit({ ...entry, ctx: true });
after -= 1;
} else if (n > 0) {
before.push(entry);
if (before.length > n) before.shift();
}
};
}
The ctx flag lets the client render context lines dimmed, so matches stand out.
Step 5 — Keep the cursor while the filter changes Permalink to this section
The SSE id is still the source offset, including for lines the filter skipped. When the user changes the filter, reopen from the current cursor:
let es, cursor = null;
function open(filter) {
es?.close();
const params = new URLSearchParams(filter);
if (cursor) params.set('from', cursor); // continue, don't restart
es = new EventSource(`/api/services/${id}/logs?${params}`, { withCredentials: true });
es.addEventListener('lines', (e) => { cursor = e.lastEventId; render(JSON.parse(e.data).lines); });
}
The server should advance the id even when a whole batch is filtered out — send an empty lines event, or a comment plus an id-only event, periodically — otherwise a highly selective filter leaves the cursor far behind and a reconnect re-scans a large range.
Validation & Monitoring Permalink to this section
# Selectivity: compare bytes received with and without a filter over 30 seconds.
timeout 30 curl -sN -b s.txt 'https://admin.example.com/api/services/api-7/logs' | wc -c
timeout 30 curl -sN -b s.txt 'https://admin.example.com/api/services/api-7/logs?level=error' | wc -c
# Hostile pattern: must be rejected or evaluated in bounded time.
curl -s -o /dev/null -w '%{http_code} %{time_total}\n' -b s.txt \
'https://admin.example.com/api/services/api-7/logs?re=(a%2B)%2B%24'
Record per stream the lines scanned, lines sent and matching time. The ratio shows how effective filters are; matching time per thousand lines is the early warning for an expensive pattern.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Why not filter in the browser to keep the server simple?
Because every filtered-out line still costs server writes, network bandwidth and client parsing. On busy services that makes filtered views as expensive as unfiltered ones, and unusable on mobile connections.
Is it safe to accept regular expressions from users?
Only with a linear-time engine such as RE2, or with strict restrictions and a time budget. A backtracking engine evaluating an arbitrary pattern on the event loop can stall every stream on the process.
What happens to the cursor when nothing matches for a long time?
The server should keep sending id-only progress events so the browser's Last-Event-ID stays close to the source's current position. Otherwise a reconnect re-scans everything since the last match.
How should the view indicate that a filter is active?
Show the active filters as removable chips above the log and a running count of lines scanned versus shown. Without it, users forget a filter is on and conclude the service has stopped logging.
Can filters be combined?
Yes. Apply the cheapest checks first — level, then substring, then regex — so most lines are rejected before the expensive test runs.