Streaming Tool Calls and Structured Output Permalink to this section
Part of Streaming AI Responses in the Browser, under Frontend Consumption & Client Patterns.
Early AI interfaces streamed one thing: text tokens. Modern ones stream a sequence of different things — reasoning status, a tool the model decided to call, the arguments it is filling in, the tool’s result, and finally an answer, sometimes as structured JSON rather than prose. Sending all of that as undifferentiated text forces the client to guess what it is looking at. This guide designs a Server-Sent Events protocol with a named event per kind of content, streams partial JSON safely, and builds a client state machine that renders each phase correctly.
Symptom & Developer Intent Permalink to this section
- Raw tool-call JSON flashes on screen before the interface realises it is not part of the answer.
- The client tries to
JSON.parsea partial structured answer on every token and throws. - The “searching the web…” indicator never clears because the end of the tool call is ambiguous.
- Two parallel tool calls interleave their argument fragments and the client mixes them up.
- A reconnect in the middle of a tool call leaves the UI in a half-rendered state.
The intent is a stream where every fragment says what it belongs to, the client knows at every moment which phase each part of the response is in, and partial structured data renders progressively without parse errors.
Root Cause Analysis Permalink to this section
Model providers stream heterogeneous content, each with its own event types. Forwarding a provider’s raw stream to the browser couples the frontend to one provider’s format; flattening it to text loses the structure. The right boundary is your own small, stable protocol between backend and browser, with a named event per kind of content and an index identifying which part of the response each fragment belongs to.
Partial JSON fails to parse because it is, by definition, incomplete. The client should either accumulate fragments and parse only on the part’s end event, or use an incremental parser that tolerates incomplete input and returns the best current view.
Step-by-Step Resolution Permalink to this section
Step 1 — Define a small, provider-neutral event protocol Permalink to this section
event: part.start
data: {"index":0,"kind":"text"}
event: part.delta
data: {"index":0,"text":"Let me check the weather"}
event: part.start
data: {"index":1,"kind":"tool_call","name":"get_weather","callId":"c_1"}
event: part.delta
data: {"index":1,"args":"{\"city\":\"Os"}
event: part.delta
data: {"index":1,"args":"lo\"}"}
event: part.end
data: {"index":1}
event: tool.result
data: {"callId":"c_1","ok":true,"result":{"tempC":7,"sky":"rain"}}
event: done
data: {"finish":"stop","usage":{"output_tokens":212}}
Every fragment carries the index of the part it belongs to, so parallel tool calls with interleaved fragments are unambiguous. The server translates whatever the provider emits into this protocol, so a change of provider does not touch the frontend.
Step 2 — Translate provider events on the server Permalink to this section
// server: map provider stream events to the browser protocol.
for await (const ev of provider.stream({ messages, tools, signal })) {
switch (ev.type) {
case 'content_block_start':
send('part.start', { index: ev.index, kind: ev.block.type === 'tool_use' ? 'tool_call' : 'text',
name: ev.block.name, callId: ev.block.id });
break;
case 'content_block_delta':
send('part.delta', ev.delta.type === 'input_json_delta'
? { index: ev.index, args: ev.delta.partial_json }
: { index: ev.index, text: ev.delta.text });
break;
case 'content_block_stop':
send('part.end', { index: ev.index });
break;
}
}
The provider-side event names above are illustrative; each provider’s SDK documents its own. What matters is that the browser sees only your protocol. When the backend runs a tool, it emits tool.result (and tool.error on failure) before continuing generation.
Step 3 — Build a client reducer keyed by part index Permalink to this section
export function responseReducer(state, { type, data }) {
switch (type) {
case 'part.start':
return { ...state, parts: { ...state.parts, [data.index]: { ...data, text: '', args: '', status: 'streaming' } } };
case 'part.delta': {
const p = state.parts[data.index];
if (!p) return state; // ignore fragments for unknown parts
return { ...state, parts: { ...state.parts, [data.index]: {
...p, text: p.text + (data.text ?? ''), args: p.args + (data.args ?? '') } } };
}
case 'part.end': {
const p = state.parts[data.index];
const status = p.kind === 'tool_call' ? 'running' : 'done';
return { ...state, parts: { ...state.parts, [data.index]: { ...p, status, parsedArgs: safeParse(p.args) } } };
}
case 'tool.result': {
const [idx, p] = Object.entries(state.parts).find(([, x]) => x.callId === data.callId) ?? [];
return idx ? { ...state, parts: { ...state.parts, [idx]: { ...p, status: 'done', result: data.result } } } : state;
}
case 'done':
return { ...state, finished: true, usage: data.usage };
default:
return state; // unknown events are ignored, never fatal
}
}
Rendering then follows the part’s kind and status: text parts render markdown incrementally, tool calls render a compact “calling get_weather…” chip while streaming or running and a result card when done. Tool arguments never appear as raw text.
Step 4 — Render partial structured output progressively Permalink to this section
When the final answer itself is JSON — a generated table, a form, a set of cards — parse it incrementally so the interface fills in as fields arrive. An incremental parser returns the best object it can build from a prefix:
import { parse as parsePartial } from 'partial-json'; // tolerant parser for incomplete JSON
function renderStructured(part) {
const view = part.status === 'done' ? JSON.parse(part.args || part.text) : parsePartial(part.args || part.text);
return <ResultCards items={view?.items ?? []} />; // cards appear one by one
}
Validate the final object against a schema on part.end, and show a clear error rather than a half-rendered view if validation fails.
Step 5 — Resume mid-call safely Permalink to this section
If the connection drops during a tool call, the resumed stream — replayed from Last-Event-ID as described in resuming an interrupted AI stream — re-sends the fragments after the cursor. Because the reducer appends by index and events carry ids, the client rebuilds exactly the same state. Give every event an id for this reason, not only text deltas.
Validation & Monitoring Permalink to this section
Record real sessions as fixtures (the raw event sequence) and replay them through the reducer in tests, asserting final state and intermediate snapshots. Include a fixture with two parallel tool calls whose fragments interleave, one with a failing tool, and one that ends mid-JSON.
In production, log unknown event types and unparseable final payloads with the model and prompt version; both are early signs of a provider or prompt change.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Should I forward the model provider's stream directly to the browser?
It works for prototypes, but it couples the frontend to one provider's event format and may expose details you do not want in the browser. Translating to your own protocol keeps both sides stable.
How do I show that a tool is running?
Render a status chip when the tool call part starts, keep it while arguments stream and the tool runs, and replace it with a result card on the tool result event.
Can I parse partial JSON with JSON.parse?
No; it throws on incomplete input. Accumulate and parse on completion, or use a tolerant incremental parser for progressive rendering.
How are parallel tool calls handled?
Each call is its own part with its own index and call id. Fragments name their index, so interleaving is unambiguous, and results match calls by call id.