Build
AI streaming
Stream model output to any number of readers and devices; each answer counts as one message.
Model providers stream tokens to the process that called them. That is enough for one tab and one request. It stops working when the agent runs in a background job, when the answer takes minutes, when a phone loses signal halfway, or when more than one person or device should watch. A run solves those: the agent writes once, every reader gets it live, and late or reconnecting readers catch up from the start.
Write an answer#
const run = await pulse.run() // opens the run lane (preset llm_agent_run)for await (const text of stream) { await run.write({ kind: 'text_delta', text })}await run.end()The first write starts the answer and end() closes it. Everything in between is part of one answer.
One stream per conversation#
pulse.run() is the session’s default run. When each conversation needs its own, name the stream:
const answer = pulse.stream({ streamId: `answer:${conversationId}`, preset: 'llm_agent_run' })await answer.write({ kind: 'text_delta', text: 'Checking the logs…' })await answer.end()Read an answer#
Any number of readers, on any device, in the browser or on a server:
answer.subscribe((event) => { if (event.body?.kind === 'text_delta') append(event.body.text) if (event.kind === 'end') done()})The AI SDK#
If your agent already produces AI SDK UI message stream parts, write them as they are; the adapter maps them to the run’s start, deltas and end.
await run.writeVercelAIStreamParts(parts)What it costs#
An answer counts as one message, however many chunks it carries (it counts again past 10,000 events or 10 MiB). Delivery to readers is free, and so is time: an agent may stream for an hour. Model usage is between you and your model provider.