JournalE.02September 30, 202611 min read

Streaming generative UI on TanStack Start

Server functions, AI SDK streams and components that render while the model thinks.

d

Dipesh Chaulagain

AI-Native Fullstack Developer

AI SDKTanStack

Stop streaming paragraphs

Most AI features ship as a text box that streams a wall of markdown. Ask where your parcel is and you get three paragraphs describing a parcel. The app that asked already has a map, a progress bar and a timeline component. The model just can't reach them.

Generative UI flips that. The model doesn't write the interface; it picks one of your components and fills in its props. Stream those props as they're produced and the component starts rendering while the model is still thinking: a chart whose line draws itself as the numbers are written, a card that appears the moment the model names the tool.

Try it. Everything below is the real pipeline: a TanStack Start server route, the AI SDK streaming protocol, useChat, and components that switch on typed message parts.

Fig. 01Components, not paragraphs
Pick a prompt

Pick a prompt. The reply streams from a real TanStack Start route.

status: ready

Tools are props

The whole trick is one reframing: a tool's input schema is a component's props. When the model “calls plotMetric”, it's really saying “render the chart, with these points.” The zod schema is the contract between a probabilistic writer and a deterministic renderer.

lib/tools.ts
import { tool } from "ai"
import type { InferUITools, UIDataTypes, UIMessage } from "ai"
import { z } from "zod"

type Shipment = { trackingId: string; eta: string; progress: number }
type Triage = { stage: "scanning" | "correlating" | "done"; events: number }

declare const shipments: { get: (id: string) => Promise<Shipment> }
declare const siem: { scan: (alertId: string) => AsyncIterable<Triage> }

export const trackShipment = tool({
  description: "Look up the live status of a shipment by tracking id.",
  inputSchema: z.object({ trackingId: z.string() }),
  execute: ({ trackingId }) => shipments.get(trackingId),
})

export const triageIncident = tool({
  description: "Triage a security alert by scanning and correlating events.",
  inputSchema: z.object({ alertId: z.string(), window: z.string() }),
  async *execute({ alertId }) {
    for await (const progress of siem.scan(alertId)) yield progress
  },
})

export const plotMetric = tool({
  description: "Render a time series chart for a metric.",
  inputSchema: z.object({
    title: z.string(),
    unit: z.string(),
    threshold: z.number(),
    points: z.array(z.object({ t: z.string(), v: z.number() })),
  }),
  execute: ({ points, threshold }) => ({
    breaches: points.filter((p) => p.v > threshold).length,
  }),
})

export const tools = { trackShipment, triageIncident, plotMetric }

export type ChatMessage = UIMessage<
  never,
  UIDataTypes,
  InferUITools<typeof tools>
>

Three kinds of tool fall out of this. trackShipment does work on the server and the UI renders its output. plotMetric is the opposite: all the interesting data is in the input the model writes, and execute barely matters. triageIncident is an async generator that reports progress while it runs. Each needs a different rendering strategy, and we'll take them one at a time.

The last line matters most. InferUITools turns the tool map into a message type where every part is a discriminated union over tool-plotMetric, tool-trackShipment and so on, each carrying its own input and output types. Change a schema and the compiler shows you every component that needs updating.

One route, one stream

On TanStack Start the endpoint is a server route: a file under routes/ with a server.handlers block. No separate API server, no framework adapter. This is the actual file serving every figure on this page:

routes/api/genui.ts
import { createFileRoute } from "@tanstack/react-router"
import { handleChat } from "@/lib/genui/chat"

export const Route = createFileRoute("/api/genui")({
  server: {
    handlers: {
      POST: ({ request }) => handleChat(request),
    },
  },
})
lib/genui/chat.ts
import {
  convertToModelMessages,
  isStepCount,
  safeValidateUIMessages,
  streamText,
} from "ai"
import { scriptedModel } from "./scripted-model"
import { tools } from "./tools"
import type { GenUIMessage } from "./tools"

const MAX_BODY = 32_000
const MAX_MESSAGES = 12

export const handleChat = async (request: Request) => {
  const raw = await request.text()
  if (raw.length > MAX_BODY) return new Response("Too large", { status: 413 })

  let body: { messages?: unknown }
  try {
    body = JSON.parse(raw) as { messages?: unknown }
  } catch {
    return new Response("Bad JSON", { status: 400 })
  }
  const parsed = await safeValidateUIMessages<GenUIMessage>({
    messages: body.messages,
    tools,
  })
  if (!parsed.success) return new Response("Bad messages", { status: 400 })

  const result = streamText({
    model: scriptedModel,
    tools,
    messages: await convertToModelMessages(parsed.data.slice(-MAX_MESSAGES)),
    stopWhen: isStepCount(3),
    abortSignal: request.signal,
  })
  return result.toUIMessageStreamResponse()
}

Four lines do most of the work. safeValidateUIMessages checks the history against the tool schemas, because a chat endpoint accepts the entire conversation from the client and the client is hostile. convertToModelMessages turns UI parts back into the provider's format. stopWhen: isStepCount(3) lets the model call a tool, read the result and write a summary, instead of stopping at the call. And abortSignal: request.signal means that when the user hits Stop, the socket closes and generation stops too, along with the bill.

What's on the wire

toUIMessageStreamResponse() returns server-sent events: one small JSON chunk per line. It's worth seeing the raw thing once, because every rendering decision you make later is a decision about these chunks.

Fig. 02What's actually on the wire
Raw SSE from /api/genui

$ curl -N -X POST /api/genui …

Text arrives as text-delta. A tool call arrives as tool-input-start, a run of tool-input-delta chunks carrying fragments of JSON, tool-input-available once the input parses and validates, then tool-output-available after execute resolves. The second start-step is the model coming back to summarise.

useChat folds that chunk stream into message parts, and each tool part walks a small state machine: input-streaming → input-available → output-available, or output-error. That state is your render switch. You never parse SSE by hand; it's only visible here so you can see where the time goes.

Render while it types

Look at the inspector again with plot latency selected. Almost the whole response is tool-input-delta. Generating twenty-four data points is the slow part, and the tool call doesn't exist until the last brace. If you render only on output-available, the user stares at a spinner for the entire generation.

During input-streaming, the AI SDK repairs the partial JSON on every delta and hands you a deep-partial input. So render it. The left pane below is the chart; the right pane is part.input exactly as the component receives it.

Fig. 03Partial JSON, live
Watch part.input fill in

Pick a prompt. The reply streams from a real TanStack Start route.

part.inputwaiting
// nothing yet
status: ready

Deep-partial means anything can be missing, including half of the last array element. { t: "14:00" } without its v is a perfectly valid moment in the stream. The renderer has to filter what's incomplete, not trust it:

components/message-view.tsx
import type { ChatMessage } from "./tools"

type Part = ChatMessage["parts"][number]
type PlotPart = Extract<Part, { type: "tool-plotMetric" }>

declare function Chart(props: {
  points: Array<{ t: string; v: number }>
  threshold?: number
  live: boolean
}): React.ReactNode
declare function ShipmentCard(props: {
  part: Extract<Part, { type: "tool-trackShipment" }>
}): React.ReactNode
declare function TriageCard(props: {
  part: Extract<Part, { type: "tool-triageIncident" }>
}): React.ReactNode

function PlotPart({ part }: { part: PlotPart }) {
  if (part.state === "output-error") return <p role="alert">{part.errorText}</p>

  const points = (part.input?.points ?? []).filter(
    (p): p is { t: string; v: number } =>
      typeof p?.t === "string" && typeof p.v === "number",
  )
  return (
    <Chart
      points={points}
      threshold={part.input?.threshold}
      live={part.state === "input-streaming"}
    />
  )
}

export function MessageView({ message }: { message: ChatMessage }) {
  return message.parts.map((part, i) => {
    switch (part.type) {
      case "text":
        return <p key={i}>{part.text}</p>
      case "tool-plotMetric":
        return <PlotPart key={part.toolCallId} part={part} />
      case "tool-trackShipment":
        return <ShipmentCard key={part.toolCallId} part={part} />
      case "tool-triageIncident":
        return <TriageCard key={part.toolCallId} part={part} />
      default:
        return null
    }
  })
}

Two more rules make partial rendering feel intentional rather than glitchy. Keep layout stable: the chart above reserves its axes up front, so new points extend a line instead of reflowing the page. And never fire side effects from partial input. It's a preview of what the model might commit to, not a command.

Tools that report back

plotMetric is slow in the input. Incident triage is slow in the execute: scan two million events, correlate, decide. A plain async execute leaves the UI with nothing to show until the very end.

Make execute an async generator instead. Every yield is sent as a preliminary output, so the same component renders live progress, and the model only ever sees the final value.

Fig. 04Preliminary outputs
A generator tool, streaming progress

Pick a prompt. The reply streams from a real TanStack Start route.

status: ready

On the client, part.preliminary tells you whether you're looking at a snapshot or the answer. The figure uses it for the pulsing live badge. Keep the snapshots self-contained, with the full list of findings so far rather than only the newest one. Then any single snapshot renders correctly on its own, and a dropped chunk costs nothing.

Not everything is a chat

A chat thread is one shape of generative UI. The more common shape in real products is a page that fills itself in: a dashboard, a report, a morning briefing. No transcript, no input box, just components arriving.

TanStack Start server functions can be async generators. Pair that with Output.array, which validates each element against a schema and emits it the moment it's complete, and you get a typed stream of objects with no route, no SSE parsing and no chat state:

lib/genui/briefing-stream.ts
import { Output, streamText } from "ai"
import { z } from "zod"
import { scriptedModel } from "./scripted-model"

export const BriefingCard = z.object({
  kind: z.enum(["metric", "alert", "note"]),
  title: z.string(),
  value: z.string(),
  delta: z.string(),
  body: z.string(),
})

export type BriefingCard = z.infer<typeof BriefingCard>

export async function* briefingCards() {
  const result = streamText({
    model: scriptedModel,
    prompt: "Write this morning's ops briefing.",
    output: Output.array({ element: BriefingCard }),
  })
  for await (const card of result.elementStream) yield card
}
lib/genui/briefing.ts
import { createServerFn } from "@tanstack/react-start"
import { briefingCards } from "./briefing-stream"

export const streamBriefing = createServerFn({ method: "GET" }).handler(
  async function* () {
    yield* briefingCards()
  },
)
Fig. 05A dashboard, not a chat
Streaming server function

await streamBriefing() → AsyncIterable<BriefingCard>

On the client that's for await (const card of await streamBriefing()). Each card is typed as BriefingCard, inferred from the zod schema through the server function boundary. Unlike partial input, elements arrive whole, so each card animates in once, fully formed. Choose this when the unit of UI is a list item; choose tool input streaming when the unit is a single rich component.

Survive a refresh

Here's what TanStack Start gives you that a client-only chat can't. A finished generative UI is just data: an array of messages whose tool parts are in output-available. Persist it, load it in a route loader, and the same components render on the server. A shared link to a conversation arrives as HTML with the charts already drawn, before any JavaScript runs.

lib/thread.ts
import { createServerFn } from "@tanstack/react-start"
import {
  convertToModelMessages,
  isStepCount,
  streamText,
  validateUIMessages,
} from "ai"
import type { LanguageModel } from "ai"
import { z } from "zod"
import { tools } from "./tools"
import type { ChatMessage } from "./tools"

declare const model: LanguageModel
declare const db: {
  load: (threadId: string) => Promise<unknown>
  save: (threadId: string, messages: Array<ChatMessage>) => Promise<void>
}

export const getThread = createServerFn({ method: "GET" })
  .inputValidator(z.object({ threadId: z.string() }))
  .handler(async ({ data }) => {
    const messages = await validateUIMessages<ChatMessage>({
      messages: await db.load(data.threadId),
      tools,
    })
    return JSON.stringify(messages)
  })

export const readThread = (json: string) =>
  JSON.parse(json) as Array<ChatMessage>

export const chat = async (threadId: string, messages: Array<ChatMessage>) => {
  const result = streamText({
    model,
    tools,
    messages: await convertToModelMessages(messages),
    stopWhen: isStepCount(3),
  })
  return result.toUIMessageStreamResponse<ChatMessage>({
    originalMessages: messages,
    onEnd: ({ messages: finished }) => db.save(threadId, finished),
  })
}

originalMessages plus onEnd hands you the full, updated conversation when the stream closes. Save it there, not on the client, which can close the tab mid-stream.

Loading has a sharp edge. Return Array<UIMessage> straight from a server function and Start's serialization check rejects it: dynamic tool parts carry an unknown input, and unknown can't be proven serializable. The compiler is right to be suspicious. Stored messages are old input from a client you don't trust, written against schemas that may have changed since. So validate them against today's tools on the server, cross the boundary as a string, and hand the result to useChat({ messages }).

Production notes

The component set is an allowlist. Never let the model emit markup, JSX or component names you look up dynamically. It picks from the tools you registered, and a default: branch renders nothing. That's the difference between generative UI and an injection vector.

Cap everything the client sends. The route above rejects bodies over 32 KB and sends only the last twelve messages to the model. A public chat endpoint without limits is a free, unmetered proxy to your API key.

Throttle re-renders. A fast model can emit a hundred deltas a second, and each one re-renders the thread. Pass throttle (in milliseconds) to useChat and memoise finished messages so only the streaming one updates.

Design the error state. output-error is a first-class part state, not an exception. Render it inside the component's frame, and use onError on the stream response to turn stack traces into something safe to show.

Announce politely. Put the thread in an aria-live="polite" region, and label the streaming state rather than announcing each token. A screen reader reading every delta aloud is unusable.

Mind the clock. Serverless functions have duration limits, and a three-step tool run can outlive them. For work longer than a request, make the agent durable (see E.01) and stream its progress instead.

The checklist

  • Every tool's input schema is designed as component props.
  • The message type comes from InferUITools, not hand-written.
  • Incoming messages are validated, size-capped and truncated.
  • request.signal is passed to streamText.
  • Components render input-streaming and filter partial items.
  • Slow tools are generators that yield their final value.
  • List-shaped UIs use Output.array through a server function.
  • Finished threads are saved in onEnd and render on the server.
  • Unknown parts render nothing. The model never writes markup.

The model is good at choosing and filling in. Your components are good at being fast, accessible and on-brand. Generative UI works when each side does only its own job, and the stream in between is typed.