DEV Community

myra haroon
myra haroon

Posted on

Streaming an LLM response in Next.js 15 without the UI feeling broken

Streaming is table stakes for an AI app now. People expect the answer to appear word by word, not after a 20-second spinner. Here's the setup I use in Next.js 15 with the Vercel AI SDK, plus the two things that quietly trip people up.

The route handler

// app/api/chat/route.ts
import { openai } from '@ai-sdk/openai'
import { streamText, convertToCoreMessages } from 'ai'

export async function POST(req: Request) {
  const { messages } = await req.json()

  const result = streamText({
    model: openai('gpt-4o-mini'),
    messages: convertToCoreMessages(messages),
    onFinish: async ({ text }) => {
      await saveMessage({ role: 'assistant', content: text })
    },
  })

  return result.toDataStreamResponse()
}
Enter fullscreen mode Exit fullscreen mode

The provider is one line. Swapping GPT for Claude is anthropic('claude-...') and nothing else in the handler changes.

The client

'use client'
import { useChat } from '@ai-sdk/react'

export function Chat() {
  const { messages, input, handleInputChange, handleSubmit, status } = useChat()
  // render messages, wire the form to handleSubmit
}
Enter fullscreen mode Exit fullscreen mode

useChat handles the streaming, the optimistic user message, and the loading state for you. You render messages and you're basically done.

Gotcha 1: save on the server, not the client

Don't persist the assistant message from the browser. The stream can be cut off (the user closes the tab) and you end up with half a message, or none. Save it on the server in onFinish — that fires once, with the complete text, whether or not the client is still listening.

Gotcha 2: spend the credit BEFORE you stream

If you meter usage, check and spend the credit before the stream starts, not in onFinish. Charge after and someone who disconnects mid-stream got a free generation. Spend first (atomically, so two requests can't both pass the check — I wrote about the race-safe version here), and refund on a rare hard failure if you want to be generous.

The polish that makes it feel real

Model output is markdown, so render it with react-markdown plus a syntax highlighter (rehype-highlight). Add a copy button on code blocks and a stop button wired to useChat's abort. Small touches, but they're the gap between "demo" and "product."


I bundled this streaming layer plus auth, the credit system, and Stripe billing into a Next.js 15 starter so I stop rebuilding it: live demo at https://ai-saas-starter-ashen.vercel.app, code at https://venturionai.gumroad.com/l/ai-saas-starter (40% off the first 10 with LAUNCH40). The patterns above work on their own though, starter or not.

Top comments (0)