DEV Community

Cover image for If Your Spring AI Streams Die After 60 Seconds, It's Not Your Network
Feezan Khattak
Feezan Khattak

Posted on Originally published at feezankhattak.com AI-assisted

If Your Spring AI Streams Die After 60 Seconds, It's Not Your Network

If you're on Spring AI 2.0.1 and your streaming calls die after about a minute with OpenAIIoException: Stream failed, stop debugging your proxy.

Since 2.0.1, every request carried a hard-coded 60-second per-call timeout that overrode whatever you configured. spring.ai.openai.timeout and spring.ai.openai.chat.timeout were ignored. Any streaming turn longer than a minute was killed, no matter what you set.

Spring AI 2.1.0-M1, released 25 September 2026, fixes it. That alone is worth the ten minutes — but the release also lands the change that the next two versions are built on.

TL;DR

  • The 60-second timeout regression is fixed. Your configured timeout is honoured again.
  • Message parts: messages are now an ordered list of typed parts, so reasoning, tool calls and media round-trip in the order the model produced them.
  • OpenAI Responses endpoint: one property. Needed if you want tool calling and reasoning effort on a current OpenAI flagship.
  • VectorStore.upsert(): write embeddings you computed somewhere else, with stable ids that make re-runs safe.
  • It's a milestone. APIs can change, and it's built against Spring Boot 4.2.0-M2.

The fix, first

The generic lesson is one I keep relearning in payments work: the shortest timeout in the chain wins, and it's usually one you didn't set. A load balancer, a gateway, an SDK default — or, here, a framework regression.

When a call dies at a suspiciously round number of seconds, that's rarely a coincidence. 60 seconds, 29 seconds, 30 seconds: those are configured limits, not network weather.

Three other fixes in the same release:

  • OpenAI strict-mode tool schemas are backfilled with additionalProperties: false at every object level, which fixes strict mode for MCP or hand-written schemas that didn't come from Spring AI's own generator.
  • Bedrock messages carrying media but no text no longer send an empty text block, which the Converse API rejected with a 400.
  • TextReader closes its stream instead of leaking a file descriptor per document.

Message parts: the actual headline

Until now, a Spring AI message was text plus side lists for tool calls and media. That can't represent what current models return: reasoning interleaved with tool calls, text between two images, or provider-specific blocks that have to go back verbatim next turn.

Now AssistantMessage, UserMessage and ToolResponseMessage hold an ordered list of MessagePart:

Part Holds
TextPart plain text
ReasoningPart the model's reasoning
ToolCallPart a requested tool call
ToolResultPart the result you sent back
MediaPart an image or file
UnknownPart raw JSON for a block the adapter doesn't model yet
UserMessage question = UserMessage.builder()
    .part(TextPart.of("What changed between these two charts?"))
    .part(MediaPart.of(firstChart))
    .part(MediaPart.of(secondChart))
    .build();

AssistantMessage answer = chatModel.call(new Prompt(question)).getResult().getOutput();
for (ReasoningPart reasoning : answer.getReasoning()) {
    System.out.println("Reasoning: " + reasoning.text());
}
System.out.println(answer.getText());
Enter fullscreen mode Exit fullscreen mode

Two details matter more than they look:

OpaquePayload on reasoning and tool-call parts holds things that must be replayed unchanged — an Anthropic thinking signature, a Gemini thought signature. Without a defined home for those, a framework either drops them (and the model loses its train of thought mid tool loop) or grows a provider-specific hack.

UnknownPart keeps the raw JSON of blocks the adapter doesn't understand yet. Nothing is silently dropped when a provider ships a new block type faster than the framework models it.

Your existing code keeps working: getText(), getMedia() and getToolCalls() are now views over the parts, and the old constructors still produce the legacy order. Streaming subscribers reading getText() see the same deltas as before.

Two caveats the release is explicit about: only the new OpenAiResponsesChatModel produces and consumes parts natively in M1 (the rest get refactored in RC1), and chat memory repositories don't persist parts yet — so only InMemoryChatMemoryRepository preserves reasoning across turns. With a persistent repository the conversation works, but the model reasons from scratch each turn.

The OpenAI Responses endpoint

# chat-completions (default) | responses
spring.ai.openai.chat.api=responses
Enter fullscreen mode Exit fullscreen mode

The rest of your spring.ai.openai config stays where it is.

The reason isn't novelty, it's correctness: from GPT-5.4 onward, Chat Completions doesn't support tool calling combined with a reasoning effort other than none. Responses does. Building an agent on a current OpenAI flagship? That's the endpoint.

It's built on message parts, so encrypted reasoning rides in a ReasoningPart and goes back verbatim across tool calls — the model keeps reasoning through the loop. It's deliberately stateless (every call sends the whole Prompt), so ChatMemory, advisors and RAG behave exactly as before. HostedTool switches on OpenAI's server-side tools: web search, file search, code interpreter, remote MCP, image generation.

Pre-computed embeddings

VectorStore.add() always computed embeddings with the store's own model. That's wrong whenever you already have vectors — a batch API ran overnight at a discount, another team owns the pipeline, a multimodal model embedded an image, or you're migrating from a system that exports text and vectors together.

float[] embedding = ...; // computed elsewhere
Document document = new Document("8f14e45f-ceea-467a-9a3e-5b1c2d6f7a90", "some text",
        Map.of("source", "manual"));

vectorStore.upsert(List.of(new EmbeddedDocument(document, embedding)));
Enter fullscreen mode Exit fullscreen mode

The name is the feature. Writing the same id again replaces the row, so an ingestion job with stable ids can be re-run after a failure instead of duplicating everything — the difference between an idempotent pipeline and a nightly job nobody dares restart.

Supported here: pgvector, Redis, Elasticsearch, Qdrant. Other stores throw until they opt in.

Should you take it?

Take it now if Wait for GA if
You're hitting the 60-second stream timeout 2.0.x is stable for you with no symptoms
You need tool calling and reasoning effort on a current OpenAI model You need reasoning persisted across turns (RC1)
You want to write vectors computed elsewhere Your vector store isn't one of the four
You can move to Spring Boot 4.2 You're pinned to an earlier Boot line

What RC1 is building on this

Message parts are the foundation for most of what's next: parts across all providers, session management moving into core (with repositories that persist the full part structure, so reasoning finally survives persistent storage), and MCP 2026-07-28 support. Spring AI Agents ships as a separate experimental project in November 2026, planned to merge into Spring AI 3.0 mid-2027.


The one thing to do after reading this: grep your logs for Stream failed. If it's there, you've just found out why.

I write about backend engineering and payments at feezankhattak.com — the longer version of this post has the full gotchas list, and I build free in-browser developer tools, no sign-up and nothing uploaded.

Top comments (0)