If you're on Spring AI 2.0.1 and your streaming calls die after about a minute with OpenAIIoException: Stream failed, stop debugging your proxy.
Since 2.0.1, every request carried a hard-coded 60-second per-call timeout that overrode whatever you configured. spring.ai.openai.timeout and spring.ai.openai.chat.timeout were ignored. Any streaming turn longer than a minute was killed, no matter what you set.
Spring AI 2.1.0-M1, released 25 September 2026, fixes it. That alone is worth the ten minutes — but the release also lands the change that the next two versions are built on.
TL;DR
- The 60-second timeout regression is fixed. Your configured timeout is honoured again.
- Message parts: messages are now an ordered list of typed parts, so reasoning, tool calls and media round-trip in the order the model produced them.
- OpenAI Responses endpoint: one property. Needed if you want tool calling and reasoning effort on a current OpenAI flagship.
-
VectorStore.upsert(): write embeddings you computed somewhere else, with stable ids that make re-runs safe. - It's a milestone. APIs can change, and it's built against Spring Boot 4.2.0-M2.
The fix, first
The generic lesson is one I keep relearning in payments work: the shortest timeout in the chain wins, and it's usually one you didn't set. A load balancer, a gateway, an SDK default — or, here, a framework regression.
When a call dies at a suspiciously round number of seconds, that's rarely a coincidence. 60 seconds, 29 seconds, 30 seconds: those are configured limits, not network weather.
Three other fixes in the same release:
- OpenAI strict-mode tool schemas are backfilled with
additionalProperties: falseat every object level, which fixes strict mode for MCP or hand-written schemas that didn't come from Spring AI's own generator. - Bedrock messages carrying media but no text no longer send an empty text block, which the Converse API rejected with a 400.
-
TextReadercloses its stream instead of leaking a file descriptor per document.
Message parts: the actual headline
Until now, a Spring AI message was text plus side lists for tool calls and media. That can't represent what current models return: reasoning interleaved with tool calls, text between two images, or provider-specific blocks that have to go back verbatim next turn.
Now AssistantMessage, UserMessage and ToolResponseMessage hold an ordered list of MessagePart:
| Part | Holds |
|---|---|
TextPart |
plain text |
ReasoningPart |
the model's reasoning |
ToolCallPart |
a requested tool call |
ToolResultPart |
the result you sent back |
MediaPart |
an image or file |
UnknownPart |
raw JSON for a block the adapter doesn't model yet |
UserMessage question = UserMessage.builder()
.part(TextPart.of("What changed between these two charts?"))
.part(MediaPart.of(firstChart))
.part(MediaPart.of(secondChart))
.build();
AssistantMessage answer = chatModel.call(new Prompt(question)).getResult().getOutput();
for (ReasoningPart reasoning : answer.getReasoning()) {
System.out.println("Reasoning: " + reasoning.text());
}
System.out.println(answer.getText());
Two details matter more than they look:
OpaquePayload on reasoning and tool-call parts holds things that must be replayed unchanged — an Anthropic thinking signature, a Gemini thought signature. Without a defined home for those, a framework either drops them (and the model loses its train of thought mid tool loop) or grows a provider-specific hack.
UnknownPart keeps the raw JSON of blocks the adapter doesn't understand yet. Nothing is silently dropped when a provider ships a new block type faster than the framework models it.
Your existing code keeps working: getText(), getMedia() and getToolCalls() are now views over the parts, and the old constructors still produce the legacy order. Streaming subscribers reading getText() see the same deltas as before.
Two caveats the release is explicit about: only the new OpenAiResponsesChatModel produces and consumes parts natively in M1 (the rest get refactored in RC1), and chat memory repositories don't persist parts yet — so only InMemoryChatMemoryRepository preserves reasoning across turns. With a persistent repository the conversation works, but the model reasons from scratch each turn.
The OpenAI Responses endpoint
# chat-completions (default) | responses
spring.ai.openai.chat.api=responses
The rest of your spring.ai.openai config stays where it is.
The reason isn't novelty, it's correctness: from GPT-5.4 onward, Chat Completions doesn't support tool calling combined with a reasoning effort other than none. Responses does. Building an agent on a current OpenAI flagship? That's the endpoint.
It's built on message parts, so encrypted reasoning rides in a ReasoningPart and goes back verbatim across tool calls — the model keeps reasoning through the loop. It's deliberately stateless (every call sends the whole Prompt), so ChatMemory, advisors and RAG behave exactly as before. HostedTool switches on OpenAI's server-side tools: web search, file search, code interpreter, remote MCP, image generation.
Pre-computed embeddings
VectorStore.add() always computed embeddings with the store's own model. That's wrong whenever you already have vectors — a batch API ran overnight at a discount, another team owns the pipeline, a multimodal model embedded an image, or you're migrating from a system that exports text and vectors together.
float[] embedding = ...; // computed elsewhere
Document document = new Document("8f14e45f-ceea-467a-9a3e-5b1c2d6f7a90", "some text",
Map.of("source", "manual"));
vectorStore.upsert(List.of(new EmbeddedDocument(document, embedding)));
The name is the feature. Writing the same id again replaces the row, so an ingestion job with stable ids can be re-run after a failure instead of duplicating everything — the difference between an idempotent pipeline and a nightly job nobody dares restart.
Supported here: pgvector, Redis, Elasticsearch, Qdrant. Other stores throw until they opt in.
Should you take it?
| Take it now if | Wait for GA if |
|---|---|
| You're hitting the 60-second stream timeout | 2.0.x is stable for you with no symptoms |
| You need tool calling and reasoning effort on a current OpenAI model | You need reasoning persisted across turns (RC1) |
| You want to write vectors computed elsewhere | Your vector store isn't one of the four |
| You can move to Spring Boot 4.2 | You're pinned to an earlier Boot line |
What RC1 is building on this
Message parts are the foundation for most of what's next: parts across all providers, session management moving into core (with repositories that persist the full part structure, so reasoning finally survives persistent storage), and MCP 2026-07-28 support. Spring AI Agents ships as a separate experimental project in November 2026, planned to merge into Spring AI 3.0 mid-2027.
The one thing to do after reading this: grep your logs for Stream failed. If it's there, you've just found out why.
I write about backend engineering and payments at feezankhattak.com — the longer version of this post has the full gotchas list, and I build free in-browser developer tools, no sign-up and nothing uploaded.
Top comments (0)