DEV Community

Francis Oyakhire
Francis Oyakhire

Posted on

Ollama v1 OpenAI Compat Drops Think Toggle

This week, the open-source community buzzed with excitement around LiveKit’s new WebRTC stack and cloud offerings. As developers and engineers, we’re always on the lookout for tools that simplify real-time communication and AI integration. But even with these advancements, we often run into subtle but impactful issues in the stack that require deep debugging. One such incident recently occurred in our work with Ollama and the Qwen3-family models.

We use Ollama’s /v1 endpoint to interface with multiple LLMs, including the Qwen3 series, in a production setting. Recently, we noticed that for certain models, the response content was consistently empty, despite the model being correctly invoked. This wasn’t a matter of timeouts or errors - it was a complete absence of output. After extensive logging and profiling, we traced the issue to a seemingly innocuous parameter: think: false.

Ollama’s /v1 endpoint is designed to be OpenAI-compatible, which makes it easy to integrate with existing tooling and workflows. However, we discovered that for the Qwen3-family models, the think: false toggle was being silently ignored. This meant that even when we explicitly set think: false, the model would still consume its entire num_predict budget on internal reasoning, leaving no tokens for the actual response. The result? Empty content, no errors, and a lot of confusion.

To confirm this, we compared the payloads sent to the /v1 endpoint versus the native /api/chat endpoint. Here’s a simplified version of what we sent to the /v1 endpoint:

{
  "model": "qwen3",
  "prompt": "What is the capital of France?",
  "options": {
    "num_predict": 100,
    "think": false
  }
}
Enter fullscreen mode Exit fullscreen mode

And here’s what we sent to the native endpoint:

{
  "model": "qwen3",
  "prompt": "What is the capital of France?",
  "options": {
    "num_predict": 100
  }
}
Enter fullscreen mode Exit fullscreen mode

The difference? The absence of the think: false toggle in the native endpoint. Surprisingly, the native endpoint didn’t require it to produce a response. This suggests that the /v1 endpoint’s compatibility layer was either misinterpreting or ignoring the think parameter for this specific model family.

The tradeoff here is clear: while the /v1 endpoint offers a familiar interface for those coming from the OpenAI ecosystem, it’s not always fully compatible with all models. In our case, it introduced a silent failure mode that was difficult to diagnose without deep logging and testing.

We’ve since moved our Qwen3-family integrations to the native /api/chat endpoint, where the behavior is predictable and consistent. However, this incident highlights a broader challenge: as we work with increasingly diverse model ecosystems, compatibility layers can introduce subtle, hard-to-detect issues.

Looking ahead, we’re exploring ways to build more robust, model-agnostic interfaces that can handle these edge cases without requiring users to switch endpoints. We’re also considering whether to contribute a patch or a workaround back to the Ollama community to improve the /v1 endpoint’s compatibility with Qwen3 and similar models. What do you think - should compatibility layers be more transparent about their limitations, or is it better to rely on native APIs for critical workflows?

Top comments (0)