DEV Community

Aman Kumar
Aman Kumar

Posted on

How to Use Claude in Cursor and Cline With a Custom OpenAI Base URL (2026 Setup Guide)

Cursor and Cline both let you plug Claude in with an Anthropic key. That's the simplest path if you're happy paying Anthropic directly. But plenty of people reach Claude through something else: a company LiteLLM proxy, OpenRouter, or a flat-rate gateway. In those cases the endpoint speaks the OpenAI wire format, and you have to wire it in through each tool's "custom OpenAI base URL" settings.

I set this up for a lot of people (I build one of these gateways, so there's a disclosure at the end), and the same few mistakes come up every time. This guide covers the exact settings as of October 2026, plus a five-minute curl check that catches most problems before you open the editor.

My examples use my own endpoint, https://apiclaw.biz/v1. Swap in your gateway's base URL, key and model IDs, and the steps are the same.

What you need

  • A base URL that ends in /v1 (or whatever your gateway documents). It shouldn't end in /chat/completions, because the tools append that themselves.
  • An API key for that gateway. This is not an Anthropic key.
  • The exact model ID the gateway uses for Claude. Gateways name models differently, for example anthropic/claude-..., claude-..., or a prefixed alias. Copy it from the gateway's model list instead of guessing.

Step 0: sanity-check the endpoint with curl

If curl can't talk to the endpoint, neither can Cursor, and Cursor's error messages are a lot less helpful. Set two variables:

export BASE_URL="https://apiclaw.biz/v1"
export API_KEY="paste-your-key-here"
Enter fullscreen mode Exit fullscreen mode

1. List models and find the Claude IDs:

curl -s "$BASE_URL/models" \
  -H "Authorization: Bearer $API_KEY" | jq -r '.data[].id' | grep -i claude
Enter fullscreen mode Exit fullscreen mode

If you get a 401 here, the key is wrong. If you get a 404, the base URL is wrong.

2. Send one chat completion:

curl -s "$BASE_URL/chat/completions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-5-5","messages":[{"role":"user","content":"Reply with the word ok"}]}' \
  | jq -r '.choices[0].message.content'
Enter fullscreen mode Exit fullscreen mode

3. Check streaming. Both editors stream, so you should see data: chunks arrive one at a time, not all at once at the end:

curl -N "$BASE_URL/chat/completions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"claude-sonnet-5-5","stream":true,"messages":[{"role":"user","content":"Count to five"}]}'
Enter fullscreen mode Exit fullscreen mode

4. Check a tool call. Agent modes depend on this, and it's the step most likely to break on a half-compatible endpoint:

curl -s "$BASE_URL/chat/completions" \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet-5-5",
    "messages": [{"role":"user","content":"What is the weather in Pune?"}],
    "tools": [{"type":"function","function":{
      "name":"get_weather",
      "parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}
    }}]
  }' | jq '.choices[0].message.tool_calls'
Enter fullscreen mode Exit fullscreen mode

You want a get_weather call with {"city":"Pune"} in the arguments. If all four checks pass, the problem is never "the endpoint". It's the editor config.

Cursor: Override OpenAI Base URL

Cursor's own API key docs put provider keys under Cursor Settings > Models. The custom endpoint goes in the OpenAI section:

  1. Open Cursor Settings → Models and expand the API keys area.
  2. Turn on OpenAI API Key and paste your gateway key, not an OpenAI one. Cursor sends this key to whatever base URL you set next.
  3. Turn on Override OpenAI Base URL and paste the base URL, for example https://apiclaw.biz/v1.
  4. Click Add Custom Model and enter the Claude model ID exactly as your gateway lists it.
  5. Click Verify, then select the model in the chat model picker with Auto turned off.

Four things that trip people up:

Name collisions. If your custom model name matches one of Cursor's built-in names, Cursor may treat it as its own model and refuse to route it through your key. Many gateways offer a prefixed alias for this reason. On my gateway it's the apiclaw/ prefix (for example apiclaw/claude-sonnet-5-5), and other gateways use anthropic/...-style IDs.

The override isn't per model. While it's on, Cursor sends OpenAI-family requests to your URL, including some of its built-in models. Disable the built-in models you aren't using, or switch the override off when you want Cursor's own models back.

Agent mode and request shape. Cursor's agent can send request bodies that a plain /chat/completions handler doesn't expect, so tool calls fail even though chat works. Some gateways run a Cursor-specific endpoint for this. OpenRouter documents /api/v1/cursor, and mine has /cursor. Check your gateway's Cursor page if chat works but agent mode doesn't.

What the key doesn't cover. According to Cursor's docs, custom keys only work with chat models, and Tab completion stays on Cursor's built-in models. On Teams and Enterprise plans, Cursor still charges its own token rate ($0.25 per million tokens) for requests made with your key. Cursor's Zero Data Retention policy doesn't apply to requests made with your own key, and those requests still go through Cursor's servers to build the final prompt.

Cline: the OpenAI Compatible provider

Cline's OpenAI Compatible docs describe the setup. In practice:

  1. Open the Cline panel in VS Code and click the ⚙️ settings icon.
  2. Set API Provider to OpenAI Compatible.
  3. Base URL: https://apiclaw.biz/v1 (or your gateway's). Again, don't add /chat/completions.
  4. API Key: your gateway key.
  5. Model ID: the exact Claude ID from Step 0. Once the URL and key are in, Cline can often fetch the model list for you.
  6. Open Model Configuration and fill in the real numbers:
Field What to put
Context Window Size The model's actual input limit from your gateway's model list
Max Output Tokens The model's actual output limit
Image Support On, if the model accepts images
Input / Output Price Optional. These only drive Cline's cost display

Don't skip the context window. Cline uses it to decide when to trim or summarise the conversation, and the generic default for an unknown model is usually wrong for Claude. On my gateway's public model list, the Claude models currently show a 1,000,000-token input window and 128,000 max output tokens, but copy the numbers from yours.

If you use Cline's Plan and Act modes, set the model for both. Otherwise one mode can quietly fall back to a different model.

Troubleshooting map

Symptom Usual cause
404 on every request Base URL missing /v1, or /chat/completions added twice
401 / "invalid API key" Pasted a vendor key instead of the gateway key, or the key is disabled
"Model not found" Typo in the model ID, or a built-in name collision in Cursor
Chat works, agent mode fails The tool-call request shape. Try the gateway's Cursor-specific endpoint
Cline loses track in long tasks Context Window Size left at the default
Works in curl, not in Cursor Cursor's Auto mode is still picking a built-in model

Final check

Run one real task: ask the agent to read a file and propose a small edit. Then open your gateway's request log and confirm the model, the token counts, and the endpoint path you expect. Two minutes of reading logs here saves an afternoon of wondering why your bill or quota looks odd.

Disclosure: I build APIClaw, the endpoint used in the examples. The Cursor and Cline steps are the same for any OpenAI-compatible gateway, and only the URL, key and model IDs change.

Top comments (0)