Cursor and Cline both let you plug Claude in with an Anthropic key. That's the simplest path if you're happy paying Anthropic directly. But plenty of people reach Claude through something else: a company LiteLLM proxy, OpenRouter, or a flat-rate gateway. In those cases the endpoint speaks the OpenAI wire format, and you have to wire it in through each tool's "custom OpenAI base URL" settings.
I set this up for a lot of people (I build one of these gateways, so there's a disclosure at the end), and the same few mistakes come up every time. This guide covers the exact settings as of October 2026, plus a five-minute curl check that catches most problems before you open the editor.
My examples use my own endpoint, https://apiclaw.biz/v1. Swap in your gateway's base URL, key and model IDs, and the steps are the same.
What you need
-
A base URL that ends in
/v1(or whatever your gateway documents). It shouldn't end in/chat/completions, because the tools append that themselves. - An API key for that gateway. This is not an Anthropic key.
-
The exact model ID the gateway uses for Claude. Gateways name models differently, for example
anthropic/claude-...,claude-..., or a prefixed alias. Copy it from the gateway's model list instead of guessing.
Step 0: sanity-check the endpoint with curl
If curl can't talk to the endpoint, neither can Cursor, and Cursor's error messages are a lot less helpful. Set two variables:
export BASE_URL="https://apiclaw.biz/v1"
export API_KEY="paste-your-key-here"
1. List models and find the Claude IDs:
curl -s "$BASE_URL/models" \
-H "Authorization: Bearer $API_KEY" | jq -r '.data[].id' | grep -i claude
If you get a 401 here, the key is wrong. If you get a 404, the base URL is wrong.
2. Send one chat completion:
curl -s "$BASE_URL/chat/completions" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5-5","messages":[{"role":"user","content":"Reply with the word ok"}]}' \
| jq -r '.choices[0].message.content'
3. Check streaming. Both editors stream, so you should see data: chunks arrive one at a time, not all at once at the end:
curl -N "$BASE_URL/chat/completions" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"claude-sonnet-5-5","stream":true,"messages":[{"role":"user","content":"Count to five"}]}'
4. Check a tool call. Agent modes depend on this, and it's the step most likely to break on a half-compatible endpoint:
curl -s "$BASE_URL/chat/completions" \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet-5-5",
"messages": [{"role":"user","content":"What is the weather in Pune?"}],
"tools": [{"type":"function","function":{
"name":"get_weather",
"parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}
}}]
}' | jq '.choices[0].message.tool_calls'
You want a get_weather call with {"city":"Pune"} in the arguments. If all four checks pass, the problem is never "the endpoint". It's the editor config.
Cursor: Override OpenAI Base URL
Cursor's own API key docs put provider keys under Cursor Settings > Models. The custom endpoint goes in the OpenAI section:
- Open Cursor Settings → Models and expand the API keys area.
- Turn on OpenAI API Key and paste your gateway key, not an OpenAI one. Cursor sends this key to whatever base URL you set next.
- Turn on Override OpenAI Base URL and paste the base URL, for example
https://apiclaw.biz/v1. - Click Add Custom Model and enter the Claude model ID exactly as your gateway lists it.
- Click Verify, then select the model in the chat model picker with Auto turned off.
Four things that trip people up:
Name collisions. If your custom model name matches one of Cursor's built-in names, Cursor may treat it as its own model and refuse to route it through your key. Many gateways offer a prefixed alias for this reason. On my gateway it's the apiclaw/ prefix (for example apiclaw/claude-sonnet-5-5), and other gateways use anthropic/...-style IDs.
The override isn't per model. While it's on, Cursor sends OpenAI-family requests to your URL, including some of its built-in models. Disable the built-in models you aren't using, or switch the override off when you want Cursor's own models back.
Agent mode and request shape. Cursor's agent can send request bodies that a plain /chat/completions handler doesn't expect, so tool calls fail even though chat works. Some gateways run a Cursor-specific endpoint for this. OpenRouter documents /api/v1/cursor, and mine has /cursor. Check your gateway's Cursor page if chat works but agent mode doesn't.
What the key doesn't cover. According to Cursor's docs, custom keys only work with chat models, and Tab completion stays on Cursor's built-in models. On Teams and Enterprise plans, Cursor still charges its own token rate ($0.25 per million tokens) for requests made with your key. Cursor's Zero Data Retention policy doesn't apply to requests made with your own key, and those requests still go through Cursor's servers to build the final prompt.
Cline: the OpenAI Compatible provider
Cline's OpenAI Compatible docs describe the setup. In practice:
- Open the Cline panel in VS Code and click the ⚙️ settings icon.
- Set API Provider to OpenAI Compatible.
-
Base URL:
https://apiclaw.biz/v1(or your gateway's). Again, don't add/chat/completions. - API Key: your gateway key.
- Model ID: the exact Claude ID from Step 0. Once the URL and key are in, Cline can often fetch the model list for you.
- Open Model Configuration and fill in the real numbers:
| Field | What to put |
|---|---|
| Context Window Size | The model's actual input limit from your gateway's model list |
| Max Output Tokens | The model's actual output limit |
| Image Support | On, if the model accepts images |
| Input / Output Price | Optional. These only drive Cline's cost display |
Don't skip the context window. Cline uses it to decide when to trim or summarise the conversation, and the generic default for an unknown model is usually wrong for Claude. On my gateway's public model list, the Claude models currently show a 1,000,000-token input window and 128,000 max output tokens, but copy the numbers from yours.
If you use Cline's Plan and Act modes, set the model for both. Otherwise one mode can quietly fall back to a different model.
Troubleshooting map
| Symptom | Usual cause |
|---|---|
| 404 on every request | Base URL missing /v1, or /chat/completions added twice |
| 401 / "invalid API key" | Pasted a vendor key instead of the gateway key, or the key is disabled |
| "Model not found" | Typo in the model ID, or a built-in name collision in Cursor |
| Chat works, agent mode fails | The tool-call request shape. Try the gateway's Cursor-specific endpoint |
| Cline loses track in long tasks | Context Window Size left at the default |
| Works in curl, not in Cursor | Cursor's Auto mode is still picking a built-in model |
Final check
Run one real task: ask the agent to read a file and propose a small edit. Then open your gateway's request log and confirm the model, the token counts, and the endpoint path you expect. Two minutes of reading logs here saves an afternoon of wondering why your bill or quota looks odd.
Disclosure: I build APIClaw, the endpoint used in the examples. The Cursor and Cline steps are the same for any OpenAI-compatible gateway, and only the URL, key and model IDs change.
Top comments (0)