DEV Community

LaoDeng Learns AI
LaoDeng Learns AI

Posted on

How to connect 12 different AI clients to any OpenAI-compatible endpoint

If you self-host a model, run an API gateway, or use a provider in your region, you've hit this wall: every AI client wants your endpoint entered differently.

Cline calls it openAiBaseUrl. Continue calls it apiBase. Aider reads OPENAI_API_BASE from the environment. Cursor hides it behind a toggle. SillyTavern only shows the custom endpoint field after you pick the right chat completion source.

So you google "how to set base url in Cline". Then "how to set base url in Continue". Then you do it all again three months later because you forgot.

This is the cheat sheet I wanted to exist. Twelve clients, plus the gotchas that cost me the most time.

First: what does "OpenAI-compatible" actually mean?

An endpoint is OpenAI-compatible if it implements at least two routes:

POST /v1/chat/completions
GET  /v1/models
Enter fullscreen mode Exit fullscreen mode

That's the whole bar. If both work, almost every client on this list will talk to it.

The one question that trips everyone up: does the base URL include /v1?

The OpenAI SDK convention is that base_url includes it, and the SDK appends the rest:

base_url = "https://api.example.com/v1"
→ requests go to https://api.example.com/v1/chat/completions
Enter fullscreen mode Exit fullscreen mode

Most clients follow that convention. Cursor is the notable exception — it appends the path itself, so what you enter depends on how your provider documents it.

Rule of thumb: if in doubt, include /v1. If you get a 404, drop it.

Test any endpoint before you touch a client:

curl https://api.example.com/v1/models \
  -H "Authorization: Bearer $YOUR_KEY"
Enter fullscreen mode Exit fullscreen mode

A JSON list of models means you're compatible. A 401 means your key is wrong. A 404 means your base URL is wrong.

The 12 clients

GUI apps

SillyTavern — API: Chat Completion → Chat Completion Source: Custom (OpenAI-compatible). The custom endpoint field only appears after you select that source. The API key field is hidden behind a toggle.

Cherry Studio — Settings → Model Providers → Add Provider → type OpenAI. Then use "Manage Models" to add the model id by hand; it does not auto-discover.

Jan — Settings → Providers → Add Provider, then add the model id manually.

Cursor — Cursor Settings → Models → enable "OpenAI API Key" → enable "Override OpenAI Base URL". Remember that Cursor appends the path itself.

VS Code extensions

Cline — three keys in your VS Code settings.json:

{
  "cline.apiProvider": "openai",
  "cline.openAiBaseUrl": "https://api.example.com/v1",
  "cline.openAiApiKey": "sk-...",
  "cline.openAiModelId": "your-model"
}
Enter fullscreen mode Exit fullscreen mode

Continue — ~/.continue/config.json:

{
  "models": [
    {
      "title": "My Endpoint",
      "provider": "openai",
      "model": "your-model",
      "apiBase": "https://api.example.com/v1",
      "apiKey": "sk-..."
    }
  ]
}
Enter fullscreen mode Exit fullscreen mode

CLI

Aider — reads standard environment variables. Put this in .env at your project root:

OPENAI_API_BASE=https://api.example.com/v1
OPENAI_API_KEY=sk-...
Enter fullscreen mode Exit fullscreen mode

then run:

aider --model openai/your-model
Enter fullscreen mode Exit fullscreen mode

Gateways and self-hosted UIs

LiteLLM — config.yaml:

model_list:
  - model_name: my-endpoint
    litellm_params:
      model: openai/your-model
      api_base: "https://api.example.com/v1"
      api_key: "sk-..."
Enter fullscreen mode Exit fullscreen mode

LibreChat — a custom endpoint block in librechat.yaml.

Open WebUI — environment variables at container start, or Admin Settings → Connections.

SDKs

Python

from openai import OpenAI

client = OpenAI(
    base_url="https://api.example.com/v1",
    api_key="sk-...",
)
Enter fullscreen mode Exit fullscreen mode

Node

import OpenAI from 'openai';

const client = new OpenAI({
  baseURL: 'https://api.example.com/v1',
  apiKey: 'sk-...',
});
Enter fullscreen mode Exit fullscreen mode

The gotchas nobody documents

These five cost me the most time.

1. The /v1 suffix is inconsistent. Covered above. Include it by default and only drop it if you get a 404.

2. Some clients populate their model dropdown from /v1/models, and some don't. If your endpoint doesn't implement that route, or returns an empty list, the dropdown stays empty and you have to type the model id by hand. Cline, Continue, and Cherry Studio all allow manual entry. Some clients don't, and those will simply not work.

3. Test with streaming on. Most clients default to streaming responses. An endpoint can implement /v1/chat/completions correctly but not stream: true — in which case curl returns a nice response and the client just hangs. Always test streaming before you conclude the endpoint is broken.

4. Context length is guessed client-side. Clients infer the context window from the model name. Serve a model under a custom name and the client may assume 4k and silently truncate your prompts. Most clients let you override this in the model config — do it.

5. SillyTavern's hidden API key field. It exists, it's just collapsed until you reveal it. People miss it and then spend an hour debugging 401s.

Generating all of this at once

I got tired of writing these out, so I built a small generator: llm-endpoint-setup.

It came out of running an OpenAI-compatible gateway — haotogen — where I had to test against every one of these clients anyway.

You enter your base URL, key, and model id once. It produces the exact file or click path for each client above. There's a web version that runs entirely in your browser (the key never leaves the page), a CLI, and a library you can import:

npx llm-endpoint-setup --base-url https://api.example.com/v1 --model your-model --client cline
Enter fullscreen mode Exit fullscreen mode

MIT licensed, and adding a client is about 15 lines if yours is missing.

Wrapping up

The pattern never changes: a base URL, a key, and a model id. The only thing that varies is where each client wants them and what it calls them.

Check /v1/models first. Include /v1 unless you have a reason not to. Test with streaming on. Everything else is just finding the right text box.

Top comments (0)