<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: LaoDeng Learns AI</title>
    <description>The latest articles on DEV Community by LaoDeng Learns AI (@dafeiai918).</description>
    <link>https://dev.to/dafeiai918</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4148962%2Fd3ee37fc-430b-4a96-94ba-c290f4be14c2.png</url>
      <title>DEV Community: LaoDeng Learns AI</title>
      <link>https://dev.to/dafeiai918</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dafeiai918"/>
    <language>en</language>
    <item>
      <title>How to connect 12 different AI clients to any OpenAI-compatible endpoint</title>
      <dc:creator>LaoDeng Learns AI</dc:creator>
      <pubDate>Tue, 29 Sep 2026 08:01:39 +0000</pubDate>
      <link>https://dev.to/dafeiai918/how-to-connect-12-different-ai-clients-to-any-openai-compatible-endpoint-2hkf</link>
      <guid>https://dev.to/dafeiai918/how-to-connect-12-different-ai-clients-to-any-openai-compatible-endpoint-2hkf</guid>
      <description>&lt;p&gt;If you self-host a model, run an API gateway, or use a provider in your region, you've hit this wall: &lt;strong&gt;every AI client wants your endpoint entered differently.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Cline calls it &lt;code&gt;openAiBaseUrl&lt;/code&gt;. Continue calls it &lt;code&gt;apiBase&lt;/code&gt;. Aider reads &lt;code&gt;OPENAI_API_BASE&lt;/code&gt; from the environment. Cursor hides it behind a toggle. SillyTavern only shows the custom endpoint field &lt;em&gt;after&lt;/em&gt; you pick the right chat completion source.&lt;/p&gt;

&lt;p&gt;So you google "how to set base url in Cline". Then "how to set base url in Continue". Then you do it all again three months later because you forgot.&lt;/p&gt;

&lt;p&gt;This is the cheat sheet I wanted to exist. Twelve clients, plus the gotchas that cost me the most time.&lt;/p&gt;

&lt;h2&gt;
  
  
  First: what does "OpenAI-compatible" actually mean?
&lt;/h2&gt;

&lt;p&gt;An endpoint is OpenAI-compatible if it implements at least two routes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight http"&gt;&lt;code&gt;&lt;span class="err"&gt;POST /v1/chat/completions
GET  /v1/models
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's the whole bar. If both work, almost every client on this list will talk to it.&lt;/p&gt;

&lt;p&gt;The one question that trips everyone up: &lt;strong&gt;does the base URL include &lt;code&gt;/v1&lt;/code&gt;?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The OpenAI SDK convention is that &lt;code&gt;base_url&lt;/code&gt; includes it, and the SDK appends the rest:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;base_url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.example.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="err"&gt;→&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt; &lt;span class="n"&gt;go&lt;/span&gt; &lt;span class="n"&gt;to&lt;/span&gt; &lt;span class="n"&gt;https&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;//&lt;/span&gt;&lt;span class="n"&gt;api&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;example&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;com&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;v1&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Most clients follow that convention. Cursor is the notable exception — it appends the path itself, so what you enter depends on how your provider documents it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Rule of thumb: if in doubt, include &lt;code&gt;/v1&lt;/code&gt;.&lt;/strong&gt; If you get a 404, drop it.&lt;/p&gt;

&lt;p&gt;Test any endpoint before you touch a client:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.example.com/v1/models &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$YOUR_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A JSON list of models means you're compatible. A 401 means your key is wrong. A 404 means your base URL is wrong.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 12 clients
&lt;/h2&gt;

&lt;h3&gt;
  
  
  GUI apps
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;SillyTavern&lt;/strong&gt; — API: &lt;code&gt;Chat Completion&lt;/code&gt; → Chat Completion Source: &lt;code&gt;Custom (OpenAI-compatible)&lt;/code&gt;. The custom endpoint field only appears &lt;em&gt;after&lt;/em&gt; you select that source. The API key field is hidden behind a toggle.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cherry Studio&lt;/strong&gt; — Settings → Model Providers → Add Provider → type &lt;code&gt;OpenAI&lt;/code&gt;. Then use "Manage Models" to add the model id by hand; it does not auto-discover.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Jan&lt;/strong&gt; — Settings → Providers → Add Provider, then add the model id manually.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Cursor&lt;/strong&gt; — Cursor Settings → Models → enable "OpenAI API Key" → enable "Override OpenAI Base URL". Remember that Cursor appends the path itself.&lt;/p&gt;

&lt;h3&gt;
  
  
  VS Code extensions
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Cline&lt;/strong&gt; — three keys in your VS Code &lt;code&gt;settings.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cline.apiProvider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cline.openAiBaseUrl"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.example.com/v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cline.openAiApiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sk-..."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cline.openAiModelId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"your-model"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Continue&lt;/strong&gt; — &lt;code&gt;~/.continue/config.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"My Endpoint"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"your-model"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"apiBase"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.example.com/v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"apiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sk-..."&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  CLI
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Aider&lt;/strong&gt; — reads standard environment variables. Put this in &lt;code&gt;.env&lt;/code&gt; at your project root:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;OPENAI_API_BASE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;https://api.example.com/v1
&lt;span class="nv"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;sk-...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;then run:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;aider &lt;span class="nt"&gt;--model&lt;/span&gt; openai/your-model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Gateways and self-hosted UIs
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;LiteLLM&lt;/strong&gt; — &lt;code&gt;config.yaml&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;model_list&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;model_name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;my-endpoint&lt;/span&gt;
    &lt;span class="na"&gt;litellm_params&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai/your-model&lt;/span&gt;
      &lt;span class="na"&gt;api_base&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.example.com/v1"&lt;/span&gt;
      &lt;span class="na"&gt;api_key&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-..."&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;LibreChat&lt;/strong&gt; — a custom endpoint block in &lt;code&gt;librechat.yaml&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Open WebUI&lt;/strong&gt; — environment variables at container start, or Admin Settings → Connections.&lt;/p&gt;

&lt;h3&gt;
  
  
  SDKs
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Python&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.example.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Node&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;OpenAI&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://api.example.com/v1&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;sk-...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  The gotchas nobody documents
&lt;/h2&gt;

&lt;p&gt;These five cost me the most time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;1. The &lt;code&gt;/v1&lt;/code&gt; suffix is inconsistent.&lt;/strong&gt; Covered above. Include it by default and only drop it if you get a 404.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Some clients populate their model dropdown from &lt;code&gt;/v1/models&lt;/code&gt;, and some don't.&lt;/strong&gt; If your endpoint doesn't implement that route, or returns an empty list, the dropdown stays empty and you have to type the model id by hand. Cline, Continue, and Cherry Studio all allow manual entry. Some clients don't, and those will simply not work.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Test with streaming on.&lt;/strong&gt; Most clients default to streaming responses. An endpoint can implement &lt;code&gt;/v1/chat/completions&lt;/code&gt; correctly but not &lt;code&gt;stream: true&lt;/code&gt; — in which case curl returns a nice response and the client just hangs. Always test streaming before you conclude the endpoint is broken.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;4. Context length is guessed client-side.&lt;/strong&gt; Clients infer the context window from the model &lt;em&gt;name&lt;/em&gt;. Serve a model under a custom name and the client may assume 4k and silently truncate your prompts. Most clients let you override this in the model config — do it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;5. SillyTavern's hidden API key field.&lt;/strong&gt; It exists, it's just collapsed until you reveal it. People miss it and then spend an hour debugging 401s.&lt;/p&gt;

&lt;h2&gt;
  
  
  Generating all of this at once
&lt;/h2&gt;

&lt;p&gt;I got tired of writing these out, so I built a small generator: &lt;a href="https://github.com/dafeiai918/llm-endpoint-setup" rel="noopener noreferrer"&gt;llm-endpoint-setup&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;It came out of running an OpenAI-compatible gateway — &lt;a href="https://haotogen.com/" rel="noopener noreferrer"&gt;haotogen&lt;/a&gt; — where I had to test against every one of these clients anyway.&lt;/p&gt;

&lt;p&gt;You enter your base URL, key, and model id once. It produces the exact file or click path for each client above. There's a web version that runs entirely in your browser (the key never leaves the page), a CLI, and a library you can import:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx llm-endpoint-setup &lt;span class="nt"&gt;--base-url&lt;/span&gt; https://api.example.com/v1 &lt;span class="nt"&gt;--model&lt;/span&gt; your-model &lt;span class="nt"&gt;--client&lt;/span&gt; cline
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;MIT licensed, and adding a client is about 15 lines if yours is missing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;The pattern never changes: a base URL, a key, and a model id. The only thing that varies is where each client wants them and what it calls them.&lt;/p&gt;

&lt;p&gt;Check &lt;code&gt;/v1/models&lt;/code&gt; first. Include &lt;code&gt;/v1&lt;/code&gt; unless you have a reason not to. Test with streaming on. Everything else is just finding the right text box.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>opensource</category>
      <category>tutorial</category>
      <category>tooling</category>
    </item>
  </channel>
</rss>
