<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: lbase-novaapi</title>
    <description>The latest articles on DEV Community by lbase-novaapi (@lbase-novaapi).</description>
    <link>https://dev.to/lbase-novaapi</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4119778%2Fa2730a73-8e1b-46a6-a1bd-4cffcf69cf09.jpg</url>
      <title>DEV Community: lbase-novaapi</title>
      <link>https://dev.to/lbase-novaapi</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/lbase-novaapi"/>
    <language>en</language>
    <item>
      <title>One Endpoint, Your Whole Toolchain: Running Chinese Frontier Models in Claude Code, Cursor, Cline, Aider and More</title>
      <dc:creator>lbase-novaapi</dc:creator>
      <pubDate>Fri, 11 Sep 2026 21:00:39 +0000</pubDate>
      <link>https://dev.to/lbase-novaapi/one-endpoint-your-whole-toolchain-running-chinese-frontier-models-in-claude-code-cursor-cline-271k</link>
      <guid>https://dev.to/lbase-novaapi/one-endpoint-your-whole-toolchain-running-chinese-frontier-models-in-claude-code-cursor-cline-271k</guid>
      <description>&lt;p&gt;The awkward part of using a Chinese frontier model isn't the model. It's wiring the same model into the seven different tools you already have open.&lt;/p&gt;

&lt;p&gt;This is the copy-paste version of that setup, for every client I've configured over the last few months. One endpoint, many tools. Scroll to the one you use.&lt;/p&gt;

&lt;h2&gt;
  
  
  First: the part that matters (base URLs)
&lt;/h2&gt;

&lt;p&gt;Nearly every tool below asks for a "base URL". Getting this wrong is the #1 setup failure, and the rule is annoyingly inconsistent:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Client type&lt;/th&gt;
&lt;th&gt;What to enter&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;OpenAI-compatible clients (Cursor, Cline, LangChain, most GUIs)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;https://api.lbase.com/v1&lt;/code&gt; — &lt;strong&gt;with&lt;/strong&gt; &lt;code&gt;/v1&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Code / Anthropic SDK clients&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;https://api.lbase.com&lt;/code&gt; — &lt;strong&gt;without&lt;/strong&gt; &lt;code&gt;/v1&lt;/code&gt; (the SDK appends &lt;code&gt;/v1/messages&lt;/code&gt; itself)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Put &lt;code&gt;/v1&lt;/code&gt; on an Anthropic client and you'll get a 404 on &lt;code&gt;/v1/v1/messages&lt;/code&gt;. That's the whole bug 90% of the time.&lt;/p&gt;

&lt;h2&gt;
  
  
  Claude Code
&lt;/h2&gt;

&lt;p&gt;Two environment variables, no config file:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.lbase.com"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_AUTH_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-your-key"&lt;/span&gt;
claude
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Want a specific model? Claude Code respects the model env var too:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"deepseek-v4-flash"&lt;/span&gt;     &lt;span class="c"&gt;# or a claude-* name if your gateway maps them&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Tip:&lt;/strong&gt; keep a separate shell profile (or a small wrapper script) for the cheap profile and another for the flagship one. Switching is then one command, and you stop hand-editing configs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cursor
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;Settings → Models → OpenAI API Key&lt;/code&gt;, then add a custom model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Base URL&lt;/strong&gt;: &lt;code&gt;https://api.lbase.com/v1&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model name&lt;/strong&gt;: &lt;code&gt;deepseek-v4-flash&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Toggle off the models you don't want shown in the picker (Cursor lists a lot of noise by default)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cursor is chat/completion oriented; for long agentic runs, Claude Code or Cline handle tool-call loops better on non-Anthropic models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cline / Roo Code (VS Code)
&lt;/h2&gt;

&lt;p&gt;In the extension settings, choose &lt;strong&gt;OpenAI Compatible&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Base URL&lt;/strong&gt;: &lt;code&gt;https://api.lbase.com/v1&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API Key&lt;/strong&gt;: your &lt;code&gt;sk-...&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model ID&lt;/strong&gt;: &lt;code&gt;deepseek-v4-flash&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Cline's agentic loops are token-hungry. This is exactly where prompt caching pays off — the same system prompt and file context get re-sent every step, and on a cached-input rate the input side is nearly free.&lt;/p&gt;

&lt;h2&gt;
  
  
  Aider (terminal)
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENAI_API_BASE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.lbase.com/v1"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-your-key"&lt;/span&gt;
aider &lt;span class="nt"&gt;--model&lt;/span&gt; openai/deepseek-v4-flash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Aider also supports Anthropic-style endpoints if you prefer the Claude model aliases.&lt;/p&gt;

&lt;h2&gt;
  
  
  Continue.dev
&lt;/h2&gt;

&lt;p&gt;&lt;code&gt;~/.continue/config.json&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"models"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"title"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"DeepSeek V4.1 (NovaAPI)"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"provider"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"deepseek-v4-flash"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"apiBase"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://api.lbase.com/v1"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"apiKey"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sk-your-key"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Dify / FastGPT / LangChain
&lt;/h2&gt;

&lt;p&gt;All three want an OpenAI-compatible provider. Take any provider you're not using, or add a custom one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# LangChain
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;langchain_openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;ChatOpenAI&lt;/span&gt;

&lt;span class="n"&gt;llm&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;ChatOpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deepseek-v4-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.lbase.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-your-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For Dify: &lt;strong&gt;Settings → Model Provider → OpenAI-API-compatible&lt;/strong&gt;, paste the base URL and key, then add &lt;code&gt;deepseek-v4-flash&lt;/code&gt; as a model name.&lt;/p&gt;

&lt;h2&gt;
  
  
  Desktop GUIs (Cherry Studio, ChatBox, Open WebUI)
&lt;/h2&gt;

&lt;p&gt;Same pattern every time: pick "OpenAI" as the provider type, then override the API host with &lt;code&gt;https://api.lbase.com/v1&lt;/code&gt;, paste the key, and fetch the model list (they'll pull whatever the endpoint exposes).&lt;/p&gt;

&lt;p&gt;One caveat with reasoning models: some GUIs render the model's &lt;code&gt;reasoning_content&lt;/code&gt; field inline. That's a client rendering quirk, not a broken response — worth knowing before you file a bug.&lt;/p&gt;

&lt;h2&gt;
  
  
  The three gotchas that cost the most time
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;1. Model names change under you.&lt;/strong&gt; Providers rename models (DeepSeek turned &lt;code&gt;deepseek-v4-flash&lt;/code&gt; into &lt;code&gt;deepseek-flash&lt;/code&gt; the day V4.1 shipped) and start auto-routing traffic between tiers. If a tool hardcodes an old name, it breaks silently. Gateways help here because the mapping lives server-side: your client keeps sending the name it always sent.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;2. Peak/off-peak pricing is real.&lt;/strong&gt; If your provider bills peak windows (DeepSeek's are Mon–Fri 01:00–04:00 and 06:00–10:00 UTC, at 2x), then "the same request" costs double depending on when a batch runs. Schedule overnight jobs outside those windows; on a large batch that's a straight 50% saving with no code changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;3. Caching is the biggest lever nobody configures.&lt;/strong&gt; Cached-input rates are an order of magnitude below cache-miss rates. Agentic tools re-send a stable prefix constantly, so enabling caching changes your bill more than switching models does.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to verify a setup in 30 seconds
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.lbase.com/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-your-key"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{"model":"deepseek-v4-flash","messages":[{"role":"user","content":"reply with OK"}],"max_tokens":5}'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A 200 with &lt;code&gt;"choices"&lt;/code&gt; means your endpoint, key and model name are all fine — any client problem after that is client-side config, not the gateway.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Full disclosure: I build &lt;a href="https://novaapi.lbase.com" rel="noopener noreferrer"&gt;NovaAPI&lt;/a&gt;, the gateway used in these examples — OpenAI/Anthropic-compatible access to Chinese frontier models, PayPal/USDT billing. Every snippet above works with any compatible gateway; the base-URL rules and gotchas are the same regardless of which one you pick.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>tutorial</category>
      <category>claude</category>
      <category>api</category>
    </item>
    <item>
      <title>DeepSeek V4.1 Flash Is Here — and You Can Use It From Claude Code Today</title>
      <dc:creator>lbase-novaapi</dc:creator>
      <pubDate>Thu, 10 Sep 2026 19:26:30 +0000</pubDate>
      <link>https://dev.to/lbase-novaapi/deepseek-v41-flash-is-here-and-you-can-use-it-from-claude-code-today-bn4</link>
      <guid>https://dev.to/lbase-novaapi/deepseek-v41-flash-is-here-and-you-can-use-it-from-claude-code-today-bn4</guid>
      <description>&lt;p&gt;If you follow Chinese AI at all, you know DeepSeek's pricing announcements have been a wild ride this quarter: peak/off-peak billing in August, a big Flash price cut in September, and now — literally today — &lt;strong&gt;V4.1 Flash&lt;/strong&gt; is rolling out, with DeepSeek routing all &lt;code&gt;deepseek-v4-pro&lt;/code&gt; traffic to the new model at the new (lower) rate.&lt;/p&gt;

&lt;p&gt;Here's the short version of today's news, and then the part that matters for developers outside China: &lt;strong&gt;how to actually use it from your existing tools.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed today
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;V4.1 Flash&lt;/strong&gt; officially launches (2026-09-10, 12:00 Beijing time).&lt;/li&gt;
&lt;li&gt;Internally tested to beat &lt;strong&gt;V4 Pro&lt;/strong&gt; on performance, speed, cost, and total response time.&lt;/li&gt;
&lt;li&gt;Until V4.1 Pro ships, requests to &lt;code&gt;deepseek-v4-pro&lt;/code&gt; are &lt;strong&gt;auto-routed to V4.1 Flash&lt;/strong&gt; and billed at V4.1 Flash rates.&lt;/li&gt;
&lt;li&gt;New pricing (per 1M tokens, off-peak / peak): input &lt;strong&gt;¥1 / ¥2&lt;/strong&gt;, output &lt;strong&gt;¥4 / ¥8&lt;/strong&gt; — roughly &lt;strong&gt;30% cheaper than V4 Flash&lt;/strong&gt;, and far below V4 Pro's old ¥4.5 / ¥13.5.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That's a serious price-to-performance jump. The models are frontier-level, the context window is 1M tokens, and the API speaks both OpenAI &lt;em&gt;and&lt;/em&gt; Anthropic dialects natively.&lt;/p&gt;

&lt;h2&gt;
  
  
  The one annoying thing
&lt;/h2&gt;

&lt;p&gt;DeepSeek's own platform is built for the domestic market. If you're outside China, signing up means dealing with a &lt;strong&gt;Chinese phone number&lt;/strong&gt; for verification, Chinese payment rails, and a docs experience that assumes you live there. A lot of us just want to point our existing clients at the model and go.&lt;/p&gt;

&lt;p&gt;That's exactly the gap I built &lt;a href="https://novaapi.lbase.com" rel="noopener noreferrer"&gt;NovaAPI&lt;/a&gt; for: a gateway that gives you the &lt;strong&gt;same DeepSeek models (V4.1 line included)&lt;/strong&gt; through a normal international flow — sign up with any email, pay with &lt;strong&gt;PayPal or USDT&lt;/strong&gt;, no Chinese phone number anywhere. It also aggregates Claude-compatible endpoints, so Claude Code talks to it natively. (Full disclosure: this is my project — but the tutorial below works for any OpenAI/Anthropic-compatible gateway.)&lt;/p&gt;

&lt;h2&gt;
  
  
  5-minute setup: DeepSeek V4.1 from Claude Code
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Step 1 — Create an account and a key&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Sign up at &lt;a href="https://api.lbase.com/register" rel="noopener noreferrer"&gt;api.lbase.com/register&lt;/a&gt; (any email works).&lt;/li&gt;
&lt;li&gt;Top up with PayPal or USDT — the minimum is small, just enough to try it.&lt;/li&gt;
&lt;li&gt;Go to &lt;strong&gt;API Keys&lt;/strong&gt; and create a key. Pick the &lt;strong&gt;DeepSeek&lt;/strong&gt; channel when asked.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;You'll get something like &lt;code&gt;sk-...&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 2 — Point Claude Code at it&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Claude Code respects the &lt;code&gt;ANTHROPIC_BASE_URL&lt;/code&gt; and &lt;code&gt;ANTHROPIC_AUTH_TOKEN&lt;/code&gt; environment variables:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.lbase.com"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_AUTH_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-your-novaapi-key"&lt;/span&gt;
claude
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That's it. Claude Code will use the Anthropic-compatible endpoint, and NovaAPI maps Claude model names (&lt;code&gt;claude-sonnet-4-6&lt;/code&gt;, &lt;code&gt;claude-fable-5&lt;/code&gt;, etc.) onto the DeepSeek V4.1 line automatically.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 3 — Or use any OpenAI-compatible client (Cursor, Cherry Studio, code)&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-your-novaapi-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.lbase.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deepseek-v4-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# serving V4.1 Flash as of today
&lt;/span&gt;    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write a quick Python snippet for a retry-with-backoff helper.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cursor: Settings → Models → OpenAI-compatible → Base URL &lt;code&gt;https://api.lbase.com/v1&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Step 4 — Verify&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.lbase.com/v1/models &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer sk-your-novaapi-key"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;You'll see the model list. Run a couple of prompts and check your usage page — billing is transparent, pay-as-you-go, in USD.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why not just call DeepSeek directly?
&lt;/h2&gt;

&lt;p&gt;If you already have a Chinese phone number and a way to pay in CNY, absolutely go direct — it'll be marginally cheaper. NovaAPI is for everyone else, and for people who want:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;One key for multiple Chinese models&lt;/strong&gt; (DeepSeek today, more families coming),&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code / Anthropic-SDK compatibility&lt;/strong&gt; without shims,&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PayPal/USDT billing&lt;/strong&gt; instead of Chinese payment apps,&lt;/li&gt;
&lt;li&gt;No phone verification, no VPN, no WeChat.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  A note on pricing and peak hours
&lt;/h2&gt;

&lt;p&gt;DeepSeek bills peak/off-peak (peak = Mon–Fri 01:00–04:00 &amp;amp; 06:00–10:00 UTC). NovaAPI passes the official &lt;strong&gt;off-peak price × a transparent multiplier&lt;/strong&gt; to you, and like DeepSeek, peak-hour requests cost more. The price table on &lt;a href="https://novaapi.lbase.com" rel="noopener noreferrer"&gt;novaapi.lbase.com&lt;/a&gt; is always current — check it before you commit a big batch job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Wrapping up
&lt;/h2&gt;

&lt;p&gt;V4.1 Flash looks like the best price-to-performance point in the DeepSeek lineup right now, and with the Pro auto-routing in place, "deepseek-v4-flash" is secretly a very fast, very cheap flagship. If you've been wanting to try Chinese frontier models but bounced off the phone-number wall — the door is open.&lt;/p&gt;

&lt;p&gt;Questions, corrections, or war stories from wiring this into your stack? Drop a comment. I read all of them.&lt;/p&gt;

</description>
      <category>ai</category>
      <category>deepseek</category>
      <category>tutorial</category>
      <category>claude</category>
    </item>
    <item>
      <title>How to Cut Your LLM API Bill by ~80% Without Switching Your Tools</title>
      <dc:creator>lbase-novaapi</dc:creator>
      <pubDate>Thu, 10 Sep 2026 19:12:46 +0000</pubDate>
      <link>https://dev.to/lbase-novaapi/how-to-cut-your-llm-api-bill-by-80-without-switching-your-tools-29b4</link>
      <guid>https://dev.to/lbase-novaapi/how-to-cut-your-llm-api-bill-by-80-without-switching-your-tools-29b4</guid>
      <description>&lt;p&gt;I run Claude Code all day. For a while I assumed the API cost was just the price of doing business — the tooling is excellent, the model is excellent, and the bill is… the bill.&lt;/p&gt;

&lt;p&gt;Then I actually did the arithmetic on what a coding assistant costs at frontier-model prices, and it turned out I was paying a premium for capability I don't always need. Here's the breakdown, with real numbers, and where I landed.&lt;/p&gt;

&lt;h2&gt;
  
  
  The price gap, in one table
&lt;/h2&gt;

&lt;p&gt;Published list prices, per 1M tokens (USD):&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Anthropic Claude Sonnet (list)&lt;/td&gt;
&lt;td&gt;$3.00&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4.1 Flash (off-peak)&lt;/td&gt;
&lt;td&gt;$0.22&lt;/td&gt;
&lt;td&gt;$0.66&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4.1 Flash (peak)&lt;/td&gt;
&lt;td&gt;$0.44&lt;/td&gt;
&lt;td&gt;$1.32&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That's roughly &lt;strong&gt;13–14× cheaper on input and 11–23× cheaper on output&lt;/strong&gt;, depending on the hour. DeepSeek bills peak/off-peak (peak = Mon–Fri 01:00–04:00 and 06:00–10:00 UTC), and off-peak is half price — so batching heavy jobs into off-peak hours cuts the bill in half again.&lt;/p&gt;

&lt;h2&gt;
  
  
  What that means for a real workload
&lt;/h2&gt;

&lt;p&gt;Say you're a solo developer running a coding assistant for a month:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;5M input tokens&lt;/strong&gt; (system prompts, file context, diffs that keep getting re-sent)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;1M output tokens&lt;/strong&gt; (the model's actual work)&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Setup&lt;/th&gt;
&lt;th&gt;Monthly cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet at list price&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$30.00&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude-compatible endpoint on a Chinese frontier model (e.g. NovaAPI, Sonnet tier)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$6.16&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Native DeepSeek V4.1 Flash, off-peak&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$1.76&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Same tools. Same &lt;code&gt;ANTHROPIC_BASE_URL&lt;/code&gt; env var. Same request shape. &lt;strong&gt;~80% cheaper&lt;/strong&gt; on the managed route, ~94% cheaper if you go direct and stay off-peak.&lt;/p&gt;

&lt;p&gt;Scale that to a small team (50M in / 10M out per month) and the difference is ~$300/month vs ~$62/month.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why is it so much cheaper?
&lt;/h2&gt;

&lt;p&gt;Three reasons, none of them magic:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Chinese lab pricing is aggressive by design.&lt;/strong&gt; DeepSeek in particular has been cutting Flash-series prices this quarter (an August restructure, a September cut, and now V4.1 Flash, which the vendor says beats their own V4 Pro while costing less).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Efficiency at the architecture level.&lt;/strong&gt; These are MoE-style models with aggressive KV-cache and context handling. The token price reflects a cheaper serving stack, not a cheaper product decision.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Prompt caching is nearly free.&lt;/strong&gt; DeepSeek's cached-input rate is $0.007/1M off-peak — effectively nothing. If your client re-sends a stable prefix (Claude Code does), the effective input cost collapses.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  The honest caveats
&lt;/h2&gt;

&lt;p&gt;I'm not going to pretend this is a free lunch:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agentic, long-horizon tasks still favor the Western flagships.&lt;/strong&gt; If your workflow is "let the model run 40 tool calls and recover from its own mistakes," Claude Opus / GPT-5.x still earn their price. Test on &lt;em&gt;your&lt;/em&gt; workload before migrating anything critical.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality is workload-dependent.&lt;/strong&gt; Benchmarks are vendor claims. V4.1 Flash is genuinely strong at code, summarization, and long-context work in my testing; for subtle reasoning chains, verify before you trust.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Peak-hour pricing is real.&lt;/strong&gt; If your workload runs during Chinese business hours (UTC 01–04, 06–10), you pay 2×.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data path matters.&lt;/strong&gt; Your prompts go to a Chinese model provider. Don't send anything you wouldn't send to any third-party API.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The setup that actually works
&lt;/h2&gt;

&lt;p&gt;The practical trick is that &lt;strong&gt;you don't have to choose one provider.&lt;/strong&gt; Because the clients we already use speak configurable endpoints, you can point them anywhere:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Claude Code, pointed at an Anthropic-compatible endpoint&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_BASE_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"https://api.lbase.com"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;ANTHROPIC_AUTH_TOKEN&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-your-key"&lt;/span&gt;
claude
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# Or plain OpenAI SDK
&lt;/span&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-your-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.lbase.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;deepseek-v4-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;   &lt;span class="c1"&gt;# served by V4.1 Flash today
&lt;/span&gt;    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Refactor this function for readability.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then use whichever profile makes sense per task: cheap-and-fast for refactors, summaries, test generation and bulk jobs; flagship for hairy architecture work. I keep two shells open with different env vars and that's the whole system.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I actually spend now
&lt;/h2&gt;

&lt;p&gt;Real numbers from my own usage this month, running a coding assistant and a couple of batch jobs through a Chinese frontier model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A full evening of interactive coding: &lt;strong&gt;23 requests, 18,174 tokens, ~$0.008&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;A single complex Claude-compatible request: &lt;strong&gt;~$0.003&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;A 10k-document summarization batch: &lt;strong&gt;a few cents&lt;/strong&gt;, off-peak&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Compare that to a month of Sonnet at list price and the gap is not marginal — it's the difference between "careful with tokens" and "stop thinking about tokens."&lt;/p&gt;

&lt;h2&gt;
  
  
  When to switch, when not to
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Switch (or add a second profile) if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;You're cost-sensitive, or you run high-volume batch/summarization work&lt;/li&gt;
&lt;li&gt;Your tasks are code edits, refactors, tests, docs, extraction — not long autonomous chains&lt;/li&gt;
&lt;li&gt;You want to experiment widely without watching a meter&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Don't switch if:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your product depends on the strongest available reasoning under long tool-use loops&lt;/li&gt;
&lt;li&gt;You need a specific vendor's compliance posture or data residency guarantees&lt;/li&gt;
&lt;li&gt;You can't tolerate a third-party gateway between you and the model&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Most developers I know end up in the middle: &lt;strong&gt;flagship for the hard 10%, cheap frontier models for the other 90%.&lt;/strong&gt; That split is where the 80% number comes from — it isn't a magic trick, just arithmetic plus tooling that lets you route per task.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Full disclosure: I built &lt;a href="https://novaapi.lbase.com" rel="noopener noreferrer"&gt;NovaAPI&lt;/a&gt;, an OpenAI/Anthropic-compatible gateway for Chinese frontier models (PayPal/USDT billing, no Chinese phone number needed). The prices above are DeepSeek's published rates and Anthropic's list rates — check both before you commit to a migration. Questions about specific workloads? Ask in the comments and I'll run the numbers.&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>claude</category>
      <category>tutorial</category>
      <category>api</category>
    </item>
  </channel>
</rss>
