<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hassann</title>
    <description>The latest articles on DEV Community by Hassann (@hassann).</description>
    <link>https://dev.to/hassann</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3890506%2F89a141f2-4995-48b3-b5f2-e00ba5055afb.png</url>
      <title>DEV Community: Hassann</title>
      <link>https://dev.to/hassann</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hassann"/>
    <language>en</language>
    <item>
      <title>How to Use Claude for Free in 2026: Every Option That Works</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Tue, 25 Aug 2026 08:06:07 +0000</pubDate>
      <link>https://dev.to/hassann/how-to-use-claude-for-free-in-2026-every-option-that-works-58j4</link>
      <guid>https://dev.to/hassann/how-to-use-claude-for-free-in-2026-every-option-that-works-58j4</guid>
      <description>&lt;p&gt;Claude is free to use in 2026. The permanent free plan at &lt;a href="https://claude.ai" rel="noopener noreferrer"&gt;claude.ai&lt;/a&gt; includes Claude Sonnet 5—one of Anthropic’s strongest models—without a trial period or credit card. Free options also exist for Claude Code and API testing with &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This guide covers the practical free routes: the official plan, Claude Code guest passes, API starter credits, third-party tools, and cloud-provider trials. For deciding whether to upgrade later, see our &lt;a href="http://apidog.com/blog/claude-free-vs-pro?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Free vs Pro comparison&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is Claude free?
&lt;/h2&gt;

&lt;p&gt;Yes. The free plan at &lt;a href="http://claude.ai" rel="noopener noreferrer"&gt;claude.ai&lt;/a&gt; never expires and requires no credit card. As of August 2026, it includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Sonnet 5, the default free model since July 2026&lt;/li&gt;
&lt;li&gt;Claude Haiku 4.5 for faster, lighter responses&lt;/li&gt;
&lt;li&gt;Web search, file uploads, and image inputs&lt;/li&gt;
&lt;li&gt;Web, desktop, and mobile access&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It does &lt;strong&gt;not&lt;/strong&gt; include Claude Fable 5 or other Opus-tier models. Anthropic’s &lt;a href="https://support.claude.com/en/articles/15424964-claude-fable-5-on-your-plan" rel="noopener noreferrer"&gt;help center&lt;/a&gt; lists Fable 5 as paid-plan only: Pro, Max, Team, and Enterprise.&lt;/p&gt;

&lt;p&gt;A short free Fable 5 promotion ran in June 2026 and ended on July 19, 2026. See &lt;a href="http://apidog.com/blog/how-to-use-claude-fable-5-for-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;How to Use Claude Fable 5 for Free&lt;/a&gt; for the details.&lt;/p&gt;

&lt;p&gt;Free usage is metered. Anthropic does not publish fixed limits, but community testing suggests roughly 15–40 messages per rolling five-hour window. Limits can shrink during peak demand because paid accounts receive priority.&lt;/p&gt;

&lt;p&gt;Use the free plan for daily questions, document work, and short coding sessions—not extended back-and-forth agent runs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Option 1: Use the &lt;a href="http://claude.ai" rel="noopener noreferrer"&gt;claude.ai&lt;/a&gt; free plan
&lt;/h2&gt;

&lt;p&gt;This is the simplest path:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Go to &lt;a href="https://claude.ai" rel="noopener noreferrer"&gt;claude.ai&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Sign up with email or Google.&lt;/li&gt;
&lt;li&gt;Start using Sonnet 5.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Free Sonnet 5 is the same model available to paying users, not a reduced version. It handles code generation, document analysis, and multi-step reasoning well; most free users hit quota limits before quality limits.&lt;/p&gt;

&lt;p&gt;For more detail, read our guide to &lt;a href="http://apidog.com/blog/claude-sonnet-5-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;using Claude Sonnet 5 for free&lt;/a&gt; and our &lt;a href="http://apidog.com/blog/claude-sonnet-5-benchmarks?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Sonnet 5 benchmarks breakdown&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Stretch your free quota
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Front-load context:&lt;/strong&gt; Long threads consume more capacity because Claude reprocesses the conversation. Start a fresh chat for a new task.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use Haiku 4.5 for quick questions:&lt;/strong&gt; Select it in the chat composer when you need speed rather than deeper reasoning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Watch the rolling window:&lt;/strong&gt; When you reach the cap, Claude tells you when access resets.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Option 2: Get Claude Code through guest passes
&lt;/h2&gt;

&lt;p&gt;Claude Code is not included in the standard free plan, but guest passes can provide free temporary access.&lt;/p&gt;

&lt;p&gt;Anthropic gives Max subscribers a limited number of passes. Each pass can give a new user seven days of Pro-level access, including Claude Code. Ask a colleague with a Max subscription to run &lt;code&gt;/passes&lt;/code&gt; in Claude Code, or look for trusted developer-community giveaways.&lt;/p&gt;

&lt;p&gt;Keep these constraints in mind:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Passes are typically for new accounts.&lt;/li&gt;
&lt;li&gt;Reports indicate a card may be required at redemption, even though the seven-day period is free.&lt;/li&gt;
&lt;li&gt;After expiration, the account returns to the free plan rather than automatically upgrading to Pro—but always review the current redemption terms.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For other temporary routes, see &lt;a href="http://apidog.com/blog/use-claude-code-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;a proven way to use Claude Code for free&lt;/a&gt; and &lt;a href="http://apidog.com/blog/claude-code-free-credits?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;how to claim promotional Claude Code credits&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Option 3: Use free API credits with Apidog
&lt;/h2&gt;

&lt;p&gt;New accounts at &lt;a href="https://platform.claude.com" rel="noopener noreferrer"&gt;platform.claude.com&lt;/a&gt; receive a small starter credit balance. Third-party reports place it at about $5—enough to experiment with Sonnet 5 or Haiku 4.5 through the API.&lt;/p&gt;

&lt;p&gt;For all available credit options, read &lt;a href="http://apidog.com/blog/free-claude-api-access?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;free Claude API access&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Starter credits disappear quickly if you spend them debugging malformed requests. Use &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; to minimize avoidable API usage:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Import the Anthropic API specification&lt;/strong&gt; so endpoints, headers, and request schemas are ready to use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create one canonical &lt;code&gt;/v1/messages&lt;/code&gt; request&lt;/strong&gt; for each use case.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Inspect and save the response&lt;/strong&gt; as an example.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Create a local mock server&lt;/strong&gt; from that saved response.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Develop against the mock&lt;/strong&gt;, then call the real API only for final verification.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This approach can reduce development API spend by an order of magnitude. See &lt;a href="http://apidog.com/blog/cut-claude-api-bill?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;How to Cut Your Claude API Bill&lt;/a&gt;. When you move to production, &lt;a href="http://apidog.com/blog/what-is-prompt-caching?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;prompt caching&lt;/a&gt; can reduce costs further.&lt;/p&gt;

&lt;h2&gt;
  
  
  Option 4: Access Claude through tools you already use
&lt;/h2&gt;

&lt;p&gt;Several products expose Claude models through free or limited tiers, sometimes without requiring an Anthropic account:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;GitHub Copilot:&lt;/strong&gt; Offers Claude as a selectable model through its subscription and limited free tier.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Poe, DuckDuckGo AI Chat, and Brave Leo:&lt;/strong&gt; Provide Claude variants with small daily allowances.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cursor and Windsurf:&lt;/strong&gt; Include Claude among model options during trials.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These routes are useful for occasional usage or a short evaluation sprint. For Cursor, follow our guide on &lt;a href="http://apidog.com/blog/claude-sonnet-5-cursor?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;using Claude Sonnet 5 in Cursor&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Verify the exact model before relying on an integration. Third-party products may route requests to older or smaller Claude models and apply different system prompts than &lt;a href="http://claude.ai" rel="noopener noreferrer"&gt;claude.ai&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Option 5: Use cloud-provider trial credits
&lt;/h2&gt;

&lt;p&gt;AWS Bedrock, Google Vertex AI, and Azure AI Foundry all provide Claude models. Each cloud offers trial credits to new accounts, and existing startup or education credits may also cover usage.&lt;/p&gt;

&lt;p&gt;Setup varies by platform. Our guide to &lt;a href="http://apidog.com/blog/claude-fable-5-cloud-availability?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;running Claude Fable 5 on Bedrock, Vertex, and Foundry&lt;/a&gt; explains the setup; the same process applies to Sonnet 5.&lt;/p&gt;

&lt;p&gt;This requires more setup than the consumer free plan, but it is currently the only free route to Fable 5-class models through trial or existing cloud credits.&lt;/p&gt;

&lt;h2&gt;
  
  
  What free Claude does not include
&lt;/h2&gt;

&lt;p&gt;Know the limits before planning a workflow around the free tier:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No Fable 5:&lt;/strong&gt; It is available only on paid plans. On Pro, it uses pay-as-you-go credits rather than the flat subscription allowance. See &lt;a href="http://apidog.com/blog/how-to-access-claude-fable-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;How to Access Claude Fable 5&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Lower peak-time priority:&lt;/strong&gt; Free capacity can shrink when demand is high.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tighter session limits:&lt;/strong&gt; Large projects and long-running agent workflows will reach limits sooner. Read &lt;a href="http://apidog.com/blog/claude-fable-5-rate-limits?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Fable 5 Rate Limits Explained&lt;/a&gt; for how limits differ by tier.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you consistently hit the limit every week, the $20 Pro plan may be worthwhile. Until then, the free options above cover most individual use cases.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Which model does the free Claude plan use in 2026?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Claude Sonnet 5 is the default model, with Haiku 4.5 available for faster replies. See the &lt;a href="http://apidog.com/blog/claude-sonnet-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;full Sonnet 5 guide&lt;/a&gt; for capabilities and API pricing.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How many messages do free users get?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Anthropic does not publish a fixed number. Community estimates suggest roughly 15–40 messages in a rolling five-hour window, depending on server load and message size.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I use the Claude API for free?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
New &lt;a href="http://platform.claude.com" rel="noopener noreferrer"&gt;platform.claude.com&lt;/a&gt; accounts receive starter credit, and Bedrock, Vertex, and Foundry trial credits can also fund API usage. Use Apidog mock servers so development traffic does not consume real API credits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Claude Code free?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Not through the standard free plan. Guest passes can provide seven free days, Anthropic occasionally offers promotional credits, and starter API credit can fund a few short sessions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start free, test smart
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="http://claude.ai" rel="noopener noreferrer"&gt;claude.ai&lt;/a&gt; free plan is the easiest no-cost way to use a frontier model in 2026: Sonnet 5, no card, and no expiry. Add a guest pass to evaluate Claude Code, and use API starter credits for integration tests.&lt;/p&gt;

&lt;p&gt;Put &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; between your application and the API so mock servers handle development traffic and your free credits go toward calls that matter. &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; for free, then upgrade only when your actual usage requires it.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to Use NotebookLM in 2026: A Practical Guide to Google's Free Research Tool</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Tue, 25 Aug 2026 07:34:57 +0000</pubDate>
      <link>https://dev.to/hassann/how-to-use-notebooklm-in-2026-a-practical-guide-to-googles-free-research-tool-417k</link>
      <guid>https://dev.to/hassann/how-to-use-notebooklm-in-2026-a-practical-guide-to-googles-free-research-tool-417k</guid>
      <description>&lt;h1&gt;
  
  
  NotebookLM in 2026: A Practical Guide for Research and API Documentation
&lt;/h1&gt;

&lt;p&gt;NotebookLM turns a pile of documents into something you can question. Upload PDFs, docs, links, and videos, then chat with an AI that answers only from those sources—with clickable citations. In 2026, it can also generate podcast-style audio, narrated video explainers, and mind maps. For more on its research capabilities, see our &lt;a href="http://apidog.com/blog/notebooklm-deep-research?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Deep Research breakdown&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This guide covers setup, source limits, useful features, paid tiers, and a developer workflow for working with API documentation. It also reflects the product as of August 2026, based on &lt;a href="https://support.google.com/notebooklm/answer/16213268" rel="noopener noreferrer"&gt;Google’s official support documentation&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is NotebookLM, and is it free?
&lt;/h2&gt;

&lt;p&gt;NotebookLM is Google’s source-grounded AI research assistant. The standard tier is free and includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;100 notebooks&lt;/li&gt;
&lt;li&gt;50 sources per notebook&lt;/li&gt;
&lt;li&gt;50 chat queries per day&lt;/li&gt;
&lt;li&gt;3 Audio Overviews per day&lt;/li&gt;
&lt;li&gt;3 Video Overviews per day&lt;/li&gt;
&lt;li&gt;10 daily generations of reports, flashcards, quizzes, and mind maps&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You need a Google account and a browser, or the mobile app for &lt;a href="https://play.google.com/store/apps/details?id=com.google.android.apps.labs.language.tailwind" rel="noopener noreferrer"&gt;Android&lt;/a&gt; or iOS.&lt;/p&gt;

&lt;p&gt;“Source-grounded” is the important distinction. A general chatbot may rely on training data and the open web, which can lead to outdated or fabricated answers. NotebookLM answers from your uploaded sources and cites each claim. If the sources do not answer a question, it says so instead of guessing.&lt;/p&gt;

&lt;p&gt;Google has been folding NotebookLM into the Gemini brand during 2026. Official support pages now call it Gemini Notebook, but it remains the same product with the same limits. The app is still available at &lt;a href="https://notebooklm.google.com" rel="noopener noreferrer"&gt;notebooklm.google.com&lt;/a&gt;, and most users still call it NotebookLM.&lt;/p&gt;

&lt;h2&gt;
  
  
  Set up your first notebook
&lt;/h2&gt;

&lt;p&gt;You can create a useful notebook in about two minutes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open &lt;a href="https://notebooklm.google.com" rel="noopener noreferrer"&gt;notebooklm.google.com&lt;/a&gt; and sign in. There is no waitlist or credit card requirement.&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;Create new notebook&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Give the notebook one focused purpose: a research topic, course, API, deal, or project. NotebookLM answers from everything in the notebook, so unrelated sources can reduce answer quality.&lt;/li&gt;
&lt;li&gt;Add three or four initial sources.&lt;/li&gt;
&lt;li&gt;Wait for NotebookLM to process and index them. Processing takes seconds to a few minutes, depending on file size.&lt;/li&gt;
&lt;li&gt;Ask a specific question in the chat panel.&lt;/li&gt;
&lt;li&gt;Check the citation numbers and click them to open the relevant passage in the source.&lt;/li&gt;
&lt;li&gt;Pin useful answers to the Studio panel as notes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The core workflow is simple: add sources, ask questions, verify citations, and save useful answers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add sources: supported formats and limits
&lt;/h2&gt;

&lt;p&gt;Source quality determines answer quality, so keep these boundaries in mind.&lt;/p&gt;

&lt;h3&gt;
  
  
  Supported sources
&lt;/h3&gt;

&lt;p&gt;NotebookLM accepts:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;PDFs and text files, including scanned documents&lt;/li&gt;
&lt;li&gt;Google Docs and Google Slides from Drive&lt;/li&gt;
&lt;li&gt;Public website URLs&lt;/li&gt;
&lt;li&gt;YouTube videos with captions&lt;/li&gt;
&lt;li&gt;Audio files, such as lectures and meeting recordings&lt;/li&gt;
&lt;li&gt;Markdown files&lt;/li&gt;
&lt;li&gt;Pasted text&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For YouTube, NotebookLM reads the transcript rather than the video pixels.&lt;/p&gt;

&lt;h3&gt;
  
  
  Limits and common problems
&lt;/h3&gt;

&lt;p&gt;The free plan supports 50 sources per notebook. Individual sources have long carried a limit of roughly 500,000 words or 200 MB per file. That is enough for a 900-page PDF, and usually more than enough for a typical project.&lt;/p&gt;

&lt;p&gt;Watch for these issues:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Paywalled and login-protected pages cannot be scraped. Export them as PDFs first.&lt;/li&gt;
&lt;li&gt;YouTube videos without captions do not provide NotebookLM with readable content.&lt;/li&gt;
&lt;li&gt;Google Docs do not sync automatically. Use the sync button on the source when the underlying document changes.&lt;/li&gt;
&lt;li&gt;Mixing unrelated material in one notebook can make answers less focused.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Features worth using
&lt;/h2&gt;

&lt;p&gt;NotebookLM includes many features, but four are especially practical.&lt;/p&gt;

&lt;h3&gt;
  
  
  Audio Overviews
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://blog.google/technology/ai/notebooklm-audio-overviews/" rel="noopener noreferrer"&gt;Audio Overviews&lt;/a&gt; turn your sources into a podcast-style discussion between two AI hosts. Interactive mode lets you interrupt the episode and ask questions. Before generating an overview, you can focus it on a specific chapter or skip background material.&lt;/p&gt;

&lt;p&gt;Free users get three Audio Overviews per day.&lt;/p&gt;

&lt;h3&gt;
  
  
  Video Overviews
&lt;/h3&gt;

&lt;p&gt;Video Overviews create narrated, slide-style explainers from your sources and offer selectable visual styles. They take longer to generate than audio, but work well when you need to explain material to someone who will not read the original documents.&lt;/p&gt;

&lt;p&gt;The free tier includes three Video Overviews per day.&lt;/p&gt;

&lt;h3&gt;
  
  
  Mind maps
&lt;/h3&gt;

&lt;p&gt;Mind maps generate a branching diagram of the concepts across your sources. They are useful for quickly surveying an unfamiliar topic:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Generate a mind map.&lt;/li&gt;
&lt;li&gt;Find a branch you do not understand.&lt;/li&gt;
&lt;li&gt;Open it and ask a focused question.&lt;/li&gt;
&lt;li&gt;Follow the citations back to the source.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Reports, flashcards, and quizzes
&lt;/h3&gt;

&lt;p&gt;The Studio panel can turn sources into briefing documents, study guides, flashcards, and quizzes. Students can use these for revision, while teams can generate a briefing document before a meeting on an unfamiliar topic.&lt;/p&gt;

&lt;p&gt;These features share a pool of 10 generations per day on the free plan.&lt;/p&gt;

&lt;p&gt;In 2026, all tiers use Gemini 3 models. Paying for a higher tier increases usage limits; it does not provide a smarter NotebookLM model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Free vs. paid NotebookLM tiers
&lt;/h2&gt;

&lt;p&gt;Google offers the upper NotebookLM tiers through Google AI subscriptions rather than as a standalone NotebookLM plan. The main difference is volume:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Limit&lt;/th&gt;
&lt;th&gt;Free&lt;/th&gt;
&lt;th&gt;Google AI Plus&lt;/th&gt;
&lt;th&gt;Google AI Pro&lt;/th&gt;
&lt;th&gt;Google AI Ultra&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Notebooks&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sources per notebook&lt;/td&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;100&lt;/td&gt;
&lt;td&gt;300&lt;/td&gt;
&lt;td&gt;500–600&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chat queries per day&lt;/td&gt;
&lt;td&gt;50&lt;/td&gt;
&lt;td&gt;200&lt;/td&gt;
&lt;td&gt;500&lt;/td&gt;
&lt;td&gt;2,500–5,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Audio/Video Overviews per day&lt;/td&gt;
&lt;td&gt;3&lt;/td&gt;
&lt;td&gt;6&lt;/td&gt;
&lt;td&gt;20&lt;/td&gt;
&lt;td&gt;Up to 200 video&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Google AI Pro removes the watermark from generated videos in most regions. Ultra also unlocks the cinematic Video Overview style built on Google’s Veo video model.&lt;/p&gt;

&lt;p&gt;Google AI Pro costs $19.99 per month, while the Plus tier has been reported at $4.99. Students in several countries have received &lt;a href="http://apidog.com/blog/google-ai-pro-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Google AI Pro free for a year&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The practical approach is to stay on the free tier until a limit affects your work. Heavy users are most likely to hit the 50-chat-per-day limit first; that is when Plus or Pro may become worthwhile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Developer workflow: use NotebookLM for API documentation
&lt;/h2&gt;

&lt;p&gt;API documentation is a strong NotebookLM use case: it is long, structured, detail-heavy, and difficult to search when you need one specific answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  Build an API documentation notebook
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Create a notebook for the API you are integrating.&lt;/li&gt;
&lt;li&gt;Add the provider’s reference documentation as URLs or PDFs.&lt;/li&gt;
&lt;li&gt;Include relevant RFCs, such as OAuth 2.0’s &lt;a href="https://datatracker.ietf.org/doc/html/rfc6749" rel="noopener noreferrer"&gt;RFC 6749&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Add your OpenAPI specification as Markdown or a shareable document.&lt;/li&gt;
&lt;li&gt;Ask concrete implementation questions.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Useful prompts include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;“What is the rate limit for the batch endpoint?”&lt;/li&gt;
&lt;li&gt;“Which scopes does the refresh flow require?”&lt;/li&gt;
&lt;li&gt;“Do any endpoints in our specification return a 429 without a &lt;code&gt;Retry-After&lt;/code&gt; header?”&lt;/li&gt;
&lt;li&gt;“Which response fields are required for the create operation?”&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;NotebookLM returns citations to the relevant section of the specification, which is often faster than searching multiple documentation sites. Teams working with &lt;a href="http://apidog.com/blog/how-to-use-gemini-3-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini’s own API&lt;/a&gt; can use the same workflow to query model parameters and API behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  Know what NotebookLM cannot do
&lt;/h3&gt;

&lt;p&gt;NotebookLM reads documents; it does not call API endpoints. It may accurately cite a response schema that your production server stopped honoring several releases ago.&lt;/p&gt;

&lt;p&gt;That is where &lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; fits:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Maintain the OpenAPI specification as a living source of truth.&lt;/li&gt;
&lt;li&gt;Send real requests to the API.&lt;/li&gt;
&lt;li&gt;Validate responses against the schema.&lt;/li&gt;
&lt;li&gt;Detect differences between documentation and production.&lt;/li&gt;
&lt;li&gt;Re-export the validated specification to NotebookLM when it changes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use NotebookLM to understand an API, and Apidog to verify that the API still behaves as documented.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Does NotebookLM train on my uploaded data?
&lt;/h3&gt;

&lt;p&gt;Google states that NotebookLM does not use sources, queries, or responses from free and paid personal accounts to train its models. Workspace accounts receive additional contractual data protections. Still, apply normal caution when handling sensitive information because policies can change.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I share a notebook?
&lt;/h3&gt;

&lt;p&gt;Yes. Notebooks can be shared like Drive files with viewer and editor roles. Viewers can chat with the sources without modifying them. Public links to Audio Overviews are also supported.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is there a NotebookLM API?
&lt;/h3&gt;

&lt;p&gt;There was no public NotebookLM API as of August 2026. For programmatic source-grounded answers, use Gemini models directly. &lt;a href="http://apidog.com/blog/gemini-3-pro-ollama-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3 Pro through Ollama&lt;/a&gt; is an option for local usage, while Google’s File API supports cloud uploads.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which languages does NotebookLM support?
&lt;/h3&gt;

&lt;p&gt;Chat and generated outputs work in dozens of languages. You can set the output language independently from the source language. Audio Overviews also support many languages, although non-English voices may be less polished than the English hosts.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;NotebookLM’s free tier includes the core product: 50 sources per notebook, cited answers, audio and video generation, and mind maps. Start with one notebook focused on a real problem. Once you are comfortable with the workflow, explore these &lt;a href="http://apidog.com/blog/notebooklm-use-scenarios?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;three real-world scenarios&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If your sources are API documentation, complete the loop with &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt;. Design the specification, mock it, test it, and keep the documentation aligned with the API that actually runs.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Is DeepSeek Free? Chat, API Pricing, and Every Free Path in 2026</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Tue, 25 Aug 2026 07:13:43 +0000</pubDate>
      <link>https://dev.to/hassann/is-deepseek-free-chat-api-pricing-and-every-free-path-in-2026-4k70</link>
      <guid>https://dev.to/hassann/is-deepseek-free-chat-api-pricing-and-every-free-path-in-2026-4k70</guid>
      <description>&lt;h1&gt;
  
  
  Is DeepSeek Free? Chat, API, and Local Hosting Costs Explained
&lt;/h1&gt;

&lt;p&gt;DeepSeek’s chat app is free, its API is paid but inexpensive, and its model weights are open for self-hosting. Which option is “free” depends on how you use it. For a complete setup guide, see &lt;a href="http://apidog.com/blog/use-deepseek-v4?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;how to use DeepSeek V4&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This guide focuses on what costs money in August 2026, using rates from &lt;a href="https://api-docs.deepseek.com/quick_start/pricing" rel="noopener noreferrer"&gt;DeepSeek’s pricing page&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Is DeepSeek free?
&lt;/h2&gt;

&lt;p&gt;Yes—depending on the access method:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Web and mobile chat:&lt;/strong&gt; Free. No subscription, Plus tier, or paywall.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API:&lt;/strong&gt; Paid, with V4 Flash output starting at &lt;strong&gt;$0.66 per million tokens off-peak&lt;/strong&gt;. New accounts receive no free credit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model weights:&lt;/strong&gt; Free to download and run under an MIT license. You pay only for hardware and electricity.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In short: free for chat, cheap for API development, and free-to-self-host if you have the hardware.&lt;/p&gt;

&lt;h2&gt;
  
  
  Free chat app: what you get and its limits
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://chat.deepseek.com" rel="noopener noreferrer"&gt;DeepSeek chat app&lt;/a&gt;, including its iOS and Android apps, is free. It includes reasoning mode, file uploads, and web search.&lt;/p&gt;

&lt;p&gt;The limits are operational rather than commercial:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No published daily message quota&lt;/li&gt;
&lt;li&gt;“Server busy” errors and slower responses during traffic spikes&lt;/li&gt;
&lt;li&gt;No API access, team workspace, or uptime guarantee&lt;/li&gt;
&lt;li&gt;Prompts are processed on DeepSeek’s servers in China under DeepSeek’s privacy policy&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use the chat app for research, drafting, coding questions, and evaluating model quality. Choose local hosting if data handling is a concern.&lt;/p&gt;

&lt;h2&gt;
  
  
  DeepSeek API pricing
&lt;/h2&gt;

&lt;p&gt;As of August 16, 2026, DeepSeek uses peak and off-peak pricing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Peak:&lt;/strong&gt; 01:00–04:00 and 06:00–10:00 UTC on weekdays&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Off-peak:&lt;/strong&gt; Every other hour, at half the peak rate&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Prices are per 1 million tokens:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Cache-hit input&lt;/th&gt;
&lt;th&gt;Cache-miss input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;deepseek-v4-flash&lt;/td&gt;
&lt;td&gt;$0.007 off-peak / $0.014 peak&lt;/td&gt;
&lt;td&gt;$0.22 / $0.44&lt;/td&gt;
&lt;td&gt;$0.66 / $1.32&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;deepseek-v4-pro&lt;/td&gt;
&lt;td&gt;$0.022 / $0.044&lt;/td&gt;
&lt;td&gt;$0.66 / $1.32&lt;/td&gt;
&lt;td&gt;$1.98 / $3.96&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;deepseek-v4-flash-vision-exp&lt;/td&gt;
&lt;td&gt;$0.007 / $0.014&lt;/td&gt;
&lt;td&gt;$0.22 / $0.44&lt;/td&gt;
&lt;td&gt;$0.66 / $1.32&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three implementation details reduce spend:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Use prompt caching.&lt;/strong&gt; Repeated prefixes—such as system prompts and few-shot examples—are billed at the cache-hit rate. On Flash, that is roughly 31× cheaper than a cache miss.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Schedule batch work off-peak.&lt;/strong&gt; Moving non-urgent workloads outside peak hours cuts API cost in half without changing code.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use Flash unless Pro is justified.&lt;/strong&gt; Even at peak, Flash output is $1.32 per million tokens. Use Pro when its quality improvement is worth the 3× premium; see the &lt;a href="http://apidog.com/blog/how-to-use-deepseek-v4-pro-0813-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;V4 Pro-0813 API guide&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For example, a chatbot using 10M mostly cached input tokens and 2M output tokens per month on off-peak V4 Flash costs about &lt;strong&gt;$1.50/month&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Free and near-free API options
&lt;/h2&gt;

&lt;p&gt;DeepSeek does not offer trial credit, so a truly free API route must come from another provider.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;OpenRouter free model pool:&lt;/strong&gt; OpenRouter offers 20+ free models without a credit card. DeepSeek’s free variants were removed by mid-2026, but &lt;a href="https://openrouter.ai/deepseek" rel="noopener noreferrer"&gt;OpenRouter’s DeepSeek models&lt;/a&gt; remain paid and start at roughly $0.035 per million V4 Flash input tokens. Its 50 free daily requests on other models are useful for testing your integration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cloud credits:&lt;/strong&gt; Signup credits from cloud providers and inference hosts can pay for hosted DeepSeek models temporarily.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A small DeepSeek top-up:&lt;/strong&gt; A $2 balance can support substantial hobby-scale usage and is the simplest way to access the first-party API.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For current providers and no-cost variants, see &lt;a href="http://apidog.com/blog/how-to-get-deepseek-free-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;how to get a DeepSeek free API key&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run DeepSeek locally for free
&lt;/h2&gt;

&lt;p&gt;DeepSeek publishes weights on &lt;a href="https://huggingface.co/deepseek-ai" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt; under the MIT license, allowing commercial use, fine-tuning, and redistribution.&lt;/p&gt;

&lt;p&gt;The V4 family includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;V4 Flash:&lt;/strong&gt; 284B total parameters, with 13B active&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;V4 Pro:&lt;/strong&gt; Approximately 1.7T parameters; full weights left preview in August 2026&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Hardware is the trade-off:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;V4 Flash needs about &lt;strong&gt;33 GB VRAM&lt;/strong&gt; when heavily quantized, or one &lt;strong&gt;80 GB H100&lt;/strong&gt; at FP8.&lt;/li&gt;
&lt;li&gt;V4 Pro’s full weights approach &lt;strong&gt;900 GB&lt;/strong&gt;, making it a datacenter-scale deployment.&lt;/li&gt;
&lt;li&gt;For a desktop machine, use a smaller distilled or earlier-generation model instead.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;To run a practical reasoning model on consumer hardware, follow &lt;a href="http://apidog.com/blog/run-deepseek-r1-locally-with-ollama?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;this guide to running DeepSeek R1 locally with Ollama&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Local deployment gives you:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;No per-token fees&lt;/li&gt;
&lt;li&gt;No rate limits&lt;/li&gt;
&lt;li&gt;No data leaving your machine&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But you manage quantization, GPU capacity, and the serving stack. For production workloads, use tools such as vLLM or SGLang.&lt;/p&gt;

&lt;h2&gt;
  
  
  Chat vs. API vs. local hosting
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Choose this&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;th&gt;Cost&lt;/th&gt;
&lt;th&gt;Effort&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Chat app&lt;/td&gt;
&lt;td&gt;Questions, drafting, and model evaluation&lt;/td&gt;
&lt;td&gt;$0&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API&lt;/td&gt;
&lt;td&gt;Products and automations&lt;/td&gt;
&lt;td&gt;Cents to a few dollars monthly at hobby scale&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local weights&lt;/td&gt;
&lt;td&gt;Privacy, offline use, or existing GPUs&lt;/td&gt;
&lt;td&gt;No platform fees; hardware costs apply&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The API is OpenAI-compatible, so most SDK integrations only require changing the base URL.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test one real call, mock the rest
&lt;/h2&gt;

&lt;p&gt;Use &lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; to test DeepSeek without wasting tokens during development:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Import an OpenAI-compatible API specification.&lt;/li&gt;
&lt;li&gt;Set the base URL to &lt;code&gt;https://api.deepseek.com&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Store your API key in an environment variable.&lt;/li&gt;
&lt;li&gt;Send one real request and inspect its parsed response, timing, and streamed tokens.&lt;/li&gt;
&lt;li&gt;Capture that response and create a mock server for frontend, CI, integration, and error-state tests.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Your application can then exercise realistic model responses at $0 while only sending live requests when validating real DeepSeek behavior. This is especially useful during peak hours, when API rates double.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is DeepSeek chat unlimited?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
DeepSeek does not publish a message cap. In practice, server capacity is the limit, and you may see “Server busy” responses at high-traffic times. For workarounds, see &lt;a href="http://apidog.com/blog/how-to-use-deepseek-v4-for-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;how to use DeepSeek V4 for free&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does the DeepSeek API have a free trial?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
No. New accounts start with a zero balance and no promotional credit. Use a small top-up, an aggregator, or cloud signup credits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is DeepSeek open source?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Its weights are released under an MIT license, permitting commercial self-hosting. Its training data and full training pipeline are not published, so it is more precisely described as open-weight.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I use DeepSeek for coding without paying?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Yes. Use the free chat app for ad-hoc questions, or run DeepSeek Harness against a local model without API charges. See &lt;a href="http://apidog.com/blog/what-is-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what DeepSeek Harness is&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Bottom line
&lt;/h2&gt;

&lt;p&gt;DeepSeek is free for chat, inexpensive for API development, and free to self-host if you provide the hardware. Start with chat to evaluate it, move to the API when building, and keep costs low by caching prompts, scheduling batch jobs off-peak, and mocking non-production calls.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; to import DeepSeek’s API, validate a real request, and mock the rest for free.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to Run Any Model in DeepSeek Harness ?</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Thu, 20 Aug 2026 07:57:32 +0000</pubDate>
      <link>https://dev.to/hassann/how-to-run-any-model-in-deepseek-harness--5282</link>
      <guid>https://dev.to/hassann/how-to-run-any-model-in-deepseek-harness--5282</guid>
      <description>&lt;p&gt;DeepSeek Harness (&lt;code&gt;dsh&lt;/code&gt;) ships with DeepSeek models, but model providers are configurable. Point a provider block at an OpenAI-compatible endpoint, reference a credential, and run agent sessions against the model behind that URL. This works for local Ollama, company gateways, Qwen through DashScope compatibility mode, and catalog providers such as Anthropic and OpenAI.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This guide explains the provider configuration and includes three implementation recipes: local Ollama, a hosted OpenAI-compatible endpoint, and built-in catalog providers. The configuration details come from the official &lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/docs/user/guide/providers.md" rel="noopener noreferrer"&gt;providers guide&lt;/a&gt; on the master branch, fetched August 20, 2026.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; &lt;code&gt;dsh&lt;/code&gt; is a developer preview. The README warns about compatibility-breaking changes, so verify the documentation against the version you have installed before using these settings in production.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;If you are new to the harness, start with &lt;a href="http://apidog.com/blog/what-is-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what DeepSeek Harness is and how it works&lt;/a&gt;, then return here for provider setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why change models in an agent harness?
&lt;/h2&gt;

&lt;p&gt;An agent harness runs a loop:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The model plans.&lt;/li&gt;
&lt;li&gt;The model calls tools.&lt;/li&gt;
&lt;li&gt;The harness returns tool results.&lt;/li&gt;
&lt;li&gt;The model continues with the updated context.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The harness owns that loop. The model is replaceable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost
&lt;/h3&gt;

&lt;p&gt;Agent sessions can consume tokens quickly because tool results are added back into context. You can route routine sessions to a cheaper model, use DeepSeek V4-Flash instead of V4-Pro, and keep a frontier model configured for harder tasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Data locality
&lt;/h3&gt;

&lt;p&gt;For repositories or prompts that cannot leave your network, point a provider at a model running on your own hardware. Prompts, file contents, and tool outputs stay local while you retain the same harness workflow.&lt;/p&gt;

&lt;h3&gt;
  
  
  Local development
&lt;/h3&gt;

&lt;p&gt;A local model is useful for plugin development and loop testing. It avoids API costs and network dependency during iteration. Switch back to a larger hosted model when model quality matters.&lt;/p&gt;

&lt;p&gt;Provider routes are owned by the &lt;code&gt;dsh-llm-pi-ai&lt;/code&gt; plugin. The &lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/docs/config-catalog.md" rel="noopener noreferrer"&gt;plugin config catalog&lt;/a&gt; describes it as holding the “provider routes this instance owns.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Configure a provider block
&lt;/h2&gt;

&lt;p&gt;Custom providers live in &lt;code&gt;$DSH_HOME/settings.yaml&lt;/code&gt;. You can also create them in the web UI under &lt;strong&gt;Settings → Models&lt;/strong&gt;.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;llm-pi-ai&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;providers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;my-gateway&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;apiKeyEnv&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;GATEWAY_API_KEY&lt;/span&gt;
      &lt;span class="na"&gt;api&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai-completions&lt;/span&gt;
      &lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://gateway.example/v1&lt;/span&gt;
      &lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;legacy-chat&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;vision-preview&lt;/span&gt;
          &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;text&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Provider fields
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;my-gateway&lt;/code&gt;&lt;/strong&gt;: The provider ID. Treat it as permanent. Choose a stable identifier; the UI display name is configured separately.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;apiKeyEnv&lt;/code&gt;&lt;/strong&gt;: The environment variable containing the API key. Store only the variable name in configuration, never the secret itself.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;api&lt;/code&gt;&lt;/strong&gt;: The endpoint protocol. Use &lt;code&gt;openai-completions&lt;/code&gt; for OpenAI-compatible APIs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;baseURL&lt;/code&gt;&lt;/strong&gt;: The API root URL used by the harness.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;models&lt;/code&gt;&lt;/strong&gt;: A list of available model IDs. Each &lt;code&gt;id&lt;/code&gt; must match the identifier expected by the endpoint.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;input&lt;/code&gt;&lt;/strong&gt;: Declares supported input modalities. Custom models default to text-only. Add &lt;code&gt;input: [text, image]&lt;/code&gt; for vision models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;defaultInput&lt;/code&gt;&lt;/strong&gt;: A provider-level input fallback. A model-level &lt;code&gt;input&lt;/code&gt; setting overrides it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;compat&lt;/code&gt;&lt;/strong&gt;: Compatibility settings for endpoints that differ from standard OpenAI behavior:

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;supportsDeveloperRole: false&lt;/code&gt; for backends that reject the &lt;code&gt;developer&lt;/code&gt; role.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;maxTokensField: max_tokens&lt;/code&gt; for backends expecting the older token-limit field.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can configure &lt;code&gt;compat&lt;/code&gt; at the provider level or per model.&lt;/p&gt;

&lt;p&gt;When adding a custom provider through the UI, use &lt;strong&gt;Fetch available models&lt;/strong&gt; if the endpoint implements OpenAI-compatible &lt;code&gt;GET /models&lt;/code&gt;. The UI can populate the model list automatically.&lt;/p&gt;

&lt;h3&gt;
  
  
  Store API keys safely
&lt;/h3&gt;

&lt;p&gt;Secrets are stored write-only in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$DSH_HOME/.credentials.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After saving a key through the UI, dsh returns only a redacted descriptor. The plaintext value is not shown again.&lt;/p&gt;

&lt;p&gt;Keep &lt;code&gt;settings.yaml&lt;/code&gt; limited to references such as &lt;code&gt;apiKeyEnv&lt;/code&gt; and credential descriptors. This lets you share provider configuration without exposing secrets and rotate keys without editing model settings.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recipe 1: Run a local model with Ollama
&lt;/h2&gt;

&lt;p&gt;Ollama exposes an OpenAI-compatible endpoint at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://localhost:11434/v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;See Ollama’s &lt;a href="https://ollama.com/blog/openai-compatibility" rel="noopener noreferrer"&gt;OpenAI compatibility guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Configure it as a custom provider:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;llm-pi-ai&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;providers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;ollama-local&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;apiKeyEnv&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;OLLAMA_API_KEY&lt;/span&gt;
      &lt;span class="na"&gt;api&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai-completions&lt;/span&gt;
      &lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://localhost:11434/v1&lt;/span&gt;
      &lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;gpt-oss:20b&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;qwen3&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Local setup checklist
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Start Ollama.&lt;/li&gt;
&lt;li&gt;Pull the model you want to use:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   ollama pull gpt-oss:20b
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Confirm the installed model names:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   ollama list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Set a placeholder key for the provider schema:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   &lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;OLLAMA_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;ollama
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Ollama does not require an API key locally, but the dsh schema expects a credential reference.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Verify that Ollama serves models:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   curl http://localhost:11434/v1/models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The model ID in &lt;code&gt;settings.yaml&lt;/code&gt; must match the Ollama tag exactly, including its version tag.&lt;/p&gt;

&lt;p&gt;For the complete local model workflow, see &lt;a href="http://apidog.com/blog/run-gpt-oss-using-ollama?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;how to run GPT-OSS using Ollama&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Before configuring dsh, test this endpoint in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET http://localhost:11434/v1/models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the request returns your model list, the server and base URL are correct. If it fails, fix the Ollama setup before debugging the harness configuration.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Small local models are useful for testing plugins and agent loops, but agent workflows rely heavily on tool calling and long context. Expect weaker planning and less reliable tool use than with larger frontier models.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Recipe 2: Use a hosted OpenAI-compatible endpoint with Qwen via DashScope
&lt;/h2&gt;

&lt;p&gt;Alibaba Cloud Model Studio (DashScope) documents OpenAI compatibility through &lt;code&gt;/compatible-mode/v1&lt;/code&gt;. Its &lt;a href="https://www.alibabacloud.com/help/en/model-studio/compatibility-of-openai-with-dashscope" rel="noopener noreferrer"&gt;OpenAI compatibility page&lt;/a&gt; lists regional, workspace-specific endpoints.&lt;/p&gt;

&lt;p&gt;For example, a Singapore workspace uses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Configure Qwen in dsh:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;llm-pi-ai&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;providers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;qwen-dashscope&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;apiKeyEnv&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;DASHSCOPE_API_KEY&lt;/span&gt;
      &lt;span class="na"&gt;api&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai-completions&lt;/span&gt;
      &lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1&lt;/span&gt;
      &lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;qwen3-max&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Replace &lt;code&gt;{WorkspaceId}&lt;/code&gt; with the workspace domain from the Model Studio console.&lt;/p&gt;

&lt;p&gt;Set the credential in the environment where dsh runs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;DASHSCOPE_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;your_api_key
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Check the vendor documentation for current model IDs. For related Qwen API details, see the &lt;a href="http://apidog.com/blog/qwen-3-8-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Qwen 3.8 API guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This same configuration pattern works for any vendor with documented OpenAI compatibility:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Moonshot Kimi API&lt;/li&gt;
&lt;li&gt;OpenRouter&lt;/li&gt;
&lt;li&gt;vLLM deployments&lt;/li&gt;
&lt;li&gt;Internal company gateways&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Usually, only these fields change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiKeyEnv&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;YOUR_PROVIDER_API_KEY&lt;/span&gt;
&lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://your-provider.example/v1&lt;/span&gt;
&lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;your-model-id&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If you have configured &lt;a href="http://apidog.com/blog/codex-open-source-models-oss-mode?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;open-source models in Codex&lt;/a&gt;, the role is similar: dsh’s provider YAML serves the same purpose as Codex &lt;code&gt;model_providers&lt;/code&gt; configuration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Hosted endpoint compatibility settings
&lt;/h3&gt;

&lt;p&gt;If your provider rejects requests because of roles or token parameters, add &lt;code&gt;compat&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;llm-pi-ai&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;providers&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;qwen-dashscope&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="na"&gt;apiKeyEnv&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;DASHSCOPE_API_KEY&lt;/span&gt;
      &lt;span class="na"&gt;api&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai-completions&lt;/span&gt;
      &lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;https://{WorkspaceId}.ap-southeast-1.maas.aliyuncs.com/compatible-mode/v1&lt;/span&gt;
      &lt;span class="na"&gt;compat&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="na"&gt;supportsDeveloperRole&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
        &lt;span class="na"&gt;maxTokensField&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;max_tokens&lt;/span&gt;
      &lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
        &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;qwen3-max&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For vision models, explicitly declare image input support:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;your-vision-model&lt;/span&gt;
    &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;text&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Custom models are text-only unless configured otherwise.&lt;/p&gt;

&lt;h2&gt;
  
  
  Recipe 3: Use built-in catalog providers
&lt;/h2&gt;

&lt;p&gt;You do not need a custom provider block for mainstream cloud providers. dsh includes catalog providers for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DeepSeek&lt;/li&gt;
&lt;li&gt;Anthropic&lt;/li&gt;
&lt;li&gt;OpenAI&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For these providers, setup is primarily API-key configuration.&lt;/p&gt;

&lt;p&gt;Other catalog entries use their native authentication flows:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Authentication&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;AWS Bedrock&lt;/td&gt;
&lt;td&gt;AWS credentials&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Google Vertex&lt;/td&gt;
&lt;td&gt;ADC project configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Azure&lt;/td&gt;
&lt;td&gt;API version configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Codex&lt;/td&gt;
&lt;td&gt;OAuth&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Use catalog providers when you want the lowest-friction path to Claude, GPT, or DeepSeek models. Use custom providers for local runtimes, regional vendors, OpenAI-compatible aggregators, and internal gateways.&lt;/p&gt;

&lt;p&gt;DeepSeek V4-Pro launched alongside the harness in August 2026. Refer to &lt;a href="https://api-docs.deepseek.com" rel="noopener noreferrer"&gt;api-docs.deepseek.com&lt;/a&gt; for DeepSeek API details.&lt;/p&gt;

&lt;h2&gt;
  
  
  Select a model and understand session behavior
&lt;/h2&gt;

&lt;p&gt;Adding a provider makes its models available. Selecting a model under &lt;strong&gt;Settings → Models&lt;/strong&gt; sets the default for new sessions.&lt;/p&gt;

&lt;p&gt;Two behaviors matter:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Existing sessions keep their original model.&lt;/strong&gt; Changing the default does not alter the model used by active or previous sessions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Deleting the current default provider blocks the composer.&lt;/strong&gt; You must select another model before continuing.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This session pinning improves reproducibility. For example, when comparing harness behavior in &lt;a href="http://apidog.com/blog/deepseek-harness-vs-claude-code?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek Harness vs Claude Code&lt;/a&gt;, each transcript remains tied to the model it started with.&lt;/p&gt;

&lt;h2&gt;
  
  
  Troubleshooting provider configuration
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Wrong or unreachable &lt;code&gt;baseURL&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Confirm that the endpoint ends at the expected API root:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OpenAI-compatible: /v1
DashScope compatible mode: /compatible-mode/v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then test the model route outside dsh:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$GATEWAY_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  https://gateway.example/v1/models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; to send the same request and inspect the actual status code and response body instead of a wrapped harness error.&lt;/p&gt;

&lt;p&gt;For offline development or unstable providers, mock these endpoints in Apidog:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET  /models
POST /chat/completions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then point &lt;code&gt;baseURL&lt;/code&gt; at the mock server while developing your integration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Missing or empty environment variable
&lt;/h3&gt;

&lt;p&gt;&lt;code&gt;apiKeyEnv&lt;/code&gt; names an environment variable. It does not create it.&lt;/p&gt;

&lt;p&gt;If dsh cannot read that variable, requests are sent without valid authentication and usually fail with &lt;code&gt;401&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Check the variable in the same environment that starts dsh:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;echo&lt;/span&gt; &lt;span class="nv"&gt;$GATEWAY_API_KEY&lt;/span&gt;
dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A process started from a GUI or service manager may not inherit your shell profile.&lt;/p&gt;

&lt;h3&gt;
  
  
  Image input does not work
&lt;/h3&gt;

&lt;p&gt;Custom models are text-only by default. Add image support per model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;id&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;vision-preview&lt;/span&gt;
    &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;text&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or set it for all provider models:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;defaultInput&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="pi"&gt;[&lt;/span&gt;&lt;span class="nv"&gt;text&lt;/span&gt;&lt;span class="pi"&gt;,&lt;/span&gt; &lt;span class="nv"&gt;image&lt;/span&gt;&lt;span class="pi"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Unsupported roles or token fields
&lt;/h3&gt;

&lt;p&gt;If your backend rejects a &lt;code&gt;developer&lt;/code&gt; role or token limit field, configure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;compat&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;supportsDeveloperRole&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;maxTokensField&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;max_tokens&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  A previously working configuration stopped working
&lt;/h3&gt;

&lt;p&gt;dsh is a developer preview. Pin the version you deploy, read release notes before upgrading, and expect configuration schemas to change.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;deepseek-harness repository&lt;/a&gt; is the source of truth.&lt;/p&gt;

&lt;p&gt;Model providers are only one part of customization. You can also connect the tools that agents call. See &lt;a href="http://apidog.com/blog/apidog-cli-in-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;using Apidog CLI inside DeepSeek Harness&lt;/a&gt; for that workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Does DeepSeek Harness officially support Ollama?
&lt;/h3&gt;

&lt;p&gt;The official provider documentation does not mention Ollama by name. It supports endpoints using the &lt;code&gt;openai-completions&lt;/code&gt; protocol, and Ollama documents an OpenAI-compatible API at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://localhost:11434/v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Ollama example combines those documented pieces. Test it with your installed dsh version because the project is a developer preview.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where does dsh store API keys?
&lt;/h3&gt;

&lt;p&gt;dsh stores keys write-only in:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$DSH_HOME/.credentials.yaml
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The UI returns a redacted descriptor after saving. &lt;code&gt;settings.yaml&lt;/code&gt; stores only references such as &lt;code&gt;apiKeyEnv&lt;/code&gt;, not plaintext API keys.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can different sessions use different models?
&lt;/h3&gt;

&lt;p&gt;Yes. Changing the selected model changes the default for new sessions only. Existing sessions continue using the model they started with.&lt;/p&gt;

&lt;p&gt;For example, use &lt;a href="http://apidog.com/blog/deepseek-v4-flash-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek V4-Flash&lt;/a&gt; for routine sessions, switch to a larger model for a difficult task, and keep earlier sessions unchanged.&lt;/p&gt;

&lt;h3&gt;
  
  
  My endpoint works with curl but fails in dsh. What should I check?
&lt;/h3&gt;

&lt;p&gt;Compare the exact request payloads.&lt;/p&gt;

&lt;p&gt;The harness may send a &lt;code&gt;developer&lt;/code&gt; role or a newer token-cap field that your backend does not support. Use the documented compatibility settings:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;compat&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;supportsDeveloperRole&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;
  &lt;span class="na"&gt;maxTokensField&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;max_tokens&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Replay the harness-shaped request in an API client to identify which field the backend rejects.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to Use the Apidog CLI in DeepSeek Harness</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Thu, 20 Aug 2026 07:56:56 +0000</pubDate>
      <link>https://dev.to/hassann/how-to-use-the-apidog-cli-in-deepseek-harness-3n67</link>
      <guid>https://dev.to/hassann/how-to-use-the-apidog-cli-in-deepseek-harness-3n67</guid>
      <description>&lt;p&gt;DeepSeek Harness is a loop. The agent reads your workspace, edits files, runs commands through its bash tool, and decides what to do next based on the output. So why aren’t your API tests in that loop? They sit in Apidog behind a GUI and run when someone remembers to click. The agent never touches them.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;The fix is one config block. The Apidog CLI is an npm package, &lt;code&gt;apidog-cli&lt;/code&gt;, that runs the test scenarios you built in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; straight from a terminal. Once the CLI is installed and DeepSeek Harness knows it exists, the agent runs an Apidog scenario the same way it runs your unit tests: fire the command, read the exit code, fix the code if it is red.&lt;/p&gt;

&lt;p&gt;There is also a token argument for doing this. An agent that confirms your API still works by re-reading handler code and reasoning about response shapes burns context on every pass. An agent that runs one command gets ground truth back in a few lines. The CLI compresses “is the API correct?” into an exit code, and the agent spends its context on the fix instead.&lt;/p&gt;

&lt;p&gt;This guide covers the harness-specific part the generic install guide skips: which instructions file DeepSeek Harness actually reads, how its bash tool executes &lt;code&gt;apidog run&lt;/code&gt;, and how to keep the loop honest. If you have not installed the CLI yet, do that first. &lt;a href="http://apidog.com/blog/apidog-cli-installation-guide?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;How to install the Apidog CLI with an AI coding agent&lt;/a&gt; walks through the npm install, authentication, and the first run. This article assumes &lt;code&gt;apidog --version&lt;/code&gt; prints a number and your machine is authenticated.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which DeepSeek Harness this is about
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness, &lt;code&gt;dsh&lt;/code&gt; on the command line, is the open-source agent harness &lt;a href="https://venturebeat.com/technology/deepseek-harness-launches-as-open-source-rival-to-claude-code-alongside-v4-pro-on-api-with-higher-prices" rel="noopener noreferrer"&gt;DeepSeek released on August 13, 2026&lt;/a&gt;, alongside V4-Pro on the API. It is MIT licensed, sits at &lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;github.com/deepseek-ai/deepseek-harness&lt;/a&gt;, and had climbed past 169k stars as of August 20.&lt;/p&gt;

&lt;p&gt;Start it with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @deepseek-ai/dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This serves a local web UI at &lt;code&gt;http://127.0.0.1:3080&lt;/code&gt;. Select a workspace—the project directory where you launched it—and the agent works inside it: reading and editing files, running commands, and asking before operations that require approval under the active permission policy.&lt;/p&gt;

&lt;p&gt;Two details shape the setup below:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;DeepSeek Harness is a &lt;strong&gt;developer preview&lt;/strong&gt;. Its README warns that compatibility-breaking changes will happen. Treat the file names and config keys here as accurate for late August 2026, and re-check the &lt;a href="https://github.com/deepseek-ai/deepseek-harness/tree/master/docs" rel="noopener noreferrer"&gt;repo docs&lt;/a&gt; if something does not load.&lt;/li&gt;
&lt;li&gt;Everything in dsh is a plugin built on the Cordis architecture. That makes the key question answerable: which plugin reads project rules, and what files does it load?&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For a broader introduction, see &lt;a href="http://apidog.com/blog/what-is-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what DeepSeek Harness is&lt;/a&gt;. For a comparison with the incumbent, see &lt;a href="http://apidog.com/blog/deepseek-harness-vs-claude-code?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek Harness vs Claude Code&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 1: put the CLI in &lt;code&gt;AGENTS.md&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness reads workspace instructions through its &lt;code&gt;@deepseek-ai/dsh-agent-instructions&lt;/code&gt; plugin. According to the plugin source and the &lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/docs/config-catalog.md" rel="noopener noreferrer"&gt;config catalog&lt;/a&gt;, the loader walks upward from the session working directory to the project root, marked by &lt;code&gt;.git&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;It loads these files in each directory:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;AGENTS.md&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;CLAUDE.md&lt;/code&gt; as a fallback&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;AGENTS.local.md&lt;/code&gt; as a local overlay&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;CLAUDE.local.md&lt;/code&gt; as a local overlay&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Local overlays load after base files. A user-global &lt;code&gt;AGENTS.md&lt;/code&gt; in &lt;code&gt;$DSH_HOME&lt;/code&gt;, which defaults to &lt;code&gt;~/.dsh&lt;/code&gt;, applies across projects. Files larger than 1 MiB are ignored.&lt;/p&gt;

&lt;p&gt;If your repository already has an &lt;code&gt;AGENTS.md&lt;/code&gt; for Codex or a &lt;code&gt;CLAUDE.md&lt;/code&gt; for Claude Code, DeepSeek Harness can use it without extra setup. Add an Apidog block like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;&lt;span class="gu"&gt;## API testing with the Apidog CLI&lt;/span&gt;
&lt;span class="p"&gt;-&lt;/span&gt; To test the API, run the Apidog scenario. Do not click through the GUI.
&lt;span class="p"&gt;-&lt;/span&gt; Command: apidog run -t &lt;span class="nt"&gt;&amp;lt;scenario_id&amp;gt;&lt;/span&gt; -e &lt;span class="nt"&gt;&amp;lt;env_id&amp;gt;&lt;/span&gt; -r cli
&lt;span class="p"&gt;-&lt;/span&gt; Exit code 0 means every assertion passed. Non-zero means a failure; read the report and fix the code.
&lt;span class="p"&gt;-&lt;/span&gt; The machine is already authenticated. Never add an --access-token flag and never put a token in this file.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use the repository-level file for real scenario and environment IDs. If you work across projects, use &lt;code&gt;~/.dsh/AGENTS.md&lt;/code&gt; for a general rule such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight markdown"&gt;&lt;code&gt;Always verify API changes with the project's apidog run command.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep the project-specific IDs in each repository’s own &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 2: get the command from Apidog
&lt;/h2&gt;

&lt;p&gt;Do not guess scenario or environment IDs.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Open the test scenario in Apidog.&lt;/li&gt;
&lt;li&gt;Go to the &lt;strong&gt;CI/CD&lt;/strong&gt; tab.&lt;/li&gt;
&lt;li&gt;Copy the generated command.&lt;/li&gt;
&lt;li&gt;Paste it into your repository’s &lt;code&gt;AGENTS.md&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The generated command looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;apidog run &lt;span class="nt"&gt;-t&lt;/span&gt; 123456 &lt;span class="nt"&gt;-e&lt;/span&gt; 789012 &lt;span class="nt"&gt;-r&lt;/span&gt; cli
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The flags are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;-t&lt;/code&gt;: test scenario ID&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;-e&lt;/code&gt;: environment ID&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;-r cli&lt;/code&gt;: inline reporter output&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;cli&lt;/code&gt; reporter is important because it gives the agent output it can read directly. Put the real command from Apidog in &lt;code&gt;AGENTS.md&lt;/code&gt; so the agent runs the configured scenario rather than inventing IDs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 3: have the agent run the test
&lt;/h2&gt;

&lt;p&gt;Start a session in the dsh web UI with the correct workspace selected. Because the instruction loader has already included your &lt;code&gt;AGENTS.md&lt;/code&gt;, the agent knows the command to run.&lt;/p&gt;

&lt;p&gt;After changing API code, ask:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Run the Apidog test scenario and tell me the exit code.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The agent executes the scenario through its bash tool. According to the &lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/docs/tool-catalog.md" rel="noopener noreferrer"&gt;tool catalog&lt;/a&gt;, the default bash tool runs every command in a &lt;strong&gt;fresh shell&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That means:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Working directories do not persist between tool calls.&lt;/li&gt;
&lt;li&gt;Environment variables do not persist between calls.&lt;/li&gt;
&lt;li&gt;Shell functions do not persist between calls.&lt;/li&gt;
&lt;li&gt;Commands run from the session workspace unless the tool receives a &lt;code&gt;workdir&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This works well for a self-contained command such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;apidog run &lt;span class="nt"&gt;-t&lt;/span&gt; 123456 &lt;span class="nt"&gt;-e&lt;/span&gt; 789012 &lt;span class="nt"&gt;-r&lt;/span&gt; cli
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not rely on the agent running &lt;code&gt;cd some-directory&lt;/code&gt; in one command and &lt;code&gt;apidog run ...&lt;/code&gt; in the next. If the test must run from a subdirectory, keep the full invocation on one line in your instructions file.&lt;/p&gt;

&lt;p&gt;A non-zero command exit is returned with an explicit marker:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[exit code: N]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That preserves the pass/fail signal even if long command output is truncated. Commands can also run under a file sandbox. If the sandbox blocks an operation, dsh reports a policy denial rather than a command failure. A read-only test run should rarely trigger this, but the HTML reporter may need permission to write to &lt;code&gt;./apidog-reports&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Whether the command requires approval depends on the active permission policy. The web UI asks before operations that require approval under that policy, as described in the &lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/docs/user/guide/index.md" rel="noopener noreferrer"&gt;user guide&lt;/a&gt;. If dsh prompts for &lt;code&gt;apidog run&lt;/code&gt;, approve it when the scenario is safe for the target environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Step 4: read the report
&lt;/h2&gt;

&lt;p&gt;When a scenario fails, use the CLI output to drive the next edit.&lt;/p&gt;

&lt;p&gt;With &lt;code&gt;-r cli&lt;/code&gt;, the agent receives an inline breakdown of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Requests made by the scenario&lt;/li&gt;
&lt;li&gt;Assertions evaluated&lt;/li&gt;
&lt;li&gt;Failed assertions&lt;/li&gt;
&lt;li&gt;Expected versus actual values&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, a failure may identify a wrong status code, a missing &lt;code&gt;total&lt;/code&gt; field, or an incorrect currency code. That is usually enough for the agent to find the relevant handler and make a targeted fix.&lt;/p&gt;

&lt;p&gt;To also create a report you can open in a browser or share with a teammate, add the HTML reporter:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;apidog run &lt;span class="nt"&gt;-t&lt;/span&gt; 123456 &lt;span class="nt"&gt;-e&lt;/span&gt; 789012 &lt;span class="nt"&gt;-r&lt;/span&gt; cli,html
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;html&lt;/code&gt; reporter writes a self-contained file to &lt;code&gt;./apidog-reports&lt;/code&gt;. Keep &lt;code&gt;cli&lt;/code&gt; in the reporter list so the agent still receives inline output for its edit-test-fix loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  The loop, end to end
&lt;/h2&gt;

&lt;p&gt;Suppose the agent is editing a checkout handler.&lt;/p&gt;

&lt;p&gt;Without the CLI, the loop may end at: “the code looks right.”&lt;/p&gt;

&lt;p&gt;With the command in &lt;code&gt;AGENTS.md&lt;/code&gt;, the loop becomes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Edit the handler.&lt;/li&gt;
&lt;li&gt;Run the scenario:
&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;   apidog run &lt;span class="nt"&gt;-t&lt;/span&gt; 123456 &lt;span class="nt"&gt;-e&lt;/span&gt; 789012 &lt;span class="nt"&gt;-r&lt;/span&gt; cli
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;Read the result.&lt;/li&gt;
&lt;li&gt;If the command exits with &lt;code&gt;0&lt;/code&gt;, move on.&lt;/li&gt;
&lt;li&gt;If it returns &lt;code&gt;[exit code: 1]&lt;/code&gt;, inspect the failed assertion.&lt;/li&gt;
&lt;li&gt;Patch the handler.&lt;/li&gt;
&lt;li&gt;Run the scenario again.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The agent can now catch problems such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;code&gt;500&lt;/code&gt; response where &lt;code&gt;200&lt;/code&gt; is expected&lt;/li&gt;
&lt;li&gt;A missing &lt;code&gt;total&lt;/code&gt; field&lt;/li&gt;
&lt;li&gt;A wrong currency code&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The API contract check becomes part of the same edit-test-fix cycle used for unit tests.&lt;/p&gt;

&lt;p&gt;The agent does not need to re-read every route file to reason about whether the API works. The scenario already encodes expected behavior and can be authored visually in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;. The agent delegates verification to a deterministic command and uses its context on the fix.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verify dsh actually ran it
&lt;/h2&gt;

&lt;p&gt;Do not accept a success summary without checking the tool call. Use these three checks.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Confirm the command ran
&lt;/h3&gt;

&lt;p&gt;The dsh web UI shows agent tool calls and output in the session. Look for the literal bash call:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;apidog run ...
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the agent says it ran tests but no matching tool call appears, ask it to run the command again and show the raw output.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Confirm the exit code
&lt;/h3&gt;

&lt;p&gt;Ask directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;What was the exit code of that apidog run command?
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On failure, dsh returns an explicit &lt;code&gt;[exit code: N]&lt;/code&gt; marker. If an agent summary says “tests passed” but the command has a non-zero marker, trust the marker.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Confirm it used the real scenario
&lt;/h3&gt;

&lt;p&gt;A “scenario not found” error usually means the agent invented or misremembered an ID.&lt;/p&gt;

&lt;p&gt;Compare the &lt;code&gt;-t&lt;/code&gt; and &lt;code&gt;-e&lt;/code&gt; values against:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your &lt;code&gt;AGENTS.md&lt;/code&gt; block&lt;/li&gt;
&lt;li&gt;The command generated in Apidog’s CI/CD tab&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The IDs in the rules file are the source of truth.&lt;/p&gt;

&lt;h2&gt;
  
  
  Optional: add the Apidog MCP server for spec access
&lt;/h2&gt;

&lt;p&gt;Running scenarios verifies behavior. If you also want the agent to read your API specification while writing code, use MCP.&lt;/p&gt;

&lt;p&gt;As of late August 2026, MCP support is not documented in the DeepSeek Harness core README or user guide. A &lt;strong&gt;community plugin&lt;/strong&gt;, &lt;code&gt;hyqhyq3/dsh-mcp-manager&lt;/code&gt;, is available through the &lt;code&gt;dsh-plugin&lt;/code&gt; GitHub topic. It provides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An MCP page in Settings&lt;/li&gt;
&lt;li&gt;Support for remote HTTP and local stdio servers&lt;/li&gt;
&lt;li&gt;Tools registered as &lt;code&gt;mcp__&amp;lt;name&amp;gt;__*&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Per-project server definitions in &lt;code&gt;&amp;lt;workspace&amp;gt;/.dsh/dshmm/mcp.json&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can use it to connect the &lt;a href="http://apidog.com/blog/apidog-mcp-server?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog MCP server&lt;/a&gt;, which exposes API specifications over MCP. This lets the agent inspect an endpoint schema before writing the handler instead of discovering a mismatch only after a scenario fails.&lt;/p&gt;

&lt;p&gt;Treat this as an optional layer. Both the host harness and the community plugin can change. The CLI integration remains the load-bearing path because it only needs a shell.&lt;/p&gt;

&lt;h2&gt;
  
  
  Preview caveats, and where this goes
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness is moving quickly and explicitly warns about breaking changes. Re-check these details when upgrading:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Instructions plugin file candidates&lt;/li&gt;
&lt;li&gt;Bash tool sandbox behavior and reporting&lt;/li&gt;
&lt;li&gt;Community MCP plugin configuration&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The pattern is portable:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Put a verification command in a rules file.&lt;/li&gt;
&lt;li&gt;Use a CLI that returns a clear exit code.&lt;/li&gt;
&lt;li&gt;Require the agent to run that command after API changes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That works in dsh for the same reason it works in &lt;a href="http://apidog.com/blog/apidog-cli-in-claude-code?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Code&lt;/a&gt; and other coding harnesses: agents can read command output, but they should not be trusted without it.&lt;/p&gt;

&lt;p&gt;So, &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;download Apidog&lt;/a&gt;, build one test scenario visually, copy its &lt;code&gt;apidog run&lt;/code&gt; command from the CI/CD tab, and add it to your repository’s &lt;code&gt;AGENTS.md&lt;/code&gt;. The next time DeepSeek Harness changes API code, it can check its own work before reporting that it is done.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Does DeepSeek Harness read &lt;a href="http://AGENTS.md" rel="noopener noreferrer"&gt;AGENTS.md&lt;/a&gt; natively?
&lt;/h3&gt;

&lt;p&gt;Yes. The &lt;code&gt;@deepseek-ai/dsh-agent-instructions&lt;/code&gt; plugin loads &lt;code&gt;AGENTS.md&lt;/code&gt;, with &lt;code&gt;CLAUDE.md&lt;/code&gt; as a fallback, from the project root and directories above the session working directory. It also supports &lt;code&gt;AGENTS.local.md&lt;/code&gt; and &lt;code&gt;CLAUDE.local.md&lt;/code&gt; overlays, plus a user-global &lt;code&gt;AGENTS.md&lt;/code&gt; in &lt;code&gt;~/.dsh&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If you already maintain an &lt;code&gt;AGENTS.md&lt;/code&gt; for other agents, dsh can use it unchanged.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I need a paid DeepSeek plan to use the Apidog CLI in dsh?
&lt;/h3&gt;

&lt;p&gt;No. The harness is MIT-licensed open source, and you bring your own model. Catalog providers cover Anthropic, OpenAI, Bedrock, Vertex, and Azure. Custom gateways work through &lt;code&gt;settings.yaml&lt;/code&gt;, as covered in &lt;a href="http://apidog.com/blog/run-any-model-in-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;how to run any model in DeepSeek Harness&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The Apidog CLI is a free npm package. It requires an Apidog test scenario and authentication, not a specific model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why does the agent’s second command forget the directory the first one changed to?
&lt;/h3&gt;

&lt;p&gt;This is by design. The default dsh bash tool runs every call in a fresh shell, so &lt;code&gt;cd&lt;/code&gt; does not persist between commands.&lt;/p&gt;

&lt;p&gt;Pass the tool’s &lt;code&gt;workdir&lt;/code&gt; parameter or, more simply, keep the complete &lt;code&gt;apidog run&lt;/code&gt; invocation on one line in your rules file.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can dsh run the scenario without asking me every time?
&lt;/h3&gt;

&lt;p&gt;That depends on the active permission policy. The web UI asks before operations that require approval under that policy. The user guide does not enumerate policy levels, so check Settings in your build to see what your deployment allows.&lt;/p&gt;

&lt;p&gt;When dsh prompts, approving an &lt;code&gt;apidog run&lt;/code&gt; command against staging is a safe choice for a read-mostly verification step.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>DeepSeek Harness vs Claude Code: Which Coding Agent Fits Your Stack?</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Thu, 20 Aug 2026 06:02:04 +0000</pubDate>
      <link>https://dev.to/hassann/deepseek-harness-vs-claude-code-which-coding-agent-fits-your-stack-4ppc</link>
      <guid>https://dev.to/hassann/deepseek-harness-vs-claude-code-which-coding-agent-fits-your-stack-4ppc</guid>
      <description>&lt;p&gt;DeepSeek Harness (dsh) landed on August 13, 2026, framed as an open-source alternative to Claude Code. &lt;a href="https://venturebeat.com/technology/deepseek-harness-launches-as-open-source-rival-to-claude-code-alongside-v4-pro-on-api-with-higher-prices" rel="noopener noreferrer"&gt;VentureBeat’s launch coverage&lt;/a&gt; described it as an “open source rival to Claude Code,” released alongside DeepSeek V4-Pro on the API. One week later, the &lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;GitHub repository&lt;/a&gt; had roughly 169k stars as of August 20.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;Stars do not answer the implementation question: should you run your coding workflows with DeepSeek Harness or Claude Code?&lt;/p&gt;

&lt;p&gt;The tools make different trade-offs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek Harness&lt;/strong&gt;: MIT-licensed developer preview, plugin-kernel architecture, local web UI, and multiple model providers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Code&lt;/strong&gt;: proprietary, mature coding agent with native MCP, skills, hooks, subagents, permission controls, and multiple interfaces.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This guide compares licensing, setup, model configuration, pricing, maturity, extensibility, permissions, and MCP support. It avoids unverified performance claims. If you are new to dsh, start with &lt;a href="http://apidog.com/blog/what-is-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what DeepSeek Harness is and how it works&lt;/a&gt;.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;💡 Both agents write code against APIs, and the output quality depends on the API contract you provide. Keep your specification tested and current with Apidog, whichever agent you choose.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  The quick comparison
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;DeepSeek Harness (dsh)&lt;/th&gt;
&lt;th&gt;Claude Code&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;License&lt;/td&gt;
&lt;td&gt;MIT, source on GitHub&lt;/td&gt;
&lt;td&gt;Proprietary; “All rights reserved,” Anthropic Commercial Terms&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Age&lt;/td&gt;
&lt;td&gt;Released Aug 13, 2026; developer preview&lt;/td&gt;
&lt;td&gt;Generally available, mature product&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stability&lt;/td&gt;
&lt;td&gt;README warns of compatibility-breaking changes&lt;/td&gt;
&lt;td&gt;Stable release channels, versioned settings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Interface&lt;/td&gt;
&lt;td&gt;Local web UI at &lt;code&gt;127.0.0.1:3080&lt;/code&gt;, plus profile-based CLI modes including headless&lt;/td&gt;
&lt;td&gt;Terminal CLI, VS Code, JetBrains, desktop app, web, and mobile&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Models&lt;/td&gt;
&lt;td&gt;DeepSeek, catalog providers for Anthropic, OpenAI, Bedrock, Vertex, Azure, and OpenAI-compatible endpoints&lt;/td&gt;
&lt;td&gt;Claude models only, direct or through Bedrock, Vertex, or Foundry&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pricing&lt;/td&gt;
&lt;td&gt;Free harness; pay per token for the API you connect&lt;/td&gt;
&lt;td&gt;Claude Pro/Max subscription or usage-based API billing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extensibility&lt;/td&gt;
&lt;td&gt;Everything is a plugin through the Cordis kernel&lt;/td&gt;
&lt;td&gt;Plugins, skills, hooks, subagents, Agent SDK&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Permissions&lt;/td&gt;
&lt;td&gt;Web UI approval prompts under an active permission policy&lt;/td&gt;
&lt;td&gt;Six documented modes plus allow/deny rules&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;MCP&lt;/td&gt;
&lt;td&gt;Community plugin: &lt;code&gt;dsh-mcp-manager&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Native, first-class support&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Open source vs proprietary: what MIT changes
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness is MIT licensed, with third-party dependencies listed in &lt;code&gt;THIRD_PARTY_NOTICES.md&lt;/code&gt;. In practical terms, you can:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Inspect the agent loop.&lt;/li&gt;
&lt;li&gt;Audit what the harness sends over the network.&lt;/li&gt;
&lt;li&gt;Fork and patch the project.&lt;/li&gt;
&lt;li&gt;Embed it in commercial internal tooling.&lt;/li&gt;
&lt;li&gt;Replace parts of the implementation through plugins.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That matters when source access, auditability, or vendor independence is a requirement.&lt;/p&gt;

&lt;p&gt;Claude Code takes the opposite approach. Its public repository is for issues and documentation. Its license states:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;© Anthropic PBC. All rights reserved. Use is subject to Anthropic’s Commercial Terms of Service.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;You use a supported product and its documented extension points, but you cannot inspect or fork the core implementation.&lt;/p&gt;

&lt;p&gt;A practical decision rule:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Choose &lt;strong&gt;dsh&lt;/strong&gt; if the agent must become part of your infrastructure.&lt;/li&gt;
&lt;li&gt;Choose &lt;strong&gt;Claude Code&lt;/strong&gt; if you prefer a supported product over maintaining agent internals.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;MIT licensing gives you rights, not maintenance guarantees. dsh is still a developer preview.&lt;/p&gt;

&lt;h2&gt;
  
  
  Interface: local web UI vs multiple surfaces
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Start DeepSeek Harness
&lt;/h3&gt;

&lt;p&gt;The dsh quick start launches a local browser interface:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @deepseek-ai/dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This starts the UI at:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://127.0.0.1:3080
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To start the server without opening a browser automatically:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @deepseek-ai/dsh web &lt;span class="nt"&gt;--no-open&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;From the UI, select the workspace that maps to the directory where you launched dsh.&lt;/p&gt;

&lt;p&gt;dsh also supports profile-based CLI execution. According to the &lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/apps/cli/README.md" rel="noopener noreferrer"&gt;CLI README&lt;/a&gt;, profiles live under:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$DSH_HOME/profiles/&amp;lt;name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Launch a profile with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dsh &lt;span class="nt"&gt;--profile&lt;/span&gt; &amp;lt;name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A headless profile can run a persisted session, print the final response, and exit. This makes automation possible, although the web UI is currently the primary interface.&lt;/p&gt;

&lt;p&gt;Plugin management is also profile-scoped:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dsh plugin
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This forwards plugin operations to &lt;code&gt;pnpm&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Run Claude Code where you work
&lt;/h3&gt;

&lt;p&gt;Claude Code supports more execution surfaces. Per the &lt;a href="https://code.claude.com/docs/en/overview" rel="noopener noreferrer"&gt;official overview&lt;/a&gt;, you can use it in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Terminal&lt;/li&gt;
&lt;li&gt;VS Code&lt;/li&gt;
&lt;li&gt;JetBrains IDEs&lt;/li&gt;
&lt;li&gt;Desktop app&lt;/li&gt;
&lt;li&gt;Browser at &lt;a href="http://claude.ai/code" rel="noopener noreferrer"&gt;claude.ai/code&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Mobile&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For non-interactive scripts, CI, cron jobs, or shell pipelines:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;claude &lt;span class="nt"&gt;-p&lt;/span&gt; &lt;span class="s2"&gt;"prompt"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use dsh when a local browser workflow fits your project. Use Claude Code when you need to move between IDEs, terminals, web sessions, CI, and mobile.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model freedom: configure any supported backend in dsh
&lt;/h2&gt;

&lt;p&gt;This is the clearest architectural difference.&lt;/p&gt;

&lt;p&gt;dsh is model-agnostic. Its &lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/docs/user/guide/providers.md" rel="noopener noreferrer"&gt;provider documentation&lt;/a&gt; includes catalog providers for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Anthropic&lt;/li&gt;
&lt;li&gt;OpenAI&lt;/li&gt;
&lt;li&gt;Amazon Bedrock&lt;/li&gt;
&lt;li&gt;Google Vertex&lt;/li&gt;
&lt;li&gt;Azure&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It also supports any OpenAI-compatible endpoint through &lt;code&gt;$DSH_HOME/settings.yaml&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;A provider configuration includes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;apiKeyEnv&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;MY_PROVIDER_API_KEY&lt;/span&gt;
&lt;span class="na"&gt;api&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;openai-completions&lt;/span&gt;
&lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;http://localhost:8000/v1&lt;/span&gt;
&lt;span class="na"&gt;models&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;my-local-model&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep API credentials in a separate &lt;code&gt;.credentials.yaml&lt;/code&gt; file so you can share settings without committing secrets.&lt;/p&gt;

&lt;p&gt;This lets dsh work with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DeepSeek APIs&lt;/li&gt;
&lt;li&gt;Anthropic APIs&lt;/li&gt;
&lt;li&gt;Hosted OpenAI-compatible APIs&lt;/li&gt;
&lt;li&gt;Local model servers&lt;/li&gt;
&lt;li&gt;Quantized models running on your own GPU&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a full walkthrough, see &lt;a href="http://apidog.com/blog/run-any-model-in-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;how to run any model in DeepSeek Harness&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Claude Code runs Claude models only. You can use Amazon Bedrock, Google Cloud’s Agent Platform, or Microsoft Foundry for infrastructure and billing, but the underlying model remains Claude.&lt;/p&gt;

&lt;p&gt;Use dsh when model portability is important. Use Claude Code when you want a tightly integrated Claude-specific agent experience.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing: token billing vs subscriptions
&lt;/h2&gt;

&lt;p&gt;The DeepSeek Harness software itself is free. You pay for inference from whichever provider you configure.&lt;/p&gt;

&lt;p&gt;If you use DeepSeek’s API, pricing follows &lt;a href="https://api-docs.deepseek.com" rel="noopener noreferrer"&gt;DeepSeek’s per-token rates&lt;/a&gt;. VentureBeat reported that V4-Pro, released with dsh, had higher pricing than earlier models.&lt;/p&gt;

&lt;p&gt;To call V4-Pro directly, review &lt;a href="http://apidog.com/blog/how-to-use-deepseek-v4-pro-0813-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;the DeepSeek V4-Pro-0813 API guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Per-token billing works well when usage is light or bursty:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;No coding sessions = no inference charges
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The trade-off is that long agentic sessions can use an unpredictable number of tokens.&lt;/p&gt;

&lt;p&gt;Claude Code is commonly used through:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Pro: &lt;code&gt;$20/month&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Claude Max: &lt;code&gt;$100/month&lt;/code&gt; or &lt;code&gt;$200/month&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Usage-based API billing through the Claude Console&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Subscriptions make costs more predictable, although they include usage limits. Anthropic raised weekly limits by 50% in July 2026; see &lt;a href="http://apidog.com/blog/claude-code-weekly-limits-50-percent-increase-july-2026?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;the Claude Code weekly limits increase coverage&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;A useful rule of thumb:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Light or intermittent use&lt;/strong&gt;: per-token pricing may be cheaper.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Daily, heavy use&lt;/strong&gt;: a subscription may be easier to budget.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Because dsh can connect to Anthropic’s API, the model provider and harness decision can be evaluated independently.&lt;/p&gt;

&lt;h2&gt;
  
  
  Maturity: developer preview vs established ecosystem
&lt;/h2&gt;

&lt;p&gt;DeepSeek Harness is explicitly a developer preview. Its README states:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;THERE WILL BE COMPATIBILITY-BREAKING CHANGES.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Plan accordingly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Pin your configuration where possible.&lt;/li&gt;
&lt;li&gt;Expect settings and plugins to change.&lt;/li&gt;
&lt;li&gt;Test upgrades in a non-production workspace.&lt;/li&gt;
&lt;li&gt;Avoid building critical workflows without a maintenance plan.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Claude Code has been available to developers since early 2025. Its ecosystem includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="http://CLAUDE.md" rel="noopener noreferrer"&gt;&lt;code&gt;CLAUDE.md&lt;/code&gt;&lt;/a&gt; project memory&lt;/li&gt;
&lt;li&gt;Auto memory&lt;/li&gt;
&lt;li&gt;Skills for reusable workflows&lt;/li&gt;
&lt;li&gt;Hooks that run commands around agent actions&lt;/li&gt;
&lt;li&gt;Subagents coordinated by a lead agent&lt;/li&gt;
&lt;li&gt;Agent SDK&lt;/li&gt;
&lt;li&gt;GitHub Actions integration&lt;/li&gt;
&lt;li&gt;GitLab CI/CD integration&lt;/li&gt;
&lt;li&gt;Scheduled routines&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;These features are documented and versioned. For another comparison of ecosystem maturity, see &lt;a href="http://apidog.com/blog/claude-code-vs-codex-cli?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Code vs Codex CLI&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;dsh has substantial early community activity, with roughly 169k stars and 18.1k forks in its first week as of August 20. That can accelerate plugin development, but it does not remove preview-stage risk.&lt;/p&gt;

&lt;h2&gt;
  
  
  Extensibility: plugin kernel vs extension points
&lt;/h2&gt;

&lt;p&gt;Both tools are extensible, but at different layers.&lt;/p&gt;

&lt;h3&gt;
  
  
  DeepSeek Harness: plugins are the architecture
&lt;/h3&gt;

&lt;p&gt;dsh is built on Cordis, a plugin kernel described in the paper &lt;em&gt;A Programming Paradigm for Spatiotemporal Composability&lt;/em&gt;.&lt;/p&gt;

&lt;p&gt;In dsh, components such as these are replaceable plugins:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Model adapters&lt;/li&gt;
&lt;li&gt;Tool registries&lt;/li&gt;
&lt;li&gt;Session logging&lt;/li&gt;
&lt;li&gt;Agent loop behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, if you need custom retry behavior for failed tool calls, dsh’s architecture allows you to replace the relevant loop component instead of waiting for a vendor feature.&lt;/p&gt;

&lt;p&gt;Discover community packages through the GitHub &lt;code&gt;dsh-plugin&lt;/code&gt; topic.&lt;/p&gt;

&lt;h3&gt;
  
  
  Claude Code: stable extension seams
&lt;/h3&gt;

&lt;p&gt;Claude Code exposes supported extension points:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Plugins&lt;/li&gt;
&lt;li&gt;Skills&lt;/li&gt;
&lt;li&gt;Hooks&lt;/li&gt;
&lt;li&gt;MCP servers&lt;/li&gt;
&lt;li&gt;Subagents&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You customize behavior without replacing the core agent loop. This is usually safer for teams that prioritize upgrade stability.&lt;/p&gt;

&lt;p&gt;Choose dsh if you need to reshape the engine. Choose Claude Code if documented customization points are sufficient.&lt;/p&gt;

&lt;h2&gt;
  
  
  Permissions: approval prompts vs six documented modes
&lt;/h2&gt;

&lt;p&gt;Coding agents need controls before they edit files, execute commands, or access external services.&lt;/p&gt;

&lt;p&gt;dsh’s web UI asks for approval when an operation requires it under the active permission policy. That behavior is documented, but the public documentation does not yet define all policy levels and semantics in detail.&lt;/p&gt;

&lt;p&gt;Claude Code’s &lt;a href="https://code.claude.com/docs/en/permissions" rel="noopener noreferrer"&gt;permission system&lt;/a&gt; documents six modes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;default
acceptEdits
plan
auto
dontAsk
bypassPermissions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;It also supports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Fine-grained allow and deny rules by tool or command&lt;/li&gt;
&lt;li&gt;Working-directory boundaries&lt;/li&gt;
&lt;li&gt;Organization-managed policies&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;plan&lt;/code&gt; mode for exploration without edits&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;auto&lt;/code&gt; mode using a background classifier to review actions&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;bypassPermissions&lt;/code&gt; for sandboxed containers&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For regulated repositories, junior-heavy teams, or autonomous CI jobs, Claude Code currently has more documented controls.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP: native support vs a community plugin
&lt;/h2&gt;

&lt;p&gt;Model Context Protocol (MCP) lets coding agents access external systems such as databases, ticket trackers, and API specifications.&lt;/p&gt;

&lt;p&gt;Claude Code supports MCP natively. MCP servers are a first-class, documented integration, and their tools use the same permission model as other agent actions.&lt;/p&gt;

&lt;p&gt;In dsh, MCP is currently provided through the community-maintained &lt;code&gt;dsh-mcp-manager&lt;/code&gt; plugin. It provides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A settings page for MCP configuration&lt;/li&gt;
&lt;li&gt;Remote HTTP server support&lt;/li&gt;
&lt;li&gt;Local stdio server support&lt;/li&gt;
&lt;li&gt;OAuth or static-token authentication&lt;/li&gt;
&lt;li&gt;Per-project server configuration&lt;/li&gt;
&lt;li&gt;Tool registration with names such as:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;mcp__&amp;lt;name&amp;gt;__*
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This works, but it is not core functionality and carries community-plugin and preview-stage maintenance risk.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use Apidog as the API contract layer
&lt;/h2&gt;

&lt;p&gt;For API-focused work, connect your agent to the real API contract instead of relying on generated assumptions.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog MCP Server&lt;/a&gt; gives an MCP-capable coding agent access to your team’s API specification. That helps the agent generate client code against actual endpoints, fields, and schemas.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;In Claude Code, connect Apidog through native MCP support.&lt;/li&gt;
&lt;li&gt;In dsh, connect it through &lt;code&gt;dsh-mcp-manager&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;You can also &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;download Apidog&lt;/a&gt; and run its MCP server against the same project directory used by either agent.&lt;/p&gt;

&lt;p&gt;A practical workflow looks like this:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Maintain the API specification in Apidog.&lt;/li&gt;
&lt;li&gt;Expose it to the agent through MCP.&lt;/li&gt;
&lt;li&gt;Ask the agent to implement or update API client code.&lt;/li&gt;
&lt;li&gt;Run Apidog CLI regression tests in CI.&lt;/li&gt;
&lt;li&gt;Review generated changes before merge.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The same API test suite can validate changes produced by either agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which one should you pick?
&lt;/h2&gt;

&lt;p&gt;There is no universal winner. The right choice depends on your constraints.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pick DeepSeek Harness if
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;You need source access, MIT licensing, auditability, or forkability.&lt;/li&gt;
&lt;li&gt;You want to use multiple model providers from one harness.&lt;/li&gt;
&lt;li&gt;You need OpenAI-compatible local model support.&lt;/li&gt;
&lt;li&gt;You want to modify agent internals through plugins.&lt;/li&gt;
&lt;li&gt;You can tolerate configuration and plugin churn.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Start with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @deepseek-ai/dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then test it in a non-critical repository before standardizing on it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pick Claude Code if
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;You need a mature product now.&lt;/li&gt;
&lt;li&gt;You work across terminal, IDE, desktop, browser, mobile, and CI.&lt;/li&gt;
&lt;li&gt;You require documented enterprise permission controls.&lt;/li&gt;
&lt;li&gt;You depend on native MCP integration.&lt;/li&gt;
&lt;li&gt;You prefer predictable subscription-based billing.&lt;/li&gt;
&lt;li&gt;You are comfortable standardizing on Claude models.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Consider running both during evaluation
&lt;/h3&gt;

&lt;p&gt;A practical approach is to test both against the same repository and API contract:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Run dsh locally.&lt;/li&gt;
&lt;li&gt;Configure the model provider you want to evaluate.&lt;/li&gt;
&lt;li&gt;Connect both agents to the same Apidog MCP server.&lt;/li&gt;
&lt;li&gt;Run the same implementation tasks.&lt;/li&gt;
&lt;li&gt;Validate output with the same Apidog CLI regression suite.&lt;/li&gt;
&lt;li&gt;Compare maintenance effort, permission behavior, cost, and developer workflow fit.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Whichever agent you choose, keep the API layer reliable with &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;: one tested specification, available to either agent through MCP, and validated by CLI regression tests after every agent-generated change.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is DeepSeek Harness actually open source, unlike Claude Code?
&lt;/h3&gt;

&lt;p&gt;Yes. dsh is MIT licensed, and its source is available on GitHub, including the agent loop and plugin kernel. Claude Code’s public repository uses an all-rights-reserved notice under Anthropic’s Commercial Terms and does not provide source code you can fork.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can DeepSeek Harness use Claude models?
&lt;/h3&gt;

&lt;p&gt;Yes. dsh includes catalog providers for Anthropic, OpenAI, Bedrock, Vertex, and Azure, plus custom OpenAI-compatible endpoints through &lt;code&gt;settings.yaml&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;The reverse is not true: Claude Code runs Claude models only, whether accessed directly through Anthropic or through Bedrock, Vertex, or Foundry.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is DeepSeek Harness stable enough for daily work?
&lt;/h3&gt;

&lt;p&gt;It is a developer preview, and the README explicitly warns about compatibility-breaking changes. You can use it for real work, but expect configuration and plugin churn.&lt;/p&gt;

&lt;p&gt;Claude Code is the safer option for workflows you cannot afford to rebuild.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do both agents work with Apidog?
&lt;/h3&gt;

&lt;p&gt;Yes. Apidog’s MCP server exposes your API specification to MCP-capable agents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Code supports it natively.&lt;/li&gt;
&lt;li&gt;dsh supports it through the community &lt;code&gt;dsh-mcp-manager&lt;/code&gt; plugin.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The Apidog CLI can also run scripted regression tests from either agent’s terminal workflow. See &lt;a href="http://apidog.com/blog/apidog-cli-in-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;using Apidog CLI in DeepSeek Harness&lt;/a&gt;.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>What is DeepSeek Harness (dsh)? The Open-Source Claude Code Rival, Explained</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Thu, 20 Aug 2026 04:28:51 +0000</pubDate>
      <link>https://dev.to/hassann/what-is-deepseek-harness-dsh-the-open-source-claude-code-rival-explained-103c</link>
      <guid>https://dev.to/hassann/what-is-deepseek-harness-dsh-the-open-source-claude-code-rival-explained-103c</guid>
      <description>&lt;p&gt;DeepSeek shipped something unusual on August 13, 2026: not a model, but the machine that runs one. DeepSeek Harness (&lt;code&gt;dsh&lt;/code&gt;) is the company’s official open-source agent harness—the software layer that turns a large language model into a coding agent with a session loop, tool execution, permission checks, and a local web UI. It launched alongside DeepSeek V4-Pro on the API, and &lt;a href="https://venturebeat.com/technology/deepseek-harness-launches-as-open-source-rival-to-claude-code-alongside-v4-pro-on-api-with-higher-prices" rel="noopener noreferrer"&gt;VentureBeat framed it&lt;/a&gt; as an open-source rival to Claude Code.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;The developer community reacted quickly. As of August 20, the &lt;a href="https://github.com/deepseek-ai/deepseek-harness" rel="noopener noreferrer"&gt;deepseek-harness repository&lt;/a&gt; had roughly 169,000 stars and 18,100 forks—one week after release. That level of interest suggests developers want an agent harness they can inspect, modify, and connect to different models.&lt;/p&gt;

&lt;h2&gt;
  
  
  What DeepSeek Harness actually is
&lt;/h2&gt;

&lt;p&gt;A harness is everything around the model. The model predicts tokens; the harness decides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What context the model receives&lt;/li&gt;
&lt;li&gt;Which tools it can call&lt;/li&gt;
&lt;li&gt;How file edits and shell commands are approved&lt;/li&gt;
&lt;li&gt;How multi-step sessions are stored and resumed&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Claude Code, Codex CLI, and Gemini CLI are all harnesses wrapped around their vendors’ models. For background on two of those tools, see &lt;a href="http://apidog.com/blog/claude-code-vs-codex-cli?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Code vs Codex CLI&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;DeepSeek Harness has three defining properties:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Official:&lt;/strong&gt; It is a first-party project from DeepSeek AI, not a community wrapper around the DeepSeek API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open source:&lt;/strong&gt; It is MIT licensed, with third-party dependencies documented in the repository’s &lt;code&gt;THIRD_PARTY_NOTICES&lt;/code&gt; file. You can inspect the agent loop that operates on your codebase.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Developer preview:&lt;/strong&gt; The README warns: “THERE WILL BE COMPATIBILITY-BREAKING CHANGES.” Treat that as an operational constraint.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The launch timing also matters. &lt;code&gt;dsh&lt;/code&gt; arrived alongside DeepSeek V4-Pro on the API, so the harness and its flagship default model were introduced as a pair. For the model side, see the guide to the &lt;a href="http://apidog.com/blog/how-to-use-deepseek-v4-pro-0813-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek V4-Pro API&lt;/a&gt;, including endpoints, model IDs, and request examples.&lt;/p&gt;

&lt;h2&gt;
  
  
  The architecture: everything is a plugin
&lt;/h2&gt;

&lt;p&gt;Most coding-agent harnesses are monolithic. The agent loop, model client, tool definitions, and session store ship as one application. You may be able to configure or extend them, but you cannot replace the core components.&lt;/p&gt;

&lt;p&gt;&lt;code&gt;dsh&lt;/code&gt; takes the opposite approach. Its design principle is “everything is a plugin,” built on a framework called &lt;strong&gt;Cordis&lt;/strong&gt;. Cordis’s design is described in the paper “A Programming Paradigm for Spatiotemporal Composability.”&lt;/p&gt;

&lt;p&gt;In practical terms, components that are usually tightly coupled become replaceable modules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Model adapter:&lt;/strong&gt; The plugin that communicates with the LLM API.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool registry:&lt;/strong&gt; The tools available to the agent, such as file editing, shell execution, and search.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session log:&lt;/strong&gt; The mechanism for recording and replaying sessions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent loop:&lt;/strong&gt; The decide-act-observe cycle itself.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This matters when you want to experiment with context management, permission policies, or repository-specific tool sets. With a monolithic agent, you wait for the vendor to implement your idea. With &lt;code&gt;dsh&lt;/code&gt;, you can build a plugin.&lt;/p&gt;

&lt;p&gt;The trade-off is compatibility risk. A system where every component is replaceable has more ways to break, especially when the project explicitly warns about compatibility-breaking changes. The current trade-off is flexibility now in exchange for stability later.&lt;/p&gt;

&lt;h2&gt;
  
  
  Quick start: run a local agent
&lt;/h2&gt;

&lt;p&gt;Start the local web UI with one command:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @deepseek-ai/dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This launches the UI at &lt;code&gt;http://127.0.0.1:3080&lt;/code&gt; and opens it in your browser. To prevent the browser from opening automatically, use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npx @deepseek-ai/dsh web &lt;span class="nt"&gt;--no-open&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;There is no global install, and account creation is not required as an installation step.&lt;/p&gt;

&lt;h3&gt;
  
  
  Build from source
&lt;/h3&gt;

&lt;p&gt;To run the project from source:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;git clone https://github.com/deepseek-ai/deepseek-harness
&lt;span class="nb"&gt;cd &lt;/span&gt;deepseek-harness
pnpm &lt;span class="nb"&gt;install
&lt;/span&gt;pnpm run build
pnpm dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Complete the first-run setup
&lt;/h3&gt;

&lt;p&gt;The first-run flow has three steps:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Configure a DeepSeek API key.&lt;/strong&gt; Add it in Settings. Credentials are stored in &lt;code&gt;$DSH_HOME/.credentials.yaml&lt;/code&gt;, separate from the main settings file, which stores references to them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Choose a workspace.&lt;/strong&gt; Select the project directory where you started &lt;code&gt;dsh&lt;/code&gt;. The session composer remains unavailable until a workspace is selected. This gives the agent an explicit scope for the files it can inspect and modify.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run a task and approve operations.&lt;/strong&gt; The web UI asks for approval before operations covered by the active permission policy. File writes and shell commands appear as prompts instead of running silently.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Use profiles and headless mode
&lt;/h3&gt;

&lt;p&gt;The web UI is one entry point into the underlying profile system:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dsh web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is shorthand for:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dsh &lt;span class="nt"&gt;--profile&lt;/span&gt; web
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Profiles are stored under:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$DSH_HOME/profiles/&amp;lt;name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For scripts and CI, use headless mode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dsh &lt;span class="nt"&gt;--profile&lt;/span&gt; headless &lt;span class="s2"&gt;"job"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This runs one fresh session, prints the result, and exits.&lt;/p&gt;

&lt;p&gt;You can also manage profile plugins through pnpm:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dsh plugin
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;To inspect the composed configuration without starting the harness:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;dsh &lt;span class="nt"&gt;--dump-config&lt;/span&gt;
dsh &lt;span class="nt"&gt;--dump-default-config&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;See the &lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/apps/cli/README.md" rel="noopener noreferrer"&gt;CLI README&lt;/a&gt; for the complete command reference.&lt;/p&gt;

&lt;h2&gt;
  
  
  What models can it run?
&lt;/h2&gt;

&lt;p&gt;DeepSeek models are the default, with V4-Pro as the headline pairing. DeepSeek also made its off-peak discount permanent, which affects the cost of running an agent that uses tokens continuously. See the post on the &lt;a href="http://apidog.com/blog/deepseek-v4-pro-permanent-price-cut?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek V4-Pro price cut&lt;/a&gt;, and consult the official &lt;a href="https://api-docs.deepseek.com" rel="noopener noreferrer"&gt;DeepSeek API documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;The model adapter is a plugin, so &lt;code&gt;dsh&lt;/code&gt; supports more than DeepSeek’s models through two paths:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Catalog providers:&lt;/strong&gt; Built-in provider entries for Anthropic, OpenAI, Bedrock, Vertex, and Azure, with provider-specific credential handling.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom providers:&lt;/strong&gt; Any OpenAI-compatible endpoint can be registered in &lt;code&gt;$DSH_HOME/settings.yaml&lt;/code&gt; with a base URL, an environment variable for the API key, and a model list. This includes local runtimes and API gateways.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Selecting a model makes it the default for new sessions. Each session records the model it started with, so changing models does not obscure your project history.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/deepseek-ai/deepseek-harness/blob/master/docs/user/guide/providers.md" rel="noopener noreferrer"&gt;providers guide&lt;/a&gt; documents the configuration format. For the complete YAML setup for custom endpoints, see &lt;a href="http://apidog.com/blog/run-any-model-in-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;how to run any model in DeepSeek Harness&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The plugin ecosystem, one week in
&lt;/h2&gt;

&lt;p&gt;Plugins are discoverable through the &lt;a href="https://github.com/topics/dsh-plugin" rel="noopener noreferrer"&gt;&lt;code&gt;dsh-plugin&lt;/code&gt; GitHub topic&lt;/a&gt;. The community also coordinates through GitHub Discussions and a Discord server.&lt;/p&gt;

&lt;p&gt;One week after launch, the ecosystem already includes several categories:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Desktop wrappers:&lt;/strong&gt; Projects such as &lt;code&gt;deepseek-harness-desktop&lt;/code&gt; (Tauri) and &lt;code&gt;dsh_desktop&lt;/code&gt; (Windows) package the web UI as native applications. These are community projects, not official DeepSeek releases. Review them carefully because they may access your API keys.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Capability plugins:&lt;/strong&gt; Community repositories such as &lt;code&gt;dsh-context&lt;/code&gt; and &lt;code&gt;dsh-vision-router&lt;/code&gt; extend what sessions can see and how requests are routed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP support:&lt;/strong&gt; The core does not include native Model Context Protocol support as of this writing. The community plugin &lt;code&gt;dsh-mcp-manager&lt;/code&gt; adds an MCP Settings page supporting remote HTTP or local stdio servers, OAuth or static-token authentication, tools named with the &lt;code&gt;mcp__&amp;lt;name&amp;gt;__*&lt;/code&gt; pattern, and per-project server configuration inside the workspace’s &lt;code&gt;.dsh&lt;/code&gt; directory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The MCP distinction is important. Saying that &lt;code&gt;dsh&lt;/code&gt; “supports MCP” is only accurate through the community plugin. The core may support MCP directly in the future, but it does not as of this writing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where your API workflow fits
&lt;/h2&gt;

&lt;p&gt;An agent harness ultimately executes API calls:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The model API that powers the agent&lt;/li&gt;
&lt;li&gt;The APIs inside the project you provide as its workspace&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When &lt;code&gt;dsh&lt;/code&gt; writes code against your backend, it relies on the API behavior it can infer from the codebase. If the implementation does not match the specification, the agent may generate code against the wrong contract and fail only at runtime.&lt;/p&gt;

&lt;p&gt;A practical workflow is to verify the API surface before giving the repository to the agent:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Design or import the OpenAPI specification.&lt;/li&gt;
&lt;li&gt;Test the real endpoints against that specification.&lt;/li&gt;
&lt;li&gt;Create mock servers with stable, spec-accurate responses.&lt;/li&gt;
&lt;li&gt;Let the agent develop against those mocks while the backend changes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; covers this API design, testing, and mocking workflow. Working against a verified mock reduces the chance of hallucinated integrations caused by stale code or incomplete endpoint documentation.&lt;/p&gt;

&lt;p&gt;There is also a direct integration path. The &lt;a href="http://apidog.com/blog/apidog-mcp-server?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog MCP Server&lt;/a&gt; exposes API specifications to AI tools over MCP. In &lt;code&gt;dsh&lt;/code&gt;, connect it through the community &lt;code&gt;dsh-mcp-manager&lt;/code&gt; plugin:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Install &lt;code&gt;dsh-mcp-manager&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Register the Apidog MCP Server.&lt;/li&gt;
&lt;li&gt;Start a session in the relevant workspace.&lt;/li&gt;
&lt;li&gt;Let the agent query the API specification instead of inferring it from source code.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The complete workflow, including CLI-based test runs that the agent can trigger, is covered in &lt;a href="http://apidog.com/blog/apidog-cli-in-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;using Apidog CLI in DeepSeek Harness&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;If you want to prepare the API side first, &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; and import your specification before experimenting with &lt;code&gt;dsh&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Should you try it now or wait?
&lt;/h2&gt;

&lt;p&gt;The right answer depends on how you plan to use it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Try it now if
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You want to understand agent internals.&lt;/strong&gt; &lt;code&gt;dsh&lt;/code&gt; is highly inspectable, and reading a real agent loop is a practical way to learn how coding agents work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You need model flexibility.&lt;/strong&gt; The pluggable model adapter supports different hosted and self-hosted providers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You build developer tooling.&lt;/strong&gt; The plugin ecosystem is new, so early contributors may have an opportunity to shape it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You already use DeepSeek’s API.&lt;/strong&gt; &lt;code&gt;dsh&lt;/code&gt; provides a first-party agent experience for V4-Pro.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Wait if
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You need a stable daily driver.&lt;/strong&gt; The compatibility-breaking-change warning means configurations, plugins, and workflows may break between versions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your organization requires supported tooling.&lt;/strong&gt; A developer preview with community plugins has a different risk profile from a generally available product with a support contract.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You expect mature UX.&lt;/strong&gt; Claude Code has had a longer head start on ergonomics, and a one-week-old preview will not match it in every area.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For most developers, the pragmatic approach is to use both: keep your current agent for production work, run &lt;code&gt;dsh&lt;/code&gt; in a side project, and evaluate it before relying on it for critical repositories. For a direct comparison with Claude Code, see &lt;a href="http://apidog.com/blog/deepseek-harness-vs-claude-code?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek Harness vs Claude Code&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is DeepSeek Harness free?
&lt;/h3&gt;

&lt;p&gt;The harness is free and open source under the MIT license. The model may cost money: API usage on DeepSeek’s platform, or through whichever provider you configure, is billed by that provider.&lt;/p&gt;

&lt;p&gt;Because the adapter layer is pluggable, you can also connect &lt;code&gt;dsh&lt;/code&gt; to a locally hosted model and avoid per-token API charges. See &lt;a href="http://apidog.com/blog/run-any-model-in-deepseek-harness?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;run any model in DeepSeek Harness&lt;/a&gt; for the setup.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does dsh only work with DeepSeek models?
&lt;/h3&gt;

&lt;p&gt;No. DeepSeek models are the default, but the model adapter is a plugin. Catalog providers include Anthropic, OpenAI, Bedrock, Vertex, and Azure. You can also add any OpenAI-compatible endpoint through &lt;code&gt;$DSH_HOME/settings.yaml&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is DeepSeek Harness safe to run on my codebase?
&lt;/h3&gt;

&lt;p&gt;Safety depends on the permission model, the installed plugins, and your review process.&lt;/p&gt;

&lt;p&gt;The web UI requires a workspace before a session can run and prompts before operations that require approval under the active policy. However, &lt;code&gt;dsh&lt;/code&gt; is still a developer preview, and community plugins—including desktop wrappers—are third-party code that may handle your API keys.&lt;/p&gt;

&lt;p&gt;Review every plugin before installing it, and avoid using the preview on repositories where an accidental edit could cause serious damage.&lt;/p&gt;

&lt;h3&gt;
  
  
  How is a “harness” different from a model?
&lt;/h3&gt;

&lt;p&gt;The model is the reasoning engine; the harness is the system that lets it act.&lt;/p&gt;

&lt;p&gt;Session management, tool calls, file access, permission prompts, and context assembly all belong to the harness. Two agents using the same model can behave very differently because their harnesses implement different tools, policies, and session loops. That is why the harness layer has become a major area of coding-agent development.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Self-Hosting GLM-5.3: Get Ready for the Open-Weights Drop</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Sun, 16 Aug 2026 15:33:37 +0000</pubDate>
      <link>https://dev.to/hassann/self-hosting-glm-53-get-ready-for-the-open-weights-drop-4221</link>
      <guid>https://dev.to/hassann/self-hosting-glm-53-get-ready-for-the-open-weights-drop-4221</guid>
      <description>&lt;p&gt;Zhipu AI released GLM-5.3 on August 14, 2026. The key detail for infrastructure teams is the planned open-weights release roughly two weeks later, around August 28, on Zhipu’s &lt;a href="https://huggingface.co/zai-org" rel="noopener noreferrer"&gt;Hugging Face organization&lt;/a&gt;. Use that window to size hardware, select a serving stack, and capture a regression baseline against the hosted API before the safetensors shards are available.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;Zhipu reports major gains over GLM-5.2: 50% stronger coding capability, a Terminal-Bench 3.0 increase from 4.6 to 28.3, and agent performance described as “approaching Claude Fable 5,” according to &lt;a href="https://finance.biggo.com/news/0b571a42-9531-433c-b81b-c8468d173989" rel="noopener noreferrer"&gt;launch reporting&lt;/a&gt;. For benchmark context and known gaps versus frontier models, see the &lt;a href="http://apidog.com/blog/what-is-glm-5-3?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GLM-5.3 explainer&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This post focuses on implementation: what to prepare before the weights arrive so you can run GLM-5.3 yourself.&lt;/p&gt;

&lt;p&gt;The weights are not downloadable yet. Until they are, use the hosted API as your reference implementation. Capture responses now, then replay the same test collection against your self-hosted endpoint later with &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GLM-5.3 launched on August 14, 2026. Zhipu says open weights should arrive around August 28 on &lt;a href="https://huggingface.co/zai-org" rel="noopener noreferrer"&gt;huggingface.co/zai-org&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;According to &lt;a href="http://Z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt; documentation, the GLM-5 family uses a Mixture of Experts architecture with 744B total parameters, roughly 40B active parameters per pass, and a 200K-token context window.&lt;/li&gt;
&lt;li&gt;At BF16, weights alone are approximately 1.5 TB. At FP8, they are approximately 744 GB, excluding KV cache.&lt;/li&gt;
&lt;li&gt;Previous GLM-5 releases included BF16 and FP8 repositories. Expect a similar &lt;code&gt;GLM-5.3&lt;/code&gt; and &lt;code&gt;GLM-5.3-FP8&lt;/code&gt; release pattern, while community GGUF quants may arrive later.&lt;/li&gt;
&lt;li&gt;vLLM and SGLang are the practical day-one serving options. Both expose OpenAI-compatible APIs.&lt;/li&gt;
&lt;li&gt;Build a hosted-versus-local regression suite now using &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;: one collection, two environments, and assertions for response shape and expected content.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Zhipu is releasing, and when
&lt;/h2&gt;

&lt;p&gt;Zhipu, branded internationally as &lt;a href="http://Z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt;, launched the GLM-5.3 API with a promise to publish open weights around August 28, 2026.&lt;/p&gt;

&lt;p&gt;Zhipu says the delay supports its most extensive risk-review process to date. That matters because the model reportedly scored 84.5% on CyberGym, slightly above Claude Mythos 5 and GPT-5.6 Sol. &lt;a href="https://www.seekingalpha.com/news/4472588-chinese-openai-challenger-zhipu-is-said-to-unveil-new-open-source-model" rel="noopener noreferrer"&gt;Seeking Alpha&lt;/a&gt; describes the release as part of Zhipu’s effort to maintain its open-model position against DeepSeek.&lt;/p&gt;

&lt;p&gt;For self-hosting, two release details matter:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The base model is unchanged.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
GLM-5.3 is GLM-5 with scaled post-training. The serving architecture should therefore match the GLM-5 and GLM-5.2 family already supported by vLLM and SGLang. No new attention mechanism or tokenizer changes are expected.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;The repository pattern is established.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Zhipu’s Hugging Face organization already hosts GLM-5, GLM-5.1, and GLM-5.2, including companion FP8 repositories. GLM-5.2 alone shows 2.69M downloads. Plan for a BF16 safetensors release plus an official FP8 release.&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;License terms for GLM-5.3 were not confirmed in launch coverage. Read the model card before deploying it in a commercial product.&lt;/p&gt;

&lt;h2&gt;
  
  
  What 744B total and 40B active means for hardware
&lt;/h2&gt;

&lt;p&gt;The GLM-5 family is a Mixture of Experts model with 744B total parameters, around 40B active parameters per forward pass, and 200K context, according to &lt;a href="https://docs.z.ai/guides/llm/glm-5" rel="noopener noreferrer"&gt;Z.ai’s GLM-5 documentation&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;These are family-level specifications rather than GLM-5.3-specific claims. However, because Zhipu says the base model is unchanged, they are appropriate planning numbers.&lt;/p&gt;

&lt;p&gt;MoE models create an important memory-versus-compute split:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Compute resembles a 40B dense model.&lt;/strong&gt; Only routed experts execute for each token, so inference throughput can be substantially better than a 744B dense model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Memory resembles a 744B model.&lt;/strong&gt; All experts must remain addressable. At 2 bytes per parameter, BF16 weights require approximately 1.5 TB. At 1 byte per parameter, FP8 weights require approximately 744 GB. Neither estimate includes KV cache.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Precision&lt;/th&gt;
&lt;th&gt;Weight footprint (arithmetic)&lt;/th&gt;
&lt;th&gt;Practical deployment target&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;BF16&lt;/td&gt;
&lt;td&gt;~1.5 TB&lt;/td&gt;
&lt;td&gt;Multi-node cluster or the largest single-server GPU configurations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FP8, official&lt;/td&gt;
&lt;td&gt;~745 GB&lt;/td&gt;
&lt;td&gt;High-end multi-GPU server, potentially one node&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;INT4-class community quants&lt;/td&gt;
&lt;td&gt;~370–400 GB&lt;/td&gt;
&lt;td&gt;Smaller multi-GPU systems; validate output quality first&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If you have a single consumer GPU, full GLM-5.3 weights are not the target deployment. Instead:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Rent GPU capacity for evaluation.&lt;/li&gt;
&lt;li&gt;Wait for tested community quantizations.&lt;/li&gt;
&lt;li&gt;Keep GLM-5.3 on the hosted API while running smaller open models locally.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;a href="http://apidog.com/blog/best-local-llms-2026?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;best local LLMs in 2026 guide&lt;/a&gt; covers models that fit single-GPU and workstation hardware.&lt;/p&gt;

&lt;p&gt;Also treat the 200K context window as a memory decision. KV cache grows with both context length and batch size. Set a deployment-specific context cap before launch rather than defaulting to the model maximum.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pick a serving stack before the weights land
&lt;/h2&gt;

&lt;p&gt;Three serving paths matter, but they will not all be ready at the same time.&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 1: vLLM
&lt;/h3&gt;

&lt;p&gt;vLLM is the safest default for GLM-5.3-scale inference. It already supports the GLM-5 family and provides:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;MoE routing&lt;/li&gt;
&lt;li&gt;Tensor parallelism&lt;/li&gt;
&lt;li&gt;Expert parallelism&lt;/li&gt;
&lt;li&gt;Multi-GPU and multi-node deployment options&lt;/li&gt;
&lt;li&gt;An OpenAI-compatible API server&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A release-day command may look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;vllm serve zai-org/GLM-5.3-FP8 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--tensor-parallel-size&lt;/span&gt; 8 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--max-model-len&lt;/span&gt; 65536 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--served-model-name&lt;/span&gt; glm-5.3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Treat this as a template, not a production-ready command:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The exact repository name must be confirmed after release.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;--tensor-parallel-size&lt;/code&gt; depends on GPU count, interconnect, and available memory.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;--max-model-len&lt;/code&gt; should reflect your KV-cache budget, not just the model’s advertised maximum.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Option 2: SGLang
&lt;/h3&gt;

&lt;p&gt;SGLang is the primary alternative. It is especially relevant for agent workloads that repeatedly send long shared prompts because its radix-tree prefix caching can reduce repeated prefill work.&lt;/p&gt;

&lt;p&gt;Like vLLM, SGLang exposes an OpenAI-compatible endpoint. That means you can switch serving stacks without rewriting your application client.&lt;/p&gt;

&lt;h3&gt;
  
  
  Option 3: llama.cpp, Ollama, and LM Studio
&lt;/h3&gt;

&lt;p&gt;The llama.cpp ecosystem requires GGUF conversions. Those are usually created by the community days or weeks after safetensors weights are released.&lt;/p&gt;

&lt;p&gt;This route may eventually make more aggressive quantization practical, but verify quality against your own baseline. Do not assume a community quant behaves like the hosted model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prepare the runtime now
&lt;/h3&gt;

&lt;p&gt;Before the release window:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Install vLLM or SGLang.&lt;/li&gt;
&lt;li&gt;Test CUDA drivers, container runtime, networking, and GPU visibility.&lt;/li&gt;
&lt;li&gt;Dry-run the stack with GLM-5.2 weights if your hardware supports it.&lt;/li&gt;
&lt;li&gt;Otherwise, use another available MoE model to validate your deployment path.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Avoid debugging CUDA, NCCL, storage mounts, and model-serving configuration on release day.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use the hosted API as your baseline
&lt;/h2&gt;

&lt;p&gt;Before self-hosting, record what the reference implementation returns.&lt;/p&gt;

&lt;p&gt;The hosted API gives you a comparison point when local outputs differ. Without a baseline, you cannot distinguish between:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Quantization quality loss&lt;/li&gt;
&lt;li&gt;A serving-stack configuration issue&lt;/li&gt;
&lt;li&gt;A model-loading bug&lt;/li&gt;
&lt;li&gt;Normal sampling variance&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The hosted API is OpenAI-compatible:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;International endpoint: &lt;code&gt;https://api.z.ai/api/paas/v4/chat/completions&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Mainland China endpoint: &lt;code&gt;https://open.bigmodel.cn/api/paas/v4/chat/completions&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Authentication: &lt;code&gt;Authorization: Bearer &amp;lt;key&amp;gt;&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="http://Z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt;’s documentation currently lists &lt;code&gt;glm-5&lt;/code&gt;. Confirm the exact GLM-5.3 model ID in the &lt;a href="https://docs.z.ai/" rel="noopener noreferrer"&gt;official docs&lt;/a&gt; before running production tests.&lt;/p&gt;

&lt;p&gt;For regional setup details, use the &lt;a href="http://apidog.com/blog/how-to-use-glm-5-3-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GLM-5.3 API quickstart&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Start by capturing deterministic-ish baselines with &lt;code&gt;temperature: 0&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.z.ai/api/paas/v4/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$GLM_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "glm-5.3",
    "temperature": 0,
    "messages": [
      {
        "role": "user",
        "content": "Write a Python function that parses RFC 3339 timestamps and returns UTC datetimes. Include error handling for invalid input."
      }
    ]
  }'&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; baseline-rfc3339.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Build 20 to 50 baseline prompts from your actual workload:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Code generation and refactoring tasks&lt;/li&gt;
&lt;li&gt;Tool-calling and agent loops&lt;/li&gt;
&lt;li&gt;Structured JSON output&lt;/li&gt;
&lt;li&gt;Long-context summarization&lt;/li&gt;
&lt;li&gt;Domain-specific prompts&lt;/li&gt;
&lt;li&gt;Security-sensitive transformations&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Temperature zero does not guarantee perfectly identical output. It does reduce variance enough that meaningful quality regressions become easier to identify.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build the regression harness in Apidog
&lt;/h2&gt;

&lt;p&gt;Raw &lt;code&gt;curl&lt;/code&gt; scripts work for a single endpoint. They become difficult to manage when you compare hosted and local deployments across multiple quantization levels, context limits, and serving stacks.&lt;/p&gt;

&lt;p&gt;A structured test collection gives you repeatability. This is standard API regression testing, similar to the workflow in the &lt;a href="http://apidog.com/blog/api-testing-tool-qa-engineers?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;API testing guide for QA engineers&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Create one collection for all baseline prompts
&lt;/h3&gt;

&lt;p&gt;Create one request per baseline case against the chat completions endpoint.&lt;/p&gt;

&lt;p&gt;Because the API shape is OpenAI-compatible, you can use an OpenAI-style schema to validate request structure.&lt;/p&gt;

&lt;p&gt;Suggested collection layout:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GLM-5.3 Regression
├── code-rfc3339-parser
├── code-refactor-nested-loops
├── tool-call-weather
├── json-mode-extraction
├── long-context-summary
└── agent-planning
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Create &lt;code&gt;hosted&lt;/code&gt; and &lt;code&gt;local&lt;/code&gt; environments
&lt;/h3&gt;

&lt;p&gt;Configure two environments.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Hosted&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;base_url = https://api.z.ai/api/paas/v4
api_key = &amp;lt;your GLM API key&amp;gt;
model = glm-5.3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Local&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;base_url = http://localhost:8000/v1
api_key = local-serving
model = glm-5.3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use environment variables in every request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST {{base_url}}/chat/completions
Authorization: Bearer {{api_key}}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Switching between hosted and local then becomes an environment change rather than a request rewrite.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Assert response shape first
&lt;/h3&gt;

&lt;p&gt;Start with stable assertions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;HTTP status is &lt;code&gt;200&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;choices[0].message.content&lt;/code&gt; is not empty&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;usage&lt;/code&gt; exists and contains sensible token counts&lt;/li&gt;
&lt;li&gt;The expected response type is returned&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Then add content-level checks that tolerate wording variation.&lt;/p&gt;

&lt;p&gt;For a Python-code prompt, examples include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Response contains "def "
Response contains "datetime"
Response contains "try"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Avoid exact-output equality for generative responses unless you control all relevant decoding settings and have confirmed deterministic behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Save hosted responses as fixtures
&lt;/h3&gt;

&lt;p&gt;Save hosted responses as examples or reference fixtures.&lt;/p&gt;

&lt;p&gt;On release day:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Start the local GLM-5.3 server.&lt;/li&gt;
&lt;li&gt;Switch the collection environment from &lt;code&gt;hosted&lt;/code&gt; to &lt;code&gt;local&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Run the collection.&lt;/li&gt;
&lt;li&gt;Compare local results with hosted fixtures.&lt;/li&gt;
&lt;li&gt;Investigate failures before directing production traffic to the new deployment.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  5. Run the collection in CI or from the CLI
&lt;/h3&gt;

&lt;p&gt;Use Apidog’s runner to execute the collection headlessly.&lt;/p&gt;

&lt;p&gt;That lets you rerun the same test suite for each change:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;FP8 vs INT4 quant
vLLM vs SGLang
32K vs 64K context cap
different tensor-parallel sizes
prefix-caching changes
new model revision
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The desired output is a repeatable answer to this question:&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Does this local deployment behave sufficiently like the hosted reference for our workload?&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Your client code does not need to change
&lt;/h2&gt;

&lt;p&gt;The OpenAI-compatible API convention is the practical advantage of this deployment path.&lt;/p&gt;

&lt;p&gt;Your application can switch from hosted GLM-5.3 to local GLM-5.3 through configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="c1"&gt;# Hosted: GLM_BASE_URL=https://api.z.ai/api/paas/v4
# Local:  GLM_BASE_URL=http://localhost:8000/v1
&lt;/span&gt;
&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GLM_BASE_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GLM_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;local-serving&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;glm-5.3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Refactor this function to remove the nested loops: ...&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When starting vLLM, set a stable served model name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;--served-model-name&lt;/span&gt; glm-5.3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That keeps the model string consistent between hosted and local environments.&lt;/p&gt;

&lt;p&gt;Streaming, tool calls, and JSON mode use the same general API surface. However, explicitly regression-test tool calling and streaming. These are common areas where local serving stacks can diverge from hosted behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost framing: hosted API versus your own GPUs
&lt;/h2&gt;

&lt;p&gt;Zhipu had not published GLM-5.3-specific API pricing at launch. Check the &lt;a href="https://docs.z.ai/guides/overview/pricing" rel="noopener noreferrer"&gt;official pricing page&lt;/a&gt; before building a per-token cost model.&lt;/p&gt;

&lt;p&gt;The tradeoff is structural:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Deployment option&lt;/th&gt;
&lt;th&gt;Main cost model&lt;/th&gt;
&lt;th&gt;Best fit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Hosted API&lt;/td&gt;
&lt;td&gt;Variable per-token cost&lt;/td&gt;
&lt;td&gt;Low or variable traffic, fast evaluation, no infrastructure overhead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-hosted&lt;/td&gt;
&lt;td&gt;Fixed GPU capacity plus operational cost&lt;/td&gt;
&lt;td&gt;Sustained utilization, governance requirements, latency control&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rented GPU evaluation&lt;/td&gt;
&lt;td&gt;Temporary hourly cost&lt;/td&gt;
&lt;td&gt;Benchmarking before committing to infrastructure&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Self-hosting a 744B-class MoE can make sense when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Sustained traffic is high enough to justify GPU capacity.&lt;/li&gt;
&lt;li&gt;Prompts or data must stay inside your network.&lt;/li&gt;
&lt;li&gt;You need latency, availability, or deployment control that a shared API cannot guarantee.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For low or uncertain demand, the hosted API is usually the safer cost profile. Renting GPUs for evaluation is also safer than buying hardware for a model you have not validated.&lt;/p&gt;

&lt;p&gt;There is also a pricing hedge. Provider pricing can change, as discussed in the &lt;a href="http://apidog.com/blog/deepseek-api-price-increase-cost-optimization?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek API price increase analysis&lt;/a&gt;. Open weights give you an alternative path if hosted pricing or availability changes.&lt;/p&gt;

&lt;h2&gt;
  
  
  Drop-day checklist
&lt;/h2&gt;

&lt;p&gt;Items 1 through 6 are preparation work you can complete before weights are released.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Choose your precision target: BF16, official FP8, or community quantization.&lt;/li&gt;
&lt;li&gt;Validate that your accessible hardware can support that tier using the memory ranges above.&lt;/li&gt;
&lt;li&gt;Install vLLM or SGLang and dry-run it with GLM-5.2 or another MoE model.&lt;/li&gt;
&lt;li&gt;Create a &lt;a href="http://Z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt; API key and confirm the GLM-5.3 model ID in the &lt;a href="https://docs.z.ai/" rel="noopener noreferrer"&gt;live docs&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Capture 20 to 50 hosted API baseline responses at temperature zero.&lt;/li&gt;
&lt;li&gt;Build an Apidog collection with &lt;code&gt;hosted&lt;/code&gt; and &lt;code&gt;local&lt;/code&gt; environments plus shape and content assertions.&lt;/li&gt;
&lt;li&gt;Decide the maximum served context length for each deployment tier.&lt;/li&gt;
&lt;li&gt;Watch &lt;a href="https://huggingface.co/zai-org" rel="noopener noreferrer"&gt;huggingface.co/zai-org&lt;/a&gt; for &lt;code&gt;GLM-5.3&lt;/code&gt; and &lt;code&gt;GLM-5.3-FP8&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Read the model card and license before commercial deployment.&lt;/li&gt;
&lt;li&gt;Download the weights and start the server.&lt;/li&gt;
&lt;li&gt;Point the &lt;code&gt;local&lt;/code&gt; environment to the serving endpoint.&lt;/li&gt;
&lt;li&gt;Run the regression collection.&lt;/li&gt;
&lt;li&gt;Diff local outputs against hosted fixtures.&lt;/li&gt;
&lt;li&gt;Investigate content-level failures before scaling traffic.&lt;/li&gt;
&lt;li&gt;Only then tune quantization, parallelism, prefix caching, and context limits.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Can I download GLM-5.3 weights right now?
&lt;/h3&gt;

&lt;p&gt;No. As of August 14, 2026, only the hosted API is live. Zhipu says open weights should arrive around August 28 on the &lt;a href="https://huggingface.co/zai-org" rel="noopener noreferrer"&gt;zai-org Hugging Face page&lt;/a&gt;, where GLM-5, GLM-5.1, and GLM-5.2 are already available.&lt;/p&gt;

&lt;h3&gt;
  
  
  Will GLM-5.3 run on a single consumer GPU?
&lt;/h3&gt;

&lt;p&gt;Not at full precision or official FP8 weights. The GLM-5 family’s 744B total parameters require roughly 744 GB at FP8 before KV cache. Even INT4-class community quants are expected to require multiple GPUs.&lt;/p&gt;

&lt;p&gt;For single-GPU budgets, run smaller open models locally and keep GLM-5.3 on the hosted API. The &lt;a href="http://apidog.com/blog/best-local-llms-2026?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;local LLM roundup&lt;/a&gt; covers practical alternatives.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which serving framework should I use for GLM-5.3?
&lt;/h3&gt;

&lt;p&gt;Use vLLM as the default choice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Existing GLM-5 family support&lt;/li&gt;
&lt;li&gt;MoE-aware parallelism&lt;/li&gt;
&lt;li&gt;OpenAI-compatible server&lt;/li&gt;
&lt;li&gt;Multi-GPU and multi-node deployment options&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use SGLang when shared long prefixes are common, such as agent loops that repeatedly resend system prompts and context.&lt;/p&gt;

&lt;p&gt;Use llama.cpp, Ollama, or LM Studio later, after community GGUF conversions become available.&lt;/p&gt;

&lt;h3&gt;
  
  
  Will existing OpenAI SDK code work with self-hosted GLM-5.3?
&lt;/h3&gt;

&lt;p&gt;Yes. Point the SDK’s &lt;code&gt;base_url&lt;/code&gt; to your vLLM or SGLang server instead of &lt;code&gt;https://api.z.ai/api/paas/v4&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Keep the same request shape and, if possible, preserve the same model name with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nt"&gt;--served-model-name&lt;/span&gt; glm-5.3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Test streaming and tool calling specifically because they are more likely to vary across serving implementations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why use the hosted API if I plan to self-host?
&lt;/h3&gt;

&lt;p&gt;The hosted API is your reference implementation.&lt;/p&gt;

&lt;p&gt;Without baseline responses, you cannot tell whether a local output difference is caused by:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Aggressive quantization&lt;/li&gt;
&lt;li&gt;A server configuration issue&lt;/li&gt;
&lt;li&gt;A deployment bug&lt;/li&gt;
&lt;li&gt;Normal model behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Capture hosted responses now using the &lt;a href="http://apidog.com/blog/how-to-use-glm-5-3-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GLM-5.3 API quickstart&lt;/a&gt;, then make drop day a regression-testing exercise instead of guesswork.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where GLM-5.3 fits in your stack
&lt;/h2&gt;

&lt;p&gt;GLM-5.3 is a major open-weights coding-model announcement: Zhipu reports first-place open-model results on Terminal-Bench 3.0 and Agents’ Last Exam, a CyberGym score above two frontier models, and a scheduled public weights release.&lt;/p&gt;

&lt;p&gt;The teams that benefit first will not necessarily be the teams with the largest GPU budgets. They will be the teams that use the release window to prepare:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Serving stack installed&lt;/li&gt;
&lt;li&gt;Precision tier selected&lt;/li&gt;
&lt;li&gt;Context limits defined&lt;/li&gt;
&lt;li&gt;Hosted baselines captured&lt;/li&gt;
&lt;li&gt;Regression harness ready&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Start with the checklist. Capture hosted baselines now, then use &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; to manage one collection across &lt;code&gt;hosted&lt;/code&gt; and &lt;code&gt;local&lt;/code&gt; environments. That turns “does this deployment work?” into a repeatable pass/fail report whenever you change a quantization level or serving flag.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>How to Use the GLM-5.3 API?</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Sun, 16 Aug 2026 15:32:38 +0000</pubDate>
      <link>https://dev.to/hassann/how-to-use-the-glm-53-api-4ie9</link>
      <guid>https://dev.to/hassann/how-to-use-the-glm-53-api-4ie9</guid>
      <description>&lt;p&gt;Zhipu AI, the Chinese lab that operates internationally as &lt;a href="http://Z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt;, released GLM-5.3 on August 14, 2026. Its internal evaluations report a 50% coding-capability improvement over GLM-5.2, with Terminal-Bench 3.0 rising from 4.6 to 28.3. Zhipu describes coding and agent capability as “approaching Claude Fable 5,” according to the &lt;a href="https://finance.biggo.com/news/0b571a42-9531-433c-b81b-c8468d173989" rel="noopener noreferrer"&gt;BigGo launch report&lt;/a&gt;. Open weights are expected about two weeks later. For benchmarks and capability details, see &lt;a href="http://apidog.com/blog/what-is-glm-5-3?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what GLM-5.3 is&lt;/a&gt;; this post focuses on getting the API working.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This quickstart covers getting an API key, making a cURL request, using Python and Node.js through the OpenAI SDK, streaming output, tuning core parameters, and testing requests in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; before wiring them into your application.&lt;/p&gt;

&lt;p&gt;The API is OpenAI-compatible. If you already use OpenAI-style chat completions, you primarily need to change the base URL, API key, and model ID.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;GLM-5.3 launched recently, so verify model IDs, pricing, and limits against the live Z.ai documentation before deploying.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GLM-5.3 was released on August 14, 2026. Zhipu reports a 50% coding improvement over GLM-5.2 and a Terminal-Bench 3.0 increase from 4.6 to 28.3.&lt;/li&gt;
&lt;li&gt;International endpoint: &lt;code&gt;POST https://api.z.ai/api/paas/v4/chat/completions&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Mainland China endpoint: &lt;code&gt;POST https://open.bigmodel.cn/api/paas/v4/chat/completions&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Authenticate with &lt;code&gt;Authorization: Bearer $GLM_API_KEY&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;The &lt;a href="https://docs.z.ai/guides/llm/glm-5" rel="noopener noreferrer"&gt;GLM-5 docs&lt;/a&gt; listed &lt;code&gt;glm-5&lt;/code&gt; at launch. The expected GLM-5.3 ID is &lt;code&gt;glm-5.3&lt;/code&gt;, but confirm before hardcoding it.&lt;/li&gt;
&lt;li&gt;Test request shapes, regional endpoints, model variables, and reasoning settings in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; before writing production code.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why GLM-5.3 matters
&lt;/h2&gt;

&lt;p&gt;GLM-5.3 is a post-training update on the GLM-5 base model. Zhipu reports that Terminal-Bench 3.0 increased from 4.6 to 28.3, while SWE-Marathon roughly doubled compared with GLM-5.2.&lt;/p&gt;

&lt;p&gt;Treat these as vendor-reported results until independently reproduced. For API evaluation, the practical takeaway is simple: test GLM-5.3 on your own coding, shell, agent, and debugging prompts.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1oufl857uxi4cj6mgkbj.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F1oufl857uxi4cj6mgkbj.png" alt="GLM-5.3 benchmark results" width="799" height="654"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The GLM-5 family uses a Mixture of Experts architecture with 744B total parameters, about 40B active parameters per forward pass, and a 200K-token context window, according to &lt;a href="https://docs.z.ai/guides/llm/glm-5" rel="noopener noreferrer"&gt;Z.ai’s docs&lt;/a&gt;. These are GLM-5 family specifications, not necessarily GLM-5.3-specific changes.&lt;/p&gt;

&lt;p&gt;Zhipu says it plans to release GLM-5.3 open weights around August 28, 2026, according to &lt;a href="https://pandaily.com/zhipu-glm-5-3-release-tang-jie-sooooooon-coding-security-aug2026" rel="noopener noreferrer"&gt;Pandaily’s launch coverage&lt;/a&gt;. If self-hosting is on your roadmap, use the hosted API now to create a prompt and response regression suite. See the &lt;a href="http://apidog.com/blog/self-host-glm-5-3-open-weights?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GLM-5.3 self-hosting prep guide&lt;/a&gt; for that workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Get an API key
&lt;/h2&gt;

&lt;p&gt;Zhipu operates separate international and mainland China platforms.&lt;/p&gt;

&lt;h3&gt;
  
  
  Z.ai: international platform
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Sign up at &lt;a href="https://z.ai/" rel="noopener noreferrer"&gt;z.ai&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Open the API console.&lt;/li&gt;
&lt;li&gt;Create an API key.&lt;/li&gt;
&lt;li&gt;Use the documentation at &lt;a href="https://docs.z.ai/" rel="noopener noreferrer"&gt;docs.z.ai&lt;/a&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This guide defaults to the international endpoint.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bigmodel.cn: mainland China platform
&lt;/h3&gt;

&lt;p&gt;For mainland China traffic, use &lt;a href="https://open.bigmodel.cn/" rel="noopener noreferrer"&gt;open.bigmodel.cn&lt;/a&gt;. The request format and authentication scheme are the same, but the hostname and billing are separate.&lt;/p&gt;

&lt;p&gt;Store the key in an environment variable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GLM_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-key-from-the-console"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not commit the key to source control. Add &lt;code&gt;.env&lt;/code&gt; files to &lt;code&gt;.gitignore&lt;/code&gt;, and use your deployment platform’s secret store in production.&lt;/p&gt;

&lt;h2&gt;
  
  
  Endpoint and authentication
&lt;/h2&gt;

&lt;p&gt;Use the chat completions endpoint.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;International:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST https://api.z.ai/api/paas/v4/chat/completions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Mainland China:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST https://open.bigmodel.cn/api/paas/v4/chat/completions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Send your key in the &lt;code&gt;Authorization&lt;/code&gt; header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer $GLM_API_KEY
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The request and response formats follow the OpenAI chat completions pattern:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Request fields include &lt;code&gt;model&lt;/code&gt; and &lt;code&gt;messages&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Responses include &lt;code&gt;choices&lt;/code&gt;, &lt;code&gt;message&lt;/code&gt;, &lt;code&gt;finish_reason&lt;/code&gt;, and &lt;code&gt;usage&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;OpenAI Python and Node.js SDKs work after changing the provider base URL.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you have an existing OpenAI-compatible integration, this is usually a configuration change rather than a rewrite. The same migration pattern applies to &lt;a href="http://apidog.com/blog/how-to-use-deepseek-v4-pro-0813-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek V4 Pro’s API&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Keep the model ID configurable
&lt;/h3&gt;

&lt;p&gt;At launch, Z.ai documentation listed &lt;code&gt;glm-5&lt;/code&gt; on the GLM-5 page. Since Zhipu lists &lt;code&gt;glm-5.2&lt;/code&gt; and &lt;code&gt;glm-5.1&lt;/code&gt; separately on its pricing page, &lt;code&gt;glm-5.3&lt;/code&gt; is the expected model ID.&lt;/p&gt;

&lt;p&gt;Use an environment variable rather than hardcoding it:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GLM_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"glm-5.3"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If &lt;code&gt;glm-5.3&lt;/code&gt; is unavailable in your region, confirm the current ID in the &lt;a href="https://docs.z.ai/guides/llm/glm-5" rel="noopener noreferrer"&gt;GLM-5 docs&lt;/a&gt; and use the documented model name.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your first request with cURL
&lt;/h2&gt;

&lt;p&gt;Create a minimal smoke test before integrating an SDK:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s2"&gt;"https://api.z.ai/api/paas/v4/chat/completions"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$GLM_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "glm-5.3",
    "messages": [
      {
        "role": "system",
        "content": "You are a code reviewer. Flag issues as blocking or non-blocking."
      },
      {
        "role": "user",
        "content": "Review this shell script for safety:\n\nrm -rf $BUILD_DIR/*\ncp dist/* $DEPLOY_TARGET"
      }
    ],
    "temperature": 0.3,
    "max_tokens": 1024
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read the generated content from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;choices[0].message.content
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Track token usage from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;usage.prompt_tokens
usage.completion_tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For multi-step coding or agent tasks, enable the documented reasoning mode:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"thinking"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"enabled"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use &lt;code&gt;thinking&lt;/code&gt; for tasks that benefit from planning or deeper analysis. Skip it for simple classification, extraction, or short transformations where extra reasoning is unnecessary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Python quickstart
&lt;/h2&gt;

&lt;p&gt;Install the OpenAI SDK:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;--upgrade&lt;/span&gt; openai
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then point the client at Z.ai:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GLM_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.z.ai/api/paas/v4&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GLM_MODEL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;glm-5.3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a code reviewer. Flag issues as blocking or non-blocking.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Review this Flask route for security issues:&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;@app.route(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;/user/&amp;lt;id&amp;gt;&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;def get_user(id):&lt;/span&gt;&lt;span class="se"&gt;\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;    return db.execute(f&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;SELECT * FROM users WHERE id = {id}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;)&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
            &lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;2048&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;input tokens:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;prompt_tokens&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output tokens:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completion_tokens&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Log &lt;code&gt;usage&lt;/code&gt; from the first day of testing. Until GLM-5.3-specific pricing is published, token counts are your best input for cost projections.&lt;/p&gt;

&lt;h2&gt;
  
  
  Node.js quickstart
&lt;/h2&gt;

&lt;p&gt;Install the SDK:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install &lt;/span&gt;openai
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Create a client with the Z.ai base URL:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;OpenAI&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;GLM_API_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;baseURL&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://api.z.ai/api/paas/v4&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;GLM_MODEL&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;glm-5.3&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;system&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;You are a terminal automation agent. Return each step as a shell command with a one-line rationale.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;role&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;user&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;A Node service on port 3000 stopped responding after a deploy. Give me a diagnosis sequence.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;temperature&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;max_tokens&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;2048&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;content&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;usage:&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If your application already uses OpenAI, create a second &lt;code&gt;OpenAI&lt;/code&gt; client with the Z.ai &lt;code&gt;baseURL&lt;/code&gt;. This makes model comparisons a routing decision instead of an integration rewrite.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stream responses
&lt;/h2&gt;

&lt;p&gt;Set &lt;code&gt;stream=True&lt;/code&gt; in Python:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GLM_MODEL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;glm-5.3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain the N+1 query problem with a concrete ORM example.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;delta&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;flush&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For raw HTTP requests, send:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"stream"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then parse server-sent event (&lt;code&gt;SSE&lt;/code&gt;) &lt;code&gt;data:&lt;/code&gt; lines using the OpenAI chunk format.&lt;/p&gt;

&lt;p&gt;Two implementation details matter:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Token usage is available at or after the final chunk, so finalize accounting only after the stream closes.&lt;/li&gt;
&lt;li&gt;With &lt;code&gt;thinking&lt;/code&gt; enabled, hard prompts may have a longer time-to-first-token because the model reasons before emitting visible output.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Parameters that matter
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Parameter&lt;/th&gt;
&lt;th&gt;Type&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;max_tokens&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;integer&lt;/td&gt;
&lt;td&gt;Caps output length. This is a primary cost control.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;temperature&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;number&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;0.2&lt;/code&gt;–&lt;code&gt;0.4&lt;/code&gt; for code and extraction; use &lt;code&gt;0.7+&lt;/code&gt; for more open-ended writing.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;thinking&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;object&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;{"type": "enabled"}&lt;/code&gt; enables reasoning mode for multi-step tasks.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;stream&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;boolean&lt;/td&gt;
&lt;td&gt;Returns server-sent events instead of a single response.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;messages&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;array&lt;/td&gt;
&lt;td&gt;Standard OpenAI-compatible &lt;code&gt;system&lt;/code&gt;, &lt;code&gt;user&lt;/code&gt;, and &lt;code&gt;assistant&lt;/code&gt; messages.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Zhipu had not published GLM-5.3-specific API pricing at launch. Check the &lt;a href="https://docs.z.ai/guides/overview/pricing" rel="noopener noreferrer"&gt;official pricing page&lt;/a&gt; instead of relying on reseller estimates.&lt;/p&gt;

&lt;p&gt;At the time of writing, the page listed:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input price per 1M tokens&lt;/th&gt;
&lt;th&gt;Output price per 1M tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5.2&lt;/td&gt;
&lt;td&gt;$1.40&lt;/td&gt;
&lt;td&gt;$4.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GLM-5&lt;/td&gt;
&lt;td&gt;$1.00&lt;/td&gt;
&lt;td&gt;$3.20&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For repeated instructions, keep stable system prompts consistent so eligible requests can benefit from cached-input discounts. The cost-control patterns in this &lt;a href="http://apidog.com/blog/deepseek-api-price-increase-cost-optimization?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek API price increase postmortem&lt;/a&gt; apply here as well.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test GLM-5.3 in Apidog before writing application code
&lt;/h2&gt;

&lt;p&gt;Use &lt;a href="http://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; to validate request shape, compare models, and save fixtures before putting prompts into your application.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Create the chat completions request
&lt;/h3&gt;

&lt;p&gt;Create:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST /chat/completions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use the OpenAI-compatible request body:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"{{GLM_MODEL}}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"messages"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"content"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Explain this error message."&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"temperature"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"max_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Create regional environments
&lt;/h3&gt;

&lt;p&gt;Create two environments:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;zai-international&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;bigmodel-mainland&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Set these variables:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Variable&lt;/th&gt;
&lt;th&gt;&lt;code&gt;zai-international&lt;/code&gt;&lt;/th&gt;
&lt;th&gt;&lt;code&gt;bigmodel-mainland&lt;/code&gt;&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;BASE_URL&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;https://api.z.ai/api/paas/v4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;https://open.bigmodel.cn/api/paas/v4&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GLM_API_KEY&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Your international key&lt;/td&gt;
&lt;td&gt;Your mainland China key&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;GLM_MODEL&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;glm-5.3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;glm-5.3&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Set the request URL to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{{BASE_URL}}/chat/completions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set the authorization header at the environment level:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer {{GLM_API_KEY}}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This prevents keys from being copied into saved requests and lets you switch regions without editing the endpoint.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Test reasoning mode side by side
&lt;/h3&gt;

&lt;p&gt;Duplicate a request.&lt;/p&gt;

&lt;p&gt;In one request, add:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="nl"&gt;"thinking"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"enabled"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Compare the same prompt across both requests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;time to first token&lt;/li&gt;
&lt;li&gt;final latency&lt;/li&gt;
&lt;li&gt;response quality&lt;/li&gt;
&lt;li&gt;token usage&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use the result to decide which workloads justify reasoning mode.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Save successful responses as examples
&lt;/h3&gt;

&lt;p&gt;Save good outputs as examples or fixtures. Reuse them for request-shape and schema validation without calling the live API repeatedly.&lt;/p&gt;

&lt;p&gt;Then build regression scenarios that assert:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;HTTP status&lt;/li&gt;
&lt;li&gt;&lt;code&gt;finish_reason&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;response schema&lt;/li&gt;
&lt;li&gt;expected response content patterns&lt;/li&gt;
&lt;li&gt;token usage thresholds&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a broader workflow, see the &lt;a href="http://apidog.com/blog/api-testing-tool-qa-engineers?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;API testing guide for QA engineers&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Error handling and rate limits
&lt;/h2&gt;

&lt;p&gt;Expect OpenAI-style errors with an &lt;code&gt;error&lt;/code&gt; object containing fields such as &lt;code&gt;message&lt;/code&gt;, &lt;code&gt;type&lt;/code&gt;, and &lt;code&gt;code&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Common status codes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Typical cause&lt;/th&gt;
&lt;th&gt;What to do&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;400&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Invalid request body or unknown model&lt;/td&gt;
&lt;td&gt;Validate JSON and confirm the model ID.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;401&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Missing, invalid, or revoked key&lt;/td&gt;
&lt;td&gt;Check &lt;code&gt;GLM_API_KEY&lt;/code&gt; and authorization headers.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;429&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Rate limit reached&lt;/td&gt;
&lt;td&gt;Retry with exponential backoff and jitter.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;5xx&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Transient provider failure&lt;/td&gt;
&lt;td&gt;Retry with bounded backoff and log request IDs if available.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Implement retries for &lt;code&gt;429&lt;/code&gt; and &lt;code&gt;5xx&lt;/code&gt; errors. For example, in Python:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;APIStatusError&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;create_completion_with_retry&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;**&lt;/span&gt;&lt;span class="n"&gt;kwargs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;APIStatusError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="n"&gt;retryable&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;429&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;=&lt;/span&gt; &lt;span class="mi"&gt;500&lt;/span&gt;

            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;retryable&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="k"&gt;raise&lt;/span&gt;

            &lt;span class="n"&gt;delay&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;min&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt; &lt;span class="o"&gt;**&lt;/span&gt; &lt;span class="n"&gt;attempt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;random&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;uniform&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;delay&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Avoid hardcoding rate-limit assumptions. Check &lt;a href="https://docs.z.ai/" rel="noopener noreferrer"&gt;Z.ai’s official docs&lt;/a&gt; for current concurrency and tier limits.&lt;/p&gt;

&lt;p&gt;Keep the model ID in configuration so you can roll back to &lt;code&gt;glm-5.2&lt;/code&gt; or another documented family model without redeploying code. The debugging workflow used for &lt;a href="http://apidog.com/blog/test-debug-grok-4-6-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Grok’s API&lt;/a&gt; transfers directly because both APIs use the OpenAI-style format.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  What is the GLM-5.3 API model ID?
&lt;/h3&gt;

&lt;p&gt;The expected model ID is &lt;code&gt;glm-5.3&lt;/code&gt;, following the &lt;code&gt;glm-5.2&lt;/code&gt; and &lt;code&gt;glm-5.1&lt;/code&gt; naming pattern. However, the launch-day &lt;a href="https://docs.z.ai/guides/llm/glm-5" rel="noopener noreferrer"&gt;GLM-5 documentation&lt;/a&gt; listed &lt;code&gt;glm-5&lt;/code&gt;, so confirm the currently supported ID before pinning it in production.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does GLM-5.3 work with the OpenAI SDK?
&lt;/h3&gt;

&lt;p&gt;Yes. Use the official &lt;code&gt;openai&lt;/code&gt; packages for Python or Node.js, set the base URL to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://api.z.ai/api/paas/v4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For mainland China, use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://open.bigmodel.cn/api/paas/v4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The request and response shapes follow the chat completions standard, including streaming.&lt;/p&gt;

&lt;h3&gt;
  
  
  How much does GLM-5.3 cost?
&lt;/h3&gt;

&lt;p&gt;Zhipu had not published GLM-5.3-specific pricing at launch. Use the &lt;a href="https://docs.z.ai/guides/overview/pricing" rel="noopener noreferrer"&gt;official pricing page&lt;/a&gt; for current numbers. GLM-5.2 was listed at $1.40 per 1M input tokens and $4.40 per 1M output tokens.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does GLM-5.3 compare with Claude and GPT?
&lt;/h3&gt;

&lt;p&gt;Zhipu’s evaluations describe GLM-5.3 coding and agent capability as “approaching Claude Fable 5.” It reported CyberGym at 84.5% and ExploitBench at 54.4%. These are vendor-reported results, so validate performance using your own workload and prompt suite. For a broader comparison, see &lt;a href="http://apidog.com/blog/grok-4-6-vs-gpt-5-6-vs-claude-fable-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Grok 4.6 vs GPT-5.6 vs Claude Fable 5&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I run GLM-5.3 locally?
&lt;/h3&gt;

&lt;p&gt;Not yet. Zhipu says open weights should arrive around August 28, 2026, through its &lt;a href="https://huggingface.co/zai-org" rel="noopener noreferrer"&gt;Hugging Face organization&lt;/a&gt;. The GLM-5 family’s 744B-parameter MoE architecture is server-class infrastructure, not a typical laptop deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where GLM-5.3 fits in your stack
&lt;/h2&gt;

&lt;p&gt;GLM-5.3 is worth evaluating for agent loops, code review, terminal automation, and multi-step coding tasks. The OpenAI-compatible API means the evaluation is low-effort:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create an API key.&lt;/li&gt;
&lt;li&gt;Run the cURL smoke test.&lt;/li&gt;
&lt;li&gt;Put the endpoint, API key, and model ID behind environment variables.&lt;/li&gt;
&lt;li&gt;Test &lt;code&gt;thinking&lt;/code&gt; enabled and disabled with your production-like prompts.&lt;/li&gt;
&lt;li&gt;Save outputs as regression fixtures.&lt;/li&gt;
&lt;li&gt;Port the approved request into Python or Node.js.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; to create separate regional environments, compare model configurations, inspect token usage, and turn working prompts into reusable API tests.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>What Is GLM-5.3? Zhipu's Open-Weight Coding Model Explained</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Sun, 16 Aug 2026 15:29:50 +0000</pubDate>
      <link>https://dev.to/hassann/what-is-glm-53-zhipus-open-weight-coding-model-explained-ji3</link>
      <guid>https://dev.to/hassann/what-is-glm-53-zhipus-open-weight-coding-model-explained-ji3</guid>
      <description>&lt;p&gt;GLM-5.3 is a large language model released on August 14, 2026 by Zhipu AI, the Chinese lab that operates internationally as &lt;a href="http://Z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt;. It is a post-training upgrade of the GLM-5 base model focused on coding and agentic workloads. Zhipu plans to publish open weights on Hugging Face about two weeks after launch and reports a 50% coding-capability improvement over GLM-5.2 in its internal evaluations.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;The release matters because Zhipu—a lab that &lt;a href="https://www.seekingalpha.com/news/4472588-chinese-openai-challenger-zhipu-is-said-to-unveil-new-open-source-model" rel="noopener noreferrer"&gt;Seeking Alpha describes as a “Chinese OpenAI challenger”&lt;/a&gt;—is shipping a coding-focused model it says is “approaching Claude Fable 5,” then making the weights available. It arrived one day after DeepSeek’s latest release, covered in our &lt;a href="http://apidog.com/blog/how-to-use-deepseek-v4-pro-0813-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek V4 Pro API guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;This guide covers the model’s reported benchmarks, the open-weights plan, and how to send a first API request. You can test the request in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; without writing application code.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;GLM-5.3 shipped on August 14, 2026 from Zhipu AI (&lt;a href="http://Z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt;). It uses the same GLM-5 base model; reported gains come from scaled post-training.&lt;/li&gt;
&lt;li&gt;Terminal-Bench 3.0 increased from 4.6 to 28.3: a 6.2x improvement and first place among open-source models, according to launch reporting.&lt;/li&gt;
&lt;li&gt;It also ranks first among open models on Agents’ Last Exam.&lt;/li&gt;
&lt;li&gt;CyberGym: 84.5%, slightly above Claude Mythos 5 and GPT-5.6 Sol in the reported results. ExploitBench: 54.4%, still behind frontier models.&lt;/li&gt;
&lt;li&gt;Open weights are expected on &lt;a href="https://huggingface.co/zai-org" rel="noopener noreferrer"&gt;Hugging Face&lt;/a&gt; around August 28, 2026.&lt;/li&gt;
&lt;li&gt;Official GLM-5 family specs: Mixture of Experts, 744B total parameters, roughly 40B active per forward pass, and a 200K-token context window.&lt;/li&gt;
&lt;li&gt;The API is OpenAI-compatible at &lt;code&gt;https://api.z.ai/api/paas/v4/chat/completions&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;No GLM-5.3-specific pricing was published at launch.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What GLM-5.3 is
&lt;/h2&gt;

&lt;p&gt;GLM-5.3 is the third point release in the GLM-5 family and its most coding-focused release. Zhipu did not retrain the foundation model. The &lt;a href="https://docs.z.ai/" rel="noopener noreferrer"&gt;official docs&lt;/a&gt; describe GLM-5 as a Mixture of Experts model with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;744B total parameters&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;~40B active parameters per forward pass&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;200K-token context window&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those base specifications carry over unchanged. Zhipu attributes the improvements to scaled post-training rather than changes to the foundation weights.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fswwpnagfu6ekluy8hsmm.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fswwpnagfu6ekluy8hsmm.png" alt="GLM-5.3 benchmark and release information" width="799" height="654"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;That distinction is important when evaluating the model. The reported gains target terminal-driven agent tasks, long-horizon software engineering, and security analysis—not general chat performance.&lt;/p&gt;

&lt;p&gt;For implementation teams, GLM-5.3 is worth evaluating when you need autonomous coding behavior, terminal-tool workflows, or a future self-hosting path once the open weights are available.&lt;/p&gt;

&lt;h2&gt;
  
  
  GLM-5.3 benchmarks: results and sources
&lt;/h2&gt;

&lt;p&gt;Launch coverage from &lt;a href="https://finance.biggo.com/news/0b571a42-9531-433c-b81b-c8468d173989" rel="noopener noreferrer"&gt;BigGo Finance&lt;/a&gt; and &lt;a href="https://pandaily.com/zhipu-glm-5-3-release-tang-jie-sooooooon-coding-security-aug2026" rel="noopener noreferrer"&gt;Pandaily&lt;/a&gt; reported the following results.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;GLM-5.3 result&lt;/th&gt;
&lt;th&gt;Context&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench 3.0&lt;/td&gt;
&lt;td&gt;28.3, up from 4.6&lt;/td&gt;
&lt;td&gt;6.2x increase; first among open-source models&lt;/td&gt;
&lt;td&gt;Launch report&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agents’ Last Exam&lt;/td&gt;
&lt;td&gt;First among open-source models&lt;/td&gt;
&lt;td&gt;Score not disclosed at launch&lt;/td&gt;
&lt;td&gt;Launch report&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CyberGym&lt;/td&gt;
&lt;td&gt;84.5%&lt;/td&gt;
&lt;td&gt;Slightly above Claude Mythos 5 and GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;Launch report&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ExploitBench&lt;/td&gt;
&lt;td&gt;54.4%&lt;/td&gt;
&lt;td&gt;Trails frontier models&lt;/td&gt;
&lt;td&gt;Launch report&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SWE-Marathon&lt;/td&gt;
&lt;td&gt;Roughly 2x GLM-5.2&lt;/td&gt;
&lt;td&gt;Long-horizon software engineering&lt;/td&gt;
&lt;td&gt;Zhipu internal&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding capability aggregate&lt;/td&gt;
&lt;td&gt;+50% vs. GLM-5.2&lt;/td&gt;
&lt;td&gt;Zhipu’s headline claim&lt;/td&gt;
&lt;td&gt;Zhipu internal&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Treat the source column as part of the result.&lt;/p&gt;

&lt;p&gt;A Terminal-Bench 3.0 increase from 4.6 to 28.3 is substantial. Terminal benchmarks test whether a model can chain shell commands, inspect command output, and recover from errors across multiple steps. If those reported results hold in independent testing, GLM-5.3 becomes a stronger candidate for coding-agent pipelines than GLM-5.2.&lt;/p&gt;

&lt;p&gt;However, the most frequently quoted claims—the 50% coding gain and roughly doubled SWE-Marathon score—come from Zhipu’s internal evaluations. They are useful evaluation signals, but they are not independently reproducible until the weights are released and external evaluators rerun the suites.&lt;/p&gt;

&lt;h2&gt;
  
  
  What “approaching Claude Fable 5” means in practice
&lt;/h2&gt;

&lt;p&gt;Zhipu describes GLM-5.3’s coding and agent capability as “approaching Claude Fable 5.” The wording matters.&lt;/p&gt;

&lt;p&gt;The reported agent and terminal results are close to frontier territory:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;First among open-source models on Terminal-Bench 3.0&lt;/li&gt;
&lt;li&gt;First among open-source models on Agents’ Last Exam&lt;/li&gt;
&lt;li&gt;84.5% on CyberGym, slightly above Claude Mythos 5 and GPT-5.6 Sol in the launch report&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;But the model does not match frontier systems across every benchmark. On ExploitBench, GLM-5.3 scored 54.4% and remained behind frontier models. Exploit development requires deep, multi-step reasoning under ambiguity, and that result suggests closed frontier models still lead on the hardest reasoning-heavy work.&lt;/p&gt;

&lt;p&gt;For your stack, use GLM-5.3 as an evaluation candidate for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Terminal agents&lt;/li&gt;
&lt;li&gt;CI automation&lt;/li&gt;
&lt;li&gt;Repository-scale coding tasks&lt;/li&gt;
&lt;li&gt;Tool-using development assistants&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For difficult reasoning workloads, keep frontier closed models in your comparison set. The &lt;a href="http://apidog.com/blog/grok-4-6-vs-gpt-5-6-vs-claude-fable-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Grok 4.6 vs GPT-5.6 vs Claude Fable 5 comparison&lt;/a&gt; covers similar benchmark categories with worked examples.&lt;/p&gt;

&lt;h2&gt;
  
  
  The open-weights plan: around August 28 on Hugging Face
&lt;/h2&gt;

&lt;p&gt;Zhipu committed to publishing GLM-5.3 open weights about two weeks after launch, placing the expected release around August 28, 2026. Watch the &lt;a href="https://huggingface.co/zai-org" rel="noopener noreferrer"&gt;zai-org Hugging Face organization&lt;/a&gt;, where the company hosts previous open releases.&lt;/p&gt;

&lt;p&gt;The gap between hosted API availability and weight release is intentional. Zhipu says it built its most extensive risk review system to date for this release. That matters because a model with meaningful offensive-security capability presents a different risk profile when released as downloadable weights than when served through a monitored API.&lt;/p&gt;

&lt;h3&gt;
  
  
  Plan for self-hosting
&lt;/h3&gt;

&lt;p&gt;A 744B-parameter MoE model is still server-class hardware territory, even when roughly 40B parameters are active per token. Most teams should expect to evaluate:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The hosted API first.&lt;/li&gt;
&lt;li&gt;Quantized community variants after the weights release.&lt;/li&gt;
&lt;li&gt;A local serving stack only after establishing an API baseline.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For sizing, serving options, and a hosted-vs-local test plan, see the &lt;a href="http://apidog.com/blog/self-host-glm-5-3-open-weights?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GLM-5.3 self-hosting guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  How GLM-5.3 fits the open-model landscape
&lt;/h2&gt;

&lt;p&gt;The most obvious comparison is DeepSeek. Both labs release open weights, price aggressively, and shipped major releases within 24 hours of each other. The difference is focus:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek V4 Pro:&lt;/strong&gt; generalist flagship positioning&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GLM-5.3:&lt;/strong&gt; specialist positioning for coding agents&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If most of your workload is agentic coding, GLM-5.3’s specialization is a reason to benchmark it first.&lt;/p&gt;

&lt;p&gt;Cost is another factor. DeepSeek’s recent price increase, analyzed in the &lt;a href="http://apidog.com/blog/deepseek-api-price-increase-cost-optimization?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek API cost optimization guide&lt;/a&gt;, shows how quickly open-model API pricing can change. Open weights provide a self-hosting hedge for teams that need more control over cost or deployment.&lt;/p&gt;

&lt;p&gt;Zhipu had not published GLM-5.3-specific API pricing at launch. The &lt;a href="https://docs.z.ai/guides/overview/pricing" rel="noopener noreferrer"&gt;official pricing page&lt;/a&gt; listed GLM-5.2 at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;$1.4 per 1M input tokens&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;$4.4 per 1M output tokens&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use those figures only as a reference point. Check the live pricing page before making budget commitments.&lt;/p&gt;

&lt;p&gt;Once quantized variants are available, GLM-5.3 joins the models you can run without API dependency. See the roundup of the &lt;a href="http://apidog.com/blog/best-local-llms-2026?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;best local LLMs in 2026&lt;/a&gt; for the current open-model landscape.&lt;/p&gt;

&lt;h2&gt;
  
  
  GLM Coding Plan quotas were reset
&lt;/h2&gt;

&lt;p&gt;Zhipu reset GLM Coding Plan quotas for all users on August 14. If you subscribe through &lt;a href="https://z.ai/" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt;, your allowance restarted on release day.&lt;/p&gt;

&lt;p&gt;The Coding Plan is Zhipu’s flat-rate offering for coding tools. It is separate from pay-per-token API billing, so the reset does not change API charges.&lt;/p&gt;

&lt;p&gt;Mainland China users access the same ecosystem through &lt;a href="https://open.bigmodel.cn/" rel="noopener noreferrer"&gt;open.bigmodel.cn&lt;/a&gt;, which has separate plans and billing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Try the API in five minutes
&lt;/h2&gt;

&lt;p&gt;Z.ai’s API is OpenAI-compatible. The international endpoint is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://api.z.ai/api/paas/v4/chat/completions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The mainland China endpoint is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://open.bigmodel.cn/api/paas/v4/chat/completions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Authentication uses a Bearer token.&lt;/p&gt;

&lt;p&gt;At launch, the official docs still listed &lt;code&gt;glm-5&lt;/code&gt; as the documented model ID. The example below uses &lt;code&gt;glm-5.3&lt;/code&gt; based on the GLM-5 family naming convention. Confirm the current model ID in the &lt;a href="https://docs.z.ai/guides/llm/glm-5" rel="noopener noreferrer"&gt;model documentation&lt;/a&gt; before deploying.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GLM_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-key-from-z.ai"&lt;/span&gt;

curl https://api.z.ai/api/paas/v4/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$GLM_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "glm-5.3",
    "messages": [
      {
        "role": "user",
        "content": "Write a bash script that finds the five largest files in a git repo, excluding the .git directory."
      }
    ],
    "temperature": 0.6,
    "max_tokens": 1024
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response follows the OpenAI-style schema:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;choices&lt;/code&gt; contains generated message content.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;usage&lt;/code&gt; contains token counts.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Make the request reusable in Apidog
&lt;/h3&gt;

&lt;p&gt;Instead of repeatedly editing cURL commands:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Import the request into &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Store &lt;code&gt;GLM_API_KEY&lt;/code&gt; as an environment variable.&lt;/li&gt;
&lt;li&gt;Create a &lt;code&gt;z-ai&lt;/code&gt; environment using &lt;code&gt;https://api.z.ai&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Create a &lt;code&gt;bigmodel-cn&lt;/code&gt; environment using &lt;code&gt;https://open.bigmodel.cn&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Store successful responses as regression baselines.&lt;/li&gt;
&lt;li&gt;Switch environments and model IDs from the UI instead of modifying request text.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This gives you a repeatable test setup for comparing the hosted API against a local deployment after the weights are released.&lt;/p&gt;

&lt;p&gt;For Python and Node.js clients, streaming, and error handling, see the full &lt;a href="http://apidog.com/blog/how-to-use-glm-5-3-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GLM-5.3 API quickstart&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is GLM-5.3 open source?
&lt;/h3&gt;

&lt;p&gt;Not yet. Zhipu committed to releasing open weights on Hugging Face around August 28, 2026, roughly two weeks after API launch. Until then, access is through the hosted API.&lt;/p&gt;

&lt;h3&gt;
  
  
  How is GLM-5.3 different from GLM-5.2?
&lt;/h3&gt;

&lt;p&gt;The base model is identical. Zhipu attributes the improvements to scaled post-training and reports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;50% coding capability improvement in internal evaluations&lt;/li&gt;
&lt;li&gt;Terminal-Bench 3.0 increase from 4.6 to 28.3&lt;/li&gt;
&lt;li&gt;Roughly doubled SWE-Marathon score&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The GLM-5 family architecture, context window, and parameter count remain unchanged.&lt;/p&gt;

&lt;h3&gt;
  
  
  What does GLM-5.3 cost?
&lt;/h3&gt;

&lt;p&gt;Zhipu had not published GLM-5.3-specific API pricing at launch. GLM-5.2 was listed at $1.4 per 1M input tokens and $4.4 per 1M output tokens on the official pricing page.&lt;/p&gt;

&lt;p&gt;Check current pricing before budgeting. For general per-token cost controls, see &lt;a href="http://apidog.com/blog/deepseek-api-price-increase-cost-optimization?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;API cost optimization after the DeepSeek price increase&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I run GLM-5.3 on my own hardware?
&lt;/h3&gt;

&lt;p&gt;After the weight release, yes—with caveats. The GLM-5 family has 744B total parameters and approximately 40B active parameters per forward pass. Full-precision serving requires multi-GPU server hardware.&lt;/p&gt;

&lt;p&gt;Quantized community builds should lower the hardware requirement. Review the &lt;a href="http://apidog.com/blog/self-host-glm-5-3-open-weights?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;self-hosting preparation guide&lt;/a&gt; before planning a deployment.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is GLM-5.3 better than Claude or GPT for coding?
&lt;/h3&gt;

&lt;p&gt;It depends on your workload.&lt;/p&gt;

&lt;p&gt;The reported results show strong performance on agent and terminal benchmarks, including first place among open models on Terminal-Bench 3.0. It also reportedly edged Claude Mythos 5 and GPT-5.6 Sol on CyberGym.&lt;/p&gt;

&lt;p&gt;However, GLM-5.3 trails frontier models on ExploitBench at 54.4%. Run your own workload-specific evaluation before moving production traffic, just as you would &lt;a href="http://apidog.com/blog/api-testing-tool-qa-engineers?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;test any API before adopting it&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where GLM-5.3 fits in your stack
&lt;/h2&gt;

&lt;p&gt;GLM-5.3 is a specialist release with a clear implementation path:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Test the OpenAI-compatible hosted API now.&lt;/li&gt;
&lt;li&gt;Save representative outputs from your real workloads.&lt;/li&gt;
&lt;li&gt;Build a small regression suite for coding and agent tasks.&lt;/li&gt;
&lt;li&gt;Re-run the suite when open weights and quantized variants become available.&lt;/li&gt;
&lt;li&gt;Compare hosted API results against your own deployment before migrating traffic.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The reported benchmark story is strong for terminal and agent tasks, acknowledges an ExploitBench gap, and still needs independent verification after the weights release. That is enough evidence to justify evaluation, but not enough to justify an automatic production migration.&lt;/p&gt;

&lt;p&gt;Use &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; to organize your evaluation suite, keep &lt;a href="http://Z.ai" rel="noopener noreferrer"&gt;Z.ai&lt;/a&gt; and &lt;a href="http://bigmodel.cn" rel="noopener noreferrer"&gt;bigmodel.cn&lt;/a&gt; as switchable environments, and turn exploratory calls into a regression baseline for comparing the hosted API with future self-hosted deployments.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Gemini 3.7 Flash Pricing Explained: Lock In Rates Before They Double</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 14 Aug 2026 07:54:21 +0000</pubDate>
      <link>https://dev.to/hassann/gemini-37-flash-pricing-explained-lock-in-rates-before-they-double-e4l</link>
      <guid>https://dev.to/hassann/gemini-37-flash-pricing-explained-lock-in-rates-before-they-double-e4l</guid>
      <description>&lt;p&gt;Google shipped Gemini 3.7 Flash on August 13, 2026, three weeks after 3.6 Flash, and describes it as &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/" rel="noopener noreferrer"&gt;“our most intelligent workhorse model”&lt;/a&gt;. The implementation detail that matters most for production teams is pricing: the introductory rate of $0.75 per million input tokens and $3.75 per million output tokens ends on December 31, 2026. On January 1, 2027, both rates double.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;That is a scheduled 2x increase for every Gemini 3.7 Flash workload. For example, a chatbot costing roughly $790/month at introductory pricing becomes about $1,575/month at standard pricing—even if your code, traffic, and prompts do not change.&lt;/p&gt;

&lt;p&gt;This guide shows how to calculate both pricing tiers, estimate three common workloads, and add token-cost checks to your API workflow. If you have not sent a request yet, start with the &lt;a href="http://apidog.com/blog/how-to-use-gemini-3-7-flash-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.7 Flash API quickstart&lt;/a&gt;. Once requests are running, &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; exposes &lt;code&gt;usageMetadata&lt;/code&gt; token counts in responses so you can validate estimates against actual traffic.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Introductory pricing through December 31, 2026:&lt;/strong&gt; $0.75 per 1M input tokens and $3.75 per 1M output tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Standard pricing from January 1, 2027:&lt;/strong&gt; $1.50 input and $7.50 output per 1M tokens.&lt;/li&gt;
&lt;li&gt;Gemini 3.7 Flash's introductory price is half of Gemini 3.6 Flash's launch price.&lt;/li&gt;
&lt;li&gt;The model supports text, image, video, audio, and PDF input, with a 1M-token context window and a 64k-token output limit.&lt;/li&gt;
&lt;li&gt;A chatbot handling 10,000 requests/day costs about &lt;strong&gt;$26.25/day&lt;/strong&gt; at introductory rates and &lt;strong&gt;$52.50/day&lt;/strong&gt; at standard rates.&lt;/li&gt;
&lt;li&gt;The biggest cost controls are output caps, context caching, batching, and routing simpler requests to a smaller model.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The two price tiers
&lt;/h2&gt;

&lt;p&gt;Gemini 3.7 Flash launched with a temporary discount. For Gemini API usage billed through an AI Studio key, budget for both tiers:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tier&lt;/th&gt;
&lt;th&gt;Window&lt;/th&gt;
&lt;th&gt;Input (per 1M tokens)&lt;/th&gt;
&lt;th&gt;Output (per 1M tokens)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Introductory&lt;/td&gt;
&lt;td&gt;August 13, 2026 to December 31, 2026&lt;/td&gt;
&lt;td&gt;$0.75&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;From January 1, 2027&lt;/td&gt;
&lt;td&gt;$1.50&lt;/td&gt;
&lt;td&gt;$7.50&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two practical implications follow:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The introductory rate is half of Gemini 3.6 Flash's launch price. If you are considering an upgrade, use the &lt;a href="http://apidog.com/blog/gemini-3-6-to-3-7-flash-migration-guide?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;3.6 to 3.7 Flash migration guide&lt;/a&gt; to plan regression tests and compare behavior.&lt;/li&gt;
&lt;li&gt;Output tokens cost &lt;strong&gt;5x more&lt;/strong&gt; than input tokens in both tiers. Reducing output length usually delivers a larger saving than trimming a static prompt.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The prices above are Gemini API rates through AI Studio. Vertex AI uses Google Cloud billing and separate SKUs. Context caching and batch processing can also have separate line items. Verify current pricing on the &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;official Gemini API pricing page&lt;/a&gt; before committing to a production budget.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calculate cost from token usage
&lt;/h2&gt;

&lt;p&gt;Use this formula for each endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cost =
  (input_tokens / 1_000_000 * input_rate) +
  (output_tokens / 1_000_000 * output_rate)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the introductory tier:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cost =
  (input_tokens / 1_000_000 * 0.75) +
  (output_tokens / 1_000_000 * 3.75)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For the standard tier:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cost =
  (input_tokens / 1_000_000 * 1.50) +
  (output_tokens / 1_000_000 * 7.50)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Always calculate both. The standard-tier value is exactly 2x the introductory-tier value.&lt;/p&gt;

&lt;h2&gt;
  
  
  What real workloads cost
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Workload 1: customer support chatbot
&lt;/h3&gt;

&lt;p&gt;Assume:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;10,000 requests/day&lt;/li&gt;
&lt;li&gt;2,000 input tokens/request&lt;/li&gt;
&lt;li&gt;300 output tokens/request&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Line item&lt;/th&gt;
&lt;th&gt;Daily tokens&lt;/th&gt;
&lt;th&gt;Intro cost/day&lt;/th&gt;
&lt;th&gt;Standard cost/day&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;20M&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;td&gt;$30.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;3M&lt;/td&gt;
&lt;td&gt;$11.25&lt;/td&gt;
&lt;td&gt;$22.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;23M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$26.25&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$52.50&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Over 30 days, that is roughly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;$788/month&lt;/strong&gt; at introductory rates&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$1,575/month&lt;/strong&gt; at standard rates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A practical implementation target is to cap support responses near the expected 300–500-token range rather than allowing unconstrained output.&lt;/p&gt;

&lt;h3&gt;
  
  
  Workload 2: PDF document pipeline
&lt;/h3&gt;

&lt;p&gt;Gemini 3.7 Flash accepts PDFs natively. Its GDP.pdf benchmark score increased from 22.0% to 34.0% over 3.6 Flash, making document extraction a relevant workload.&lt;/p&gt;

&lt;p&gt;Assume:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;500 documents/day&lt;/li&gt;
&lt;li&gt;40,000 input tokens/document&lt;/li&gt;
&lt;li&gt;1,000 output tokens/document for a structured summary&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Line item&lt;/th&gt;
&lt;th&gt;Daily tokens&lt;/th&gt;
&lt;th&gt;Intro cost/day&lt;/th&gt;
&lt;th&gt;Standard cost/day&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;20M&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;td&gt;$30.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;0.5M&lt;/td&gt;
&lt;td&gt;$1.88&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20.5M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$16.88&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$33.75&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is approximately:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;$506/month&lt;/strong&gt; at introductory rates&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$1,013/month&lt;/strong&gt; at standard rates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This workload is input-heavy. Focus on caching repeated instructions and using batch processing for non-interactive document jobs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Workload 3: agent loop
&lt;/h3&gt;

&lt;p&gt;Agents amplify token use because each step often includes growing conversation context and tool results.&lt;/p&gt;

&lt;p&gt;Assume:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;200 tasks/day&lt;/li&gt;
&lt;li&gt;12 model calls/task&lt;/li&gt;
&lt;li&gt;8,000 input tokens/call&lt;/li&gt;
&lt;li&gt;400 output tokens/call&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Line item&lt;/th&gt;
&lt;th&gt;Daily tokens&lt;/th&gt;
&lt;th&gt;Intro cost/day&lt;/th&gt;
&lt;th&gt;Standard cost/day&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;19.2M&lt;/td&gt;
&lt;td&gt;$14.40&lt;/td&gt;
&lt;td&gt;$28.80&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;0.96M&lt;/td&gt;
&lt;td&gt;$3.60&lt;/td&gt;
&lt;td&gt;$7.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;20.16M&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$18.00&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$36.00&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;$540/month&lt;/strong&gt; at introductory rates&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$1,080/month&lt;/strong&gt; at standard rates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For agent workflows, track both call count and input size. Increasing a task from 12 calls to 20 calls raises cost by roughly 67% before accounting for any additional context growth.&lt;/p&gt;

&lt;p&gt;The 64k output limit also bounds worst-case output cost per request at about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;$0.24&lt;/strong&gt; at introductory rates&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$0.48&lt;/strong&gt; at standard rates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That is still expensive when retries repeatedly hit the maximum output size, so enforce endpoint-level limits.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Gemini 3.7 Flash compares
&lt;/h2&gt;

&lt;p&gt;Provider pricing changes frequently, so treat comparisons as market positioning rather than a permanent price sheet.&lt;/p&gt;

&lt;p&gt;Gemini 3.7 Flash sits in the workhorse tier: cheaper than frontier offerings such as Gemini Pro, larger Claude models, or OpenAI's flagship tier, while posting benchmark results that overlap with models previously considered flagship-class. Google reports 65.3% on DeepSWE v1.1 and a 1588 WebDev Arena Elo.&lt;/p&gt;

&lt;p&gt;The introductory rate is an opportunity to test whether the model is sufficient for your production routes before standard pricing starts. Do not assume a model's price will remain fixed: &lt;a href="http://apidog.com/blog/deepseek-api-price-increase-cost-optimization?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DeepSeek's API price increase and cost optimization guide&lt;/a&gt; shows why token budgets should include room for pricing changes.&lt;/p&gt;

&lt;p&gt;For a small prototype, doubling may be irrelevant: $3 becomes $6. For a sustained $10,000/month pipeline, the same scheduled change becomes $20,000/month.&lt;/p&gt;

&lt;h2&gt;
  
  
  Five ways to cut token spend
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Cap output tokens per endpoint
&lt;/h3&gt;

&lt;p&gt;Set &lt;code&gt;maxOutputTokens&lt;/code&gt; in &lt;code&gt;generationConfig&lt;/code&gt; to the smallest value that satisfies the endpoint's contract.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"generationConfig"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"maxOutputTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Support replies: 300–500 tokens&lt;/li&gt;
&lt;li&gt;Classification: 20–100 tokens&lt;/li&gt;
&lt;li&gt;Structured extraction: size according to the schema&lt;/li&gt;
&lt;li&gt;Agent planning: a deliberately bounded limit per step&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Output is 5x the cost of input, so this is usually the highest-return optimization.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Cache static context
&lt;/h3&gt;

&lt;p&gt;Do not repeatedly pay full input pricing for unchanged system prompts, policies, reference documents, or large tool instructions.&lt;/p&gt;

&lt;p&gt;If every request includes a static 3,000-token policy document, context caching can reduce the cost of repeatedly supplying it. Check the &lt;a href="https://ai.google.dev/gemini-api/docs" rel="noopener noreferrer"&gt;Gemini API documentation&lt;/a&gt; for current caching setup and pricing.&lt;/p&gt;

&lt;p&gt;For the chatbot example, caching a 1,500-token static prefix can remove much of that repeated input cost.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Batch non-interactive work
&lt;/h3&gt;

&lt;p&gt;Document ingestion, summarization, extraction, and overnight report generation often do not require immediate responses.&lt;/p&gt;

&lt;p&gt;Use batch processing where latency is acceptable. This is especially useful for large PDF pipelines, where input volume dominates the bill.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Route requests to the smallest capable model
&lt;/h3&gt;

&lt;p&gt;Use Gemini 3.7 Flash for work that needs multi-step planning, debugging, or tool-call chains. Route simpler operations—such as classification, routing, and short extraction—to a Flash-Lite tier model when quality testing supports it.&lt;/p&gt;

&lt;p&gt;A simple routing strategy can look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight typescript"&gt;&lt;code&gt;&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;selectModel&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;classify&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;extract&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;agent&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;|&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;debug&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;task&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;classify&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;task&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;extract&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;flash-lite&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gemini-3.7-flash&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Validate quality per route before changing production traffic.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Prototype on the free tier
&lt;/h3&gt;

&lt;p&gt;Use AI Studio's free quota to finalize prompts, schemas, and output limits before sending billed production traffic. The &lt;a href="http://apidog.com/blog/get-free-unlimited-gemini-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;free Gemini API access guide&lt;/a&gt; explains the available free path and its limitations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Track spend per endpoint with Apidog
&lt;/h2&gt;

&lt;p&gt;Estimates are useful for planning; per-request measurements are what keep a service inside budget.&lt;/p&gt;

&lt;p&gt;Gemini responses include &lt;code&gt;usageMetadata&lt;/code&gt;, including prompt and output token counts. Capture these values in your API tests and treat token growth as a regression.&lt;/p&gt;

&lt;p&gt;A typical Gemini response exposes values similar to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"usageMetadata"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"promptTokenCount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2148&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"candidatesTokenCount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;312&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"totalTokenCount"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2460&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;, set up cost checks per endpoint:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create one request for each production route: chat turn, document summary, agent step, and so on.&lt;/li&gt;
&lt;li&gt;Store the API key in an environment variable such as &lt;code&gt;GEMINI_API_KEY&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Send realistic payloads, not placeholder prompts.&lt;/li&gt;
&lt;li&gt;Extract &lt;code&gt;usageMetadata.promptTokenCount&lt;/code&gt; and &lt;code&gt;usageMetadata.candidatesTokenCount&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Add token-budget assertions.&lt;/li&gt;
&lt;li&gt;Run the scenario whenever prompts, schemas, retrieval context, or tool outputs change.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For example, enforce an input-token budget for a chat endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;inputTokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usageMetadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;promptTokenCount&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;outputTokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usageMetadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;candidatesTokenCount&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Chat prompt stays within token budget&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;inputTokens&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;be&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;below&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2500&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Chat output stays within token budget&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outputTokens&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;be&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;below&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If a prompt edit pushes input from 2,000 to 2,500 tokens, your test can fail before deployment instead of silently increasing cost by 25%.&lt;/p&gt;

&lt;p&gt;To calculate the measured request cost:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;inputTokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usageMetadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;promptTokenCount&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;outputTokens&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usageMetadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;candidatesTokenCount&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;introCost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;inputTokens&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nx"&gt;_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;0.75&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outputTokens&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nx"&gt;_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;3.75&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;standardCost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;inputTokens&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nx"&gt;_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;1.5&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outputTokens&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nx"&gt;_000_000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mf"&gt;7.5&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;introCost&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;standardCost&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Token counts can regress just like latency or error rate. The difference is that token regressions show up later as an invoice. For a broader test-suite structure, see the &lt;a href="http://apidog.com/blog/api-testing-tool-qa-engineers?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;API testing guide for QA engineers&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  When does Gemini 3.7 Flash pricing double?
&lt;/h3&gt;

&lt;p&gt;On January 1, 2027. The introductory rate of $0.75 per 1M input tokens and $3.75 per 1M output tokens runs through December 31, 2026. It then changes to $1.50 input and $7.50 output per 1M tokens.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is Gemini 3.7 Flash cheaper than Gemini 3.6 Flash?
&lt;/h3&gt;

&lt;p&gt;At launch, yes. Gemini 3.7 Flash's introductory price is half of Gemini 3.6 Flash's launch price, while reported benchmarks improved, including DeepSWE v1.1 from 49.0% to 65.3%. See the &lt;a href="http://apidog.com/blog/gemini-3-7-flash-specs-pricing-reference?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.7 Flash quick reference&lt;/a&gt; for the full comparison.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does the introductory price apply on Vertex AI?
&lt;/h3&gt;

&lt;p&gt;This guide covers Gemini API pricing billed through an AI Studio key. Vertex AI uses Google Cloud billing, separate SKUs, and enterprise terms. Confirm current Vertex pricing in your GCP billing console and on the official pricing page.&lt;/p&gt;

&lt;h3&gt;
  
  
  What counts toward input tokens?
&lt;/h3&gt;

&lt;p&gt;Everything sent in the request counts: text, images, video, audio, and PDF content are converted to input tokens. For exact post-request counts, inspect the response's &lt;code&gt;usageMetadata&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I estimate tokens before sending a request?
&lt;/h3&gt;

&lt;p&gt;Use the API's &lt;code&gt;countTokens&lt;/code&gt; endpoint for a preflight estimate, or send representative requests and inspect &lt;code&gt;usageMetadata&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Measure real payloads, including system prompts, retrieved documents, tool definitions, and conversation history. Token behavior does not necessarily match tokenizers from other providers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Gemini 3.7 Flash fits in your stack
&lt;/h2&gt;

&lt;p&gt;The practical approach is straightforward:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Move appropriate workhorse traffic to Gemini 3.7 Flash during the introductory period.&lt;/li&gt;
&lt;li&gt;Measure actual token use per endpoint.&lt;/li&gt;
&lt;li&gt;Add token budgets to API tests.&lt;/li&gt;
&lt;li&gt;Model the January 2027 standard-rate bill now.&lt;/li&gt;
&lt;li&gt;Route lighter requests to cheaper models where quality allows.&lt;/li&gt;
&lt;li&gt;Use caching, batching, and output caps before traffic scales.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; to keep Gemini requests, environments, token assertions, and cost checks in one workspace—so January's invoice is a number you predicted rather than a surprise.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>What's New in Gemini 3.7 Flash? Features, Benchmarks, and API Access</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 14 Aug 2026 07:53:24 +0000</pubDate>
      <link>https://dev.to/hassann/whats-new-in-gemini-37-flash-features-benchmarks-and-api-access-3hj7</link>
      <guid>https://dev.to/hassann/whats-new-in-gemini-37-flash-features-benchmarks-and-api-access-3hj7</guid>
      <description>&lt;p&gt;Gemini 3.7 Flash is Google’s newest workhorse AI model, released on August 13, 2026, three weeks after Gemini 3.6 Flash. It keeps the 1M token context window and multimodal input of its predecessor while posting large gains on coding and agent benchmarks, and it launches at an introductory API price of $0.75 per 1M input tokens. Google calls it “our most intelligent workhorse model.”&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;That three-week gap between releases is the story. Google is shipping Flash updates on a sprint cadence while &lt;a href="https://www.axios.com/2026/08/13/google-gemini-37-flash" rel="noopener noreferrer"&gt;Gemini 3.5 Pro stays delayed&lt;/a&gt;, which means the mid-tier model is now where the interesting engineering lands first. The benchmark deltas back that up: DeepSWE jumps 16 points, AutomationBench nearly doubles, and the intro pricing undercuts the 3.6 Flash launch rate by half.&lt;/p&gt;

&lt;p&gt;This post covers what changed, the benchmark results, the pricing windows to calendar, where to access the model, and a cURL request you can run in five minutes. For a deeper endpoint walkthrough, see the &lt;a href="http://apidog.com/blog/how-to-use-gemini-3-7-flash-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.7 Flash API quickstart&lt;/a&gt;. You can also use &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; to test the new model against your current prompts before switching production traffic.&lt;/p&gt;

&lt;h2&gt;
  
  
  TL;DR
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Gemini 3.7 Flash launched August 13, 2026, three weeks after 3.6 Flash. Model ID: &lt;code&gt;gemini-3.7-flash&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Coding and agent benchmarks improved: DeepSWE v1.1 went from 49.0% to 65.3%, AutomationBench from 17.0% to 30.4%, and WebDev Arena from 1538 to 1588 Elo.&lt;/li&gt;
&lt;li&gt;Intro API pricing is $0.75 per 1M input tokens and $3.75 per 1M output tokens—half of 3.6 Flash’s launch price. Standard rates ($1.50 / $7.50) start January 1, 2027.&lt;/li&gt;
&lt;li&gt;Specs are unchanged: 1M-token input context, 64k output limit, multimodal input, function calling, search as a tool, and computer use.&lt;/li&gt;
&lt;li&gt;It is available in the Gemini API, AI Studio, Gemini app, Gemini Spark, Google Antigravity, Android Studio, and Gemini Enterprise across 160+ countries.&lt;/li&gt;
&lt;li&gt;Gemini 3.5 Pro is still delayed, while Google continues shipping Flash improvements.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Gemini 3.7 Flash is
&lt;/h2&gt;

&lt;p&gt;Flash is the middle tier of the Gemini line: cheaper and faster than Pro, smarter than Flash-Lite, and tuned for high-volume production workloads. Gemini 3.7 Flash is the third Flash release in the 3.x line.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/" rel="noopener noreferrer"&gt;official announcement&lt;/a&gt; positions it as a coding and agent model first, and a chat model second.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyz5bxakqovt9d06iter4.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fyz5bxakqovt9d06iter4.png" alt="Gemini 3.7 Flash overview" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;The spec sheet carries over from 3.6 Flash:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Context:&lt;/strong&gt; 1M-token input window and 64k-token output limit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Input modalities:&lt;/strong&gt; text, image, video, audio, and PDF. Output is text.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tooling:&lt;/strong&gt; function calling, search as a tool, and computer use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safety:&lt;/strong&gt; updated CBRN and cyber safeguards ship with the release.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;What changed is behavior rather than architecture. Google says 3.7 Flash debugs better, produces deployable code on the first try more often, adapts when it hits roadblocks, asks clarifying questions for ambiguous intent, and follows instructions more reliably.&lt;/p&gt;

&lt;p&gt;The benchmarks are the part you can validate with your own prompt suite.&lt;/p&gt;

&lt;h2&gt;
  
  
  Benchmark improvements: 3.6 vs. 3.7
&lt;/h2&gt;

&lt;p&gt;Google spread these results across a blog post and model card. Here are the key deltas:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Gemini 3.6 Flash&lt;/th&gt;
&lt;th&gt;Gemini 3.7 Flash&lt;/th&gt;
&lt;th&gt;Change&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE v1.1 (agentic coding)&lt;/td&gt;
&lt;td&gt;49.0%&lt;/td&gt;
&lt;td&gt;65.3%&lt;/td&gt;
&lt;td&gt;+16.3 pts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FrontierCode 1.1 Main&lt;/td&gt;
&lt;td&gt;34.4%&lt;/td&gt;
&lt;td&gt;43.6%&lt;/td&gt;
&lt;td&gt;+9.2 pts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;WebDev Arena (Elo)&lt;/td&gt;
&lt;td&gt;1538&lt;/td&gt;
&lt;td&gt;1588&lt;/td&gt;
&lt;td&gt;+50 Elo&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GDP.pdf (document reasoning)&lt;/td&gt;
&lt;td&gt;22.0%&lt;/td&gt;
&lt;td&gt;34.0%&lt;/td&gt;
&lt;td&gt;+12.0 pts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AutomationBench (agent tasks)&lt;/td&gt;
&lt;td&gt;17.0%&lt;/td&gt;
&lt;td&gt;30.4%&lt;/td&gt;
&lt;td&gt;+13.4 pts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two standalone scores add context:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;90.7%&lt;/strong&gt; on Harvey LAB-AA, a legal reasoning benchmark.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;97.0%&lt;/strong&gt; on 128k-needle long-context retrieval.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The largest gains are on multi-step agent tasks. AutomationBench moving from 17.0% to 30.4% may change which workflows are viable for a mid-tier model. &lt;a href="https://9to5google.com/2026/08/13/gemini-3-7-flash-launch/" rel="noopener noreferrer"&gt;Launch coverage&lt;/a&gt; highlighted the DeepSWE result for the same reason: a 16.3-point gain in agentic coding over three weeks is notable.&lt;/p&gt;

&lt;p&gt;For cross-model benchmark context, see &lt;a href="http://apidog.com/blog/grok-4-6-vs-gpt-5-6-vs-claude-fable-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Grok 4.6 vs GPT 5.6 vs Claude Fable 5&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Coding and agent improvements
&lt;/h2&gt;

&lt;p&gt;Use the benchmark changes to define a targeted evaluation instead of relying on generic chat prompts.&lt;/p&gt;

&lt;p&gt;Test these areas:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Debugging:&lt;/strong&gt; Give the model failing tests and multi-file repositories. DeepSWE is the closest benchmark signal because it measures multi-file bug fixes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;First-try deployable code:&lt;/strong&gt; Ask for implementation changes with explicit acceptance criteria, then run the generated code. FrontierCode improved by 9.2 points.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Roadblock handling:&lt;/strong&gt; Simulate failed tool calls, missing files, or permission errors. Check whether the model changes its plan instead of repeating the same failed step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Intent clarification:&lt;/strong&gt; Send ambiguous requirements. Measure whether it asks a useful question before making assumptions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Instruction fidelity:&lt;/strong&gt; Test output schemas, length limits, negative constraints, and formatting requirements across longer conversations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The tool stack is also unchanged from 3.6 Flash:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Function calling for structured integrations.&lt;/li&gt;
&lt;li&gt;Search as a tool for retrieval workflows.&lt;/li&gt;
&lt;li&gt;Computer use for browser-driven tasks when an API does not exist.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Computer use can be useful, but it is slower and less deterministic than a structured API integration. See &lt;a href="http://apidog.com/blog/computer-use-vs-structured-apis?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;computer use vs. structured APIs&lt;/a&gt; to decide which approach fits a workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing: half price until December 31, 2026
&lt;/h2&gt;

&lt;p&gt;Gemini 3.7 Flash has a two-phase Gemini API pricing schedule:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Period&lt;/th&gt;
&lt;th&gt;Input per 1M tokens&lt;/th&gt;
&lt;th&gt;Output per 1M tokens&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Through December 31, 2026&lt;/td&gt;
&lt;td&gt;$0.75&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;From January 1, 2027&lt;/td&gt;
&lt;td&gt;$1.50&lt;/td&gt;
&lt;td&gt;$7.50&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The introductory price is half of Gemini 3.6 Flash’s launch price. However, the cost doubles on January 1, 2027.&lt;/p&gt;

&lt;p&gt;Before deploying:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Verify current pricing on the &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;official Gemini API pricing page&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Forecast long-lived workloads using the standard 2027 rate.&lt;/li&gt;
&lt;li&gt;Record input and output tokens for each request from API usage metadata.&lt;/li&gt;
&lt;li&gt;Compare cost per successful task, not only cost per token.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For worked token-cost examples, see the &lt;a href="http://apidog.com/blog/gemini-3-7-flash-pricing-explained?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.7 Flash pricing breakdown&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where you can use it today
&lt;/h2&gt;

&lt;p&gt;Google shipped 3.7 Flash across its product surface area in 160+ countries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemini API:&lt;/strong&gt; Get a key from AI Studio and call &lt;code&gt;gemini-3.7-flash&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google AI Studio:&lt;/strong&gt; Browser playground for prompt testing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini app:&lt;/strong&gt; Consumer chat app access.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini Spark:&lt;/strong&gt; Available to AI Pro and Ultra subscribers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Google Antigravity:&lt;/strong&gt; Google’s agentic development environment.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Android Studio:&lt;/strong&gt; IDE code assistance.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini Enterprise:&lt;/strong&gt; Managed organizational offering.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For API access at scale, Vertex AI provides a path through &lt;code&gt;aiplatform.googleapis.com&lt;/code&gt; with OAuth, IAM, and audit logging. The model and request schema are the same, but authentication and hosting guarantees differ.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Google shipped Flash before Gemini 3.5 Pro
&lt;/h2&gt;

&lt;p&gt;The release order is unusual: 3.6 Flash arrived in late July, 3.7 Flash followed three weeks later, and Gemini 3.5 Pro still has no date.&lt;/p&gt;

&lt;p&gt;Axios reports that Google is shipping Flash updates while its next flagship remains delayed. For developers, the practical takeaway is straightforward: do not treat Flash as the trailing edge of the model lineup. Agent capabilities are landing there first.&lt;/p&gt;

&lt;p&gt;If you deferred a migration from 3.6 Flash while waiting for Pro, start with the &lt;a href="http://apidog.com/blog/gemini-3-6-to-3-7-flash-migration-guide?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;3.6 to 3.7 Flash migration guide&lt;/a&gt;. For most codebases, the migration is a model ID change followed by regression testing.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to try the API in five minutes
&lt;/h2&gt;

&lt;p&gt;Create an API key in &lt;a href="https://aistudio.google.com/apikey" rel="noopener noreferrer"&gt;AI Studio&lt;/a&gt;, export it, then send a request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GEMINI_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"AIza..."&lt;/span&gt;

curl &lt;span class="s2"&gt;"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.7-flash:generateContent"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-goog-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$GEMINI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "contents": [{
      "parts": [{
        "text": "Review this function for bugs: def dedupe(items): return list(set(items))"
      }]
    }],
    "generationConfig": {
      "temperature": 0.4,
      "maxOutputTokens": 1024
    }
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The response includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;code&gt;candidates&lt;/code&gt; array containing generated content in &lt;code&gt;content.parts&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;A &lt;code&gt;usageMetadata&lt;/code&gt; object with exact input and output token counts.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Capture &lt;code&gt;usageMetadata&lt;/code&gt; from day one. It is your cost meter.&lt;/p&gt;

&lt;p&gt;To stream output, replace &lt;code&gt;:generateContent&lt;/code&gt; with &lt;code&gt;:streamGenerateContent?alt=sse&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s2"&gt;"https://generativelanguage.googleapis.com/v1beta/models/gemini-3.7-flash:streamGenerateContent?alt=sse"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-goog-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$GEMINI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "contents": [{
      "parts": [{ "text": "Explain server-sent events in three bullet points." }]
    }]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Once the request works, move comparisons out of the terminal. In &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Import the Generative Language API specification.&lt;/li&gt;
&lt;li&gt;Store &lt;code&gt;GEMINI_API_KEY&lt;/code&gt; as an environment variable.&lt;/li&gt;
&lt;li&gt;Bind it to the &lt;code&gt;x-goog-api-key&lt;/code&gt; header.&lt;/li&gt;
&lt;li&gt;Save the model ID as a path variable.&lt;/li&gt;
&lt;li&gt;Run identical requests against &lt;code&gt;gemini-3.6-flash&lt;/code&gt; and &lt;code&gt;gemini-3.7-flash&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Save outputs as examples for regression testing.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This gives you a direct comparison of output quality, instruction following, streaming behavior, latency, and token usage using your own workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Gemini 3.7 Flash free to use?
&lt;/h3&gt;

&lt;p&gt;The Gemini API has a free tier through AI Studio with daily quota for prototyping. Paid usage starts at $0.75 per 1M input tokens through December 31, 2026.&lt;/p&gt;

&lt;p&gt;For guidance on working within quota limits, see &lt;a href="http://apidog.com/blog/get-free-unlimited-gemini-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;free Gemini API access&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  What’s the difference between Gemini 3.6 Flash and 3.7 Flash?
&lt;/h3&gt;

&lt;p&gt;The model specs are the same: context window, output limit, modalities, and tool support remain unchanged.&lt;/p&gt;

&lt;p&gt;The differences are benchmark and behavior improvements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;DeepSWE: +16.3 points&lt;/li&gt;
&lt;li&gt;AutomationBench: +13.4 points&lt;/li&gt;
&lt;li&gt;WebDev Arena: +50 Elo&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Google also claims improved debugging, instruction following, and multi-step planning.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Gemini 3.7 Flash replace Gemini 3.5 Pro?
&lt;/h3&gt;

&lt;p&gt;No. Pro remains the flagship tier for the hardest reasoning tasks, and Gemini 3.5 Pro is still in development.&lt;/p&gt;

&lt;p&gt;Gemini 3.7 Flash targets high-volume production workloads where cost and latency matter alongside quality.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens to pricing on January 1, 2027?
&lt;/h3&gt;

&lt;p&gt;The introductory price ends. Standard pricing becomes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input:&lt;/strong&gt; $1.50 per 1M tokens&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output:&lt;/strong&gt; $7.50 per 1M tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Budget at the standard rate for workloads that will continue beyond 2026.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can Gemini 3.7 Flash process images and PDFs?
&lt;/h3&gt;

&lt;p&gt;Yes. It accepts text, image, video, audio, and PDF inputs in the same &lt;code&gt;contents&lt;/code&gt; array. Output is text only.&lt;/p&gt;

&lt;p&gt;Google’s GDP.pdf result increased from 22.0% to 34.0%, which it presents as evidence of improved document reasoning.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where 3.7 Flash fits in your stack
&lt;/h2&gt;

&lt;p&gt;Gemini 3.7 Flash combines agent-focused benchmark improvements with unchanged integration specs and a temporary pricing discount.&lt;/p&gt;

&lt;p&gt;The implementation path is simple:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Run your existing prompt suite against &lt;code&gt;gemini-3.7-flash&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Compare outputs with your current model.&lt;/li&gt;
&lt;li&gt;Track token usage and latency.&lt;/li&gt;
&lt;li&gt;Test structured output, tool failures, long-context retrieval, and ambiguous instructions.&lt;/li&gt;
&lt;li&gt;Make the migration decision from measured results.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Use &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; to save 3.6 Flash responses as examples, replay the same requests against 3.7 Flash, and validate whether the benchmark improvements hold for your workload.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
