<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Mason Reed</title>
    <description>The latest articles on DEV Community by Mason Reed (@masonreed1).</description>
    <link>https://dev.to/masonreed1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4113593%2F5718bbeb-e76f-4666-b75c-87fb4df12464.png</url>
      <title>DEV Community: Mason Reed</title>
      <link>https://dev.to/masonreed1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/masonreed1"/>
    <language>en</language>
    <item>
      <title>Changing Gemini CLI Directories: Workspace, Project Settings, and Global Config</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Tue, 22 Sep 2026 02:14:24 +0000</pubDate>
      <link>https://dev.to/masonreed1/changing-gemini-cli-directories-workspace-project-settings-and-global-config-15fl</link>
      <guid>https://dev.to/masonreed1/changing-gemini-cli-directories-workspace-project-settings-and-global-config-15fl</guid>
      <description>&lt;p&gt;When I need to change a Gemini CLI directory, I first separate three tasks: starting the agent in another repository, adding repositories to its workspace, and relocating its persistent configuration. Each uses a different mechanism.&lt;/p&gt;

&lt;p&gt;Gemini CLI is Google’s open-source terminal agent. Alongside code inspection and assistance, it can execute shell commands with safeguards and integrate tools such as Google Search and Model Context Protocol (MCP) extensions. Those capabilities make directory scope matter: the files available to the agent and the location of its credentials are separate concerns.&lt;/p&gt;

&lt;p&gt;Here’s how I approach each case, including the version-dependent parts I would check before relying on them.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the directory you actually want to change
&lt;/h2&gt;

&lt;p&gt;The default user configuration directory is &lt;code&gt;.gemini&lt;/code&gt; under your home directory:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scope&lt;/th&gt;
&lt;th&gt;Typical location&lt;/th&gt;
&lt;th&gt;Purpose&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;User, Linux/macOS&lt;/td&gt;
&lt;td&gt;&lt;code&gt;~/.gemini/&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;User settings and persistent state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User, Windows&lt;/td&gt;
&lt;td&gt;&lt;code&gt;%USERPROFILE%\.gemini&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Windows user settings and persistent state&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Project&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;.gemini/settings.json&lt;/code&gt; in the project&lt;/td&gt;
&lt;td&gt;Project-specific settings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;System&lt;/td&gt;
&lt;td&gt;OS-specific locations, such as under &lt;code&gt;/etc/&lt;/code&gt; or &lt;code&gt;%PROGRAMDATA%&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;System configuration, when applicable&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Typical contents include &lt;code&gt;settings.json&lt;/code&gt;, &lt;code&gt;GEMINI.md&lt;/code&gt;, &lt;code&gt;commands/&lt;/code&gt;, cached credentials, telemetry identifiers, and other local state. Project settings override corresponding user settings when operating in that project.&lt;/p&gt;

&lt;p&gt;Historically, the user configuration path has been tied to the home directory and the &lt;code&gt;.gemini&lt;/code&gt; name. I would therefore check the installed release before assuming a configuration-directory environment variable works.&lt;/p&gt;

&lt;p&gt;For ordinary repository work, I start Gemini from the intended folder:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;cd&lt;/span&gt; /path/to/your-project
gemini
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That selects the starting working directory. It does not relocate the user configuration or credential cache.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add another repository without moving configuration
&lt;/h2&gt;

&lt;p&gt;If I only need the agent to inspect a second repository, workspace inclusion is the direct solution.&lt;/p&gt;

&lt;p&gt;At startup:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gemini &lt;span class="nt"&gt;--include-directories&lt;/span&gt; /path/to/repo
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or inside an interactive session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/directory add /path/to/another/project
/directory list
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These commands extend the workspace context. They do not move &lt;code&gt;~/.gemini&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;This is useful when a change spans an application and a shared library, or when another repository provides reference material. Moving configuration would not solve that access problem.&lt;/p&gt;

&lt;p&gt;If a directory still appears unavailable, I check whether the CLI process can read it. Network mounts and filesystem permissions can prevent access even when the directory has been added successfully.&lt;/p&gt;

&lt;h3&gt;
  
  
  Don’t rely on shell-mode &lt;code&gt;cd&lt;/code&gt; to switch the session
&lt;/h3&gt;

&lt;p&gt;Some platforms have had issues where &lt;code&gt;cd&lt;/code&gt; inside Gemini’s shell mode does not change the working directory as expected.&lt;/p&gt;

&lt;p&gt;My practical workaround is to change directories in the parent terminal before launching the CLI. For additional context within an existing session, I use &lt;code&gt;/directory add&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep repository settings with the repository
&lt;/h2&gt;

&lt;p&gt;For project-specific behavior, I use a project &lt;code&gt;.gemini&lt;/code&gt; directory instead of redirecting the global one:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;your-project/
├── .gemini/
│   ├── settings.json
│   └── GEMINI.md
└── src/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run &lt;code&gt;gemini&lt;/code&gt; from the project directory so it can discover the project configuration and context. Context discovery can search upward through the directory tree.&lt;/p&gt;

&lt;p&gt;This gives a repository its own settings while leaving user-wide state in the normal location. It also makes the intended scope easier to inspect: a repository override lives alongside the code it affects.&lt;/p&gt;

&lt;p&gt;The distinction matters when debugging authentication. Adding project settings does not, by itself, move the user credential cache.&lt;/p&gt;

&lt;h2&gt;
  
  
  Relocate global configuration only when necessary
&lt;/h2&gt;

&lt;p&gt;For a centralized configuration store, another drive, or a restricted home directory, there are two main options: a supported native override or a filesystem redirect.&lt;/p&gt;

&lt;h3&gt;
  
  
  Check environment-variable support for your release
&lt;/h3&gt;

&lt;p&gt;Several similarly named settings serve different purposes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;GEMINI_API_KEY&lt;/code&gt; supplies a key for Gemini API authentication.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;GEMINI_MODEL&lt;/code&gt; selects a model where supported.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;GEMINI_CLI_SYSTEM_SETTINGS_PATH&lt;/code&gt; overrides the system settings file path where supported.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;GEMINI_CONFIG_DIR&lt;/code&gt; has appeared as a code constant and in community proposals for a configurable directory.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A constant named &lt;code&gt;GEMINI_CONFIG_DIR&lt;/code&gt; does not establish that the CLI reads an environment variable with that name. Support and behavior have varied, with Windows issues reported in particular.&lt;/p&gt;

&lt;p&gt;If your installed version documents the directory override, the shell configuration looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GEMINI_CONFIG_DIR&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/custom_gemini_dir"&lt;/span&gt;
gemini
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;PowerShell:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;GEMINI_CONFIG_DIR&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;'C:\Users\you\CustomGemini'&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;gemini&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Treat those as conditional examples. Before building CI around them, check the release documentation and verify where that version actually writes its state.&lt;/p&gt;

&lt;p&gt;The system-settings override is narrower:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GEMINI_CLI_SYSTEM_SETTINGS_PATH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"/etc/my-gemini/system.settings.json"&lt;/span&gt;
gemini
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Even when supported, that selects a system settings file; it is not a general relocation mechanism for credentials, commands, and caches.&lt;/p&gt;

&lt;h3&gt;
  
  
  Use a symlink or junction when native relocation is unavailable
&lt;/h3&gt;

&lt;p&gt;A filesystem redirect lets the CLI continue addressing its expected path while the directory contents live elsewhere.&lt;/p&gt;

&lt;p&gt;On Linux or macOS, assuming the current configuration exists and the destination parent is available:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Preserve the current configuration.&lt;/span&gt;
&lt;span class="nb"&gt;mv&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.gemini"&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/gemini_backup"&lt;/span&gt;

&lt;span class="c"&gt;# Populate the destination before linking it.&lt;/span&gt;
&lt;span class="nb"&gt;mkdir&lt;/span&gt; &lt;span class="nt"&gt;-p&lt;/span&gt; /path/to/central/gemini-config
&lt;span class="nb"&gt;cp&lt;/span&gt; &lt;span class="nt"&gt;-a&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/gemini_backup/."&lt;/span&gt; /path/to/central/gemini-config/

&lt;span class="nb"&gt;ln&lt;/span&gt; &lt;span class="nt"&gt;-s&lt;/span&gt; /path/to/central/gemini-config &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$HOME&lt;/span&gt;&lt;span class="s2"&gt;/.gemini"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Windows, a directory junction provides a similar approach. From an appropriately privileged PowerShell session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;Move-Item&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Path&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;USERPROFILE&lt;/span&gt;&lt;span class="s2"&gt;\.gemini"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Destination&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;'C:\GeminiConfigBackup'&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="n"&gt;New-Item&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-ItemType&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;Directory&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Path&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;'C:\CentralGeminiConfig'&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Force&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="n"&gt;Copy-Item&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Path&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;'C:\GeminiConfigBackup\*'&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Destination&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;'C:\CentralGeminiConfig'&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Recurse&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-Force&lt;/span&gt;&lt;span class="w"&gt;

&lt;/span&gt;&lt;span class="n"&gt;New-Item&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nt"&gt;-ItemType&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;Junction&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;-Path&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$&lt;/span&gt;&lt;span class="nn"&gt;env&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="nv"&gt;USERPROFILE&lt;/span&gt;&lt;span class="s2"&gt;\.gemini"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="se"&gt;`
&lt;/span&gt;&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="nt"&gt;-Target&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;'C:\CentralGeminiConfig'&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I keep the backup until the CLI has loaded the expected settings and authentication state.&lt;/p&gt;

&lt;p&gt;Symlinks and junctions depend on filesystem privileges and can behave differently across Windows and container environments. Some CLI versions also restrict following certain symlinks for security, so I verify both &lt;code&gt;settings.json&lt;/code&gt; and context discovery after the move.&lt;/p&gt;

&lt;h3&gt;
  
  
  Treat home-directory changes as an isolated-runtime option
&lt;/h3&gt;

&lt;p&gt;In CI or containers, controlling the process’s effective home directory can influence where the default &lt;code&gt;.gemini&lt;/code&gt; path resolves. The source environment may use &lt;code&gt;HOME&lt;/code&gt; on Unix or profile-related values on Windows.&lt;/p&gt;

&lt;p&gt;That change has a broader effect than a configuration override: other tools and authentication flows, including Google OAuth caches, may also depend on the home directory. I would reserve this approach for an isolated runtime whose filesystem and authentication setup I control.&lt;/p&gt;

&lt;h2&gt;
  
  
  Diagnose failures by scope
&lt;/h2&gt;

&lt;p&gt;The error usually tells me which layer to inspect.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;What I check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Another repository is invisible&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;--include-directories&lt;/code&gt;, &lt;code&gt;/directory add&lt;/code&gt;, and read permissions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Project settings seem ignored&lt;/td&gt;
&lt;td&gt;Launch directory and project &lt;code&gt;.gemini/settings.json&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Relocated settings are missing&lt;/td&gt;
&lt;td&gt;Link target, copied contents, and version-specific symlink behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;GEMINI_CONFIG_DIR&lt;/code&gt; appears ignored&lt;/td&gt;
&lt;td&gt;Whether that release supports it as an environment variable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Windows reports &lt;code&gt;EPERM&lt;/code&gt; creating &lt;code&gt;.gemini&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Write permissions on &lt;code&gt;%USERPROFILE%&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For Windows &lt;code&gt;EPERM&lt;/code&gt;, adjusting folder permissions or using an appropriately elevated terminal can resolve the underlying access problem. A supported alternate location or junction may also help when the default profile directory is unsuitable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep API routing separate from directory configuration
&lt;/h2&gt;

&lt;p&gt;Sometimes the underlying goal is to use Gemini from a script or CI job. In that case, a direct API request may be sufficient, without configuring the terminal agent.&lt;/p&gt;

&lt;p&gt;A unified multi-model API such as CometAPI exposes an OpenAI-style chat-completions endpoint using bearer authentication:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;COMET_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"sk-xxxx"&lt;/span&gt;

curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://api.cometapi.com/v1/chat/completions"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$COMET_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "gemini-2.5-pro",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "Summarize the 3 key benefits of unit tests."}
    ],
    "max_tokens": 300
  }'&lt;/span&gt; | jq &lt;span class="nb"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This calls the model directly. It does not provide the CLI’s workspace, file inspection, or shell-tool behavior.&lt;/p&gt;

&lt;p&gt;Connecting the official CLI to a gateway requires checking both custom-base-URL support and API compatibility. Some releases or proposals have introduced overrides such as &lt;code&gt;GOOGLE_GEMINI_BASE_URL&lt;/code&gt;, but a base URL alone does not translate Gemini requests into OpenAI-style requests. A compatible endpoint or translating proxy is still necessary.&lt;/p&gt;

&lt;p&gt;For model-ID mismatches, inspect the gateway’s &lt;code&gt;/v1/models&lt;/code&gt; response and use the exact identifier. A variant such as &lt;code&gt;gemini-2.5-flash-preview-04-17&lt;/code&gt; should not be assumed interchangeable with a shorter family name.&lt;/p&gt;

&lt;p&gt;For my own directory setup, I keep the choice narrow: project settings for repository behavior, workspace commands for additional files, and a verified relocation mechanism only when persistent state needs another home.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-change-gemini-cli-directory/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-change-gemini-cli-directory"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Building an AI Companion: Persona, Memory, Tools, and the Tests That Matter</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Mon, 21 Sep 2026 07:58:58 +0000</pubDate>
      <link>https://dev.to/masonreed1/building-an-ai-companion-persona-memory-tools-and-the-tests-that-matter-4j6k</link>
      <guid>https://dev.to/masonreed1/building-an-ai-companion-persona-memory-tools-and-the-tests-that-matter-4j6k</guid>
      <description>&lt;p&gt;I’d start an AI companion with a narrow role, an explicit memory policy, and a set of difficult conversations to test. Voice and avatars come later. They make the experience more expressive, but they also make inconsistent behavior more noticeable.&lt;/p&gt;

&lt;p&gt;The engineering work sits across several layers: personality, conversation state, persistent memory, retrieval, tool execution, and presentation. Customization means deciding how those layers work together.&lt;/p&gt;

&lt;p&gt;Here’s how I’d approach a build in 2026, including where existing platforms make sense and where I’d want control over the backend.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose the relationship before the stack
&lt;/h2&gt;

&lt;p&gt;“AI companion” covers several products with different requirements. I’d pick one primary role before choosing a model or writing a system prompt.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Role&lt;/th&gt;
&lt;th&gt;Main behavior&lt;/th&gt;
&lt;th&gt;What the implementation needs&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Study coach or mentor&lt;/td&gt;
&lt;td&gt;Structured guidance and accountability&lt;/td&gt;
&lt;td&gt;Goals, task tracking, reminders&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Social companion&lt;/td&gt;
&lt;td&gt;Warmth, banter, conversational continuity&lt;/td&gt;
&lt;td&gt;Preference memory and consistent tone&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wellness companion&lt;/td&gt;
&lt;td&gt;Reflection, journaling, habit support&lt;/td&gt;
&lt;td&gt;Clear boundaries for sensitive conversations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Creative partner&lt;/td&gt;
&lt;td&gt;Brainstorming, stories, character development&lt;/td&gt;
&lt;td&gt;Project context and flexible generation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Roleplay character&lt;/td&gt;
&lt;td&gt;Consistent characterization and scene progression&lt;/td&gt;
&lt;td&gt;Backstory, scene state, narrative memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Productivity assistant&lt;/td&gt;
&lt;td&gt;Planning and task execution&lt;/td&gt;
&lt;td&gt;Calendar, document, and application integrations&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Trying to cover all these roles immediately makes prompt design and evaluation harder. A study coach that gives concise guidance has a different conversational rhythm from a roleplay character.&lt;/p&gt;

&lt;p&gt;This choice also determines what to remember. A mentor needs recurring goals. A creative partner needs project decisions. A social companion needs continuity without turning every casual statement into a permanent preference.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decide how much infrastructure you want to own
&lt;/h2&gt;

&lt;p&gt;Consumer platforms such as Replika, Character.AI, Kindroid, Nomi, and Kalon offer an accessible route to personality and visual customization. For work-oriented use cases, Zoom AI Companion, Microsoft Copilot, and custom GPTs fit more naturally into existing workflows.&lt;/p&gt;

&lt;p&gt;A custom implementation gives you control over memory, model selection, tool permissions, and the interface. It also makes you responsible for keeping those components consistent.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Customization&lt;/th&gt;
&lt;th&gt;Best fit&lt;/th&gt;
&lt;th&gt;Cost model&lt;/th&gt;
&lt;th&gt;Main advantage&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Consumer apps: Kalon, Kindroid, Nomi&lt;/td&gt;
&lt;td&gt;Personality, visuals, backstory, long-term memory&lt;/td&gt;
&lt;td&gt;Personal and social use&lt;/td&gt;
&lt;td&gt;Freemium or subscription&lt;/td&gt;
&lt;td&gt;Fast setup and immersion&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Zoom Custom AI Companion&lt;/td&gt;
&lt;td&gt;Agents, knowledge, avatars&lt;/td&gt;
&lt;td&gt;Enterprise workflows&lt;/td&gt;
&lt;td&gt;Add-on, approximately $12/user/month&lt;/td&gt;
&lt;td&gt;Workflow integration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom GPTs / Copilot&lt;/td&gt;
&lt;td&gt;Prompts and platform-dependent memory&lt;/td&gt;
&lt;td&gt;Productivity&lt;/td&gt;
&lt;td&gt;Subscription&lt;/td&gt;
&lt;td&gt;Existing ecosystem&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Custom API implementation&lt;/td&gt;
&lt;td&gt;Application-controlled persona, memory, and tools&lt;/td&gt;
&lt;td&gt;Custom products and scaling&lt;/td&gt;
&lt;td&gt;Pay per use&lt;/td&gt;
&lt;td&gt;Backend flexibility&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Open-source models such as Llama&lt;/td&gt;
&lt;td&gt;Fine-tuning and self-hosted behavior&lt;/td&gt;
&lt;td&gt;Advanced deployments and privacy requirements&lt;/td&gt;
&lt;td&gt;Hosting and operation costs&lt;/td&gt;
&lt;td&gt;Ownership and control&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For comparing models behind one interface, a unified gateway such as &lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt; offers an OpenAI-compatible API covering 500+ models, including GPT-5, Claude, Grok, and open-source options; its advertised 20–40% savings and 1M free testing tokens are terms I’d verify against the models and workload I actually intend to use.&lt;/p&gt;

&lt;p&gt;I’d choose the platform based on the controls the product requires. Visual customization alone is a different requirement from owning retrieval, deletion, and tool execution.&lt;/p&gt;

&lt;h2&gt;
  
  
  Write a persona you can evaluate
&lt;/h2&gt;

&lt;p&gt;“Be helpful and friendly” leaves too much behavior unspecified.&lt;/p&gt;

&lt;p&gt;My persona specification would include a name, role, speaking style, values, interests, emotional range, relationship dynamic, and prohibited behaviors. For a fictional character, I’d add age and backstory.&lt;/p&gt;

&lt;p&gt;The useful details are observable. “Short answers first” is testable. “A great personality” is not.&lt;/p&gt;

&lt;p&gt;A compact study-coach instruction might look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Be a calm, empathetic study coach who gives short answers first,
adds examples only when asked, avoids slang, and checks in with
the user after stressful topics.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a character-driven companion, the source’s Elara example provides a different starting point:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;You are Elara, a witty 28-year-old astrophysicist companion who
loves sci-fi and deep conversations. You respond warmly but
directly, using analogies from space exploration.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Neither prompt defines the whole product. Memory rules, sensitive-topic behavior, and tool access still need their own specifications.&lt;/p&gt;

&lt;p&gt;I’d compare prompt variants against the same conversations and then compare models. Claims about one model being better at empathy or creativity are useful hypotheses; the actual persona needs evaluation on its own workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make memory a product feature with user controls
&lt;/h2&gt;

&lt;p&gt;I’d separate memory into three layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Session context:&lt;/strong&gt; the current conversation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Persistent user memory:&lt;/strong&gt; stable preferences, recurring goals, and relevant history.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Retrieved context:&lt;/strong&gt; facts or documents fetched for the current response.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a custom build, persistent recall can use a memory layer such as mem0 or a custom implementation with Upstash Redis. RAG can bring in documents and user data when needed. A user profile can hold preferences such as communication style, favorite topics, and goals.&lt;/p&gt;

&lt;p&gt;The important design choice is what deserves persistence. Storing everything makes it harder to distinguish a lasting preference from a passing comment.&lt;/p&gt;

&lt;p&gt;ChatGPT illustrates the controls users increasingly expect: it can reference past chats, saved memories, and, where available, files and connected Gmail. Users can delete or clear memories, disable memory, or use Temporary Chat to prevent new memories from being created. Memory-source controls can also show what context contributed to personalization.&lt;/p&gt;

&lt;p&gt;For my own implementation, I’d make remembered information inspectable and correctable. The critical test is a preference change: when the user contradicts an older preference, does the companion follow the updated one?&lt;/p&gt;

&lt;h2&gt;
  
  
  Specify boundaries alongside the personality
&lt;/h2&gt;

&lt;p&gt;A warm persona still needs explicit limits. I’d define what it can discuss, when it should refuse, when it should redirect, and how it should handle sensitive emotional situations.&lt;/p&gt;

&lt;p&gt;That specification should cover unsafe advice, privacy, disallowed content, and behavior that encourages emotional dependency.&lt;/p&gt;

&lt;p&gt;Human-like presentation increases the chance that users attribute understanding or authority to the system. The companion should remain clear about its capabilities while responding with care.&lt;/p&gt;

&lt;p&gt;These rules belong in evaluation from the beginning. A character that follows its backstory perfectly but mishandles a distressed user is not behaving consistently with the product requirements.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ground answers before adding actions
&lt;/h2&gt;

&lt;p&gt;Knowledge retrieval and tool execution solve different problems, and I’d configure them separately.&lt;/p&gt;

&lt;p&gt;Knowledge sources can include internal documents, project notes, or web search. Custom dictionaries help with domain terminology; Zoom’s Custom AI Companion supports knowledge bases and custom dictionaries for this kind of adaptation.&lt;/p&gt;

&lt;p&gt;Tools let the companion act through calendars, email, or other APIs. For custom implementations, function calling or model tool use provides the integration mechanism.&lt;/p&gt;

&lt;p&gt;Model selection matters here as much as it does for conversation. A model that produces appealing dialogue still needs evaluation on tool selection and reasoning.&lt;/p&gt;

&lt;p&gt;I’d also keep behavioral tuning explicit:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Temperature and top-p affect sampling behavior.&lt;/li&gt;
&lt;li&gt;Response templates and custom dictionaries can improve consistency.&lt;/li&gt;
&lt;li&gt;User ratings provide feedback for evaluation and, where implemented, retraining or RLHF-like processes.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Routine conversation should not be described as “training” unless the application actually updates model parameters. Often the change is in stored context, prompts, or retrieval.&lt;/p&gt;

&lt;h2&gt;
  
  
  Add voice and visuals once the text behavior holds up
&lt;/h2&gt;

&lt;p&gt;Text remains the largest AI companion segment, while Grand View Research identifies multimodal companions as the fastest-growing segment.&lt;/p&gt;

&lt;p&gt;That direction makes sense for product design: voice changes delivery, visual identity shapes how users perceive the character, and reactions to photos or screenshots provide additional context.&lt;/p&gt;

&lt;p&gt;I’d still introduce these features incrementally:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Stabilize text behavior and memory.&lt;/li&gt;
&lt;li&gt;Add a visual identity or scene imagery.&lt;/li&gt;
&lt;li&gt;Add voice.&lt;/li&gt;
&lt;li&gt;Evaluate more dynamic avatar behavior.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Image generation and editing options named in the source include GPT-image-2, Flux, and Midjourney. Voice can use selected or cloned TTS voices with emotional inflection; ElevenLabs integrations are common. Real-time avatars responding through emotion detection are an emerging direction in products such as Genies.&lt;/p&gt;

&lt;p&gt;Each layer needs to match the established persona. An expressive voice cannot compensate for forgotten preferences or contradictory responses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Ship against a conversation test suite
&lt;/h2&gt;

&lt;p&gt;Before deployment, I’d run the same scenarios across prompt and model changes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;What I’d check&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;The user has had a bad day&lt;/td&gt;
&lt;td&gt;Tone adapts without abandoning boundaries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Playful banter&lt;/td&gt;
&lt;td&gt;Humor stays consistent with the persona&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A sensitive emotional topic&lt;/td&gt;
&lt;td&gt;Support and redirection follow the specification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A conversation depends on older context&lt;/td&gt;
&lt;td&gt;Relevant memory is retrieved accurately&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;The user changes a preference&lt;/td&gt;
&lt;td&gt;New information takes precedence over stale context&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Track coherence, user satisfaction, and latency alongside those cases. Test model changes against the whole interaction flow, including retrieval and tools.&lt;/p&gt;

&lt;p&gt;Deployment can be a web interface, a mobile app, or an integration into an existing product. I’d keep the first release focused enough that a failure can be traced to a specific layer: persona, memory, retrieval, tools, or presentation. That gives each subsequent iteration a concrete problem to solve.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-customize-an-ai-companion/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-customize-an-ai-companion"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Top 6 OpenClaw Skills you can't afford to miss in 2026</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Mon, 21 Sep 2026 06:58:43 +0000</pubDate>
      <link>https://dev.to/masonreed1/top-6-openclaw-skills-you-cant-afford-to-miss-in-2026-29ng</link>
      <guid>https://dev.to/masonreed1/top-6-openclaw-skills-you-cant-afford-to-miss-in-2026-29ng</guid>
      <description>&lt;p&gt;OpenClaw has emerged as one of the most transformative open-source projects of 2026, powering autonomous AI agents that don't just chat—they act. Running locally on your machine or VPS, OpenClaw connects large language models (like Claude, GPT, or local alternatives) with your files, apps, browser, terminal, and messaging platforms (WhatsApp, Telegram, Discord, etc.). It handles real tasks: clearing inboxes, managing calendars, executing workflows, and running 24/7 via heartbeat schedulers.&lt;/p&gt;

&lt;p&gt;At the heart of OpenClaw's power are &lt;strong&gt;Skills&lt;/strong&gt;—modular Markdown files (typically &lt;code&gt;SKILL.md&lt;/code&gt;) that package instructions, prompts, tool calls, and workflows. These reusable components turn a generic agent into a specialized digital coworker. With thousands available on ClawHub and community repos, selecting the right ones is critical.&lt;/p&gt;

&lt;h3&gt;
  
  
  What Is OpenClaw and Why Skills Matter in 2026
&lt;/h3&gt;

&lt;p&gt;OpenClaw (formerly Clawdbot/Moltbot) is a self-hosted agent runtime. It runs on Mac, Windows, or Linux, connects to any LLM (OpenAI, Anthropic, local models via Ollama, etc.), and uses messaging apps as the primary interface. It features persistent local memory (Markdown files), browser automation, shell execution, and proactive scheduling.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Skills&lt;/strong&gt; are the extensibility layer. Defined primarily via &lt;code&gt;SKILL.md&lt;/code&gt; (natural language instructions + tool calls), they allow the LLM to interpret and execute complex, multi-step tasks reliably. Community contributions exploded in 2026, with high-quality skills vetted on ClawHub.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Benefits&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Modularity&lt;/strong&gt;: Install only what you need; chain them for complex workflows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extensibility&lt;/strong&gt;: Community and self-created skills allow custom behaviors.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Persistence&lt;/strong&gt;: Combined with memory systems (e.g., MEMORY.md, SOUL.md), skills enable long-term learning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safety &amp;amp; Control&lt;/strong&gt;: Local execution keeps data private; vet skills carefully.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Data Support&lt;/strong&gt;: Analyses of ClawHub show native/bundled tools cover ~70% of calls, but top community skills handle high-value tasks like email, browsing, and project management. Users report 90-day reliability improvements and significant time savings.&lt;/p&gt;

&lt;h3&gt;
  
  
  Installation Basics (General for Most Skills):
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Ensure OpenClaw is installed and running (Docker, direct install, or VPS recommended).&lt;/li&gt;
&lt;li&gt;Use the ClawHub CLI or manual placement in the skills directory.&lt;/li&gt;
&lt;li&gt;Restart/reload the agent and test via your preferred chat app.&lt;/li&gt;
&lt;li&gt;Configure API keys (e.g., for external services) in environment variables or config files.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Pro Recommendation&lt;/strong&gt;: Power OpenClaw with &lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;&lt;strong&gt;CometAPI&lt;/strong&gt;&lt;/a&gt; . This single OpenAI-compatible endpoint provides access to 500+ models (GPT-5 series, Claude Opus/Sonnet variants, Grok, DeepSeek, Llama, multimodal, etc.) at 20–40% lower costs, with free starter tokens. It eliminates multiple API keys, offers enterprise analytics/privacy controls, and ensures high uptime—perfect for always-on OpenClaw agents. Integrate once and route models dynamically for optimal cost/performance (e.g., cheaper models for routine tasks, frontier for complex reasoning).&lt;/p&gt;

&lt;h2&gt;
  
  
  1. GOG (Google Workspace Integration) — The Productivity Powerhouse
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What it is&lt;/strong&gt;: GOG (often steipete/gog or similar wrappers) provides unified access to Gmail, Calendar, Drive, Docs, Sheets, and Contacts via Google’s APIs/CLI.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Importance&lt;/strong&gt;: Email and calendar management consume ~28% of knowledge workers’ time. GOG automates triage, scheduling, and data synthesis. It ranks among the most-installed skills (tens of thousands of downloads) and powers “AI employee” workflows.&lt;/p&gt;

&lt;h3&gt;
  
  
  How to install:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;clawhub install gog (or official variants).&lt;/li&gt;
&lt;li&gt;Authenticate via OAuth (use dedicated/scoped accounts for safety).&lt;/li&gt;
&lt;li&gt;Add to workspace and test with “Summarize my unread emails.”&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Key Functions:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Intelligent inbox triage, auto-archive, replies/drafts.&lt;/li&gt;
&lt;li&gt;Calendar conflict detection, meeting scheduling, reminders.&lt;/li&gt;
&lt;li&gt;Drive/Docs/Sheets: Search, summarize, update data, generate reports.&lt;/li&gt;
&lt;li&gt;Proactive briefings (e.g., morning digest combining email + calendar + Drive files).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Use Cases &amp;amp; Data&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Founders: Auto-coordinate meetings and update Notion/Sheets CRMs.&lt;/li&gt;
&lt;li&gt;Teams: Weekly status reports pulled from emails/Docs.&lt;/li&gt;
&lt;li&gt;Personal: Flight check-ins or expense tracking from receipts in Drive. Real-world impact: Users achieve inbox zero and reclaim hours; integration with CometAPI allows cheaper models for high-volume email processing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;CometAPI Tip&lt;/strong&gt;: Route routine summarization to cost-effective models while using premium ones for sensitive drafting.&lt;/p&gt;

&lt;h2&gt;
  
  
  2. Agent Browser / Web Automation Skill — Autonomous Internet Agent
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What it is&lt;/strong&gt;: Tools like Agent Browser or Playwright-based skills enable headless browsing, form filling, scraping, screenshots, and interaction with JS-heavy sites.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Importance&lt;/strong&gt;: Web tasks (research, monitoring, transactions) are fragmented. This skill turns OpenClaw into a true agent, with high adoption for research and ops automation.&lt;/p&gt;

&lt;h3&gt;
  
  
  How to install:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;clawhub install agent-browser (or top-rated equivalents).&lt;/li&gt;
&lt;li&gt;Configure in sandbox (Docker recommended due to power).&lt;/li&gt;
&lt;li&gt;Test: “Check flight status and summarize prices.”&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Key Functions:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Navigate sites, handle logins (with care), extract structured data.&lt;/li&gt;
&lt;li&gt;Automated check-ins, lead gen, price monitoring.&lt;/li&gt;
&lt;li&gt;Screenshots + OCR for visual confirmation.&lt;/li&gt;
&lt;li&gt;Multi-step workflows (e.g., research → fill form → confirm).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Use Cases&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Competitive intelligence: Daily SERP/competitor monitoring.&lt;/li&gt;
&lt;li&gt;E-commerce: Price alerts, order tracking.&lt;/li&gt;
&lt;li&gt;Research: Compile reports from multiple sources. Data shows web skills among top installs; combined with CometAPI’s fast models, it enables real-time loops without rate limits.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Security&lt;/strong&gt;: Sandbox heavily; use approval for actions involving logins.&lt;/p&gt;

&lt;h2&gt;
  
  
  3. Self-Improving Agent / Capability Evolver — The Meta-Skill
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What it is&lt;/strong&gt;: Skills like Self-Improving Agent or Capability Evolver log interactions, errors, and preferences to refine behavior autonomously.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Importance&lt;/strong&gt;: Static agents plateau; these create compounding intelligence. Highest-rated on ClawHub with strong community backing.&lt;/p&gt;

&lt;h3&gt;
  
  
  How to install:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;clawhub install self-improving-agent or capability-evolver.&lt;/li&gt;
&lt;li&gt;Point to memory folders; enable in SOUL.md.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Key Functions:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Persistent learning: Update preferences, avoid repeated mistakes.&lt;/li&gt;
&lt;li&gt;Auto-generate or refine other skills.&lt;/li&gt;
&lt;li&gt;Memory ontology for long-term context.&lt;/li&gt;
&lt;li&gt;Error logging and self-correction loops.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Use Cases&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Personalization: Learns your style for emails/content.&lt;/li&gt;
&lt;li&gt;Workflow evolution: Turns ad-hoc tasks into reusable automations.&lt;/li&gt;
&lt;li&gt;Long-running agents: Improves over weeks/months. Users report significant gains in reliability; pair with CometAPI for diverse model routing to accelerate learning.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. GitHub Integration — Developer and Team Workflow Accelerator
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What it is&lt;/strong&gt;: Official/community GitHub skills for repo management, PRs, issues, and commits.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Importance&lt;/strong&gt;: Dev teams spend heavily on context-switching. This skill automates reviews, notifications, and maintenance—critical as AI coding scales in 2026.&lt;/p&gt;

&lt;h3&gt;
  
  
  How to install:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;clawhub install github.&lt;/li&gt;
&lt;li&gt;OAuth setup with scoped tokens.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Key Functions:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Monitor PRs/issues, auto-summarize, suggest reviewers.&lt;/li&gt;
&lt;li&gt;Create branches, draft PRs, run basic CI checks.&lt;/li&gt;
&lt;li&gt;Daily digests and triage from chat.&lt;/li&gt;
&lt;li&gt;Code review assistance.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Use Cases&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Solo devs: “Fix failing tests” → autonomous loops.&lt;/li&gt;
&lt;li&gt;Teams: Auto-close stale issues, generate release notes.&lt;/li&gt;
&lt;li&gt;Integration with browser skill for external research. High download counts; CometAPI supports strong coding models (e.g., specialized coders) at lower cost.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  5. Summarize Skill — Knowledge Distiller
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What it is&lt;/strong&gt;: Universal summarization across URLs, YouTube, podcasts, docs, and files.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Importance&lt;/strong&gt;: Information overload is constant. This skill (10k+ downloads) delivers concise insights fast.&lt;/p&gt;

&lt;h3&gt;
  
  
  How to install:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;clawhub install summarize.&lt;/li&gt;
&lt;li&gt;Simple setup; works with local files too.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Key Functions&lt;/strong&gt;:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Multi-format input → structured output (key points, action items).&lt;/li&gt;
&lt;li&gt;Custom rubrics (e.g., “business implications”).&lt;/li&gt;
&lt;li&gt;Batch processing for newsletters/research.&lt;/li&gt;
&lt;li&gt;Integration with other skills (e.g., summarize then act).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Use Cases&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Daily news/podcast digests via heartbeats.&lt;/li&gt;
&lt;li&gt;Meeting prep: Summarize related docs.&lt;/li&gt;
&lt;li&gt;Research pipelines. Essential baseline skill; efficient with CometAPI’s balanced models.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  6. Project Management Integrations (e.g., Linear, Notion) — Ops Orchestrator
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What it is&lt;/strong&gt;: Skills for Linear, Notion, Asana, etc., syncing tasks across tools.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Importance&lt;/strong&gt;: Fragmented tools kill productivity. These unify execution.&lt;/p&gt;

&lt;h3&gt;
  
  
  How to install:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;e.g., clawhub install linear or Notion equivalents.&lt;/li&gt;
&lt;li&gt;API key/OAuth.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Key Functions:
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Create/update tickets from chat/emails.&lt;/li&gt;
&lt;li&gt;Status sync and cross-tool reports.&lt;/li&gt;
&lt;li&gt;Auto-triage bugs from logs/emails.&lt;/li&gt;
&lt;li&gt;Weekly digests and reminders.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Use Cases&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Founders: Link emails → tasks → Notion.&lt;/li&gt;
&lt;li&gt;Teams: Standup automation.&lt;/li&gt;
&lt;li&gt;Personal: Life admin tracking. Combines powerfully with GOG and self-improving skills.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  How to choose the right OpenClaw skill
&lt;/h2&gt;

&lt;p&gt;Choose skills based on repeated pain, not novelty. If a task happens every day, starts in chat, and ends with a tool action, it is a skill candidate. If it needs memory, timing, or strict guardrails, it is an even better candidate. OpenClaw’s own docs emphasize that skills teach the agent how and when to use tools, while plugins and tools provide the raw capability.&lt;/p&gt;

&lt;p&gt;A good rule for 2026 is to start with the six skills above and then add custom workspace skills only after you have measured the pain point. OpenClaw supports local overrides, workspace skills, and precedence rules, so you do not need to keep editing the same repo copy to customize behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  Comparison Table: Top 6 OpenClaw Skills
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Skill&lt;/th&gt;
&lt;th&gt;Installs/Popularity&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;th&gt;Complexity&lt;/th&gt;
&lt;th&gt;Risk Level&lt;/th&gt;
&lt;th&gt;CometAPI Synergy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GOG (Google)&lt;/td&gt;
&lt;td&gt;Very High (top-ranked)&lt;/td&gt;
&lt;td&gt;Productivity, Email/Calendar&lt;/td&gt;
&lt;td&gt;Low-Medium&lt;/td&gt;
&lt;td&gt;Medium (OAuth)&lt;/td&gt;
&lt;td&gt;High (volume tasks)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent Browser&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Research, Automation&lt;/td&gt;
&lt;td&gt;Medium-High&lt;/td&gt;
&lt;td&gt;High (sandbox)&lt;/td&gt;
&lt;td&gt;High (real-time)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-Improving&lt;/td&gt;
&lt;td&gt;High (top-rated)&lt;/td&gt;
&lt;td&gt;Long-term Autonomy&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Medium (learning loops)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GitHub&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Dev Workflows&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;High (coding models)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Summarize&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Knowledge Mgmt&lt;/td&gt;
&lt;td&gt;Very Low&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;High (efficiency)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Project Mgmt (Linear/Notion)&lt;/td&gt;
&lt;td&gt;Medium-High&lt;/td&gt;
&lt;td&gt;Ops/Teams&lt;/td&gt;
&lt;td&gt;Low-Medium&lt;/td&gt;
&lt;td&gt;Low-Medium&lt;/td&gt;
&lt;td&gt;High (orchestration)&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Skill/Category&lt;/th&gt;
&lt;th&gt;Use Case&lt;/th&gt;
&lt;th&gt;Install Difficulty&lt;/th&gt;
&lt;th&gt;Popularity (Est.)&lt;/th&gt;
&lt;th&gt;Key Benefit&lt;/th&gt;
&lt;th&gt;CometAPI Synergy&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GitHub&lt;/td&gt;
&lt;td&gt;Repo management, PRs&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Very High&lt;/td&gt;
&lt;td&gt;Autonomous dev workflows&lt;/td&gt;
&lt;td&gt;Reliable coding models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent Browser&lt;/td&gt;
&lt;td&gt;Web automation&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Browser actions without manual&lt;/td&gt;
&lt;td&gt;Vision/ multimodal models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Web Search&lt;/td&gt;
&lt;td&gt;Real-time research&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Fresh data synthesis&lt;/td&gt;
&lt;td&gt;Fast, cheap inference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Summarize/Notion&lt;/td&gt;
&lt;td&gt;Content &amp;amp; knowledge mgmt&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Structured output&lt;/td&gt;
&lt;td&gt;Long-context models (GPT-5.4)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Self-Improving&lt;/td&gt;
&lt;td&gt;Agent evolution&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Growing&lt;/td&gt;
&lt;td&gt;Reduced errors over time&lt;/td&gt;
&lt;td&gt;Consistent model perf via CometAPI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Calendar/Email&lt;/td&gt;
&lt;td&gt;Daily productivity&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;Very High&lt;/td&gt;
&lt;td&gt;Proactive scheduling&lt;/td&gt;
&lt;td&gt;Low-latency for frequent calls&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Advanced Tips, Security, and Scaling in 2026
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Memory &amp;amp; Heartbeats&lt;/strong&gt;: Combine skills with persistent memory and scheduled runs for proactive agents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Security Best Practices&lt;/strong&gt;: Dedicated user/sandbox, VirusTotal checks on ClawHub, approval gates, read-only defaults, regular audits. Consider NVIDIA NemoClaw for added guardrails.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Multi-Agent Setups&lt;/strong&gt;: Run specialized OpenClaw instances (e.g., one for coding, one for personal).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CometAPI Integration&lt;/strong&gt;: Set as primary provider in OpenClaw config. Use model routing for cost optimization (e.g., via their dashboard analytics). Benefits: Single key, broad model access (including latest releases), lower latency/cost, privacy focus. Ideal for high-token agents.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Building Custom Skills&lt;/strong&gt;: OpenClaw can help generate them—start simple with &lt;code&gt;SKILL.md&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Future Outlook&lt;/strong&gt;: By late 2026, expect deeper multimodal skills, better enterprise controls, and even more seamless integrations. Skills like these position you at the forefront.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion: Level Up Your OpenClaw Today
&lt;/h2&gt;

&lt;p&gt;These top 6 skills—GOG, Agent Browser, Self-Improving/Capability Evolver, GitHub, Summarize, and Project Management—form a robust foundation for a truly autonomous AI teammate in 2026. Start with core productivity ones (GOG + Summarize), then layer on automation and self-improvement.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Ready to deploy?&lt;/strong&gt; Head to &lt;a href="https://openclaw.ai/" rel="noopener noreferrer"&gt;openclaw.ai&lt;/a&gt;, install via the one-liner, and power it with &lt;strong&gt;CometAPI&lt;/strong&gt; at cometapi.com for seamless, affordable access to the best models. Experiment safely, iterate with your agent, and watch productivity soar.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/top-6-openclaw-skills-you-can-t-afford-to-miss-in-2026/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=top-6-openclaw-skills-you-can-t-afford-to-miss-in-2026"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How to Add AI Video Generation to a SaaS App</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Mon, 21 Sep 2026 04:40:33 +0000</pubDate>
      <link>https://dev.to/masonreed1/how-to-add-ai-video-generation-to-a-saas-app-1emk</link>
      <guid>https://dev.to/masonreed1/how-to-add-ai-video-generation-to-a-saas-app-1emk</guid>
      <description>&lt;p&gt;Adding video generation to your app is not the same as adding image generation. The API call returns immediately — but the video isn't ready yet. You get a task ID, and you have to keep asking "is it done?" until it is.&lt;/p&gt;

&lt;p&gt;Most developers hit this the first time they call a video API, wait for a response body with a video URL, and get back a task ID instead. This guide walks through the full flow: submitting a task, polling for results, handling failures, and storing the output before the URL expires.&lt;/p&gt;

&lt;h2&gt;
  
  
  What you'll build
&lt;/h2&gt;

&lt;p&gt;A backend service that accepts a text prompt or image, submits a video generation task, polls until it's complete, and returns the final video URL. You'll work with four models — Veo 3 Fast, Sora 2, Kling Video, and Runway — all through a single API key.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Prerequisites:&lt;/strong&gt;
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Python 3.8+ or Node.js 18+&lt;/li&gt;
&lt;li&gt;A &lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt; key&lt;/li&gt;
&lt;li&gt;Basic familiarity with REST APIs&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Understand why video generation is different
&lt;/h2&gt;

&lt;p&gt;With image generation, you send a request and get the image back in the same response. Video generation uses an async task queue:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Submit&lt;/strong&gt; a generation request → get back a &lt;code&gt;task_id&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Poll&lt;/strong&gt; a status endpoint every few seconds&lt;/li&gt;
&lt;li&gt;When status reaches a terminal state, you get the video URL&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Download and store&lt;/strong&gt; the video — the URL is temporary&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you treat video generation like image generation and wait for the first response to contain your video, your request will time out every time.&lt;/p&gt;

&lt;p&gt;In a production web service, this polling loop should run in a background worker (Celery, Bull, or similar), not in your request handler. The examples below use synchronous polling — fine for scripts and prototypes, but not for handling concurrent users.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose a model
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Provider&lt;/th&gt;
&lt;th&gt;Max duration&lt;/th&gt;
&lt;th&gt;Price (via CometAPI)&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Veo 3 Fast&lt;/td&gt;
&lt;td&gt;Google&lt;/td&gt;
&lt;td&gt;8 sec&lt;/td&gt;
&lt;td&gt;$0.05/sec&lt;/td&gt;
&lt;td&gt;Fast prototyping, social clips&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sora 2&lt;/td&gt;
&lt;td&gt;OpenAI (via CometAPI model ID)&lt;/td&gt;
&lt;td&gt;~10 sec&lt;/td&gt;
&lt;td&gt;$0.08/sec&lt;/td&gt;
&lt;td&gt;High-quality creative shorts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kling Video&lt;/td&gt;
&lt;td&gt;Kuaishou&lt;/td&gt;
&lt;td&gt;10 sec&lt;/td&gt;
&lt;td&gt;$0.13–$2.64/task&lt;/td&gt;
&lt;td&gt;Marketing content, granular control&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Runway Gen-3A Turbo&lt;/td&gt;
&lt;td&gt;Runway&lt;/td&gt;
&lt;td&gt;5 or 10 sec&lt;/td&gt;
&lt;td&gt;$0.32/task&lt;/td&gt;
&lt;td&gt;Image-to-video, commercial content&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Source**: CometAPI model pages, May 2026. Note: "Sora 2" is CometAPI's model&lt;/em&gt; &lt;em&gt;identifier&lt;/em&gt; &lt;em&gt;— refer to their&lt;/em&gt; &lt;a href="https://www.cometapi.com/models/openai/sora-2/" rel="noopener noreferrer"&gt;&lt;em&gt;model page&lt;/em&gt;&lt;/a&gt; &lt;em&gt;for the underlying model details.&lt;/em&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Veo 3 Fast&lt;/strong&gt; supports both text-to-video and image-to-video. Cheapest per second, good starting point.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sora 2&lt;/strong&gt; generates audio natively alongside the video — dialogue, ambient sound, and effects without a separate TTS step.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Kling Video&lt;/strong&gt; gives you &lt;code&gt;negative_prompt&lt;/code&gt;, &lt;code&gt;cfg_scale&lt;/code&gt;, camera movement settings, and a &lt;code&gt;pro&lt;/code&gt; mode. Most control of the four.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Runway&lt;/strong&gt; is image-to-video only via CometAPI. Give it a static image and a motion description, and it animates it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Submit a Veo task
&lt;/h2&gt;

&lt;p&gt;Veo uses &lt;code&gt;multipart/form-data&lt;/code&gt;. Use &lt;code&gt;files=&lt;/code&gt; in Python requests to send it correctly — &lt;code&gt;data=dict&lt;/code&gt; sends &lt;code&gt;application/x-www-form-urlencoded&lt;/code&gt;, which is not the same thing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requestsimport&lt;/span&gt; &lt;span class="n"&gt;osfrom&lt;/span&gt; &lt;span class="n"&gt;dotenv&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;load_dotenv&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt;&lt;span class="nf"&gt;load_dotenv&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;submit_veo_task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;16x9&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Submit a Veo 3 Fast text-to-video task. Returns task_id.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY environment variable is not set&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1/videos&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;files&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;veo3-fast&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;size&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;size&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="err"&gt;​​&lt;/span&gt;&lt;span class="n"&gt;task_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;submit_veo_task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;A paper kite drifting above a wheat field on a windy afternoon&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Task submitted: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;task_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Poll for the result
&lt;/h2&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;poll_veo_task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task_id&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;interval&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_wait&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;600&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Poll until Veo task completes. Returns video URL.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY environment variable is not set&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1/videos/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;task_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;elapsed&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;while&lt;/span&gt; &lt;span class="n"&gt;elapsed&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;lt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="n"&gt;max_wait&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;succeeded&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;output&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;][&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;status&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;failed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cancelled&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;RuntimeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Task &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;task_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; failed with status &lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;: &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;result&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;error&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;no error detail returned&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;interval&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;elapsed&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="n"&gt;interval&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;TimeoutError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Task &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;task_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; did not complete within &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;max_wait&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;​​&lt;/span&gt;&lt;span class="n"&gt;video_url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;poll_veo_task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;task_id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Video ready: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;video_url&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Use Kling Video for more control
&lt;/h2&gt;

&lt;p&gt;Kling has a different endpoint structure and uses JSON. Note that Kling's terminal status string is &lt;code&gt;"succeed"&lt;/code&gt; (not &lt;code&gt;"succeeded"&lt;/code&gt;) — this matches the API's actual response format:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;submit_kling_task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;std&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Submit a Kling text-to-video task. Returns task_id.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY environment variable is not set&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/kling/v1/videos/text2video&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model_name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;kling-v1-6&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;negative_prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;blurry, low quality, watermark&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;cfg_scale&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mf"&gt;0.5&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mode&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;mode&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;# "std" or "pro" &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;"aspect_ratio": "16:9", &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;"duration": duration &amp;amp;nbsp;# "5" or "10" &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;  }, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;timeout=30 &amp;amp;nbsp;  ) &amp;amp;nbsp; &amp;amp;nbsp;response.raise_for_status() &amp;amp;nbsp; &amp;amp;nbsp;return response.json()["data"]["task_id"]​​def poll_kling_task(task_id: str, interval: int = 10, max_wait: int = 600) -&amp;amp;gt; str: &amp;amp;nbsp; &amp;amp;nbsp;"""Poll Kling task until complete. Returns video URL.""" &amp;amp;nbsp; &amp;amp;nbsp;api_key = os.getenv("COMETAPI_KEY") &amp;amp;nbsp; &amp;amp;nbsp;if not api_key: &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;raise ValueError("COMETAPI_KEY environment variable is not set")​ &amp;amp;nbsp; &amp;amp;nbsp;headers = {"Authorization": f"Bearer {api_key}"} &amp;amp;nbsp; &amp;amp;nbsp;url = f"https://api.cometapi.com/kling/v1/videos/text2video/{task_id}" &amp;amp;nbsp; &amp;amp;nbsp;elapsed = 0​ &amp;amp;nbsp; &amp;amp;nbsp;while elapsed &amp;amp;lt; max_wait: &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;response = requests.get(url, headers=headers, timeout=30) &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;response.raise_for_status() &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;result = response.json() &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;status = result["data"]["task_status"]​ &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;if status == "succeed": &amp;amp;nbsp;# Kling uses "succeed", not "succeeded" &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;return result["data"]["task_result"]["videos"][0]["url"] &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;elif status == "failed": &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;error_detail = result.get("data", {}).get("task_result", "no detail") &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;raise RuntimeError( &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;f"Kling task {task_id} failed: {error_detail}" &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;  )​ &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;time.sleep(interval) &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;elapsed += interval​ &amp;amp;nbsp; &amp;amp;nbsp;raise TimeoutError(f"Kling task {task_id} timed out after {max_wait}s")
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Source**:&lt;/em&gt; &lt;a href="https://apidoc.cometapi.com/video-generation/kling-video" rel="noopener noreferrer"&gt;&lt;em&gt;CometAPI Kling Video docs&lt;/em&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Animate a static image with Runway
&lt;/h2&gt;

&lt;p&gt;Runway is image-to-video only. It also requires an extra header (X-Runway-Version):&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;submit_runway_task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;image_url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;motion_prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;duration&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;int&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Submit a Runway image-to-video task. Returns task_id.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;ValueError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY environment variable is not set&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/runwayml/v1/image_to_video&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;X-Runway-Version&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;2024-11-06&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;  &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gen3a_turbo&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;promptImage&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;image_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="c1"&gt;# must be a stable HTTPS URL &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;"promptText": motion_prompt, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;"duration": duration, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;"ratio": "1280:720", &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;"watermark": False &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;  }, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;timeout=30 &amp;amp;nbsp;  ) &amp;amp;nbsp; &amp;amp;nbsp;response.raise_for_status() &amp;amp;nbsp; &amp;amp;nbsp;return response.json()["id"]​​def poll_runway_task(task_id: str, interval: int = 5, max_wait: int = 600) -&amp;amp;gt; str: &amp;amp;nbsp; &amp;amp;nbsp;"""Poll Runway task. Returns video URL when done.""" &amp;amp;nbsp; &amp;amp;nbsp;api_key = os.getenv("COMETAPI_KEY") &amp;amp;nbsp; &amp;amp;nbsp;if not api_key: &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;raise ValueError("COMETAPI_KEY environment variable is not set")​ &amp;amp;nbsp; &amp;amp;nbsp;headers = { &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;"Authorization": f"Bearer {api_key}", &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;"X-Runway-Version": "2024-11-06" &amp;amp;nbsp;  } &amp;amp;nbsp; &amp;amp;nbsp;url = f"https://api.cometapi.com/runwayml/v1/tasks/{task_id}" &amp;amp;nbsp; &amp;amp;nbsp;elapsed = 0​ &amp;amp;nbsp; &amp;amp;nbsp;while elapsed &amp;amp;lt; max_wait: &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;response = requests.get(url, headers=headers, timeout=30) &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;response.raise_for_status() &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;result = response.json() &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;status = result.get("status")​ &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;if status == "task_not_exist": &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;# CometAPI-specific: task is still initializing, retry after a few seconds &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;time.sleep(interval) &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;elapsed += interval &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;continue &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;elif status == "succeeded": &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;return result["output"][0] &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;elif status in ("failed", "cancelled"): &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;raise RuntimeError(f"Runway task {task_id} failed: {result.get('error', 'no detail')}")​ &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;time.sleep(interval) &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;elapsed += interval​ &amp;amp;nbsp; &amp;amp;nbsp;raise TimeoutError(f"Runway task {task_id} timed out after {max_wait}s")
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;em&gt;Source**:&lt;/em&gt; &lt;a href="https://apidoc.cometapi.com/api/video/runway" rel="noopener noreferrer"&gt;&lt;em&gt;CometAPI Runway docs&lt;/em&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Save the video before the URL expires
&lt;/h2&gt;

&lt;p&gt;Video URLs from generation APIs are temporary. Download the file immediately and store it somewhere you control:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requestsimport&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;download_video&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;output_path&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;-&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt;&lt;span class="s"&gt;Download video from URL to local file using streaming.&lt;/span&gt;&lt;span class="sh"&gt;"""&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;output_path&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parent&lt;/span&gt; &lt;span class="o"&gt;!=&lt;/span&gt; &lt;span class="n"&gt;pathlib&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Path&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;parent&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;mkdir&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;parents&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;exist_ok&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;​&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;out&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;wb&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;iter_content&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk_size&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="n"&gt;f&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="n"&gt;nbsp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Saved to &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;output_path&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;​​&lt;/span&gt;&lt;span class="c1"&gt;# Full flowtask_id = submit_veo_task("A timelapse of clouds moving over a city skyline")video_url = poll_veo_task(task_id)download_video(video_url, "output/city_timelapse.mp4")
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In production, swap the local file write for an upload to S3, Cloudflare R2, or your storage of choice. The streaming pattern stays the same — pipe the bytes directly rather than loading the whole video into memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Handle failures
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;Likely cause&lt;/th&gt;
&lt;th&gt;Fix&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Task stuck in queued for 10+ min&lt;/td&gt;
&lt;td&gt;Server load or model unavailable&lt;/td&gt;
&lt;td&gt;Retry with a different model&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;task_not_exist on first Runway poll&lt;/td&gt;
&lt;td&gt;Task still initializing&lt;/td&gt;
&lt;td&gt;Wait 5 sec and retry — documented CometAPI behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;failed with no error message&lt;/td&gt;
&lt;td&gt;Prompt triggered content filter&lt;/td&gt;
&lt;td&gt;Rephrase the prompt&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video URL returns 403&lt;/td&gt;
&lt;td&gt;URL expired before download&lt;/td&gt;
&lt;td&gt;Download immediately after getting the URL&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Timeout after 10 min&lt;/td&gt;
&lt;td&gt;Generation took too long&lt;/td&gt;
&lt;td&gt;Increase max_wait or switch to Veo 3 Fast&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Kling returns "succeed" not "succeeded"&lt;/td&gt;
&lt;td&gt;Kling's API uses non-standard status string&lt;/td&gt;
&lt;td&gt;This is correct — see Kling polling code above&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Source:&lt;/em&gt; &lt;a href="https://apidoc.cometapi.com/api/video/veo3" rel="noopener noreferrer"&gt;&lt;em&gt;CometAPI video generation docs&lt;/em&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Node.js version
&lt;/h2&gt;

&lt;p&gt;Node.js 18+ includes &lt;code&gt;fetch&lt;/code&gt; and &lt;code&gt;FormData&lt;/code&gt; natively. This example covers all four models:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Node.js 18+ — no extra packages needed​const API_KEY = process.env.COMETAPI_KEY;if (!API_KEY) throw new Error('COMETAPI_KEY is not set');​// --- Veo 3 Fast ---async function submitVeoTask(prompt, size = '16x9') { &amp;amp;nbsp;const form = new FormData(); &amp;amp;nbsp;form.append('prompt', prompt); &amp;amp;nbsp;form.append('model', 'veo3-fast'); &amp;amp;nbsp;form.append('size', size);​ &amp;amp;nbsp;const res = await fetch('https://api.cometapi.com/v1/videos', { &amp;amp;nbsp; &amp;amp;nbsp;method: 'POST', &amp;amp;nbsp; &amp;amp;nbsp;headers: { 'Authorization': `Bearer ${API_KEY}` }, &amp;amp;nbsp; &amp;amp;nbsp;body: form  }); &amp;amp;nbsp;if (!res.ok) throw new Error(`Veo submit failed: ${res.status}`); &amp;amp;nbsp;return (await res.json()).id;}​async function pollVeoTask(taskId, intervalMs = 10000, maxWaitMs = 600000) { &amp;amp;nbsp;let elapsed = 0; &amp;amp;nbsp;while (elapsed &amp;amp;lt; maxWaitMs) { &amp;amp;nbsp; &amp;amp;nbsp;const res = await fetch(`https://api.cometapi.com/v1/videos/${taskId}`, { &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;headers: { 'Authorization': `Bearer ${API_KEY}` } &amp;amp;nbsp;  }); &amp;amp;nbsp; &amp;amp;nbsp;if (!res.ok) throw new Error(`Poll failed: ${res.status}`); &amp;amp;nbsp; &amp;amp;nbsp;const result = await res.json();​ &amp;amp;nbsp; &amp;amp;nbsp;if (result.status === 'succeeded') return result.output[0]; &amp;amp;nbsp; &amp;amp;nbsp;if (['failed', 'cancelled'].includes(result.status)) { &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;throw new Error(`Task ${taskId} failed: ${result.error ?? 'no detail'}`); &amp;amp;nbsp;  } &amp;amp;nbsp; &amp;amp;nbsp;await new Promise(r =&amp;amp;gt; setTimeout(r, intervalMs)); &amp;amp;nbsp; &amp;amp;nbsp;elapsed += intervalMs;  } &amp;amp;nbsp;throw new Error(`Task ${taskId} timed out`);}​// --- Kling Video ---async function submitKlingTask(prompt, duration = '5', mode = 'std') { &amp;amp;nbsp;const res = await fetch('https://api.cometapi.com/kling/v1/videos/text2video', { &amp;amp;nbsp; &amp;amp;nbsp;method: 'POST', &amp;amp;nbsp; &amp;amp;nbsp;headers: { &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;'Authorization': `Bearer ${API_KEY}`, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;'Content-Type': 'application/json' &amp;amp;nbsp;  }, &amp;amp;nbsp; &amp;amp;nbsp;body: JSON.stringify({ &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;model_name: 'kling-v1-6', &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;prompt, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;negative_prompt: 'blurry, low quality, watermark', &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;cfg_scale: 0.5, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;mode, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;aspect_ratio: '16:9', &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;duration &amp;amp;nbsp;  })  }); &amp;amp;nbsp;if (!res.ok) throw new Error(`Kling submit failed: ${res.status}`); &amp;amp;nbsp;return (await res.json()).data.task_id;}​async function pollKlingTask(taskId, intervalMs = 10000, maxWaitMs = 600000) { &amp;amp;nbsp;let elapsed = 0; &amp;amp;nbsp;while (elapsed &amp;amp;lt; maxWaitMs) { &amp;amp;nbsp; &amp;amp;nbsp;const res = await fetch( &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;`https://api.cometapi.com/kling/v1/videos/text2video/${taskId}`, &amp;amp;nbsp; &amp;amp;nbsp;  { headers: { 'Authorization': `Bearer ${API_KEY}` } } &amp;amp;nbsp;  ); &amp;amp;nbsp; &amp;amp;nbsp;if (!res.ok) throw new Error(`Kling poll failed: ${res.status}`); &amp;amp;nbsp; &amp;amp;nbsp;const result = await res.json(); &amp;amp;nbsp; &amp;amp;nbsp;const status = result.data.task_status;​ &amp;amp;nbsp; &amp;amp;nbsp;if (status === 'succeed') return result.data.task_result.videos[0].url; &amp;amp;nbsp; &amp;amp;nbsp;if (status === 'failed') { &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;throw new Error(`Kling task ${taskId} failed: ${JSON.stringify(result.data.task_result ?? 'no detail')}`); &amp;amp;nbsp;  } &amp;amp;nbsp; &amp;amp;nbsp;await new Promise(r =&amp;amp;gt; setTimeout(r, intervalMs)); &amp;amp;nbsp; &amp;amp;nbsp;elapsed += intervalMs;  } &amp;amp;nbsp;throw new Error(`Kling task ${taskId} timed out`);}​// --- Runway (image-to-video) ---async function submitRunwayTask(imageUrl, motionPrompt, duration = 5) { &amp;amp;nbsp;const res = await fetch('https://api.cometapi.com/runwayml/v1/image_to_video', { &amp;amp;nbsp; &amp;amp;nbsp;method: 'POST', &amp;amp;nbsp; &amp;amp;nbsp;headers: { &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;'Authorization': `Bearer ${API_KEY}`, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;'X-Runway-Version': '2024-11-06', &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;'Content-Type': 'application/json' &amp;amp;nbsp;  }, &amp;amp;nbsp; &amp;amp;nbsp;body: JSON.stringify({ &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;model: 'gen3a_turbo', &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;promptImage: imageUrl, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;promptText: motionPrompt, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;duration, &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;ratio: '1280:720', &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;watermark: false &amp;amp;nbsp;  })  }); &amp;amp;nbsp;if (!res.ok) throw new Error(`Runway submit failed: ${res.status}`); &amp;amp;nbsp;return (await res.json()).id;}​async function pollRunwayTask(taskId, intervalMs = 5000, maxWaitMs = 600000) { &amp;amp;nbsp;let elapsed = 0; &amp;amp;nbsp;while (elapsed &amp;amp;lt; maxWaitMs) { &amp;amp;nbsp; &amp;amp;nbsp;const res = await fetch( &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;`https://api.cometapi.com/runwayml/v1/tasks/${taskId}`, &amp;amp;nbsp; &amp;amp;nbsp;  { headers: { 'Authorization': `Bearer ${API_KEY}`, 'X-Runway-Version': '2024-11-06' } } &amp;amp;nbsp;  ); &amp;amp;nbsp; &amp;amp;nbsp;if (!res.ok) throw new Error(`Runway poll failed: ${res.status}`); &amp;amp;nbsp; &amp;amp;nbsp;const result = await res.json(); &amp;amp;nbsp; &amp;amp;nbsp;const status = result.status;​ &amp;amp;nbsp; &amp;amp;nbsp;if (status === 'task_not_exist') { &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;// CometAPI-specific: task still initializing &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;await new Promise(r =&amp;amp;gt; setTimeout(r, intervalMs)); &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;elapsed += intervalMs; &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;continue; &amp;amp;nbsp;  } &amp;amp;nbsp; &amp;amp;nbsp;if (status === 'succeeded') return result.output[0]; &amp;amp;nbsp; &amp;amp;nbsp;if (['failed', 'cancelled'].includes(status)) { &amp;amp;nbsp; &amp;amp;nbsp; &amp;amp;nbsp;throw new Error(`Runway task ${taskId} failed: ${result.error ?? 'no detail'}`); &amp;amp;nbsp;  } &amp;amp;nbsp; &amp;amp;nbsp;await new Promise(r =&amp;amp;gt; setTimeout(r, intervalMs)); &amp;amp;nbsp; &amp;amp;nbsp;elapsed += intervalMs;  } &amp;amp;nbsp;throw new Error(`Runway task ${taskId} timed out`);}​// Usage exampleconst taskId = await submitVeoTask('A paper kite drifting above a wheat field');const videoUrl = await pollVeoTask(taskId);console.log('Video ready:', videoUrl);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What's next
&lt;/h2&gt;

&lt;p&gt;You now have working code for four video models, a polling loop that handles failures, and a download step that keeps you from losing generated content.&lt;/p&gt;

&lt;p&gt;The next problem most developers hit: they've hardcoded one model, and switching to a cheaper or faster option means touching multiple files. The next article covers how to route requests across models without rewriting your code.&lt;/p&gt;

&lt;p&gt;Next: &lt;strong&gt;How to Switch Between AI Models Without Rewriting Your Code&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Q: Why do I get a task ID instead of a video in the API response?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Video generation is async — models like Veo, Sora, Kling, and Runway take 2–5 minutes to render. The API returns a task ID immediately so your request doesn't time out. You poll a separate status endpoint until the task reaches a terminal state (&lt;code&gt;succeeded&lt;/code&gt;, &lt;code&gt;succeed&lt;/code&gt;, &lt;code&gt;failed&lt;/code&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Q: How long does a generated video URL stay valid?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Video URLs from generation APIs are temporary. Download the file immediately after getting the URL and store it in your own storage (S3, Cloudflare R2, etc.). Don't store the URL and expect it to work hours later.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Q: What's the difference between Veo 3 Fast and Kling Video?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Veo 3 Fast is cheaper ($0.05/sec), faster, and simpler to call. Kling Video gives you more control: &lt;code&gt;negative_prompt&lt;/code&gt;, &lt;code&gt;cfg_scale&lt;/code&gt;, camera movement settings, and a &lt;code&gt;pro&lt;/code&gt; quality mode. If you need to fine-tune the output, use Kling. If you need speed and low cost, use Veo 3 Fast.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Q: Can I generate video from an image instead of a text prompt?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;Yes. Veo supports image-to-video by passing an &lt;code&gt;input_reference&lt;/code&gt; file. Kling supports it via the &lt;code&gt;/kling/v1/videos/image2video&lt;/code&gt; endpoint with an &lt;code&gt;image&lt;/code&gt; parameter (URL or base64). Runway is image-to-video only — it doesn't accept text-only prompts via CometAPI.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Q: Why does Runway return&lt;/strong&gt; &lt;strong&gt;&lt;code&gt;task_not_exist&lt;/code&gt;&lt;/strong&gt; &lt;strong&gt;on the first poll?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;This is documented CometAPI behavior — the task is still initializing on the backend. Wait a few seconds and retry. It's not an error. The polling code above handles this automatically.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Q: Why does Kling use&lt;/strong&gt; &lt;strong&gt;&lt;code&gt;"succeed"&lt;/code&gt;&lt;/strong&gt; &lt;strong&gt;instead of&lt;/strong&gt; &lt;strong&gt;&lt;code&gt;"succeeded"&lt;/code&gt;&lt;/strong&gt;?
&lt;/h3&gt;

&lt;p&gt;That's Kling's actual API response format. It's not a typo. Veo and Runway use &lt;code&gt;"succeeded"&lt;/code&gt; — Kling uses &lt;code&gt;"succeed"&lt;/code&gt;. If you're building a unified polling wrapper, you'll need to handle both strings.&lt;/p&gt;

&lt;h3&gt;
  
  
  &lt;strong&gt;Q: Is the synchronous polling loop safe to use in a web server?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;No. The polling loop in this guide blocks the thread for minutes at a time. In a real web service with concurrent users, run the polling in a background worker (Celery for Python, Bull for Node.js). Submit the task in the request handler, return the task ID to the client, and let the worker notify the client when the video is ready.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-add-ai-video-generation-to-a-saas-app/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-add-ai-video-generation-to-a-saas-app"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The Hidden Cost of Juggling OpenAI, Anthropic, and Google Credentials</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Mon, 21 Sep 2026 04:17:59 +0000</pubDate>
      <link>https://dev.to/masonreed1/the-hidden-cost-of-juggling-openai-anthropic-and-google-credentials-36g4</link>
      <guid>https://dev.to/masonreed1/the-hidden-cost-of-juggling-openai-anthropic-and-google-credentials-36g4</guid>
      <description>&lt;p&gt;&lt;em&gt;Multi-provider AI setups don't show their cost in the&lt;/em&gt; &lt;a href="https://www.cometapi.com/pricing/" rel="noopener noreferrer"&gt;&lt;em&gt;API bill&lt;/em&gt;&lt;/a&gt; &lt;em&gt;— they show it in developer hours. Once you put a number on it, the case for consolidation stops being a matter of taste and becomes a line item your finance team can defend.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The cost most teams never count
&lt;/h2&gt;

&lt;p&gt;Most product engineering teams running on top of three or four AI providers can tell you to the dollar what they spent on tokens last month. They can tell you which feature drove the most cost, which model is cheapest per million tokens, and whether their burn rate is on track for the quarter. What they usually cannot tell you is what the operational overhead of running three or four provider relationships actually costs them in developer time.&lt;/p&gt;

&lt;p&gt;This is not because the cost is invisible. Every engineer on the team feels it. It is because the cost is paid in increments small enough to dismiss — a credential lookup here, a debugging session there, a half-day of integration work the next time a new model ships. None of these show up in any standard cost report. The API bill captures inference cost. The cloud bill captures infrastructure cost. Engineering time spent on cross-provider operational work shows up nowhere, because no system was designed to capture it. The default reporting infrastructure has a blind spot exactly the shape of this category of work.&lt;/p&gt;

&lt;p&gt;This article is the version of that conversation that puts numbers on the table. The argument is not that &lt;a href="https://www.cometapi.com/cometapi-vs-direct-provider-apis/" rel="noopener noreferrer"&gt;multi-provider&lt;/a&gt; AI is bad — there are workloads where running multiple providers is genuinely the right architectural choice. The argument is that the operational cost of that choice is real, quantifiable, and usually larger than teams realise. Once you can name the figure, the architectural conversation becomes a real cost-benefit analysis instead of a series of competing intuitions.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The headline finding:&lt;/strong&gt; For a typical five-engineer team running three AI providers, the annual operational cost of multi-provider work — counted in developer hours alone — sits between $35,000 and $60,000. That is not a hypothetical; it is what comes out when you instrument the workflow and add up the actual time. The number doesn't appear on any budget because no system was built to capture it. The case for changing your setup is what happens when you start counting it.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  5 hidden line items
&lt;/h2&gt;

&lt;p&gt;The operational cost of multi-provider AI work breaks down into five categories, each of which can be measured if you decide to. None of them is enormous in isolation; the cost is in the aggregate. Below, each category, what it looks like in practice, and how much time it consumes per month for a representative engineering team.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Initial onboarding to each provider
&lt;/h3&gt;

&lt;p&gt;Setting up a new AI provider relationship is a multi-step process. Sign up for the account. Verify the email and any payment method. Read the rate-limit documentation. Set up secrets management for the new credential. Install the provider's SDK if it differs from what you already use. Wire the credential through your CI/CD pipeline so deploys can authenticate. Add the new provider to your secrets-rotation calendar. For a typical provider, this is 4–8 hours of engineering time, mostly done by one engineer but with at least some coordination overhead from others.&lt;/p&gt;

&lt;p&gt;This cost is paid once per provider, but the "once" matters. If your team adds one new provider per year — which is below the 2026 baseline for serious teams — you pay this cost annually. The first onboarding doesn't feel expensive because it is one engineer for one afternoon. The fourth onboarding, when the same engineer has now done it four times in eighteen months and is increasingly resistant to doing it again, is where the friction shows up.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Monthly billing reconciliation
&lt;/h3&gt;

&lt;p&gt;Every month-end, someone on the team — usually the lead engineer or technical founder — pulls usage data from each provider's dashboard, normalises the formats, attributes the costs to product features or clients, and produces a consolidated view. For a team with three providers and a clean usage pattern, this is roughly 2–4 hours per month. For a team with four or more providers, or with complex cost-attribution requirements (per-feature, per-client, or per-team), it can be 6–10 hours per month.&lt;/p&gt;

&lt;p&gt;The reconciliation work is not engineering work in any meaningful sense — it is bookkeeping done by someone who is overqualified for the task. The fact that it lands on the engineering side rather than the finance side is itself a clue that the workflow has not been designed; it has just accumulated.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Credential rotation and security hygiene
&lt;/h3&gt;

&lt;p&gt;Good security practice requires rotating API credentials periodically — quarterly for most teams, more frequently for regulated workloads. With one provider, this is a routine 30-minute task. With three or four providers, each with its own rotation interface, its own propagation timing, and its own potential failure modes, the same task expands to several hours per cycle. Add the time spent debugging when a rotated credential doesn't propagate cleanly to a production environment, and the cost rises further. A team that rotates credentials quarterly across four providers loses 8–15 hours per year to this specific category alone.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Debugging auth and integration errors across providers
&lt;/h3&gt;

&lt;p&gt;A request fails. Was it a rate limit? An auth error? A model deprecation? A content-policy refusal? On a single-provider setup, this is one debugging surface. On a multi-provider setup, it is multiple — and the error formats, status codes, and dashboard log layouts differ across each. The cognitive cost of switching between provider conventions during incident response is the friction point that bites worst, because it lands during exactly the moments when speed matters most. For a team with three providers, this category typically runs 2–4 hours per month — and spikes much higher when a provider has an outage or changes their auth model unexpectedly.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Re-evaluating model choices each time a new release lands
&lt;/h3&gt;

&lt;p&gt;In 2026, new frontier model releases happen roughly every three to six weeks. Each release triggers a small evaluation cycle: read the model card, decide whether it warrants testing against your workload, set up the integration if it's from a provider you don't already have access to, run your eval suite, compare results. On a multi-provider direct setup, this cycle is 1–2 days of engineering time per release, mostly because the setup cost is non-trivial. On a single-endpoint setup with the new model already available behind the same credential, the same evaluation is 1–2 hours. The difference, multiplied by 6–10 evaluation cycles per year, is meaningful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Putting numbers on it
&lt;/h2&gt;

&lt;p&gt;The categories above are easy to describe and easy to dismiss as small. The exercise that changes the conversation is multiplying them out for a realistic team. Below, the calculation for a five-engineer product team running three AI providers — the kind of setup that has become unremarkable for AI-native startups.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost category&lt;/th&gt;
&lt;th&gt;Hours per month&lt;/th&gt;
&lt;th&gt;Hours per year&lt;/th&gt;
&lt;th&gt;Annual cost ($)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Initial provider onboarding (1 new provider/year)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;5 hrs&lt;/td&gt;
&lt;td&gt;$675&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Monthly billing reconciliation&lt;/td&gt;
&lt;td&gt;3 hrs&lt;/td&gt;
&lt;td&gt;36 hrs&lt;/td&gt;
&lt;td&gt;$4,860&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quarterly credential rotation across 3 providers&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;12 hrs&lt;/td&gt;
&lt;td&gt;$1,620&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Debugging auth and integration errors&lt;/td&gt;
&lt;td&gt;3 hrs&lt;/td&gt;
&lt;td&gt;36 hrs&lt;/td&gt;
&lt;td&gt;$4,860&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;New model evaluations (8 releases/year)&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;120 hrs&lt;/td&gt;
&lt;td&gt;$16,200&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Daily context-switching tax (15 min/engineer)&lt;/td&gt;
&lt;td&gt;25 hrs&lt;/td&gt;
&lt;td&gt;300 hrs&lt;/td&gt;
&lt;td&gt;$40,500&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total annual operational cost&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;509 hrs&lt;/td&gt;
&lt;td&gt;$68,715&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;How the numbers are calculated.&lt;/em&gt; Hours per month for shared work (reconciliation, debugging) are total team hours, not per-engineer. The daily context-switching tax is 15 minutes per engineer per working day, multiplied by five engineers and roughly 200 working days a year. Dollar conversion uses a fully-loaded engineering cost of $135/hour, which is a conservative figure for a mid-level engineer in the US or UK once salary, benefits, taxes, and overhead are accounted for. Adjust both the team size and the hourly rate for your specific situation; the structure of the calculation is the same.&lt;/p&gt;

&lt;p&gt;Three observations about this table that matter more than the bottom-line number.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;First, the largest line is the one teams notice least.&lt;/strong&gt; The $40,500 daily context-switching tax — 15 minutes per engineer per day on dashboard checks, credential lookups, and cross-provider documentation — is paid in tiny enough increments that no one feels it as a cost. It is also, by a meaningful margin, the largest single item on the table. The accumulated effect of small daily frictions outweighs every other category combined.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Second, the model evaluation cost is the most strategically expensive.&lt;/strong&gt; $16,200 a year on evaluation cycles is significant, but the real cost is the evaluations that don't happen because the setup cost makes them not worth it. Teams running multi-provider direct setups evaluate fewer new models, take longer to migrate when a better fit appears, and end up running suboptimal model choices for longer than they should. The hidden cost of slower iteration is harder to put a number on, but it is real.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Third, the calculation is conservative.&lt;/strong&gt; The numbers above assume a team that has its multi-provider workflow working reasonably well. Teams in worse shape — with credential rotation neglected, with no consistent reconciliation cadence, with evaluation cycles that take longer because eval infrastructure isn't in place — face higher numbers. The $68,715 figure is what good operational discipline looks like; the figure for teams without it can comfortably be twice that.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this cost never appears on the budget
&lt;/h2&gt;

&lt;p&gt;If the operational cost is this large, why does no team have a line item for it? The answer is structural, not accidental. Four reasons together explain the blind spot:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;No system was built to capture this category.&lt;/strong&gt; Time-tracking systems are built for billable client work. Engineering reporting is built for feature delivery. Cost-attribution systems are built for COGS. None of them have a natural place to record "45 minutes debugging a rate-limit issue across two providers." The work happens; the recording infrastructure for it doesn't exist.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The increments are small enough to dismiss.&lt;/strong&gt; Each individual instance of this work is 5–30 minutes. That's below the threshold most engineers would treat as worth tracking. The cost shows up only when you add the increments across the year — which nobody does, because there's no system that does it automatically.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The work is invisible from outside the engineering team.&lt;/strong&gt; The CTO sees feature delivery velocity. The CFO sees the API bill. Neither sees the integration overhead in between. Unless an engineer escalates the cost explicitly — and most don't, because they have built the work into their normal routine — the category remains structurally invisible to the people making the architectural decisions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The framing is engineering culture, not finance language.&lt;/strong&gt; Engineers describe this work as "keeping the lights on" or "normal operational overhead" — language that doesn't trigger budget scrutiny. If the same work were described as "$68,715 a year of operational integration cost," the response from leadership would be immediate. The framing controls whether the cost becomes visible.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Together, these four factors create the blind spot that makes the multi-provider operational cost so persistent. The cost is real, the impact is significant, and almost nothing in the standard reporting infrastructure surfaces it. Making the case to change your setup starts with the framing — naming the cost in finance language is what brings it into the conversation.&lt;/p&gt;

&lt;h2&gt;
  
  
  The break-even calculation
&lt;/h2&gt;

&lt;p&gt;Once you have the annual operational cost named, the question becomes: at what team size or workload volume does consolidating to a single-endpoint setup pay back the migration cost? The migration itself is genuinely small — typically 4–16 engineering hours depending on how the existing codebase is structured. Below the break-even point, that migration cost outweighs the operational saving; above it, the saving accumulates from the first month onwards.&lt;/p&gt;

&lt;p&gt;Working backwards from the calculation above, the break-even for a team of five engineers running three providers is approximately one month of operational saving — about $5,700 per month of reclaimed engineering time covers the entire migration cost. For smaller teams, the break-even can be longer; for larger teams, it shortens to a few weeks. Three scenarios that bracket the typical range:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Team profile&lt;/th&gt;
&lt;th&gt;Annual operational cost (est.)&lt;/th&gt;
&lt;th&gt;Migration cost (est.)&lt;/th&gt;
&lt;th&gt;Break-even&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Solo founder, 2 providers&lt;/td&gt;
&lt;td&gt;$12,000&lt;/td&gt;
&lt;td&gt;$1,000&lt;/td&gt;
&lt;td&gt;1 month&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5-engineer startup, 3 providers&lt;/td&gt;
&lt;td&gt;$68,000&lt;/td&gt;
&lt;td&gt;$2,000&lt;/td&gt;
&lt;td&gt;2 weeks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;12-engineer scale-up, 4 providers&lt;/td&gt;
&lt;td&gt;$180,000&lt;/td&gt;
&lt;td&gt;$4,000&lt;/td&gt;
&lt;td&gt;1 week&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The pattern is consistent: the larger the team and the more providers in scope, the faster the break-even. The break-even calculation also doesn't include the secondary benefits — faster model evaluation cycles, recovered focus time, fewer credential incidents — which add to the case but are harder to quantify cleanly. The migration cost is small enough that for any team running two or more providers with non-trivial volume, it pays back within the first month.&lt;/p&gt;

&lt;h2&gt;
  
  
  The qualitative cost
&lt;/h2&gt;

&lt;p&gt;The numbers above capture the time directly spent on multi-provider operational work. They do not capture the second-order costs that show up in how the team works. These are harder to quantify but matter more in practice.&lt;/p&gt;

&lt;p&gt;Friction in the engineering loop. When even routine work requires context-switching across provider conventions, engineers ship slower. The shipping speed cost is not the literal time spent on switching; it is the cumulative effect of fragmented attention on the rest of the day. Productivity research has been clear for decades that context-switching has a residue cost that outlasts the switch itself. The engineering team that is constantly switching between provider dashboards is the same team that gets less done in a sprint than its size would suggest.&lt;/p&gt;

&lt;p&gt;Resistance to better choices. When evaluating a new model requires setting up a new provider relationship, the threshold for "is it worth trying?" rises. Engineers stop suggesting evaluations they would otherwise have run. The result is that the team's model choices drift from optimal — not because anyone made a bad decision, but because the better decisions never got made. This is the failure mode that is hardest to see in retrospect because the alternative was never tested.&lt;/p&gt;

&lt;p&gt;Burnout from administrative work. The work of managing multiple providers is genuinely tedious. Engineers tolerate it for a while, then start to resent it. The resentment shows up in standups, in slower responses to operational questions, in engineers proposing architectural changes whose real driver is escape from the credential-management overhead. The hidden cost shows up as morale, retention, and team velocity — and by the time those metrics are bad enough to notice, they have been bad for months.&lt;/p&gt;

&lt;h2&gt;
  
  
  The case to bring to your team
&lt;/h2&gt;

&lt;p&gt;If the calculation above lines up with your team's reality and you want to make the case for consolidating, here is a practical framing that works in internal conversations:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Lead with the dollar figure, not the engineering complaint.&lt;/strong&gt; "Our current multi-provider setup is costing us roughly $X in engineering time per year" lands very differently to "managing credentials is annoying." The first triggers cost-benefit analysis; the second triggers a polite acknowledgment and no action.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Show your work on the calculation.&lt;/strong&gt; Use the table structure from this article, adapted to your team's actual hours and hourly rate. The credibility of the number depends on the methodology being transparent. "Here's what we counted, here's the rate we used, here's how it adds up" is much more defensible than a single dollar figure asserted without breakdown.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Name the secondary benefits separately.&lt;/strong&gt; The break-even pays back in dollars terms within weeks for most teams. The secondary benefits — faster model evaluation, recovered focus time, reduced credential-incident risk — are presented as additional upside, not as the core case. This keeps the primary argument financially defensible while giving the team the qualitative case they care about.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Be honest about what doesn't change.&lt;/strong&gt; Aggregating to a single endpoint does not eliminate &lt;a href="https://www.cometapi.com/enterprise/" rel="noopener noreferrer"&gt;&lt;strong&gt;compliance obligations&lt;/strong&gt;&lt;/a&gt;, does not change underlying model quality, and does not solve every operational problem. Naming these limits up front is what makes the rest of the argument trustworthy. The team you're presenting to will trust your recommendation more if you have already named the trade-offs honestly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Propose a phased migration, not a big bang.&lt;/strong&gt; The most defensible proposal is to move one new feature or one experimental workload to the new setup first, measure the operational impact, then expand. This de-risks the change and gives you a real-data answer to "does this actually work for us?" within a month. Most teams that propose phased migrations get internal approval easily; teams that propose all-at-once migrations face more resistance even when the numbers are good.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Where this leaves you
&lt;/h2&gt;

&lt;p&gt;The operational cost of multi-provider AI work is real, large, and structurally invisible. Most teams pay $35,000 to $60,000 per year for a setup they assume is free because none of the cost appears on any line item. Once you start counting it, the case for consolidation moves out of "engineering preference" territory and into "defensible financial decision" territory. The numbers are the lever; the case is just letting them speak.&lt;/p&gt;

&lt;p&gt;&lt;em&gt;The practical next step:&lt;/em&gt; Run the calculation for your team. Use the structure from this article, adapt the hours to your actual setup, and produce the annual figure. The exercise takes less than an hour and produces a number that decides the question. CometAPI is one route for the single-endpoint consolidation; the practical case is the same regardless of which aggregator you choose.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Multi-provider AI doesn't cost what the API bill says it costs. The real cost includes 500+ hours of engineering time per year on integration overhead — credential rotation, billing reconciliation, dashboard navigation, daily context-switching. At realistic engineering rates, that's $35K–$60K of cost no system was built to capture. Naming it in finance language is what brings it into the conversation; running the calculation for your team is what wins the argument.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ready to integrate reliably? Head to&amp;nbsp;&lt;a href="https://cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt;&amp;nbsp;and&amp;nbsp;&lt;a href="https://apidoc.cometapi.com/" rel="noopener noreferrer"&gt;API doc&lt;/a&gt;&amp;nbsp;for seamless Claude Fable 5 access alongside other frontier models, unified billing, and enterprise-grade reliability. Sign up today and get started with generous credits for new users—your next breakthrough project awaits.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/the-hidden-cost-of-juggling-openai-anthropic-and-google-credentials/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=the-hidden-cost-of-juggling-openai-anthropic-and-google-credentials"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Seedance 2.5: What Developers Should Know Before Launch</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Mon, 21 Sep 2026 03:40:37 +0000</pubDate>
      <link>https://dev.to/masonreed1/seedance-25-what-developers-should-know-before-launch-3mo3</link>
      <guid>https://dev.to/masonreed1/seedance-25-what-developers-should-know-before-launch-3mo3</guid>
      <description>&lt;h2&gt;
  
  
  Current Status
&lt;/h2&gt;

&lt;p&gt;Seedance 2.5 is ByteDance’s next-generation video generation model in the Seedance family. ByteDance announced it on June 23, 2026, at the Volcano Engine FORCE conference. As of June 30, 2026, it was reportedly in global enterprise beta, with a public launch targeted for early July 2026.&lt;/p&gt;

&lt;p&gt;The main changes from Seedance 2.0 are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Native single-segment generation up to 30 seconds&lt;/li&gt;
&lt;li&gt;Support for up to 50 multimodal reference assets&lt;/li&gt;
&lt;li&gt;More precise local editing and continuation&lt;/li&gt;
&lt;li&gt;Improved control over motion, camera work, consistency, and pacing&lt;/li&gt;
&lt;li&gt;Native audio-video synchronization&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Seedance 2.5 accepts text alongside images, video clips, audio, and other reference material. The model is aimed at workflows that currently require stitching together several short generations and repairing inconsistencies in post-production.&lt;/p&gt;

&lt;p&gt;Independent benchmarks for 2.5 are not yet available. Seedance 2.0, however, already performs strongly in evaluations such as the Artificial Analysis Video Arena, where its text-to-video-with-audio score was approximately 1,219 Elo.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Model Adds
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Longer Native Clips
&lt;/h3&gt;

&lt;p&gt;Most video generation systems produce native clips in the 5-15 second range. Seedance 2.5 extends that to 30 seconds in one continuous generation.&lt;/p&gt;

&lt;p&gt;That matters for scenes with a setup, action, camera movement, and resolution. Keeping those elements in one pass should reduce stitching and preserve narrative rhythm. ByteDance demonstrations reportedly include multi-character interactions and spacecraft previsualization using 100k+ polygon models while maintaining structural integrity across the clip.&lt;/p&gt;

&lt;h3&gt;
  
  
  More Reference Material
&lt;/h3&gt;

&lt;p&gt;Seedance 2.5 supports up to 50 multimodal references, compared with approximately 12 in Seedance 2.0.&lt;/p&gt;

&lt;p&gt;References can include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Character sheets and style guides&lt;/li&gt;
&lt;li&gt;Images and product assets&lt;/li&gt;
&lt;li&gt;Motion-reference video&lt;/li&gt;
&lt;li&gt;Audio and music&lt;/li&gt;
&lt;li&gt;3D greybox or pre-production material&lt;/li&gt;
&lt;li&gt;Text instructions&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical benefit is tighter control over characters, environments, lighting, branding, and object identity. This is particularly relevant to teams with established asset libraries.&lt;/p&gt;

&lt;h3&gt;
  
  
  Motion and Cinematic Control
&lt;/h3&gt;

&lt;p&gt;The model is designed for more stable motion, physical interactions, lighting, and character consistency. Its control surface is intended to cover camera behavior such as dolly, tracking, and POV shots, along with performance and pacing.&lt;/p&gt;

&lt;p&gt;Seedance 2.0 already emphasized “exceptional motion stability” and “director-level control” over performance, lighting, shadows, and camera movement. Version 2.5 appears to extend those capabilities to longer and more complex scenes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Local Editing and Continuation
&lt;/h3&gt;

&lt;p&gt;Local or region-based editing allows a team to change part of a frame without regenerating the entire video. High-quality continuation is intended to extend an existing clip while preserving its style and rhythm.&lt;/p&gt;

&lt;p&gt;For e-commerce, this could mean generating a strong lifestyle video once and creating variants for different SKUs, backgrounds, languages, and promotions without rebuilding every scene.&lt;/p&gt;

&lt;h3&gt;
  
  
  Output Options
&lt;/h3&gt;

&lt;p&gt;Seedance 2.5 is expected to support resolutions up to 4K, building on Seedance 2.0’s capabilities. It also supports multiple aspect ratios, including &lt;code&gt;16:9&lt;/code&gt; and &lt;code&gt;9:16&lt;/code&gt;, for social, web, and broadcast-oriented outputs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Seedance 2.5 Compared with 2.0
&lt;/h2&gt;

&lt;p&gt;Seedance 2.0 already supports text, image, audio, and video inputs through a unified audio-video generation architecture. The newer model mainly addresses duration, reference capacity, editing, and production control.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Seedance 2.0&lt;/th&gt;
&lt;th&gt;Seedance 2.5&lt;/th&gt;
&lt;th&gt;Practical effect&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Native clip length&lt;/td&gt;
&lt;td&gt;Approximately 5-15 seconds&lt;/td&gt;
&lt;td&gt;Up to 30 seconds in one pass&lt;/td&gt;
&lt;td&gt;Fewer stitches and better continuity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Maximum references&lt;/td&gt;
&lt;td&gt;Up to approximately 12 multimodal inputs&lt;/td&gt;
&lt;td&gt;Up to 50 multimodal inputs&lt;/td&gt;
&lt;td&gt;More control over identity, style, and branding&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Editing&lt;/td&gt;
&lt;td&gt;Basic refinement&lt;/td&gt;
&lt;td&gt;Local or region editing with improved consistency&lt;/td&gt;
&lt;td&gt;More targeted iteration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resolution&lt;/td&gt;
&lt;td&gt;Up to 1080p/4K options&lt;/td&gt;
&lt;td&gt;Enhanced native 4K support&lt;/td&gt;
&lt;td&gt;Sharper production output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Continuation&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;td&gt;Higher-quality continuation with rhythmic consistency&lt;/td&gt;
&lt;td&gt;Easier longer-form work&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt adherence&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Approximately 20% better, as reported&lt;/td&gt;
&lt;td&gt;Potentially fewer retries&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical use&lt;/td&gt;
&lt;td&gt;Short clips and basic multimodal work&lt;/td&gt;
&lt;td&gt;Cinematic scenes, previsualization, and branded long-form content&lt;/td&gt;
&lt;td&gt;Broader professional use&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The reported improvement figures are vendor-reported. Generation speed, failure rates, and independent quality metrics still need to be validated after public release.&lt;/p&gt;

&lt;h2&gt;
  
  
  Release Timeline
&lt;/h2&gt;

&lt;p&gt;ByteDance previewed Seedance 2.5 in late June 2026 at the Volcano Engine conference. The expected public launch window is early July 2026, following an enterprise beta and expansion through platforms such as Dreamina, CapCut, and API providers.&lt;/p&gt;

&lt;p&gt;As of late June 2026, access remained limited. Developers should verify model availability, pricing, and exact output specifications before committing to a production workflow.&lt;/p&gt;

&lt;h2&gt;
  
  
  Calling the Video Endpoint
&lt;/h2&gt;

&lt;p&gt;For an OpenAI-compatible video workflow, the request can look like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1/videos&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]},&lt;/span&gt;
    &lt;span class="n"&gt;data&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;A slow cinematic camera push across a coastal landscape at sunrise&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;doubao-seedance-2-0&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;seconds&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;4&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;size&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;16:9&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;())&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A unified multi-model API such as CometAPI is useful here when the application needs one integration point for text, image, and video models, or when it needs to switch between Seedance versions during evaluation and rollout.&lt;/p&gt;

&lt;p&gt;For multimodal references, add an &lt;code&gt;input_reference&lt;/code&gt; file for image-guided generation. Advanced multi-reference requests can repeat the same multipart fields in upload order. The uploaded files can then be referenced sequentially as &lt;code&gt;[Image 1]&lt;/code&gt;, &lt;code&gt;[Image 2]&lt;/code&gt;, and &lt;code&gt;[Image 3]&lt;/code&gt;, with a specific role assigned to each image in the prompt.&lt;/p&gt;

&lt;p&gt;Video generation is asynchronous. Submit the task, retain its ID, and poll the video status endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GET /v1/videos/
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Seedance query endpoint returns the current task status, progress, and the signed video URL after completion. The documented polling flow works for Seedance 1.0 Pro, 1.5 Pro, and 2.0. Node.js and cURL clients use the same request shape with the appropriate base URL and credentials.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical Workloads
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Social Advertising
&lt;/h3&gt;

&lt;p&gt;A 30-second clip can contain the hook, problem, product reveal, proof point, lifestyle shot, and call to action in one generation. Teams can then create variants by audience, region, aspect ratio, product, or visual style.&lt;/p&gt;

&lt;h3&gt;
  
  
  Product and E-Commerce Video
&lt;/h3&gt;

&lt;p&gt;Reference-driven generation can turn catalog images into lifestyle scenes for seasonal campaigns, marketplace listings, product pages, and localized promotions. Human review remains necessary, but the approach can reduce the time needed for initial concepts and variant production.&lt;/p&gt;

&lt;h3&gt;
  
  
  Storyboards and Previsualization
&lt;/h3&gt;

&lt;p&gt;Film, animation, and agency teams can explore camera movement, blocking, lighting, pacing, and visual style before committing to a shoot or full production. A 30-second draft can also support pitches and client approvals.&lt;/p&gt;

&lt;h3&gt;
  
  
  Training and Education
&lt;/h3&gt;

&lt;p&gt;Short generated videos fit onboarding, safety explanations, product walkthroughs, classroom examples, and internal training. Multimodal references make it possible to match a company’s equipment, environment, or product rather than relying on generic stock footage.&lt;/p&gt;

&lt;h3&gt;
  
  
  Creative Localization
&lt;/h3&gt;

&lt;p&gt;Reference consistency and local editing are important for producing variants across languages, climates, packaging versions, and audience segments. The goal is to modify the relevant creative elements without regenerating the entire campaign from scratch.&lt;/p&gt;

&lt;h2&gt;
  
  
  Production Notes
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Build an Evaluation Set
&lt;/h3&gt;

&lt;p&gt;I would start with evaluation rather than direct customer-facing automation. Include product ads, character scenes, motion-heavy prompts, brand-safe requests, rejected prompts, and edge cases.&lt;/p&gt;

&lt;p&gt;Score at least:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompt adherence&lt;/li&gt;
&lt;li&gt;Identity and reference consistency&lt;/li&gt;
&lt;li&gt;Motion quality&lt;/li&gt;
&lt;li&gt;Visual artifacts&lt;/li&gt;
&lt;li&gt;Audio quality&lt;/li&gt;
&lt;li&gt;Moderation behavior&lt;/li&gt;
&lt;li&gt;Latency&lt;/li&gt;
&lt;li&gt;Cost&lt;/li&gt;
&lt;li&gt;Failure and retry rates&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Keep Generation Asynchronous
&lt;/h3&gt;

&lt;p&gt;A 30-second video can take substantially longer than a chat completion or image request. Use a queue, persist task IDs, poll status, and notify users when the signed output is ready. This avoids request timeouts and makes retries explicit.&lt;/p&gt;

&lt;h3&gt;
  
  
  Version Everything
&lt;/h3&gt;

&lt;p&gt;Store the prompt, model ID, parameters, reference IDs, output URL, reviewer score, and final decision. Prompt and output versioning makes comparisons between Seedance 2.5, Seedance 2.0, Veo, Kling, Runway, and other models reproducible.&lt;/p&gt;

&lt;h3&gt;
  
  
  Retain Human Review
&lt;/h3&gt;

&lt;p&gt;Commercial video introduces brand, legal, likeness, and copyright concerns. Human review is appropriate for public advertising, realistic people, influencer-like content, regulated industries, and outputs using third-party references.&lt;/p&gt;

&lt;p&gt;Avoid requests involving copyrighted characters, real people without permission, or misleading depictions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Design Fallbacks
&lt;/h3&gt;

&lt;p&gt;Do not assume one model should serve every workload. Short drafts may be better handled by Seedance 2.0 or a faster video model, while Seedance 2.5 can be reserved for final, high-value generations. Availability, latency, and cost should all be part of routing decisions.&lt;/p&gt;

&lt;h2&gt;
  
  
  Frequently Asked Questions
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Seedance 2.5 available?
&lt;/h3&gt;

&lt;p&gt;As of June 30, 2026, it was reportedly in global enterprise beta, with public availability expected in early July 2026. Developers should check the live model catalog before using it in production.&lt;/p&gt;

&lt;h3&gt;
  
  
  How long can it generate?
&lt;/h3&gt;

&lt;p&gt;The headline capability is native video generation up to 30 seconds. Seedance 2.0’s documented generation range is 4-15 seconds, so the newer model substantially increases single-pass duration.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the main difference from Seedance 2.0?
&lt;/h3&gt;

&lt;p&gt;Seedance 2.5 adds native 30-second generation, up to 50 multimodal references, and improved local editing control. Seedance 2.0 already provides text, image, audio, and video inputs, native audio-video generation, and strong motion stability.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is it best suited for?
&lt;/h3&gt;

&lt;p&gt;The strongest target workloads are 30-second social ads, product videos, cinematic storyboards, previsualization, brand-consistent creative variants, e-commerce demonstrations, training clips, and other reference-heavy generation tasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  What should developers verify before production?
&lt;/h3&gt;

&lt;p&gt;Confirm the live model ID, access status, pricing, exact resolution and duration limits, output formats, generation latency, failure behavior, moderation rules, and signed URL lifetime. These details may change as the model moves from beta to public release.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/what-is-seedance-2-5/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=what-is-seedance-2-5"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Calling Claude Through a Gateway: SDK Choices and Production Trade-offs</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Mon, 21 Sep 2026 02:39:29 +0000</pubDate>
      <link>https://dev.to/masonreed1/calling-claude-through-a-gateway-sdk-choices-and-production-trade-offs-3g6f</link>
      <guid>https://dev.to/masonreed1/calling-claude-through-a-gateway-sdk-choices-and-production-trade-offs-3g6f</guid>
      <description>&lt;p&gt;I reach for a unified gateway when an application needs several model providers but does not need several separate account, billing, and SDK integrations. It is also a way to call Claude without maintaining a direct Anthropic developer account: the application authenticates with the gateway, which routes requests upstream.&lt;/p&gt;

&lt;p&gt;CometAPI is one example, advertising Claude, GPT, and 500+ other generative AI models behind a single key and an OpenAI-compatible endpoint. Its mid-2026 model listings include Claude Opus 4.8 and Claude Sonnet 5. Those are catalog claims, so I would confirm current availability before choosing a production model.&lt;/p&gt;

&lt;p&gt;The important decision is not just which provider handles billing. It is whether the application needs a portable chat interface or Claude-specific request semantics.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose the API Contract First
&lt;/h2&gt;

&lt;p&gt;The gateway exposes two integration paths:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Concern&lt;/th&gt;
&lt;th&gt;OpenAI-compatible API&lt;/th&gt;
&lt;th&gt;Native Anthropic Messages API&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Endpoint&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/v1/chat/completions&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;/v1/messages&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SDK base URL&lt;/td&gt;
&lt;td&gt;&lt;code&gt;https://api.cometapi.com/v1&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;https://api.cometapi.com&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model scope&lt;/td&gt;
&lt;td&gt;Multiple providers, including GPT, Claude, and Gemini&lt;/td&gt;
&lt;td&gt;Claude&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authentication&lt;/td&gt;
&lt;td&gt;Bearer&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;x-api-key&lt;/code&gt; or Bearer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Response structure&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;choices&lt;/code&gt; and &lt;code&gt;message&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Content blocks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude-specific controls&lt;/td&gt;
&lt;td&gt;Not exposed through this compatibility path&lt;/td&gt;
&lt;td&gt;Thinking, caching, effort, server-side tools&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I would use chat completions for shared application code across providers. I would choose Messages when extended thinking, prompt caching, or server-side tools are part of the actual workflow. A common interface is useful, but it is not a promise that every model has identical capabilities.&lt;/p&gt;

&lt;h2&gt;
  
  
  Reuse an OpenAI Client
&lt;/h2&gt;

&lt;p&gt;Install the SDK with &lt;code&gt;pip install openai&lt;/code&gt;. Generate a gateway key, then change the client’s &lt;code&gt;base_url&lt;/code&gt; and &lt;code&gt;api_key&lt;/code&gt;. This request retains the source example’s model, temperature, and output limit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;your_gateway_key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-4.8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;What are the primary structural benefits of a unified API gateway?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For streaming, the gateway translates Claude’s native server-sent events into OpenAI-compatible chunks. Using the same client:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;stream&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-4.8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Explain the concept of latency overhead in unified APIs.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;content&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;delta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
            &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;end&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;flush&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Changing &lt;code&gt;model&lt;/code&gt; lets this integration address another supported model without replacing the client or streaming loop. I would still test parameter support, tool behavior, structured output, and error handling for each target. Request compatibility reduces integration work; it does not eliminate model-specific validation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep Claude-Specific Features on the Messages Path
&lt;/h2&gt;

&lt;p&gt;The native route uses the official Anthropic SDK, but the key comes from the gateway. Set &lt;code&gt;COMETAPI_KEY&lt;/code&gt; in the environment before running this example. Notice that the base URL does &lt;strong&gt;not&lt;/strong&gt; include &lt;code&gt;/v1&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;anthropic&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Anthropic&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;message&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-sonnet-4-6&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;system&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a helpful assistant.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Hello, world&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The SDK sends &lt;code&gt;x-api-key&lt;/code&gt; by default; the endpoint also accepts &lt;code&gt;Authorization: Bearer&lt;/code&gt;. The documented Claude-specific controls include &lt;code&gt;thinking&lt;/code&gt; with a minimum &lt;code&gt;budget_tokens&lt;/code&gt; of 1,024, &lt;code&gt;cache_control&lt;/code&gt; on content blocks, &lt;code&gt;output_config.effort&lt;/code&gt;, and server-side &lt;code&gt;web_fetch&lt;/code&gt; and &lt;code&gt;web_search&lt;/code&gt; tools. Check support against the selected model rather than assuming every Claude version accepts every control. The native example uses &lt;code&gt;claude-sonnet-4-6&lt;/code&gt;, not the Sonnet 5 model mentioned in the catalog.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Verify Before Shipping
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Latency and Compatibility
&lt;/h3&gt;

&lt;p&gt;Every gateway introduces another network hop plus authentication, routing, and potentially schema translation. The source reports typical overhead below 400 ms; I would treat that as a claim to benchmark, not a latency guarantee. Background summarization can tolerate a different budget from real-time voice or tightly timed agent workflows. Streaming does not remove the need to measure initial response latency.&lt;/p&gt;

&lt;p&gt;System prompts, temperature, structured JSON output, and tool/function calling are described as supported through the compatibility layer. Newly released or proprietary features may arrive later. For a workflow that depends on a particular feature, I would check the compatibility documentation and test the exact request. The native endpoint avoids some translation constraints, but its availability still needs verification.&lt;/p&gt;

&lt;h3&gt;
  
  
  Data Handling and Access
&lt;/h3&gt;

&lt;p&gt;The gateway describes encrypted transit, including TLS 1.3, minimized persistent logging, and enterprise zero data retention configurations. “Minimized logging” is not the same commitment as ZDR. I would review retention, zero-data-training policies, logging configuration, and contractual scope before sending sensitive inputs.&lt;/p&gt;

&lt;p&gt;Inference still happens at the upstream provider. Both the gateway’s handling and the model host’s data processing agreements matter. Likewise, using a gateway does not automatically remove regional compliance obligations, payment restrictions, or upstream access rules. A single account may simplify access, but it is not a universal compliance exemption.&lt;/p&gt;

&lt;h3&gt;
  
  
  Quotas, Cost, and Governance
&lt;/h3&gt;

&lt;p&gt;Consolidated billing is the clearest operational benefit: fewer prepaid balances, invoices, credentials, and usage dashboards to maintain. Gateway quotas can simplify capacity management, but I would verify actual limits and throttling behavior rather than assume there is no upstream constraint or production approval process.&lt;/p&gt;

&lt;p&gt;The source advertises pay-as-you-go pricing with no monthly fee, rates 20–40% below official pricing, and 1 million free signup tokens. These are commercial terms to confirm during evaluation, not architectural guarantees. Compare the price of the actual models and features the workload uses.&lt;/p&gt;

&lt;p&gt;A shared interface also makes task-based routing easier. I would reserve a flagship model such as Claude Opus 4.8 for complex reasoning or long-context analysis, then evaluate cheaper models for classification, extraction, and formatting. Switching a model identifier is easy; proving acceptable output quality is the work.&lt;/p&gt;

&lt;p&gt;The documented dashboard supports usage monitoring, project spending limits, access-token management, and budgets assigned to individual API keys. Those controls are worth testing: centralized billing only helps governance if an experimental project cannot consume the organization’s entire allocation.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Decision Rule
&lt;/h2&gt;

&lt;p&gt;I would choose a gateway when multi-model routing, A/B testing, consolidated procurement, and reduced SDK maintenance justify another dependency in the request path. I would use the OpenAI-compatible endpoint for portable chat workflows and the native Messages endpoint for Claude-specific behavior.&lt;/p&gt;

&lt;p&gt;Before rollout, I would validate model availability, feature support, latency, quotas, current pricing, and both layers of data handling. Avoiding a direct Anthropic account simplifies administration. It does not remove the engineering responsibility to understand where requests go and what contract the application relies on.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-access-the-claude-api-without-an-anthropic-account/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-access-the-claude-api-without-an-anthropic-account"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>GPT-6 Astra Access: Verify the Route, Then Make One Small Call</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Mon, 21 Sep 2026 01:53:33 +0000</pubDate>
      <link>https://dev.to/masonreed1/gpt-6-astra-access-verify-the-route-then-make-one-small-call-24h9</link>
      <guid>https://dev.to/masonreed1/gpt-6-astra-access-verify-the-route-then-make-one-small-call-24h9</guid>
      <description>&lt;p&gt;I treat API access as three separate checks: the gateway lists the model, the credential works, and an authenticated request completes. Creating a key only covers part of that.&lt;/p&gt;

&lt;p&gt;For a unified multi-model API such as CometAPI, the credential must belong to the gateway receiving the request. An OpenAI SDK client can target that gateway, but it still needs the gateway’s key and base URL.&lt;/p&gt;

&lt;p&gt;Here’s the sequence I’d use to verify &lt;code&gt;gpt-6-astra&lt;/code&gt; access before wiring it into an application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check the model catalog first
&lt;/h2&gt;

&lt;p&gt;The source guide reported &lt;code&gt;gpt-6-astra&lt;/code&gt; as available, with &lt;code&gt;upcoming: false&lt;/code&gt; and support for both &lt;code&gt;/v1/responses&lt;/code&gt; and &lt;code&gt;/v1/chat/completions&lt;/code&gt;. Treat that as a dated observation: check the live catalog before integration.&lt;/p&gt;

&lt;p&gt;The public catalog endpoint requires no Authorization header:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://api.cometapi.com/api/models &lt;span class="se"&gt;\&lt;/span&gt;
  | jq &lt;span class="s1"&gt;'.data[] | select(.id == "gpt-6-astra") | {
      id,
      provider,
      upcoming,
      endpoints
    }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Look for the exact ID, &lt;code&gt;upcoming&lt;/code&gt; set to &lt;code&gt;false&lt;/code&gt;, and the route you intend to use in &lt;code&gt;endpoints&lt;/code&gt;. That establishes public routing availability; it does not validate your account or key.&lt;/p&gt;

&lt;p&gt;If the command prints nothing, check the spelling and the &lt;a href="https://www.cometapi.com/models/openai/gpt-6-astra/" rel="noopener noreferrer"&gt;live model page&lt;/a&gt;. Don’t substitute a display name, prepend a provider name, or borrow an alias from another gateway.&lt;/p&gt;

&lt;p&gt;I’d start with &lt;code&gt;/v1/responses&lt;/code&gt; for the reasoning workflow described here. Use &lt;code&gt;/v1/chat/completions&lt;/code&gt; when the application specifically needs the messages-based interface.&lt;/p&gt;

&lt;h2&gt;
  
  
  Create a credential with a limited quota
&lt;/h2&gt;

&lt;p&gt;Create an account, sign in, and open the &lt;a href="https://www.cometapi.com/console/token" rel="noopener noreferrer"&gt;API Keys page&lt;/a&gt;. Select &lt;strong&gt;Create API Key&lt;/strong&gt;, name it something recognizable such as &lt;code&gt;astra-dev&lt;/code&gt;, choose an appropriate quota, and copy the value.&lt;/p&gt;

&lt;p&gt;You also need sufficient account credit or quota and a server-side environment for storing the credential. For a local Bash session:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;read&lt;/span&gt; &lt;span class="nt"&gt;-rsp&lt;/span&gt; &lt;span class="s2"&gt;"Gateway API key: "&lt;/span&gt; COMETAPI_KEY
&lt;span class="nb"&gt;printf&lt;/span&gt; &lt;span class="s1"&gt;'\n'&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;COMETAPI_KEY
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Keep development, staging, and production keys separate. Explicit quotas limit the damage from accidental loops or leaked credentials, while descriptive names make rotation easier.&lt;/p&gt;

&lt;p&gt;Store deployed credentials in a secrets manager or server-side environment variable. Keep them out of browser JavaScript, mobile bundles, repositories, screenshots, and support tickets. Store any direct OpenAI credentials separately so configuration changes cannot silently mix providers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the smallest useful request
&lt;/h2&gt;

&lt;p&gt;My first request would contain no tools, files, streaming, or long context. A small payload makes authentication and routing failures easier to isolate.&lt;/p&gt;

&lt;p&gt;Send the gateway key to the gateway endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;--fail-with-body&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  https://api.cometapi.com/v1/responses &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$COMETAPI_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "gpt-6-astra",
    "input": "Reply with exactly: GPT-6 Astra access confirmed.",
    "reasoning": {"effort": "low"},
    "max_output_tokens": 40
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The target is HTTP 200 with a response object containing an &lt;code&gt;id&lt;/code&gt;, a completed status, the expected model, output content, and usage data. Keep the response ID during testing; it helps investigate routing or platform problems.&lt;/p&gt;

&lt;p&gt;An HTML page or redirect does not establish API access. Check the hostname and path before changing the payload.&lt;/p&gt;

&lt;h3&gt;
  
  
  Use the same configuration in Python
&lt;/h3&gt;

&lt;p&gt;The OpenAI Python SDK can make the same request with an explicit &lt;code&gt;api_key&lt;/code&gt;, &lt;code&gt;base_url&lt;/code&gt;, and model selection. A separate gateway SDK is unnecessary.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;responses&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-6-astra&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Reply with exactly: GPT-6 Astra access confirmed.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;reasoning&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;effort&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;low&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;max_output_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;40&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nb"&gt;id&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;output_text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use &lt;code&gt;https://api.cometapi.com/v1&lt;/code&gt; as the SDK base URL and &lt;code&gt;https://api.cometapi.com/v1/responses&lt;/code&gt; for the direct request. Neither URL has a trailing period. Don’t substitute &lt;code&gt;api.openai.com&lt;/code&gt; while retaining the gateway credential.&lt;/p&gt;

&lt;p&gt;Once the minimal call succeeds, add application instructions and features individually.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decide what counts as a successful access test
&lt;/h2&gt;

&lt;p&gt;I’d keep four checks in the integration notes:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check&lt;/th&gt;
&lt;th&gt;What it establishes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Catalog returns &lt;code&gt;gpt-6-astra&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;The public catalog recognizes the model and lists its routes.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authenticated endpoint returns JSON&lt;/td&gt;
&lt;td&gt;The request reaches an API endpoint rather than a login page or redirect.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Response completes with the expected model&lt;/td&gt;
&lt;td&gt;The key, account, route, payload, and model routing worked together for that request.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Usage appears in the dashboard&lt;/td&gt;
&lt;td&gt;The request is recorded against the intended account.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A successful catalog lookup cannot prove private access. Likewise, a key’s existence cannot guarantee sufficient quota, current availability, a valid payload, or acceptance under rate limits.&lt;/p&gt;

&lt;p&gt;Check usage or billing records after the first call. This also gives you a starting point for cost monitoring.&lt;/p&gt;

&lt;h2&gt;
  
  
  Diagnose failures before retrying
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Authentication and access: 401 or 403
&lt;/h3&gt;

&lt;p&gt;For &lt;strong&gt;401 Unauthorized&lt;/strong&gt;, check for a missing, malformed, or expired key and verify the destination host. The header should be exactly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer $COMETAPI_KEY
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Reload the environment variable or rotate the credential if needed. Repeating the same invalid credential will not help.&lt;/p&gt;

&lt;p&gt;For &lt;strong&gt;403 Forbidden&lt;/strong&gt;, return to the minimal payload, remove optional fields, and check account status, quota, and model access. Treat it as an access or request problem before assuming a temporary outage.&lt;/p&gt;

&lt;h3&gt;
  
  
  Routing: HTML, redirects, or model-not-found errors
&lt;/h3&gt;

&lt;p&gt;Recheck the base URL and endpoint path. During debugging, avoid silently following redirects on the authenticated request: a wrong path can otherwise appear to be a successful connection.&lt;/p&gt;

&lt;p&gt;For a model-not-found response, query the catalog again and compare the exact &lt;code&gt;gpt-6-astra&lt;/code&gt; identifier.&lt;/p&gt;

&lt;h3&gt;
  
  
  Rate limits and server failures
&lt;/h3&gt;

&lt;p&gt;For &lt;strong&gt;429 Too Many Requests&lt;/strong&gt;, inspect the error body and account limits. Reduce burst concurrency and use exponential backoff with jitter for rate-limit retries. Monitor usage by model and route.&lt;/p&gt;

&lt;p&gt;For &lt;strong&gt;500, 503, 504, or 524&lt;/strong&gt;, retain any request or response ID and use bounded backoff for temporary platform or timeout failures. Inspect the body first: if it reports &lt;code&gt;invalid_request&lt;/code&gt;, fix the payload before retrying.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put production controls around the working call
&lt;/h2&gt;

&lt;p&gt;After access is confirmed, issue a production-specific key with the smallest practical quota. Store it in a secrets manager, restrict who can view or rotate it, and document rotation.&lt;/p&gt;

&lt;p&gt;Keep model calls on your server. Browser and mobile clients should call your authenticated backend, where you can enforce user permissions, rate limits, and spending controls.&lt;/p&gt;

&lt;p&gt;Log the route, model ID, HTTP status, latency, response ID, and token usage. Redact credentials and sensitive prompt data. Track development and production separately so test traffic does not obscure production failures.&lt;/p&gt;

&lt;p&gt;For payload changes, consult the &lt;a href="https://apidoc.cometapi.com/api/text/responses" rel="noopener noreferrer"&gt;Responses API documentation&lt;/a&gt;. The source also links a &lt;a href="https://developers.openai.com/api/docs/models/gpt-6-astra" rel="noopener noreferrer"&gt;first-party model reference&lt;/a&gt;; check current documentation and the live catalog before relying on a capability or endpoint.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/get-gpt-6-astra-api-access-cometapi/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=get-gpt-6-astra-api-access-cometapi"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>GPT-5.6 API Costs: What I’d Budget Beyond the Token Price</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Fri, 18 Sep 2026 17:10:24 +0000</pubDate>
      <link>https://dev.to/masonreed1/gpt-56-api-costs-what-id-budget-beyond-the-token-price-1m7e</link>
      <guid>https://dev.to/masonreed1/gpt-56-api-costs-what-id-budget-beyond-the-token-price-1m7e</guid>
      <description>&lt;h2&gt;
  
  
  Start with the route, not the family name
&lt;/h2&gt;

&lt;p&gt;The first thing I’d check in a GPT-5.6 integration is the model ID. According to OpenAI’s &lt;a href="https://developers.openai.com/api/docs/guides/latest-model" rel="noopener noreferrer"&gt;model guidance&lt;/a&gt;, &lt;code&gt;gpt-5.6&lt;/code&gt; routes to &lt;code&gt;gpt-5.6-sol&lt;/code&gt;. It does not automatically select the cheapest tier that can handle a request.&lt;/p&gt;

&lt;p&gt;If I want Terra or Luna pricing, I’d use &lt;code&gt;gpt-5.6-terra&lt;/code&gt; or &lt;code&gt;gpt-5.6-luna&lt;/code&gt; explicitly.&lt;/p&gt;

&lt;p&gt;The July 30, 2026 pricing update reduced Terra rates by &lt;strong&gt;20%&lt;/strong&gt; and Luna rates by &lt;strong&gt;80%&lt;/strong&gt;, while leaving Sol unchanged. Here are the updated Standard rates for requests with &lt;strong&gt;at most 272,000 input tokens&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;All token prices below are USD per 1 million tokens.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Cached input&lt;/th&gt;
&lt;th&gt;Cache write&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-5.6-sol&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$6.25&lt;/td&gt;
&lt;td&gt;$30.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-5.6-terra&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$2.50&lt;/td&gt;
&lt;td&gt;$12.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-5.6-luna&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$0.02&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;td&gt;$1.20&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Source: &lt;a href="https://developers.openai.com/api/docs/pricing" rel="noopener noreferrer"&gt;OpenAI API pricing&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;OpenAI positions Sol as the flagship, Terra as the balanced option, and Luna as the low-cost tier for high-volume work. I’d treat those descriptions as evaluation starting points, not routing rules.&lt;/p&gt;

&lt;p&gt;The number that matters most for ordinary short-context traffic: &lt;strong&gt;output costs 6× as much as uncached input&lt;/strong&gt; across all three tiers. Trimming unnecessary response text can matter more than shaving a few tokens off a system prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put the rates through an actual workload
&lt;/h2&gt;

&lt;p&gt;A rate table is useful, but I prefer turning it into request costs immediately. The examples here use the updated July rates consistently, without caching, tools, retries, or regional adjustments.&lt;/p&gt;

&lt;h3&gt;
  
  
  A 1,000-input, 500-output request
&lt;/h3&gt;

&lt;p&gt;For Sol:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Input:  1,000 / 1,000,000 × $5  = $0.005
Output:   500 / 1,000,000 × $30 = $0.015
Total:                           $0.020
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Applying the same request shape to each model:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input cost&lt;/th&gt;
&lt;th&gt;Output cost&lt;/th&gt;
&lt;th&gt;Total per request&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;$0.005&lt;/td&gt;
&lt;td&gt;$0.015&lt;/td&gt;
&lt;td&gt;$0.020&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terra&lt;/td&gt;
&lt;td&gt;$0.002&lt;/td&gt;
&lt;td&gt;$0.006&lt;/td&gt;
&lt;td&gt;$0.008&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Luna&lt;/td&gt;
&lt;td&gt;$0.0002&lt;/td&gt;
&lt;td&gt;$0.0006&lt;/td&gt;
&lt;td&gt;$0.0008&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Output accounts for &lt;strong&gt;75%&lt;/strong&gt; of the token bill even though the request contains twice as many input tokens as output tokens.&lt;/p&gt;

&lt;p&gt;For chat, code generation, and agents, I’d look for unnecessary verbosity before spending much time optimizing a small prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  One million requests per month
&lt;/h3&gt;

&lt;p&gt;Now change the workload to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;1 million requests per month&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;2,000 input tokens per request&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;500 output tokens per request&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That produces &lt;strong&gt;2 billion input tokens&lt;/strong&gt; and &lt;strong&gt;500 million output tokens&lt;/strong&gt; monthly.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Monthly input&lt;/th&gt;
&lt;th&gt;Monthly output&lt;/th&gt;
&lt;th&gt;Monthly total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;$10,000&lt;/td&gt;
&lt;td&gt;$15,000&lt;/td&gt;
&lt;td&gt;$25,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terra&lt;/td&gt;
&lt;td&gt;$4,000&lt;/td&gt;
&lt;td&gt;$6,000&lt;/td&gt;
&lt;td&gt;$10,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Luna&lt;/td&gt;
&lt;td&gt;$400&lt;/td&gt;
&lt;td&gt;$600&lt;/td&gt;
&lt;td&gt;$1,000&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For this workload, the price cuts move Terra from &lt;strong&gt;$12,500 to $10,000&lt;/strong&gt; and Luna from &lt;strong&gt;$5,000 to $1,000&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Those differences justify routing experiments. They do not prove that Luna is the cheapest way to complete the work. More retries, failed tool calls, or manual review can consume the apparent savings.&lt;/p&gt;

&lt;p&gt;My preferred metric is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cost per successful task = total workflow cost / successful tasks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I’d track human review alongside that metric, rather than pretending it disappears because it is absent from the API invoice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat 272K input tokens as a billing boundary
&lt;/h2&gt;

&lt;p&gt;GPT-5.6 supports a &lt;strong&gt;1.05M-token context window&lt;/strong&gt;, but context capacity and short-context pricing are different limits.&lt;/p&gt;

&lt;p&gt;When input exceeds &lt;strong&gt;272,000 tokens&lt;/strong&gt;, the higher rates apply to the &lt;strong&gt;entire request&lt;/strong&gt;, not just the excess tokens:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Input: &lt;strong&gt;2×&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Cached input: &lt;strong&gt;2×&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Cache writes: &lt;strong&gt;2×&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Output: &lt;strong&gt;1.5×&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Long input&lt;/th&gt;
&lt;th&gt;Long cached input&lt;/th&gt;
&lt;th&gt;Long cache write&lt;/th&gt;
&lt;th&gt;Long output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sol&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$1.00&lt;/td&gt;
&lt;td&gt;$12.50&lt;/td&gt;
&lt;td&gt;$45.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terra&lt;/td&gt;
&lt;td&gt;$4.00&lt;/td&gt;
&lt;td&gt;$0.40&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$18.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Luna&lt;/td&gt;
&lt;td&gt;$0.40&lt;/td&gt;
&lt;td&gt;$0.04&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$1.80&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;This is a pricing cliff, not a gradual overage charge. It also means the &lt;strong&gt;6× output-to-input ratio applies to short-context Standard pricing&lt;/strong&gt;; long-context pricing has a different ratio because the multipliers differ.&lt;/p&gt;

&lt;p&gt;Near the boundary, I’d inspect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Duplicate retrieval chunks&lt;/li&gt;
&lt;li&gt;Stale conversation history&lt;/li&gt;
&lt;li&gt;Repository files unrelated to the task&lt;/li&gt;
&lt;li&gt;Tool output that could be reduced before the next model call&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That cleanup can do more than reduce token count: it can keep the whole request in the lower pricing band. OpenAI’s &lt;a href="https://developers.openai.com/api/docs/guides/cost-optimization" rel="noopener noreferrer"&gt;cost optimization guide&lt;/a&gt; covers additional approaches.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choose a processing tier separately from the model
&lt;/h2&gt;

&lt;p&gt;Model selection is only one pricing decision. Eligible GPT-5.6 text workloads also have different processing options.&lt;/p&gt;

&lt;p&gt;These are short-context input/output rates per 1 million tokens:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Processing option&lt;/th&gt;
&lt;th&gt;Sol&lt;/th&gt;
&lt;th&gt;Terra&lt;/th&gt;
&lt;th&gt;Luna&lt;/th&gt;
&lt;th&gt;Where I’d consider it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;$5 / $30&lt;/td&gt;
&lt;td&gt;$2 / $12&lt;/td&gt;
&lt;td&gt;$0.20 / $1.20&lt;/td&gt;
&lt;td&gt;Normal synchronous traffic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch&lt;/td&gt;
&lt;td&gt;$2.50 / $15&lt;/td&gt;
&lt;td&gt;$1 / $6&lt;/td&gt;
&lt;td&gt;$0.10 / $0.60&lt;/td&gt;
&lt;td&gt;Offline asynchronous jobs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flex&lt;/td&gt;
&lt;td&gt;$2.50 / $15&lt;/td&gt;
&lt;td&gt;$1 / $6&lt;/td&gt;
&lt;td&gt;$0.10 / $0.60&lt;/td&gt;
&lt;td&gt;Work that tolerates slower processing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fast mode&lt;/td&gt;
&lt;td&gt;$10 / $60&lt;/td&gt;
&lt;td&gt;$4 / $24&lt;/td&gt;
&lt;td&gt;$0.40 / $2.40&lt;/td&gt;
&lt;td&gt;Latency-sensitive traffic&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The listed Batch and Flex rates are &lt;strong&gt;50% below Standard&lt;/strong&gt;. I’d check their operational constraints before moving any production workload, rather than treating them as interchangeable discounts.&lt;/p&gt;

&lt;p&gt;Priority Processing was renamed &lt;strong&gt;Fast mode on July 30, 2026&lt;/strong&gt;. Existing requests using &lt;code&gt;service_tier: "priority"&lt;/code&gt; remain compatible.&lt;/p&gt;

&lt;p&gt;For &lt;strong&gt;Sol&lt;/strong&gt;, Fast mode offers &lt;strong&gt;up to 2.5× faster processing&lt;/strong&gt; at &lt;strong&gt;2× the Standard token price&lt;/strong&gt;, without changing model intelligence. That is a latency trade-off, not a quality upgrade. I’d pay for it only where the user-facing benefit justifies the premium.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cache reuse, not merely long prompts
&lt;/h2&gt;

&lt;p&gt;Cache writes cost &lt;strong&gt;1.25× normal input&lt;/strong&gt;, while matching cached reads use the discounted cached-input rate.&lt;/p&gt;

&lt;p&gt;For a reusable &lt;strong&gt;100,000-token prefix&lt;/strong&gt; on short-context Sol:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Operation&lt;/th&gt;
&lt;th&gt;Prefix input cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One uncached use&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One cache write&lt;/td&gt;
&lt;td&gt;$0.625, approximately $0.63&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One matching cached read&lt;/td&gt;
&lt;td&gt;$0.05&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two uncached uses cost &lt;strong&gt;$1.00&lt;/strong&gt;. A cache write followed by one matching read costs &lt;strong&gt;$0.675&lt;/strong&gt;, saving &lt;strong&gt;$0.325&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That comparison covers only the reusable prefix. It excludes output, other uncached input, tools, and retries.&lt;/p&gt;

&lt;p&gt;The practical distinction is between a long prefix and a &lt;strong&gt;long, stable, reused prefix&lt;/strong&gt;. A write that never gets reused costs more than ordinary input processing.&lt;/p&gt;

&lt;p&gt;I’d measure actual cache hits and read the &lt;a href="https://developers.openai.com/api/docs/guides/prompt-caching" rel="noopener noreferrer"&gt;prompt caching documentation&lt;/a&gt; for matching requirements, explicit breakpoints, and TTL behavior before building savings into a forecast.&lt;/p&gt;

&lt;h2&gt;
  
  
  Budget for tokens you don’t see and calls outside the model
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Reasoning is billed output
&lt;/h3&gt;

&lt;p&gt;Reasoning tokens are charged as output tokens even when they do not appear in the visible response. A concise answer does not necessarily imply a small output-token bill.&lt;/p&gt;

&lt;p&gt;Where supported, &lt;code&gt;reasoning.mode = "pro"&lt;/code&gt; is a setting I’d benchmark rather than enable by default. OpenAI does not list a separate fixed Pro surcharge; its cost impact comes from resulting token usage.&lt;/p&gt;

&lt;p&gt;The comparison should include task success, total output tokens, latency, and retries—not just whether one response looks better.&lt;/p&gt;

&lt;h3&gt;
  
  
  Search has its own bill
&lt;/h3&gt;

&lt;p&gt;Standard web search is listed at &lt;strong&gt;$10 per 1,000 calls&lt;/strong&gt;, plus search-content tokens billed at the selected model rate.&lt;/p&gt;

&lt;p&gt;At that price, two searches can cost more than the model tokens for a small Luna request.&lt;/p&gt;

&lt;p&gt;There is also a separate &lt;strong&gt;web search preview&lt;/strong&gt; rate for non-reasoning models: &lt;strong&gt;$25 per 1,000 calls&lt;/strong&gt;, with search-content tokens free. I would not apply one search rate across every tool and endpoint. Check the exact combination on the &lt;a href="https://developers.openai.com/api/docs/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Residency can add an uplift
&lt;/h3&gt;

&lt;p&gt;Eligible regional processing or data-residency endpoints for models released on or after &lt;strong&gt;March 5, 2026&lt;/strong&gt; carry a &lt;strong&gt;10% uplift&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If residency is a requirement, that belongs in the initial estimate—not in an explanation for why the first invoice exceeded it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the older models in the comparison
&lt;/h2&gt;

&lt;p&gt;A new family does not make previous routes irrelevant. These are the listed Standard short-context rates for several other OpenAI text models:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Cached input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-5.5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$30.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-5.4&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$2.50&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-5.4-mini&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.75&lt;/td&gt;
&lt;td&gt;$0.075&lt;/td&gt;
&lt;td&gt;$4.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;gpt-5.4-nano&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$0.02&lt;/td&gt;
&lt;td&gt;$1.25&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Luna now matches GPT-5.4 nano’s &lt;strong&gt;$0.20 input rate&lt;/strong&gt;, with slightly cheaper output: &lt;strong&gt;$1.20 versus $1.25&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is a useful pricing comparison, not evidence that the models behave identically. I’d keep existing routes in the evaluation set until the replacement earns its place.&lt;/p&gt;

&lt;p&gt;Also, ChatGPT subscriptions and API billing are separate products. “ChatGPT API pricing” does not identify a rate; the model and usage type do.&lt;/p&gt;

&lt;h2&gt;
  
  
  My production cost checklist
&lt;/h2&gt;

&lt;p&gt;For a basic uncached request, the arithmetic is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Token cost =
    (input tokens / 1,000,000 × input rate)
  + (output tokens / 1,000,000 × output rate)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For production, I’d expand that into a workflow estimate covering uncached input, cached reads, cache writes, output and reasoning tokens, tool fees, applicable processing-tier rates, regional uplift, retries, and fallbacks.&lt;/p&gt;

&lt;p&gt;Then I’d run the same representative tasks through candidate routes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Use explicit model IDs.&lt;/strong&gt; Avoid paying Sol rates accidentally through &lt;code&gt;gpt-5.6&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Start with the cheapest plausible model.&lt;/strong&gt; Promote difficult tasks based on measured results.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Record billed output, not just visible text.&lt;/strong&gt; Include reasoning usage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Monitor the 272K boundary.&lt;/strong&gt; Aggregate averages can hide expensive individual requests.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Test Batch or Flex for non-urgent work.&lt;/strong&gt; Verify operational fit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure cache reuse, tool calls, and retries.&lt;/strong&gt; Do not assume their costs cancel out.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Compare cost per successful task against latency and reliability.&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For cross-provider evaluations, a unified API such as CometAPI can be useful for running the same workload through multiple OpenAI-compatible routes.&lt;/p&gt;

&lt;p&gt;The cheapest token rate is a starting hypothesis. The route I’d ship is the one that meets the quality and latency requirements at the lowest measured workflow cost.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/gpt-5-6-pricing/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=gpt-5-6-pricing"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Stop Reconstructing Client AI Spend From Logs</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Fri, 18 Sep 2026 16:53:33 +0000</pubDate>
      <link>https://dev.to/masonreed1/stop-reconstructing-client-ai-spend-from-logs-5ban</link>
      <guid>https://dev.to/masonreed1/stop-reconstructing-client-ai-spend-from-logs-5ban</guid>
      <description>&lt;p&gt;At the end of a billing period, I want to answer a simple question: how much did each client’s AI work cost?&lt;/p&gt;

&lt;p&gt;Most provider dashboards answer a different question: how much did the account spend in total? That number is useful, but it doesn’t tell me how to divide the bill across clients.&lt;/p&gt;

&lt;p&gt;The usual workaround is unpleasant. I export request logs, match timestamps to projects, estimate ambiguous records, and maintain a spreadsheet until the numbers look plausible. That process is slow, difficult to audit, and not something I want to use as the basis for an invoice.&lt;/p&gt;

&lt;p&gt;The underlying issue is structural: provider billing is organized around my account, while my business is organized around clients.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use API keys as cost centres
&lt;/h2&gt;

&lt;p&gt;The clean fix is to issue a separate API key for each client—or for each client workflow—and track usage independently per key.&lt;/p&gt;

&lt;p&gt;Every request already carries the identity of the key that made it. If each key maps to one client, attribution happens when the request runs instead of during month-end reconciliation. Cost per client becomes a dashboard value rather than something I have to reconstruct from raw records.&lt;/p&gt;

&lt;p&gt;This is essentially a cost-centre model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One client or workflow gets one clearly named key.&lt;/li&gt;
&lt;li&gt;Requests made with that key accumulate in its usage bucket.&lt;/li&gt;
&lt;li&gt;The bucket reports the client’s spend and activity.&lt;/li&gt;
&lt;li&gt;The invoice is built from those totals.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is particularly useful when multiple models are accessed through the same unified endpoint. I can keep all client keys and their usage in one account while still getting separate reporting. A unified multi-model API such as &lt;a href="https://www.cometapi.com/cometapi-vs-direct-provider-apis/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt; is relevant here because the alternative is consolidating usage across several provider accounts.&lt;/p&gt;

&lt;h3&gt;
  
  
  The data available per key
&lt;/h3&gt;

&lt;p&gt;A useful per-key report should expose the dimensions needed to support an invoice:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Total spend:&lt;/strong&gt; the dollar cost of requests made with the key during the billing period.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Request volume:&lt;/strong&gt; the number of calls, which helps validate activity.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Token usage:&lt;/strong&gt; input and output token counts behind the charge.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model breakdown:&lt;/strong&gt; the models used and the cost associated with each.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That gives me more than a single number. If a client’s workload uses an inexpensive model for bulk processing and a frontier model for difficult tasks, the model split explains the resulting total. If a client questions the bill, I can provide usage and token context rather than presenting an unexplained portion of a larger account charge.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this is better than the usual workarounds
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Approach&lt;/th&gt;
&lt;th&gt;Operating model&lt;/th&gt;
&lt;th&gt;Failure mode&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Single key, parsed logs&lt;/td&gt;
&lt;td&gt;One key serves every client; usage is reconstructed at invoice time.&lt;/td&gt;
&lt;td&gt;Slow, error-prone, and approximate whenever records are ambiguous.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Separate provider accounts&lt;/td&gt;
&lt;td&gt;Each client gets separate accounts with each provider.&lt;/td&gt;
&lt;td&gt;Credentials, dashboards, and invoices multiply quickly.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Manual spreadsheet tracking&lt;/td&gt;
&lt;td&gt;Usage is recorded by hand as work happens.&lt;/td&gt;
&lt;td&gt;It depends on sustained discipline, becomes stale, and compounds errors.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-key tracking&lt;/td&gt;
&lt;td&gt;Each client gets a key on one account; usage is metered automatically.&lt;/td&gt;
&lt;td&gt;Attribution is captured at request time and reported directly.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The important difference is when attribution occurs. The first three options push the work to invoice time. Per-key tracking captures the relationship between request and client at the source.&lt;/p&gt;

&lt;p&gt;For two clients, manually parsing logs may be tolerable. For fifteen, it becomes recurring operational work. With fifty clients, it is effectively a part-time job. Per-key tracking requires roughly the same setup pattern at each scale: issue a key, use it consistently, and read the resulting totals.&lt;/p&gt;

&lt;h2&gt;
  
  
  A setup pattern that holds up
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Choose the reporting boundary
&lt;/h3&gt;

&lt;p&gt;Start with one key per client unless the client needs separate billing streams. If a client has multiple projects or workflows that should be reported independently, use one key per client-project or per workflow.&lt;/p&gt;

&lt;p&gt;Finer-grained keys produce finer-grained reports, but they also increase the number of credentials to manage. I use the smallest number of keys that matches the invoice structure I actually need.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Name keys for humans
&lt;/h3&gt;

&lt;p&gt;A dashboard full of opaque key IDs is not a reporting system. Name keys with the client or project they represent so the usage list is immediately readable.&lt;/p&gt;

&lt;p&gt;This matters at invoice time: the report should look like a client list, not like a collection of tokens that needs another mapping table.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Configure each deployment with its own key
&lt;/h3&gt;

&lt;p&gt;Point each client integration at the corresponding key. In most cases, this is an environment or deployment configuration change rather than an application rewrite.&lt;/p&gt;

&lt;p&gt;The important operational rule is that a shared service must not accidentally use a default key for every client. The key selection needs to follow the same client or workflow boundary as the work being billed.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Read the totals at the billing boundary
&lt;/h3&gt;

&lt;p&gt;At the end of the billing period, inspect each key’s:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Spend&lt;/li&gt;
&lt;li&gt;Request count&lt;/li&gt;
&lt;li&gt;Input and output tokens&lt;/li&gt;
&lt;li&gt;Model usage and model-level cost&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those values form the per-client usage report. There is no need to infer ownership from timestamps or manually split an aggregate total.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Rotate keys independently
&lt;/h3&gt;

&lt;p&gt;Per-client keys also reduce the blast radius of credential operations.&lt;/p&gt;

&lt;p&gt;If a client leaves, revoke that client’s key. If one key is exposed, rotate that key without interrupting every other client. The control boundary and the billing boundary are the same, which is much safer than relying on one credential for the entire operation.&lt;/p&gt;

&lt;p&gt;Usage is metered per token against the same published rates regardless of which key made the request. That means the per-key total maps directly to the underlying pricing: the amount charged to the client traces back to the amount charged by the platform, with any agency margin added transparently.&lt;/p&gt;

&lt;h2&gt;
  
  
  The operational benefits go beyond invoicing
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Client-level profitability
&lt;/h3&gt;

&lt;p&gt;Once I can see AI cost per client, I can compare it with the client’s fee. That exposes engagements with healthy margins and those quietly consuming them.&lt;/p&gt;

&lt;p&gt;This is useful for repricing, changing the workflow, or deciding whether a particular service should be packaged differently.&lt;/p&gt;

&lt;h3&gt;
  
  
  Faster detection of runaway usage
&lt;/h3&gt;

&lt;p&gt;A bad loop, misconfigured integration, or unexpected traffic surge is easier to spot when usage is attached to a client key. The increase is visible in the relevant bucket instead of being hidden inside the account aggregate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Fewer billing disputes
&lt;/h3&gt;

&lt;p&gt;A client asking what they paid for can receive an explanation based on request volume, token counts, and model usage. That is more defensible than assigning them an estimated share of a lump-sum provider bill.&lt;/p&gt;

&lt;h3&gt;
  
  
  Better estimates for future work
&lt;/h3&gt;

&lt;p&gt;Historical per-client usage gives me real data for quoting similar projects. Instead of guessing how much model usage a new engagement might require, I can use comparable workflows and their actual costs.&lt;/p&gt;

&lt;p&gt;At larger agencies, the same account-level controls—team access, spending visibility, and administrative oversight—turn this from a billing convenience into an operational control. Per-key attribution is useful on its own, but it becomes more valuable when it is part of the broader account governance model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The practical pattern
&lt;/h2&gt;

&lt;p&gt;Provider dashboards generally make total account spend easy to see. They do not automatically know which client should receive each portion of that total.&lt;/p&gt;

&lt;p&gt;The solution is to make client identity part of the request path:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Issue one key per client or billable workflow.&lt;/li&gt;
&lt;li&gt;Give each key a clear name.&lt;/li&gt;
&lt;li&gt;Configure the matching integration to use it.&lt;/li&gt;
&lt;li&gt;Review per-key usage at the invoice boundary.&lt;/li&gt;
&lt;li&gt;Revoke or rotate keys independently when necessary.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;That changes month-end reconciliation from log analysis into reporting. It also creates a useful operating dataset for profitability, anomaly detection, client conversations, and future estimates.&lt;/p&gt;

&lt;p&gt;The dashboard behavior and per-key workflow described here were checked against platform documentation in June 2026. Platform features evolve, so specific reporting and administrative capabilities should be confirmed against current documentation before making them part of a billing process.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/per-key-usage-tracking-how-agencies-attribute-ai-spend-to-individual-clients/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=per-key-usage-tracking-how-agencies-attribute-ai-spend-to-individual-clients"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>Consolidating AI API Keys Is a Migration, Not an Endless Cleanup Task</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Thu, 17 Sep 2026 15:34:35 +0000</pubDate>
      <link>https://dev.to/masonreed1/consolidating-ai-api-keys-is-a-migration-not-an-endless-cleanup-task-4gb8</link>
      <guid>https://dev.to/masonreed1/consolidating-ai-api-keys-is-a-migration-not-an-endless-cleanup-task-4gb8</guid>
      <description>&lt;p&gt;I would scope AI credential consolidation like any other migration: inventory the dependencies, prove one workload, move the rest, and retire the old path. “Reduce credential sprawl” is an indefinite backlog item. “Route these workloads through one endpoint and revoke these keys” has an acceptance test.&lt;/p&gt;

&lt;p&gt;The distinction matters because multi-provider setups rarely come from a deliberate architecture decision. OpenAI serves the first feature, Claude gets added for another, Gemini covers a specific task, and image or audio experiments bring more accounts. Each addition makes sense locally. Together, they leave a collection of secrets, dashboards, invoices, and configuration that somebody has to maintain.&lt;/p&gt;

&lt;p&gt;The useful question is not whether fewer keys would look tidier. It is whether a bounded migration can remove enough recurring work to justify another infrastructure dependency.&lt;/p&gt;

&lt;h2&gt;
  
  
  Define “Done” Before Picking a Gateway
&lt;/h2&gt;

&lt;p&gt;My acceptance criteria would be explicit: every model currently required is reachable through the consolidated endpoint, every migrated workload passes its checks, the base URL and key live in environment variables, and obsolete provider credentials are revoked.&lt;/p&gt;

&lt;p&gt;That makes consolidation a finishable project. It does &lt;strong&gt;not&lt;/strong&gt; mean authentication, billing, or integration maintenance disappears forever. The migration can end; operating the resulting dependency still takes work.&lt;/p&gt;

&lt;p&gt;The source’s estimate is a single sprint for most teams, sometimes a couple of focused days. I would treat that as a planning hypothesis until the inventory and first workload establish how much compatibility work is involved. Likewise, a claimed payback measured in weeks needs to be checked against actual maintenance costs, not assumed.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Operational concern&lt;/th&gt;
&lt;th&gt;Separate provider integrations&lt;/th&gt;
&lt;th&gt;Consolidated access&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Credentials&lt;/td&gt;
&lt;td&gt;A set for each provider&lt;/td&gt;
&lt;td&gt;One key for covered workloads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Spend visibility&lt;/td&gt;
&lt;td&gt;Multiple provider dashboards&lt;/td&gt;
&lt;td&gt;A consolidated dashboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Billing&lt;/td&gt;
&lt;td&gt;Separate accounts and invoices&lt;/td&gt;
&lt;td&gt;One billing relationship&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model changes&lt;/td&gt;
&lt;td&gt;May require account and integration setup&lt;/td&gt;
&lt;td&gt;Can be a model-name change for compatible calls&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Migration completion&lt;/td&gt;
&lt;td&gt;No consolidation target&lt;/td&gt;
&lt;td&gt;Workloads moved and old keys retired&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For this use case, a unified multi-model API such as CometAPI is relevant because it puts models from multiple providers behind a common access layer. The meaningful coverage number is not the catalog headline; it is how many models from your own inventory are actually supported.&lt;/p&gt;

&lt;h2&gt;
  
  
  Evaluate the Dependency Before Migrating
&lt;/h2&gt;

&lt;p&gt;I would address concentration risk before scheduling the rollout. One endpoint and one credential reduce the number of independent configurations, but they also concentrate access behind a shared dependency. Fewer places to misconfigure authentication does not automatically mean lower overall operational risk.&lt;/p&gt;

&lt;p&gt;OpenAI-compatible access helps make a migration reversible when the workload already uses compatible request and response formats. Moving back may be the same base-URL-and-key change in reverse. I would still verify that against the actual destination rather than treating “compatible” as a guarantee that every provider feature is interchangeable.&lt;/p&gt;

&lt;p&gt;The trade-off depends on the workload. Several providers supporting a mix of features make consolidation more attractive. A single-provider, single-model workload at ultra-high volume may have less to gain from an intermediary.&lt;/p&gt;

&lt;p&gt;That is the decision I want made deliberately: does the common access layer save more integration and operational effort than its dependency costs? Reversibility helps, but only when the return path has been checked.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run the Migration in Five Steps
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Inventory What Production Actually Calls
&lt;/h3&gt;

&lt;p&gt;List every provider, active key, consuming workload, and model. Include environment-specific configuration and experiments that still have live credentials.&lt;/p&gt;

&lt;p&gt;This is where forgotten accounts and lightly used integrations become visible. The inventory also prevents an incomplete migration from being declared finished because the main application moved while a background job still calls a provider directly.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Verify Coverage With the New Account
&lt;/h3&gt;

&lt;p&gt;Create the consolidated account, generate the key, and confirm that every required model is accessible through it.&lt;/p&gt;

&lt;p&gt;Do this before bulk configuration changes. A catalog entry is useful for discovery; your required calls are the migration scope. Any workload that cannot move needs an explicit exception, and that exception changes the claim that everything now runs through one key.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Prove One Low-Stakes Workload
&lt;/h3&gt;

&lt;p&gt;Choose a workload with limited consequences if something fails. Change its base URL and key, then run real requests through the complete application path.&lt;/p&gt;

&lt;p&gt;The goal is more than successful authentication. Confirm that the request works and that the response still satisfies the downstream code. This pilot establishes whether the remaining migration is configuration work or whether adapters need attention.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Move the Remaining Workloads
&lt;/h3&gt;

&lt;p&gt;For integrations with unchanged request and response formats, repeat the proven base-URL-and-key change. Downstream code should not need modification merely because the endpoint changed.&lt;/p&gt;

&lt;p&gt;Keep the endpoint and credential in environment variables so later routing changes remain configuration changes. Where a workload differs from the pilot, verify that difference instead of assuming the same pattern applies everywhere.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Retire the Old Access Paths
&lt;/h3&gt;

&lt;p&gt;After every intended workload runs through the new endpoint, revoke obsolete provider keys and close accounts that are no longer needed.&lt;/p&gt;

&lt;p&gt;This is part of the migration, not optional cleanup. Leaving the old credentials active preserves much of the secret-management burden you were trying to remove. The finish line includes retirement, not just successful traffic through the new service.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Expect to Improve
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Credential-related incidents:&lt;/strong&gt; Separate credentials create separate opportunities for expiration, misconfiguration, limits, and drift between environments. Consolidation reduces that configuration surface. The source reports fewer integration incidents among teams that consolidate, but provides no quantified incident data; I would measure the effect locally.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model evaluation speed:&lt;/strong&gt; For compatible models behind the same endpoint, switching can become a model-name configuration change rather than a new account and integration exercise. That removes setup friction from trying alternatives. It does not eliminate the need to evaluate their behavior.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Billing operations:&lt;/strong&gt; One account can replace multiple invoices, payment arrangements, balances, and spend dashboards. The described pay-as-you-go offering has no minimum and credits that do not expire. Those are account-specific terms to verify, not properties inherent to every unified API.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Day-to-day context switching:&lt;/strong&gt; One authentication pattern, documentation entry point, and dashboard reduce the need to remember which provider owns which operational detail. This is harder to quantify than invoice count, but it is a practical benefit when developers repeatedly move between workloads.&lt;/p&gt;

&lt;p&gt;I would put the inventory on the next sprint’s backlog first, with the pilot immediately behind it. That produces enough evidence to scope the rest honestly. A consolidation migration is complete when its acceptance criteria pass, not when someone promises there will never be another credential task.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/your-last-api-key-setup-consolidate-500-models-before-your-next-sprint/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=your-last-api-key-setup-consolidate-500-models-before-your-next-sprint"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Seedance 2.5 Costs: Separate Token Pricing From Your Production Budget</title>
      <dc:creator>Mason Reed</dc:creator>
      <pubDate>Thu, 17 Sep 2026 14:31:37 +0000</pubDate>
      <link>https://dev.to/masonreed1/seedance-25-costs-separate-token-pricing-from-your-production-budget-4lhn</link>
      <guid>https://dev.to/masonreed1/seedance-25-costs-separate-token-pricing-from-your-production-budget-4lhn</guid>
      <description>&lt;h2&gt;
  
  
  Two Billing Models, Two Different Estimates
&lt;/h2&gt;

&lt;p&gt;I would not put Seedance 2.5 into a budget spreadsheet until I had picked the API route. The published BytePlus token prices and the gateway's per-second prices are different billing models, not interchangeable ways to quote the same request.&lt;/p&gt;

&lt;p&gt;ByteDance Seed / BytePlus ModelArk lists the official model ID as &lt;code&gt;dreamina-seedance-2-5-260628&lt;/code&gt;. Its published pricing is &lt;strong&gt;$10.70 per 1 million tokens without video input&lt;/strong&gt; and &lt;strong&gt;$6.40 per 1 million tokens with video input&lt;/strong&gt;. The published output tiers are 480p and 720p.&lt;/p&gt;

&lt;p&gt;For five-second, 16:9 outputs, BytePlus gives these examples:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Input category&lt;/th&gt;
&lt;th&gt;480p&lt;/th&gt;
&lt;th&gt;720p&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Without video input&lt;/td&gt;
&lt;td&gt;$0.514&lt;/td&gt;
&lt;td&gt;$1.156&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;With video input&lt;/td&gt;
&lt;td&gt;$0.553–$2.152&lt;/td&gt;
&lt;td&gt;$1.244–$4.838&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The lower video-input token rate does not guarantee a cheaper request: input-video duration also contributes to estimated token consumption. BytePlus says it charges only for successfully generated videos. Published pricing and model information, however, do not by themselves confirm broad API access.&lt;/p&gt;

&lt;p&gt;The per-second figures below come from the source's gateway pricing snapshots: August 27, 2026 for Seedance duration calculations and August 28, 2026 for the cross-model comparison. I would treat them as dated budgeting inputs, not a guarantee of current pricing or account access.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Verify Before Integrating
&lt;/h2&gt;

&lt;p&gt;Seedance 2.5 is a multimodal video-generation model aimed at storytelling, reference-driven generation, and targeted editing. ByteDance describes support for up to &lt;strong&gt;30 seconds of audio-video output in one pass&lt;/strong&gt;, &lt;strong&gt;30 image references&lt;/strong&gt;, &lt;strong&gt;10 video references&lt;/strong&gt;, and &lt;strong&gt;10 audio references&lt;/strong&gt;, plus timestamp-level editing and multi-round extensions.&lt;/p&gt;

&lt;p&gt;Those model capabilities should not be confused with the fields exposed by a particular API route. The documented gateway route uses &lt;code&gt;seedance-2-5&lt;/code&gt;, supports text-to-video and image-to-video, and accepts durations from &lt;strong&gt;4–30 seconds&lt;/strong&gt;, defaulting to &lt;strong&gt;5 seconds&lt;/strong&gt;. Its listed output tiers are &lt;strong&gt;480p and 720p&lt;/strong&gt;; at 16:9, the exact sizes are &lt;code&gt;854x480&lt;/code&gt; and &lt;code&gt;1280x720&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;For unified multi-model access, CometAPI exposes Seedance 2.5, Vidu Q3, and Wan 3.0 through the same base URL and &lt;code&gt;POST /v1/videos&lt;/code&gt;, allowing shared authentication, billing, queueing, polling, and monitoring. That is useful for comparisons and fallbacks, but I would still verify each route's sizes, duration limits, input fields, and callback support before shipping.&lt;/p&gt;

&lt;p&gt;The asynchronous lifecycle matters as much as the submission endpoint: persist the returned &lt;code&gt;id&lt;/code&gt; or &lt;code&gt;task_id&lt;/code&gt;, poll &lt;code&gt;GET /v1/videos/{id}&lt;/code&gt; until a documented terminal state, and download completed media promptly.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Per-Second Budget
&lt;/h2&gt;

&lt;p&gt;For the documented Seedance gateway rates, the calculation is straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;generation_cost = requested_seconds * resolution_rate
effective_generation_cost_per_clip = total_generation_spend / accepted_clips
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The first equation estimates an attempt. The second tells me whether the workflow is affordable. Neither automatically includes storage, delivery bandwidth, moderation review, editing, or reruns.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Requested duration&lt;/th&gt;
&lt;th&gt;480p at $0.103/s&lt;/th&gt;
&lt;th&gt;720p at $0.231/s&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;4 seconds&lt;/td&gt;
&lt;td&gt;$0.41&lt;/td&gt;
&lt;td&gt;$0.92&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;5 seconds&lt;/td&gt;
&lt;td&gt;$0.52&lt;/td&gt;
&lt;td&gt;$1.16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;10 seconds&lt;/td&gt;
&lt;td&gt;$1.03&lt;/td&gt;
&lt;td&gt;$2.31&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;15 seconds&lt;/td&gt;
&lt;td&gt;$1.55&lt;/td&gt;
&lt;td&gt;$3.47&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;30 seconds&lt;/td&gt;
&lt;td&gt;$3.09&lt;/td&gt;
&lt;td&gt;$6.93&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At these rates, 720p costs about &lt;strong&gt;2.24 times&lt;/strong&gt; as much as 480p for the same duration. I would use 480p for early composition, motion, and reference tests, then increase resolution once those decisions are stable. The listed route does not publish 1080p or 4K pricing; neither belongs in a budget until the live model directory and API schema support it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Volume Changes the Rounding
&lt;/h3&gt;

&lt;p&gt;Using the underlying rates rather than rounded per-clip display prices, &lt;strong&gt;100 five-second generations&lt;/strong&gt; cost &lt;strong&gt;$51.50 at 480p&lt;/strong&gt; or &lt;strong&gt;$115.50 at 720p&lt;/strong&gt;. For &lt;strong&gt;1,000&lt;/strong&gt;, the baseline is &lt;strong&gt;$515&lt;/strong&gt; or &lt;strong&gt;$1,155&lt;/strong&gt;. These estimates assume every task completes once and every output is accepted.&lt;/p&gt;

&lt;p&gt;Regenerating 10% of the 100-clip 720p batch raises generation spend to &lt;strong&gt;$127.05&lt;/strong&gt;. Acceptance changes the result further: if only &lt;strong&gt;70 of 100&lt;/strong&gt; outputs pass review, &lt;strong&gt;$115.50 ÷ 70&lt;/strong&gt; is approximately &lt;strong&gt;$1.65 per accepted clip&lt;/strong&gt;, before review or editing costs. That is the number I would put beside a competing model's measured result.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compare Routes at the Same Resolution
&lt;/h2&gt;

&lt;p&gt;The comparison below uses the source's August 28, 2026 gateway snapshot. All listed rates are per generated second; five-second totals are rounded. “Not listed” means that exact resolution tier was not published for the route.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Resolution&lt;/th&gt;
&lt;th&gt;Model ID&lt;/th&gt;
&lt;th&gt;Rate per second&lt;/th&gt;
&lt;th&gt;Five-second cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;480p&lt;/td&gt;
&lt;td&gt;&lt;code&gt;wan3.0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.05&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;480p&lt;/td&gt;
&lt;td&gt;&lt;code&gt;seedance-2-5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.103&lt;/td&gt;
&lt;td&gt;$0.52&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;480p&lt;/td&gt;
&lt;td&gt;&lt;code&gt;viduq3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Not listed&lt;/td&gt;
&lt;td&gt;Not applicable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;720p&lt;/td&gt;
&lt;td&gt;&lt;code&gt;wan3.0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;720p&lt;/td&gt;
&lt;td&gt;&lt;code&gt;viduq3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.1232&lt;/td&gt;
&lt;td&gt;$0.62&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;720p&lt;/td&gt;
&lt;td&gt;&lt;code&gt;seedance-2-5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.231&lt;/td&gt;
&lt;td&gt;$1.16&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1080p&lt;/td&gt;
&lt;td&gt;&lt;code&gt;viduq3&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.1232&lt;/td&gt;
&lt;td&gt;$0.62&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1080p&lt;/td&gt;
&lt;td&gt;&lt;code&gt;wan3.0&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1080p&lt;/td&gt;
&lt;td&gt;&lt;code&gt;seedance-2-5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Not listed&lt;/td&gt;
&lt;td&gt;Not applicable&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Wan 3.0 has the lowest listed native 480p and 720p rates in this comparison. Vidu Q3's lowest published tier is &lt;strong&gt;540p at $0.056/s&lt;/strong&gt;, or &lt;strong&gt;$0.28 for five seconds&lt;/strong&gt;; I would not relabel that as a 480p offer. Its 720p tier includes synchronized audio, and its listed native 1080p rate is lower than Wan 3.0's.&lt;/p&gt;

&lt;p&gt;Seedance is not the cheapest 720p option by unit price. Its longer duration and reference controls may justify the difference for particular shots, especially where they reduce stitching or reruns, but that needs testing against accepted output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I Would Put Cost Controls
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Recover existing tasks before retrying submissions.&lt;/strong&gt; A polling timeout does not prove generation failed. The server may already have accepted the task, and resubmitting can create another billable job. Keep automatic retries bounded and retrieve the existing task whenever its ID is available.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Keep billing failures separate from quality failures.&lt;/strong&gt; BytePlus's successful-generation billing statement is not a universal policy for every route. I would store task ID, model ID, requested seconds, size, terminal status, reported usage or charge, and rejection reason. An HTTP success is not evidence of a usable video; a terminal failure needs its own billing record.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Validate short drafts before committing to long output.&lt;/strong&gt; My starting budget would use 480p at &lt;strong&gt;4–5 seconds&lt;/strong&gt;, promoting shots only after composition, motion, and reference handling pass review. A &lt;strong&gt;30-second 720p attempt costs $6.93&lt;/strong&gt;; repeating the whole clip because of an issue near the end gets expensive quickly. Higher resolution can also increase transfer, storage, review, and post-production costs without fixing an unstable prompt.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Benchmark the workflow, not just the rate card.&lt;/strong&gt; For a fixed prompt set, I would record first-pass acceptance, rerun rate, generation time, and editing time. Cheaper routes can handle concept tests; Seedance can be reserved for validated shots that benefit from its capabilities. The decision metric is total API spend divided by accepted clips, with reviewer and editing costs added when material.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/seedance-2-5-api-pricing/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=seedance-2-5-api-pricing"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
  </channel>
</rss>
