<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Ryan Cole</title>
    <description>The latest articles on DEV Community by Ryan Cole (@ryancole1).</description>
    <link>https://dev.to/ryancole1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4113545%2Fb1c19de8-a032-4f70-9bb5-dff18bb47f9e.png</url>
      <title>DEV Community: Ryan Cole</title>
      <link>https://dev.to/ryancole1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/ryancole1"/>
    <language>en</language>
    <item>
      <title>Running GLM-5.3-Flash Locally: Memory Budgets, Runtime Choices, and Working Commands</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Thu, 24 Sep 2026 08:33:20 +0000</pubDate>
      <link>https://dev.to/ryancole1/running-glm-53-flash-locally-memory-budgets-runtime-choices-and-working-commands-fao</link>
      <guid>https://dev.to/ryancole1/running-glm-53-flash-locally-memory-budgets-runtime-choices-and-working-commands-fao</guid>
      <description>&lt;p&gt;The first number I would check before deploying GLM-5.3-Flash is &lt;strong&gt;306 GiB&lt;/strong&gt;: the approximate size of its native FP8 weights. Its 18B active parameters per token describe compute usage, but the full model has about 320B parameters that still need somewhere to live.&lt;/p&gt;

&lt;p&gt;The weights are available under the MIT license. Local deployment is possible; the useful question is which combination of RAM, VRAM, quantization, and runtime makes sense for your workload.&lt;/p&gt;

&lt;p&gt;My starting point:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Server GPUs:&lt;/strong&gt; vLLM or SGLang.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Hundreds of gigabytes of RAM plus consumer GPUs:&lt;/strong&gt; KTransformers.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quantized workstation experiments:&lt;/strong&gt; llama.cpp or Ollama.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A 24 GB or 32 GB GPU cannot hold the model alone. Single-GPU deployment relies on system RAM, offloading, and possibly quantization.&lt;/p&gt;

&lt;h2&gt;
  
  
  Budget for the checkpoint and the requests
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://recipes.vllm.ai/zai-org/GLM-5.3-Flash" rel="noopener noreferrer"&gt;official vLLM recipe&lt;/a&gt; puts native FP8 weights at approximately 306 GiB. KTransformers recommends at least &lt;strong&gt;350 GB of available system memory&lt;/strong&gt; for its native FP8 CPU-GPU path.&lt;/p&gt;

&lt;p&gt;GGUF gives you more sizing options. These are approximate model-file sizes, not total runtime requirements:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Quantization&lt;/th&gt;
&lt;th&gt;Model size&lt;/th&gt;
&lt;th&gt;How I would budget&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;BF16&lt;/td&gt;
&lt;td&gt;642 GB&lt;/td&gt;
&lt;td&gt;Server-class memory&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q8_0&lt;/td&gt;
&lt;td&gt;341 GB&lt;/td&gt;
&lt;td&gt;Large-memory server or workstation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q6_K_XL&lt;/td&gt;
&lt;td&gt;292 GB&lt;/td&gt;
&lt;td&gt;High-memory workstation/server&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q5_K_XL&lt;/td&gt;
&lt;td&gt;240 GB&lt;/td&gt;
&lt;td&gt;256 GB RAM is likely too tight after overhead&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q4_K_XL&lt;/td&gt;
&lt;td&gt;200 GB&lt;/td&gt;
&lt;td&gt;Roughly 256 GB+ system-memory class&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IQ4_XS&lt;/td&gt;
&lt;td&gt;157 GB&lt;/td&gt;
&lt;td&gt;More realistically a 192–256 GB-class system&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q3_K_XL&lt;/td&gt;
&lt;td&gt;148 GB&lt;/td&gt;
&lt;td&gt;Large-memory workstation, with a growing quality trade-off&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Q2_K_XL&lt;/td&gt;
&lt;td&gt;109 GB&lt;/td&gt;
&lt;td&gt;128 GB leaves little room beyond the file&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IQ2_XXS&lt;/td&gt;
&lt;td&gt;102 GB&lt;/td&gt;
&lt;td&gt;Aggressive compression&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;IQ1_S&lt;/td&gt;
&lt;td&gt;93.1 GB&lt;/td&gt;
&lt;td&gt;Extreme compression; task-specific validation required&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Those planning notes are not official minimum specifications. Runtime buffers, metadata, multimodal components, KV cache, and the operating system all need memory. Context length, concurrency, GPU offload, and the quantization implementation change the actual fit.&lt;/p&gt;

&lt;p&gt;For a machine with 32–64 GB of system RAM, I would choose a smaller model or hosted access. At 128 GB, the smallest GGUF builds approach the available capacity, but predictable quality and long context are difficult targets. Around 256 GB or more, Q4_K_XL becomes a more practical starting point.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why 18B active parameters does not solve the memory problem
&lt;/h3&gt;

&lt;p&gt;GLM-5.3-Flash routes each token through only part of its expert capacity. Its hybrid linear and sparse attention, including IndexPool, also reduces long-context costs.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://autoclaw.z.ai/blog/model/glm-5.3-flash/" rel="noopener noreferrer"&gt;Z.ai reports lower attention compute and KV-cache usage than GLM-5.3&lt;/a&gt;. That helps, but cache still grows with context and concurrent requests. A deployment that works at 8K can run out of memory on a much longer conversation.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://huggingface.co/zai-org/GLM-5.3-Flash" rel="noopener noreferrer"&gt;official model card&lt;/a&gt; lists the following:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Property&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Architecture&lt;/td&gt;
&lt;td&gt;Native multimodal Mixture-of-Experts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total / active parameters&lt;/td&gt;
&lt;td&gt;320B / 18B per token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Language-model layers&lt;/td&gt;
&lt;td&gt;45&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Attention&lt;/td&gt;
&lt;td&gt;Hybrid linear + sparse attention with IndexPool&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Maximum context&lt;/td&gt;
&lt;td&gt;1,048,576 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training corpus&lt;/td&gt;
&lt;td&gt;30T multimodal tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Inputs / output&lt;/td&gt;
&lt;td&gt;Text, images, video, files / text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Weights and license&lt;/td&gt;
&lt;td&gt;Open weights, MIT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Official model ID&lt;/td&gt;
&lt;td&gt;&lt;code&gt;zai-org/GLM-5.3-Flash&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning effort&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;low&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;max&lt;/code&gt;; default &lt;code&gt;max&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I would configure context around the application’s actual input size. Allocating the full supported window during initial setup makes memory debugging harder.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with GGUF for workstation experiments
&lt;/h2&gt;

&lt;p&gt;For a local experiment, I would start here unless matching the native checkpoint is a requirement. Unsloth publishes GGUF builds ranging from 1-bit through BF16.&lt;/p&gt;

&lt;p&gt;llama.cpp offers CPU-GPU offloading; Ollama provides a short launch command. Both still need enough total memory for the selected build.&lt;/p&gt;

&lt;h3&gt;
  
  
  llama.cpp
&lt;/h3&gt;

&lt;p&gt;Install on macOS or Linux:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-LsSf&lt;/span&gt; https://llama.app/install.sh | sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Windows:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight powershell"&gt;&lt;code&gt;&lt;span class="n"&gt;winget&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;install&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;llama.cpp&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Start the Q4_K_XL server:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama serve &lt;span class="nt"&gt;-hf&lt;/span&gt; unsloth/GLM-5.3-Flash-GGUF:UD-Q4_K_XL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or run it directly in the terminal:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;llama cli &lt;span class="nt"&gt;-hf&lt;/span&gt; unsloth/GLM-5.3-Flash-GGUF:UD-Q4_K_XL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This build is about &lt;strong&gt;200 GB&lt;/strong&gt;. Before starting the download, check the quantization tag, repository file size, and available disk space.&lt;/p&gt;

&lt;h3&gt;
  
  
  Ollama
&lt;/h3&gt;

&lt;p&gt;Unsloth documents direct Hugging Face loading:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;ollama run hf.co/unsloth/GLM-5.3-Flash-GGUF:UD-Q4_K_XL
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The memory budget is the same underlying problem regardless of how short the command is.&lt;/p&gt;

&lt;p&gt;If Q4 does not fit, 3-bit, 2-bit, and 1-bit builds exist. I would select the highest-quality quantization that leaves room for runtime buffers and cache, then compare it against a native or hosted reference.&lt;/p&gt;

&lt;p&gt;Aggressive quantization can affect reasoning reliability, generated code, tool-call formatting, and multimodal behavior. A model that loads has passed only the capacity check.&lt;/p&gt;

&lt;h2&gt;
  
  
  Use KTransformers when the weights belong in system RAM
&lt;/h2&gt;

&lt;p&gt;KTransformers reads the official FP8 weights directly and distributes expert inference across CPU and GPU. This is the route I would investigate for a workstation with very large RAM capacity and one or more consumer GPUs.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://github.com/kvcache-ai/ktransformers/blob/main/doc/en/kt-kernel/GLM-5.3-Flash-Tutorial.md" rel="noopener noreferrer"&gt;GLM-5.3-Flash tutorial&lt;/a&gt; documents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Approximately 306 GiB of FP8 weights and a recommendation for 350 GB of available system memory.&lt;/li&gt;
&lt;li&gt;NVIDIA SM89 and SM120 support, including RTX 40- and 50-series GPUs.&lt;/li&gt;
&lt;li&gt;AVX-512 FP8 CPU expert kernels.&lt;/li&gt;
&lt;li&gt;Single-GPU and four-GPU configurations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;An RTX 4090 or RTX 5090 can participate, but most model state still sits outside VRAM. CPU capability, RAM bandwidth, and memory placement matter considerably.&lt;/p&gt;

&lt;h3&gt;
  
  
  Install and prepare the checkpoint
&lt;/h3&gt;

&lt;p&gt;Create a Python 3.11 environment and install the SGLang integration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;conda create &lt;span class="nt"&gt;-n&lt;/span&gt; glm53flash &lt;span class="nv"&gt;python&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;3.11 &lt;span class="nt"&gt;-y&lt;/span&gt;
conda activate glm53flash

pip &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="s2"&gt;"ktransformers[sglang]"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Download &lt;a href="https://huggingface.co/zai-org/GLM-5.3-Flash" rel="noopener noreferrer"&gt;&lt;code&gt;zai-org/GLM-5.3-Flash&lt;/code&gt;&lt;/a&gt; to local storage. Allow sufficient disk space for the checkpoint and available RAM for the server configuration.&lt;/p&gt;

&lt;h3&gt;
  
  
  Launch the documented single-GPU configuration
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;MODEL_PATH&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/path/to/GLM-5.3-Flash

&lt;span class="nv"&gt;CUDA_VISIBLE_DEVICES&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;0 python &lt;span class="nt"&gt;-m&lt;/span&gt; sglang.launch_server &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--model-path&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MODEL_PATH&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--kt-weight-path&lt;/span&gt; &lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MODEL_PATH&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--served-model-name&lt;/span&gt; GLM-5.3-flash &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--tp-size&lt;/span&gt; 1 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--context-length&lt;/span&gt; 501025 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--mem-fraction-static&lt;/span&gt; 0.65 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--chunked-prefill-size&lt;/span&gt; 2048 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--kt-method&lt;/span&gt; FP8 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--kt-cpuinfer&lt;/span&gt; 64 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--kt-threadpool-count&lt;/span&gt; 2 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--kt-num-gpu-experts&lt;/span&gt; 0 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--kt-gpu-prefill-token-threshold&lt;/span&gt; 2048 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--cuda-graph-bs&lt;/span&gt; 1 2 4 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--limit-mm-data-per-request&lt;/span&gt; &lt;span class="s1"&gt;'{"image":8,"video":1}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--mm-process-config&lt;/span&gt; &lt;span class="s1"&gt;'{"image":{"max_pixels":1254400}}'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--tool-call-parser&lt;/span&gt; glm47 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--reasoning-parser&lt;/span&gt; glm45
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The tutorial validates this &lt;strong&gt;501,025-token&lt;/strong&gt; configuration. The model’s 1,048,576-token limit does not mean every deployment should allocate that much context.&lt;/p&gt;

&lt;p&gt;Check discovery:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:30000/v1/models
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The OpenAI-compatible chat endpoint is &lt;code&gt;http://localhost:30000/v1/chat/completions&lt;/code&gt;. This launch command serves the model as &lt;code&gt;GLM-5.3-flash&lt;/code&gt;, so use that identifier for chat requests.&lt;/p&gt;

&lt;p&gt;If generation is extremely slow, investigate NUMA placement, RAM bandwidth, CPU instruction support, and storage behavior during loading. With substantial expert work on the CPU, memory capacity alone does not guarantee usable interactive performance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Serve native weights on GPU infrastructure
&lt;/h2&gt;

&lt;p&gt;For production serving, I would start with vLLM when throughput and ecosystem compatibility dominate. SGLang is also worth testing for agent workloads, structured generation, multimodal requests, and concurrency.&lt;/p&gt;

&lt;p&gt;Both serve the native checkpoint and support strong multi-GPU scaling. Their focus on CPU offload is more limited than KTransformers or llama.cpp.&lt;/p&gt;

&lt;h3&gt;
  
  
  vLLM
&lt;/h3&gt;

&lt;p&gt;Use Linux, a supported NVIDIA stack, sufficient aggregate GPU memory, and a recent supported vLLM build or the container specified in the current recipe.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;vllm

vllm serve &lt;span class="s2"&gt;"zai-org/GLM-5.3-Flash"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--tensor-parallel-size&lt;/span&gt; 8 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--served-model-name&lt;/span&gt; zai-org/GLM-5.3-Flash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a reference launch configuration. Eight GPUs do not automatically imply enough usable memory or support for every feature.&lt;/p&gt;

&lt;p&gt;Verify the endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:8000/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "zai-org/GLM-5.3-Flash",
    "messages": [
      {"role": "user", "content": "Reply with OK"}
    ]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://recipes.vllm.ai/zai-org/GLM-5.3-Flash" rel="noopener noreferrer"&gt;vLLM recipe&lt;/a&gt; also covers FP8 KV cache on supported Blackwell systems, MTP speculative decoding, tool and reasoning parsers, and prefill/decode disaggregation. I would take those flags from the current recipe because support changes quickly.&lt;/p&gt;

&lt;h3&gt;
  
  
  SGLang
&lt;/h3&gt;

&lt;p&gt;Install:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;pip &lt;span class="nb"&gt;install &lt;/span&gt;sglang
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Launch:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;python3 &lt;span class="nt"&gt;-m&lt;/span&gt; sglang.launch_server &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--model-path&lt;/span&gt; &lt;span class="s2"&gt;"zai-org/GLM-5.3-Flash"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--host&lt;/span&gt; 0.0.0.0 &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--port&lt;/span&gt; 30000
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Verify:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"http://localhost:30000/v1/chat/completions"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--data&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "zai-org/GLM-5.3-Flash",
    "messages": [{"role": "user", "content": "Give me three local deployment checks."}]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The official model card includes multimodal SGLang examples. For tool calling, check the parser configuration for the current integration. Parser flags from an older GLM release are not a reliable template.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the application contract before tuning throughput
&lt;/h2&gt;

&lt;p&gt;I would validate a deployment in this order: text generation, intended context length, tool schemas, multimodal inputs, and concurrent traffic.&lt;/p&gt;

&lt;p&gt;For the vLLM endpoint above:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl http://localhost:8000/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "zai-org/GLM-5.3-Flash",
    "messages": [
      {"role": "user", "content": "Return exactly: LOCAL_OK"}
    ],
    "reasoning_effort": "low"
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Adjust the port and served model name for your runtime.&lt;/p&gt;

&lt;p&gt;The model defaults to &lt;code&gt;reasoning_effort: "max"&lt;/code&gt;. Keep &lt;code&gt;max&lt;/code&gt; when reproducing benchmarks. For workstation iteration, &lt;code&gt;low&lt;/code&gt; or &lt;code&gt;high&lt;/code&gt; can be more practical.&lt;/p&gt;

&lt;p&gt;My acceptance checks would include:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Check&lt;/th&gt;
&lt;th&gt;What to exercise&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Context&lt;/td&gt;
&lt;td&gt;A document or repository-sized prompt near the application’s intended limit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tools&lt;/td&gt;
&lt;td&gt;Argument JSON, tool selection, repeated calls, recovery from tool errors&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multimodal&lt;/td&gt;
&lt;td&gt;Actual image/video formats and resolution ranges&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Concurrency&lt;/td&gt;
&lt;td&gt;Latency and memory with multiple active requests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quantization&lt;/td&gt;
&lt;td&gt;Identical prompts against the selected GGUF and a native or hosted reference&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h3&gt;
  
  
  Separate memory tuning from latency tuning
&lt;/h3&gt;

&lt;p&gt;To reduce memory pressure, lower configured context and concurrency first. A small concurrency target with an explicit queue can suit a workstation better than server-style parallelism.&lt;/p&gt;

&lt;p&gt;Quantization reduces weight memory. CPU offload moves substantial model state into RAM, shifting performance pressure toward the CPU and memory bandwidth.&lt;/p&gt;

&lt;p&gt;Lowering &lt;code&gt;reasoning_effort&lt;/code&gt; can shorten generated reasoning, reduce latency, and consume fewer tokens. It &lt;strong&gt;does not reduce the memory needed to load the weights&lt;/strong&gt;. Any cache savings are indirect, from generating a shorter sequence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Diagnose failures by when they happen
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Symptom&lt;/th&gt;
&lt;th&gt;First checks&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Loads, then crashes on a long prompt&lt;/td&gt;
&lt;td&gt;KV-cache headroom; reduce context and concurrency, then increase gradually&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GGUF fits on disk but fails in RAM&lt;/td&gt;
&lt;td&gt;Runtime buffers, cache, operating-system headroom&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single-GPU KTransformers is too slow&lt;/td&gt;
&lt;td&gt;NUMA, RAM bandwidth, CPU support, CPU expert workload&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tool calls contain malformed JSON&lt;/td&gt;
&lt;td&gt;Current runtime parser configuration&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Download is hundreds of gigabytes&lt;/td&gt;
&lt;td&gt;Expected file size and selected quantization tag&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I would monitor both system RAM and GPU memory while increasing workload size. Otherwise, it is easy to mistake a successful checkpoint load for a viable serving configuration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decide whether the capability justifies the deployment
&lt;/h2&gt;

&lt;p&gt;The benchmark case for this model is mainly coding, tool-driven automation, long-context document work, and multimodal workflows.&lt;/p&gt;

&lt;p&gt;These are &lt;a href="https://autoclaw.z.ai/blog/model/glm-5.3-flash/" rel="noopener noreferrer"&gt;Z.ai’s reported scores&lt;/a&gt;, not measurements from a local quantized deployment:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;GLM-5.3-Flash&lt;/th&gt;
&lt;th&gt;GLM-5.2&lt;/th&gt;
&lt;th&gt;Difference&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench 2.1&lt;/td&gt;
&lt;td&gt;84.3&lt;/td&gt;
&lt;td&gt;81.0&lt;/td&gt;
&lt;td&gt;+3.3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE v1.1&lt;/td&gt;
&lt;td&gt;63.4&lt;/td&gt;
&lt;td&gt;46.2&lt;/td&gt;
&lt;td&gt;+17.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NL2Repo&lt;/td&gt;
&lt;td&gt;56.3&lt;/td&gt;
&lt;td&gt;48.9&lt;/td&gt;
&lt;td&gt;+7.4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Toolathlon Verified&lt;/td&gt;
&lt;td&gt;78.4&lt;/td&gt;
&lt;td&gt;59.9&lt;/td&gt;
&lt;td&gt;+18.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AutomationBench v1.0.6&lt;/td&gt;
&lt;td&gt;48.8&lt;/td&gt;
&lt;td&gt;26.2&lt;/td&gt;
&lt;td&gt;+22.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agents' Last Exam&lt;/td&gt;
&lt;td&gt;26.3&lt;/td&gt;
&lt;td&gt;20.4&lt;/td&gt;
&lt;td&gt;+5.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HLE with Tools&lt;/td&gt;
&lt;td&gt;55.3&lt;/td&gt;
&lt;td&gt;54.7&lt;/td&gt;
&lt;td&gt;+0.6&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GDPval-AA v2&lt;/td&gt;
&lt;td&gt;1773&lt;/td&gt;
&lt;td&gt;1504&lt;/td&gt;
&lt;td&gt;+269 Elo&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I would use those results to decide what to evaluate locally, then let workload-specific tests determine whether the hardware and operating costs are justified.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where hosted access fits
&lt;/h3&gt;

&lt;p&gt;Self-hosting gives control over data residency, offline operation, quantization, and inference settings. It also makes hardware capacity, maintenance, and scaling your responsibility.&lt;/p&gt;

&lt;p&gt;Hosted access removes the upfront hardware requirement and provider-side maintenance, but sends data to the selected service, uses provider-selected quantization, and scales within provider limits.&lt;/p&gt;

&lt;p&gt;If I needed a unified multi-model API for reference comparisons or a fallback during deployment work, CometAPI provides an OpenAI-compatible endpoint using model ID &lt;code&gt;glm-5.3-flash&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;glm-5.3-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Reply with OK&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Local hosting can make financial sense with suitable owned hardware, consistently high utilization, or a firm requirement to keep data inside your infrastructure. For intermittent workloads, I would weigh the fixed hardware and operational burden carefully.&lt;/p&gt;

&lt;p&gt;My deployment gate would be concrete: the chosen checkpoint fits with headroom, the runtime handles the required tools and inputs, and latency stays acceptable at the intended context and concurrency. Until those checks pass, a running server is still an experiment.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-run-glm-5-3-flash-locally/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-run-glm-5-3-flash-locally"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Building Short Videos with Sora 2’s Native Audio</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Tue, 22 Sep 2026 04:51:43 +0000</pubDate>
      <link>https://dev.to/ryancole1/building-short-videos-with-sora-2s-native-audio-2n92</link>
      <guid>https://dev.to/ryancole1/building-short-videos-with-sora-2s-native-audio-2n92</guid>
      <description>&lt;p&gt;Sora 2 is OpenAI’s second-generation text-to-video model, but its most useful change is not purely visual. Audio is generated as part of the video workflow: dialogue, ambience, music cues, and sound effects can be described alongside the scene and synchronized with what happens on screen.&lt;/p&gt;

&lt;p&gt;That matters because the traditional workflow is split across several tools: generate the video, record or synthesize dialogue, find sound effects, add ambience, then manually align everything in an editor. Sora 2 attempts to produce those layers together from the first render.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Sora 2 generates
&lt;/h2&gt;

&lt;p&gt;Sora 2 can produce several practical audio layers:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Synchronized dialogue&lt;/strong&gt; aligned with character lip movement and timing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sound effects&lt;/strong&gt; connected to visible events such as footsteps, impacts, doors, and object movement.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Environmental audio&lt;/strong&gt; including rain, wind, room tone, crowds, and distant traffic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Music cues&lt;/strong&gt;, such as short stings or background loops. Licensing and style restrictions may still apply.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A basic mix&lt;/strong&gt; combining the generated elements. For detailed balancing or mastering, export stems when the workflow supports it and finish in a DAW.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The important distinction is that audio is not necessarily a post-process. The model simulates the sound and image together, which can improve synchronization between speech, physical actions, and the surrounding environment.&lt;/p&gt;

&lt;h2&gt;
  
  
  The audio capabilities worth testing
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Dialogue and lip-sync in the generation pass
&lt;/h3&gt;

&lt;p&gt;Sora 2 can generate speech that matches a generated face or animated mouth. This is different from taking a finished video and running a separate lip-sync process over it: the timing and prosody are part of the generation step.&lt;/p&gt;

&lt;p&gt;That makes short dialogue-driven content practical without recording actors. Product micro-ads, instructional clips, social cameos, and quick narrative prototypes are obvious use cases.&lt;/p&gt;

&lt;h3&gt;
  
  
  Sound effects that follow the scene
&lt;/h3&gt;

&lt;p&gt;The model can associate sound with visible physical events:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A cup clinks when it hits a table.&lt;/li&gt;
&lt;li&gt;Footsteps reflect the apparent environment.&lt;/li&gt;
&lt;li&gt;A door creaks or slams at the moment it moves.&lt;/li&gt;
&lt;li&gt;An object impact produces a corresponding sound.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those details carry a surprising amount of perceived realism. Room tone gives a location scale, while a well-timed impact or sudden thud supplies an emotional cue.&lt;/p&gt;

&lt;h3&gt;
  
  
  Audio continuity across shots
&lt;/h3&gt;

&lt;p&gt;For sequences made from multiple shots or stitched clips, Sora 2 attempts to maintain characteristics such as:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Similar room reverb&lt;/li&gt;
&lt;li&gt;Consistent voice timbre for recurring characters&lt;/li&gt;
&lt;li&gt;Matching environmental noise&lt;/li&gt;
&lt;li&gt;More coherent audio across cuts&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It is not a replacement for a sound editor, but it can reduce the amount of manual EQ and room-tone matching needed during early iterations.&lt;/p&gt;

&lt;h2&gt;
  
  
  Accessing Sora 2
&lt;/h2&gt;

&lt;p&gt;There are two main access paths:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Sora app or web app&lt;/strong&gt; — OpenAI announced Sora 2 with an app for creating videos without code. Availability is staged by region, app store, and access window. Recent wider-access periods have included the US, Canada, Japan, and South Korea, with quotas and other limitations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI Video API&lt;/strong&gt; — The video API exposes &lt;code&gt;sora-2&lt;/code&gt; and &lt;code&gt;sora-2-pro&lt;/code&gt;. Requests can include a prompt, duration in seconds, output size, and input references. &lt;code&gt;sora-2&lt;/code&gt; is positioned for faster iteration, while &lt;code&gt;sora-2-pro&lt;/code&gt; targets higher fidelity and more complex scenes.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A unified multi-model gateway such as &lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt; can expose the same Sora 2 API call style and endpoints, with pricing listed below.&lt;/p&gt;

&lt;h3&gt;
  
  
  Minimal API request
&lt;/h3&gt;

&lt;p&gt;The &lt;code&gt;v1/videos&lt;/code&gt; endpoint accepts &lt;code&gt;model=sora-2&lt;/code&gt; or &lt;code&gt;sora-2-pro&lt;/code&gt;. This example asks for dialogue, applause, and a sustained piano note:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.cometapi.com/v1/videos &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OPENAI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-F&lt;/span&gt; &lt;span class="s2"&gt;"model=sora-2"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-F&lt;/span&gt; &lt;span class="s2"&gt;"prompt=A calico cat playing a piano on stage. Audio: single speaker narrator says 'At last, the show begins'. Add applause and piano sustain after the final chord."&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-F&lt;/span&gt; &lt;span class="s2"&gt;"seconds=8"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-F&lt;/span&gt; &lt;span class="s2"&gt;"size=1280x720"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The request creates a video job. Once processing finishes, the result is an MP4 with audio baked into it; the API returns a job ID and, when ready, a download URL.&lt;/p&gt;

&lt;h3&gt;
  
  
  Listed API pricing
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Sora-2&lt;/td&gt;
&lt;td&gt;$0.08 per second&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sora-2-pro&lt;/td&gt;
&lt;td&gt;$0.24 per second&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  A practical prompt-to-video workflow
&lt;/h2&gt;

&lt;p&gt;I generally treat the audio plan as part of the shot plan rather than adding it after writing the visual prompt.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Define the scene, characters, dialogue, mood, and whether the sound is diegetic or musical.&lt;/li&gt;
&lt;li&gt;Describe the audio explicitly: speakers, delivery, pacing, effects, and ambience.&lt;/li&gt;
&lt;li&gt;Start with a short render, usually 4–8 seconds for rapid iteration.&lt;/li&gt;
&lt;li&gt;Check speech timing, lip-sync, background noise, and event-based effects.&lt;/li&gt;
&lt;li&gt;Export the mixed result or stems, if available, and finish the mix externally when exact control matters.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Sora 2 is aimed at short cinematic clips. Longer sequences can be created through multi-shot or stitching workflows, but they typically require more iteration.&lt;/p&gt;

&lt;h2&gt;
  
  
  One-step audio versus a separate narration asset
&lt;/h2&gt;

&lt;p&gt;Use the video endpoint when the goal is a single prompt that produces video and audio together.&lt;/p&gt;

&lt;p&gt;A separate speech asset makes more sense when:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The narrator must remain consistent across many videos.&lt;/li&gt;
&lt;li&gt;You need to audition different voices.&lt;/li&gt;
&lt;li&gt;Voice timbre and prosody require tighter control.&lt;/li&gt;
&lt;li&gt;The same narration will be reused in multiple edits.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The speech endpoint is &lt;code&gt;/v1/audio/speech&lt;/code&gt;. You can generate an MP3, import it into Final Cut or Premiere, and replace or layer the generated audio. Where supported, the audio can also be supplied as an input reference for a remix workflow.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.openai.com/v1/audio/speech &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OPENAI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "gpt-speech-1",
    "voice": "alloy",
    "input": "Welcome to our product demo. Today we show fast AI video generation."
  }'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--output&lt;/span&gt; narration.mp3
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The primary Sora 2 workflow already generates audio. Separate speech generation is for voice consistency, external reuse, or more controlled production pipelines.&lt;/p&gt;

&lt;h2&gt;
  
  
  Node.js example
&lt;/h2&gt;

&lt;p&gt;The official SDK can create the video job directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;OpenAI&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;openai&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;apiKey&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;OPENAI_API_KEY&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;video&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;openai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;videos&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;sora-2&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`
A friendly robot waters plants on a balcony at sunrise.
Audio: soft morning birds, one speaker voiceover says
"Good morning, little world." Include distant city ambience.
Style: gentle, warm.
  `&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;seconds&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;8&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
  &lt;span class="na"&gt;size&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;1280x720&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="c1"&gt;// Poll the job and download the result after completion.&lt;/span&gt;
&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Video job created:&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;video&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The SDK call creates the job; polling and downloading the completed result follow the video API’s job lifecycle.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prompt structure that works better
&lt;/h2&gt;

&lt;p&gt;I get more predictable results when the prompt has a clear visual section followed by an explicit audio section:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Visual scene and action.

Audio:
- Number of speakers
- Voice, tone, and pacing
- Dialogue
- Sound effects and timing
- Environmental ambience
- Optional mix perspective
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A rainy evening on a narrow city alley. A woman in a red coat hurries across wet cobblestones toward a flickering neon sign.

Audio: Two speakers. Speaker A, the woman, breathes slightly and sounds hurried. Speaker B, an offscreen street vendor, calls out once. Add steady rain on a roof, a distant car, and the clatter of an empty can when she kicks it.

Dialogue:
Speaker A: "I'm late. I can't believe I missed it."
Speaker B, muffled, one line: "You better run!"

Style: cinematic, shallow depth of field, close-up when she speaks. Sync dialogue to lip movement and use naturalistic reverb.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Putting the audio instructions after the visual description helps bind sounds to the events already established in the scene. For timing-sensitive clips, include explicit timecodes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;[SFX: door_close @00:01]
[SFX: impact @00:04]
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;I also separate camera instructions from audio instructions instead of mixing everything into one sentence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Remixing and targeted edits
&lt;/h2&gt;

&lt;p&gt;Sora 2 supports remix-style workflows for changes such as extending a scene or replacing its background. Audio changes should be stated in the remix request too:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Replace the music with sparse piano. Keep the dialogue identical,
but move the second line to 2.5 seconds. Preserve the room ambience.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is useful when the visual result is close and rebuilding the entire shot would be wasteful.&lt;/p&gt;

&lt;h2&gt;
  
  
  Troubleshooting audio problems
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Lip-sync is drifting
&lt;/h3&gt;

&lt;p&gt;Make the dialogue timing more explicit and simplify competing background noise. Strong ambience can mask speech or affect perceived timing.&lt;/p&gt;

&lt;h3&gt;
  
  
  The voice sounds muffled or too reverberant
&lt;/h3&gt;

&lt;p&gt;Specify the acoustic target:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Dry voice, minimal reverb, close microphone perspective.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or, for a location-based sound:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Natural room reverb, voice slightly distant from the camera.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Effects are too loud or buried
&lt;/h3&gt;

&lt;p&gt;Use relative balance instructions:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Soft door close. Dialogue should be 3 dB louder than ambience.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  The generation contains unwanted artifacts
&lt;/h3&gt;

&lt;p&gt;Regenerate with slightly different wording. Alternate phrasing can produce cleaner audio even when the underlying scene description remains the same.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three prompt patterns
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Product reveal: 7–12 seconds
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;7s, studio product shot: small espresso machine on a counter.
Visual: slow 3/4 pan in.
Dialogue: "Perfect crema, every time."
Voice: confident, friendly, male, medium tempo.
SFX: steam release at 0:04, small metallic click at 0:06.
Ambient: low cafe murmur.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A short spoken hook paired with a recognizable product sound gives the clip an immediate sensory identity. A brand jingle can still be added in post.&lt;/p&gt;

&lt;h3&gt;
  
  
  Instructional clip: 10 seconds
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;10s, overhead kitchen shot.
Visual: hands sprinkle salt into a bowl, then whisk.
Audio: step narration, female and calm:
"One pinch of sea salt."
SFX: salt sprinkle at the start, whisking texture under narration.
Ambient: quiet kitchen.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This combines instructional narration with diegetic sounds that reinforce each action.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tension beat: 6 seconds
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;6s, alleyway at dusk.
Visual: quick low-angle shot of a bicyclist's tire skidding.
Audio: sudden metallic screech at 00:02 synced to the skid,
heartbeat-like low bass underlay, distant thunder.
No dialogue.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Short tension scenes depend on precise effects and low-frequency cues. Keeping the sound plan compact makes the timing easier to evaluate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where I would not use Sora 2 alone
&lt;/h2&gt;

&lt;p&gt;Long-form narrative work with complex dialogue, many scenes, and detailed mixes still benefits from human performers and dedicated sound design.&lt;/p&gt;

&lt;p&gt;I also would not treat synthetic media as a substitute for authenticated recordings in legal proceedings, evidence workflows, or other strict compliance contexts.&lt;/p&gt;

&lt;p&gt;For ordinary short-form production, though, native audio changes the iteration loop. I can describe the action, the sound, and their relationship in one request, then spend time refining the creative result instead of manually assembling every layer from scratch.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-create-video-using-sora-2s-audio-tool/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-create-video-using-sora-2s-audio-tool"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Building Short Midjourney V1 Videos with an Image-to-Video API</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Tue, 22 Sep 2026 02:00:39 +0000</pubDate>
      <link>https://dev.to/ryancole1/building-short-midjourney-v1-videos-with-an-image-to-video-api-5gme</link>
      <guid>https://dev.to/ryancole1/building-short-midjourney-v1-videos-with-an-image-to-video-api-5gme</guid>
      <description>&lt;p&gt;Midjourney’s V1 video model is designed around a straightforward workflow: start with one still image and turn it into a short animated clip. The source image can come from Midjourney or from an externally hosted URL.&lt;/p&gt;

&lt;p&gt;The default output is roughly five seconds long. Clips can be extended in four-second increments to approximately 21 seconds, and the result is delivered as an MP4. This makes V1 useful for animated stills, product loops, stylized social content, and short visual experiments rather than long-form cinematic sequences.&lt;/p&gt;

&lt;h2&gt;
  
  
  What V1 Actually Does
&lt;/h2&gt;

&lt;p&gt;Midjourney V1 is an image-to-video model. It prioritizes preserving the source image’s visual identity, including its color palette, brushwork, subject, and overall mood.&lt;/p&gt;

&lt;p&gt;The main characteristics are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Base clip length of approximately five seconds.&lt;/li&gt;
&lt;li&gt;Extensions in four-second increments, up to a documented limit of roughly 21 seconds.&lt;/li&gt;
&lt;li&gt;Automatic or manual animation modes.&lt;/li&gt;
&lt;li&gt;Low- and high-motion controls through &lt;code&gt;--motion low&lt;/code&gt; and &lt;code&gt;--motion high&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Configurable batch size, looping, and end frames.&lt;/li&gt;
&lt;li&gt;MP4 output.&lt;/li&gt;
&lt;li&gt;Primarily SD output at 480p, with HD output at 720p available through the appropriate model or video-type parameter.&lt;/li&gt;
&lt;li&gt;A quality and resolution tradeoff aimed at quick iteration, web content, and social media rather than full cinematic production.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical implication is that the starting image matters more than it would in a text-to-video workflow. A clear subject, deliberate composition, and usable aspect ratio give the model less ambiguity to resolve.&lt;/p&gt;

&lt;p&gt;For programmatic access, I use CometAPI’s unified REST interface, which exposes the Midjourney V1 video capability through &lt;code&gt;/mj/submit/video&lt;/code&gt;. The request accepts an image URL in &lt;code&gt;prompt&lt;/code&gt;, a &lt;code&gt;videoType&lt;/code&gt; such as &lt;code&gt;vid_1.1_i2v_480&lt;/code&gt;, a generation &lt;code&gt;mode&lt;/code&gt;, and an &lt;code&gt;animateMode&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decide the Parameters Before Submitting
&lt;/h2&gt;

&lt;p&gt;There are a few decisions worth making before writing the request.&lt;/p&gt;

&lt;h3&gt;
  
  
  Starting image
&lt;/h3&gt;

&lt;p&gt;External images need to be available at publicly reachable URLs. Choose an image with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A clear primary subject.&lt;/li&gt;
&lt;li&gt;Enough visual separation between foreground and background.&lt;/li&gt;
&lt;li&gt;A composition that works at the intended output aspect ratio.&lt;/li&gt;
&lt;li&gt;A subject whose expected movement is physically plausible.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Starting aspect ratio affects the final video dimensions and whether the result maps cleanly to SD or HD output.&lt;/p&gt;

&lt;h3&gt;
  
  
  Motion
&lt;/h3&gt;

&lt;p&gt;Use low motion for subtle camera movement, restrained subject animation, or loop-friendly clips. High motion is more appropriate when the subject needs to travel or the scene contains more visible movement.&lt;/p&gt;

&lt;p&gt;You can either let the model infer motion automatically or describe it explicitly in manual mode.&lt;/p&gt;

&lt;h3&gt;
  
  
  Duration and batch size
&lt;/h3&gt;

&lt;p&gt;The default clip is approximately five seconds. Extensions add four seconds at a time, up to roughly 21 seconds.&lt;/p&gt;

&lt;p&gt;The default batch size is four variants. For production jobs or cost-sensitive iteration, request one or two variants instead. A batch size of one is represented by &lt;code&gt;bs: 1&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Resolution
&lt;/h3&gt;

&lt;p&gt;V1 is primarily oriented around 480p SD output. HD output is 720p and requires the relevant documented video type or parameter. The examples below use &lt;code&gt;vid_1.1_i2v_480&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Submit a Video Job
&lt;/h2&gt;

&lt;p&gt;The smallest useful request includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;prompt&lt;/code&gt;: an image URL, optionally followed by a motion description.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;videoType&lt;/code&gt;: for example, &lt;code&gt;vid_1.1_i2v_480&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;mode&lt;/code&gt;: usually &lt;code&gt;"fast"&lt;/code&gt;, or &lt;code&gt;"relax"&lt;/code&gt; when supported by the plan.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;animateMode&lt;/code&gt;: &lt;code&gt;"automatic"&lt;/code&gt; or &lt;code&gt;"manual"&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Here is a complete request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;--location&lt;/span&gt; &lt;span class="nt"&gt;--request&lt;/span&gt; POST &lt;span class="s1"&gt;'https://api.cometapi.com/mj/submit/video'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s1"&gt;'Authorization: Bearer sk-YOUR_COMETAPI_KEY'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s1"&gt;'Content-Type: application/json'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--data-raw&lt;/span&gt; &lt;span class="s1"&gt;'{
    "prompt": "https://cdn.midjourney.com/example/0_0.png A peaceful seaside scene — camera slowly zooms out and a gull flies by",
    "videoType": "vid_1.1_i2v_480",
    "mode": "fast",
    "animateMode": "manual",
    "motion": "low",
    "bs": 1
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The image URL and motion instruction are combined in the &lt;code&gt;prompt&lt;/code&gt; field. In this example, manual animation is paired with low motion, a single output variant, and a slow zoom.&lt;/p&gt;

&lt;h2&gt;
  
  
  Submit and Poll from Python
&lt;/h2&gt;

&lt;p&gt;The API is asynchronous, so the basic integration pattern is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Submit the generation request.&lt;/li&gt;
&lt;li&gt;Extract the returned job identifier.&lt;/li&gt;
&lt;li&gt;Poll the status endpoint.&lt;/li&gt;
&lt;li&gt;Download the video once the job reaches &lt;code&gt;completed&lt;/code&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A minimal &lt;code&gt;requests&lt;/code&gt; implementation looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;

&lt;span class="n"&gt;API_KEY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;sk-YOUR_COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;BASE&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="n"&gt;HEADERS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;API_KEY&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;application/json&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;payload&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://cdn.midjourney.com/example/0_0.png A calm city street — camera pans left, rain falling&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;videoType&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;vid_1.1_i2v_480&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;mode&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;fast&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;animateMode&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;manual&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;motion&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;low&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;bs&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="c1"&gt;# Submit job
&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;BASE&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/mj/submit/video&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;payload&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;HEADERS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;job&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;job_id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;job_id&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Poll for completion (example polling)
&lt;/span&gt;&lt;span class="n"&gt;status_url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;BASE&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;/mj/status/&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;job_id&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;_&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;  &lt;span class="c1"&gt;# poll up to ~60 times
&lt;/span&gt;    &lt;span class="n"&gt;s&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;status_url&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;HEADERS&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="n"&gt;st&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;s&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;completed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;download_url&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;result&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;video_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Video ready:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;download_url&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;break&lt;/span&gt;
    &lt;span class="k"&gt;elif&lt;/span&gt; &lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;status&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;failed&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;error&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;raise&lt;/span&gt; &lt;span class="nc"&gt;RuntimeError&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Video generation failed: &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;st&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
    &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a production worker, I would move the polling loop into a job system, add request timeouts and retry handling, and persist the provider job ID rather than keeping it only in process memory.&lt;/p&gt;

&lt;h2&gt;
  
  
  Writing Motion Prompts
&lt;/h2&gt;

&lt;p&gt;Motion prompts are natural-language instructions. I get more predictable results by keeping the movement description short and separating camera movement from subject movement.&lt;/p&gt;

&lt;p&gt;Useful patterns include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;"camera dolly left while the subject walks forward"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;"leaf falls from tree and drifts toward camera"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;"slow zoom in, slight parallax, 2x speed"&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;"subtle motion, loopable, cinematic rhythm"&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A more structured prompt can be written like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;start_frame_url animate: "slow spiral camera, subject bobs gently, loopable", style: "film grain, cinematic, 2 fps tempo"
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The useful progression is to state the action first, then add timing and stylistic constraints. Small iterations are more productive than stacking a long list of unrelated instructions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Automatic versus manual animation
&lt;/h3&gt;

&lt;p&gt;Automatic animation is useful when plausible motion is enough. The model infers how the scene should move and is usually the fastest way to test an image.&lt;/p&gt;

&lt;p&gt;Manual animation is better when camera direction, subject movement, or choreography needs to be repeatable. It is also the better choice when the clip must align with live-action footage or a predefined sequence.&lt;/p&gt;

&lt;h2&gt;
  
  
  Extending and Looping Clips
&lt;/h2&gt;

&lt;p&gt;A generated five-second clip can be extended by four seconds per operation, up to approximately 21 seconds. In the UI, this is exposed through an &lt;strong&gt;Extend&lt;/strong&gt; control. Programmatically, wrappers generally expose an &lt;code&gt;extend&lt;/code&gt; flag or a separate extend job referencing the original clip. The exact parameterized endpoints and controls depend on the API documentation.&lt;/p&gt;

&lt;p&gt;Extensions should be treated as additional generation work, with costs similar to an initial generation.&lt;/p&gt;

&lt;p&gt;For loops, reuse the starting frame as the ending frame or use the &lt;code&gt;--loop&lt;/code&gt; parameter. For a different ending, provide another image URL through &lt;code&gt;end&lt;/code&gt; and keep its aspect ratio compatible with the starting image. Midjourney also supports a &lt;code&gt;--end&lt;/code&gt; parameter. Manual extension can help adjust the motion prompt when continuity starts to drift.&lt;/p&gt;

&lt;p&gt;Batch size is another simple cost control. Since the default is four variants, setting &lt;code&gt;bs: 1&lt;/code&gt; is useful when the goal is a single production candidate rather than exploration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Adding Audio After Generation
&lt;/h2&gt;

&lt;p&gt;Midjourney V1 produces silent MP4 video. Audio is not generated natively, so voice, music, and effects need to be added afterward.&lt;/p&gt;

&lt;p&gt;A practical audio pipeline usually has three parts:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Generate narration with a text-to-speech service such as ElevenLabs, Replica, or another voice-cloning/TTS provider.&lt;/li&gt;
&lt;li&gt;Generate or source music and sound effects using tools such as MM Audio, Magicshot, or specialized SFX generators.&lt;/li&gt;
&lt;li&gt;Mix the tracks in DaVinci Resolve, Premiere, Audacity, or a similar editor.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For timing-sensitive work, I would do the final mix in a DAW or video editor. That provides better control over dialogue timing, sound effects, background levels, and synchronization than trying to assemble everything through isolated API calls.&lt;/p&gt;

&lt;p&gt;For a simple silent video plus speech track, &lt;code&gt;ffmpeg&lt;/code&gt; is enough:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Normalize audio length (optional), then combine:&lt;/span&gt;
ffmpeg &lt;span class="nt"&gt;-i&lt;/span&gt; video.mp4 &lt;span class="nt"&gt;-i&lt;/span&gt; speech.mp3 &lt;span class="nt"&gt;-c&lt;/span&gt;:v copy &lt;span class="nt"&gt;-c&lt;/span&gt;:a aac &lt;span class="nt"&gt;-shortest&lt;/span&gt; output_with_audio.mp4
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For dialogue, music, and effects, render one mixed audio track first, then mux that track into the MP4 using the same approach.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where V1 Fits
&lt;/h2&gt;

&lt;p&gt;V1 is deliberately constrained: short clips, image-driven motion, and relatively modest control over long sequences. Those constraints are also what make it useful for fast iteration.&lt;/p&gt;

&lt;p&gt;The strongest use cases are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Animated illustrations.&lt;/li&gt;
&lt;li&gt;Product hero loops.&lt;/li&gt;
&lt;li&gt;Short character movements.&lt;/li&gt;
&lt;li&gt;Stylized social clips.&lt;/li&gt;
&lt;li&gt;Loopable background visuals.&lt;/li&gt;
&lt;li&gt;Motion studies from existing still images.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It is less suitable for long continuous scenes, complex multi-shot narratives, or workflows that require precise cinematic camera rigs. The model is expected to improve over time in sequence length, fidelity, and camera control, but the current workflow is best treated as concise motion design built from a strong still frame.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-create-a-video-in-midjourney-api/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-create-a-video-in-midjourney-api"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How I’d Choose Between GPT-Image-2.5 and Nano Banana 2 for Production</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Tue, 22 Sep 2026 01:23:01 +0000</pubDate>
      <link>https://dev.to/ryancole1/how-id-choose-between-gpt-image-25-and-nano-banana-2-for-production-55pa</link>
      <guid>https://dev.to/ryancole1/how-id-choose-between-gpt-image-25-and-nano-banana-2-for-production-55pa</guid>
      <description>&lt;p&gt;I’d start with Sunburst for edits that must preserve a product or person, Flare for rapid iteration, and Nano Banana 2 for generation that needs search context, unusual canvas shapes, or predictable batch costs.&lt;/p&gt;

&lt;p&gt;That split is more useful than declaring an overall winner. The reported Arena results favor GPT-Image-2.5, especially for editing, but a preference leaderboard doesn’t measure retrieval integration, output constraints, or the cost of producing thousands of acceptable assets.&lt;/p&gt;

&lt;p&gt;This comparison uses the September 2026 release and leaderboard snapshot discussed below. The OpenAI benchmark entries are still preliminary; I’d treat them as a reason to prioritize an evaluation, rather than a substitute for one.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the API contract
&lt;/h2&gt;

&lt;p&gt;GPT-Image-2.5 has two API variants. OpenAI launched both on September 8, 2026:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Flare&lt;/strong&gt; targets everyday generation, fast iteration, and volume.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sunburst&lt;/strong&gt; spends more generation time on precision editing and tighter creative control.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Nano Banana 2 is Google’s &lt;strong&gt;Gemini 3.1 Flash Image&lt;/strong&gt;, introduced on February 26, 2026. Google’s release notes place general availability on May 28, 2026, with video-to-image context support added. The older preview model ID was deprecated, although some integration URLs still contain the preview slug.&lt;/p&gt;

&lt;p&gt;Here’s the contract I’d compare before looking at sample images:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Capability&lt;/th&gt;
&lt;th&gt;GPT-Image-2.5 Flare&lt;/th&gt;
&lt;th&gt;GPT-Image-2.5 Sunburst&lt;/th&gt;
&lt;th&gt;Nano Banana 2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;API model ID&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gpt-image-2.5-flare&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gpt-image-2.5-sunburst&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemini-3.1-flash-image&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;Text, image&lt;/td&gt;
&lt;td&gt;Text, image&lt;/td&gt;
&lt;td&gt;Text, image, video context&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;Image&lt;/td&gt;
&lt;td&gt;Image&lt;/td&gt;
&lt;td&gt;Image, text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main role&lt;/td&gt;
&lt;td&gt;Fast generation and controlled edits&lt;/td&gt;
&lt;td&gt;Precision editing and premium assets&lt;/td&gt;
&lt;td&gt;Grounded multimodal generation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quality settings&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;low&lt;/code&gt;, &lt;code&gt;medium&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;xhigh&lt;/code&gt;, &lt;code&gt;max&lt;/code&gt;, &lt;code&gt;auto&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Same&lt;/td&gt;
&lt;td&gt;Primarily resolution-driven&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Common output sizes&lt;/td&gt;
&lt;td&gt;1024×1024, 1536×1024, 1024×1536&lt;/td&gt;
&lt;td&gt;Same&lt;/td&gt;
&lt;td&gt;0.5K, 1K, 2K, 4K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Maximum edge&lt;/td&gt;
&lt;td&gt;3840 px&lt;/td&gt;
&lt;td&gt;3840 px&lt;/td&gt;
&lt;td&gt;Up to 4K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Aspect ratios&lt;/td&gt;
&lt;td&gt;Approximately 1:3 through 3:1&lt;/td&gt;
&lt;td&gt;Same&lt;/td&gt;
&lt;td&gt;Includes 1:4, 4:1, 1:8, 8:1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transparent backgrounds&lt;/td&gt;
&lt;td&gt;PNG and WebP&lt;/td&gt;
&lt;td&gt;PNG and WebP&lt;/td&gt;
&lt;td&gt;Not a headline API capability&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Native search grounding&lt;/td&gt;
&lt;td&gt;Not documented at image-model level&lt;/td&gt;
&lt;td&gt;Not documented at image-model level&lt;/td&gt;
&lt;td&gt;Google Web and Image Search&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thinking&lt;/td&gt;
&lt;td&gt;Not exposed as a core capability&lt;/td&gt;
&lt;td&gt;Not exposed as a core capability&lt;/td&gt;
&lt;td&gt;Supported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch API&lt;/td&gt;
&lt;td&gt;Not a headline 2.5 feature&lt;/td&gt;
&lt;td&gt;Not a headline 2.5 feature&lt;/td&gt;
&lt;td&gt;Supported&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;OpenAI’s model documentation describes &lt;a href="https://developers.openai.com/api/docs/models/gpt-image-2.5-flare" rel="noopener noreferrer"&gt;Flare&lt;/a&gt; as the faster everyday option and &lt;a href="https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst" rel="noopener noreferrer"&gt;Sunburst&lt;/a&gt; as the variant optimized for editing precision. Both expose the six quality settings above.&lt;/p&gt;

&lt;p&gt;For Google, I’d check the current &lt;a href="https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-image" rel="noopener noreferrer"&gt;Gemini 3.1 Flash Image documentation&lt;/a&gt;, particularly when migrating an integration that still references the preview model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Editing is where I’d test Sunburst first
&lt;/h2&gt;

&lt;p&gt;The useful editing question is how much unrelated content changes after each revision.&lt;/p&gt;

&lt;p&gt;A product workflow might replace a background, adjust lighting, add seasonal styling, revise copy, and then change the aspect ratio. Each individual render can look good while the bottle, label, or person slowly drifts away from the reference.&lt;/p&gt;

&lt;p&gt;OpenAI’s &lt;a href="https://openai.com/index/introducing-chatgpt-images-2-5/" rel="noopener noreferrer"&gt;Images 2.5 announcement&lt;/a&gt; emphasizes preservation of reference subjects, natural lighting, richer textures, and more consistent multi-turn edits. Sunburst specifically targets the workflows where preserving those details matters more than minimum latency.&lt;/p&gt;

&lt;p&gt;That makes it my first evaluation candidate for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Product photographs where packaging and geometry must survive revisions.&lt;/li&gt;
&lt;li&gt;Campaign assets with repeated changes to copy, placement, and styling.&lt;/li&gt;
&lt;li&gt;Character or subject transformations driven by reference images.&lt;/li&gt;
&lt;li&gt;Structured posters, branded layouts, and transparent design elements.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Flare occupies the faster lane in the same family. OpenAI claims higher-quality output with &lt;strong&gt;up to 50% lower generation latency than GPT Image 2&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That figure is an OpenAI generation-to-generation comparison. It says nothing conclusive about Flare’s speed relative to Nano Banana 2. I’d benchmark both with the same prompts, output dimensions, and concurrency before making a latency commitment.&lt;/p&gt;

&lt;p&gt;Google’s model also supports conversational editing and subject consistency. Its stronger differentiator is the surrounding Gemini workflow: reasoning, search, multilingual text, and multiple references.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Arena numbers actually support
&lt;/h2&gt;

&lt;p&gt;Arena aggregates blind side-by-side human preferences. The reported snapshot looks like this:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Arena metric&lt;/th&gt;
&lt;th&gt;Sunburst&lt;/th&gt;
&lt;th&gt;Flare&lt;/th&gt;
&lt;th&gt;Nano Banana 2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Text-to-Image score&lt;/td&gt;
&lt;td&gt;1421±13&lt;/td&gt;
&lt;td&gt;1399±13&lt;/td&gt;
&lt;td&gt;1261±5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text-to-Image rank&lt;/td&gt;
&lt;td&gt;#1&lt;/td&gt;
&lt;td&gt;#2&lt;/td&gt;
&lt;td&gt;#9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text-to-Image votes&lt;/td&gt;
&lt;td&gt;3,149&lt;/td&gt;
&lt;td&gt;2,856&lt;/td&gt;
&lt;td&gt;41,957&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single Image Edit score&lt;/td&gt;
&lt;td&gt;1520±9&lt;/td&gt;
&lt;td&gt;1491±9&lt;/td&gt;
&lt;td&gt;1387±4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single Image Edit rank&lt;/td&gt;
&lt;td&gt;#1&lt;/td&gt;
&lt;td&gt;#2&lt;/td&gt;
&lt;td&gt;#12&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single Image Edit votes&lt;/td&gt;
&lt;td&gt;6,704&lt;/td&gt;
&lt;td&gt;5,676&lt;/td&gt;
&lt;td&gt;157,693&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sunburst leads Nano Banana 2 by &lt;strong&gt;160 points&lt;/strong&gt; for text-to-image generation and &lt;strong&gt;133 points&lt;/strong&gt; for single-image editing. Flare’s corresponding leads are &lt;strong&gt;138&lt;/strong&gt; and &lt;strong&gt;104 points&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That is a strong early preference signal, and the editing result is consistent with OpenAI’s stated priorities. It makes Sunburst a sensible first test for controlled revisions.&lt;/p&gt;

&lt;p&gt;The sample sizes matter, though. Both OpenAI entries are marked &lt;strong&gt;Preliminary&lt;/strong&gt;, with only a few thousand votes. Google’s entry has 41,957 generation votes and 157,693 editing votes. The new entries have more room to move as comparisons accumulate.&lt;/p&gt;

&lt;p&gt;Arena points also aren’t percentages. A 160-point lead doesn’t translate into a fixed percentage improvement in image quality, and the ranking reflects Arena’s prompt distribution. It doesn’t establish which model handles your packaging, typography, references, or acceptance criteria best.&lt;/p&gt;

&lt;p&gt;I’d use the leaderboard to choose evaluation order, then measure performance on the application’s actual workload.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Nano Banana 2 changes the implementation
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Search can be part of generation
&lt;/h3&gt;

&lt;p&gt;Google documents both Web Search and Image Search grounding for Gemini 3.1 Flash Image. Retrieved text and images can inform generation using current web information.&lt;/p&gt;

&lt;p&gt;This applies when the Google Search tool is enabled and its attribution requirements are followed. Merely selecting the model doesn’t make every output grounded.&lt;/p&gt;

&lt;p&gt;For a travel application, educational diagram, visual search interface, or information graphic, that integration can reduce the work needed to assemble context. A landmark illustration can use retrieved references; a localized graphic can combine factual context with rendered text.&lt;/p&gt;

&lt;p&gt;GPT-Image-2.5 doesn’t document equivalent native grounding at the image-model level. An application can supply its own retrieved context and reference images, which narrows the practical difference. If the application already uses Gemini, keeping retrieval and generation in that workflow may be simpler.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reference scale is explicitly documented
&lt;/h3&gt;

&lt;p&gt;Google highlights identity preservation for &lt;strong&gt;up to five characters&lt;/strong&gt; and fidelity for &lt;strong&gt;up to 14 objects&lt;/strong&gt; in a workflow.&lt;/p&gt;

&lt;p&gt;Those figures are useful when scoping compositions with several recurring subjects. They are documented capability claims, rather than a guarantee that every arrangement will preserve every detail.&lt;/p&gt;

&lt;p&gt;OpenAI emphasizes strong reference preservation, especially with Sunburst, but the comparison doesn’t provide equivalent character and object counts for its models. I’d separate documented reference scale from observed fidelity during testing.&lt;/p&gt;

&lt;h3&gt;
  
  
  Localization has a clear place in the feature set
&lt;/h3&gt;

&lt;p&gt;Nano Banana 2 emphasizes international text rendering and translation of text inside an existing visual. That is useful for adapting one campaign across markets.&lt;/p&gt;

&lt;p&gt;OpenAI emphasizes infographic accuracy, hierarchy, and layout. For a static advertisement or UI concept where exact spatial structure dominates, I’d try GPT-Image-2.5 first. For multilingual graphics or visuals whose copy depends on retrieved information, I’d start with Google.&lt;/p&gt;

&lt;p&gt;Both still need output verification when numerical or legal copy matters. A convincing layout can contain incorrect text.&lt;/p&gt;

&lt;h2&gt;
  
  
  Canvas requirements can decide the model before quality does
&lt;/h2&gt;

&lt;p&gt;Nano Banana 2 supports &lt;strong&gt;0.5K, 1K, 2K, and 4K&lt;/strong&gt; output, spanning the documented 512px-to-4K range. Its extreme aspect ratios include &lt;strong&gt;1:4, 4:1, 1:8, and 8:1&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That gives it a straightforward advantage for panoramic backgrounds, narrow mobile creatives, tall product displays, and campaigns needing many canvas shapes.&lt;/p&gt;

&lt;p&gt;GPT-Image-2.5 supports custom dimensions within an approximately &lt;strong&gt;1:3 to 3:1&lt;/strong&gt; aspect-ratio range. Its documented limits are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Maximum edge: &lt;strong&gt;3840 pixels&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Maximum output area: &lt;strong&gt;8,294,400 pixels&lt;/strong&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those constraints should stay explicit in application validation. “Supports 4K” would obscure the difference between OpenAI’s limits and Nano Banana 2’s 4096×4096 4K tier.&lt;/p&gt;

&lt;p&gt;OpenAI has a separate practical advantage: &lt;strong&gt;transparent PNG and WebP backgrounds are explicitly supported&lt;/strong&gt;. That matters for product cutouts, composited design elements, and assets that need an alpha channel.&lt;/p&gt;

&lt;p&gt;My routing decision here would be mechanical: check canvas dimensions and transparency requirements before spending time comparing creative quality.&lt;/p&gt;

&lt;h2&gt;
  
  
  Photorealism needs its own evaluation
&lt;/h2&gt;

&lt;p&gt;I wouldn’t infer a universal photorealism winner from the Arena ranking.&lt;/p&gt;

&lt;p&gt;OpenAI describes improvements to natural lighting, texture, and recognizable reference subjects. Google makes similar claims about detail, texture, lighting, and photographic quality.&lt;/p&gt;

&lt;p&gt;Some early subjective comparisons have favored Nano Banana 2 for camera feel, materials, and natural product scenes, while favoring GPT-Image-2.5 for controlled edits and text-heavy layouts. Those observations are useful leads, but they aren’t standardized cross-provider measurements.&lt;/p&gt;

&lt;p&gt;For a production evaluation, I’d include skin and hair, glossy and matte products, transparent materials, fabric, food, architecture, indoor lighting, shallow depth of field, and identity preservation from reference photos.&lt;/p&gt;

&lt;p&gt;The metric I care about is &lt;strong&gt;cost per accepted image&lt;/strong&gt;. A model that produces one exceptional sample can still be expensive if most outputs need another render or manual correction.&lt;/p&gt;

&lt;h3&gt;
  
  
  One shared product prompt
&lt;/h3&gt;

&lt;p&gt;The supplied comparison uses this exact prompt:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Photorealistic product photograph of a matte black water bottle standing on a pale concrete ledge. Soft morning light from the left, gentle reflections, shallow depth of field. Centered composition, Include ONLY this text (verbatim): headline "YOURS TO CREATE" in bold sans-serif across the top, subhead "Limited Edition" smaller at the bottom. No other text or logos.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;GPT-Image-2.5 Sunburst&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftqwxzkfztwkljeur4l0t.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Ftqwxzkfztwkljeur4l0t.png" alt="Sunburst output for the matte black water bottle prompt" width="800" height="533"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Nano Banana 2&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkdgtcyaej94at1fsghwj.jpeg" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fkdgtcyaej94at1fsghwj.jpeg" alt="Nano Banana 2 output for the same water bottle prompt" width="800" height="447"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;I’d score this prompt separately for exact copy, text placement, bottle geometry, material appearance, lighting direction, and unwanted logos. It exercises several requirements at once, so a single overall preference can hide the reason an image fails acceptance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compare billing at the image level
&lt;/h2&gt;

&lt;p&gt;The pricing models encourage different budgeting approaches.&lt;/p&gt;

&lt;p&gt;GPT-Image-2.5 uses token billing, with consumption affected by quality settings. Nano Banana 2 publishes image-output costs by resolution, making that part of the budget easier to estimate.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Official billing unit&lt;/th&gt;
&lt;th&gt;Price&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-Image-2.5 text input&lt;/td&gt;
&lt;td&gt;$5 / 1M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-Image-2.5 image input&lt;/td&gt;
&lt;td&gt;$8 / 1M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-Image-2.5 image output&lt;/td&gt;
&lt;td&gt;$30 / 1M image tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nano Banana 2, 0.5K&lt;/td&gt;
&lt;td&gt;$0.045 / image&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nano Banana 2, 1K&lt;/td&gt;
&lt;td&gt;$0.067 / image&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nano Banana 2, 2K&lt;/td&gt;
&lt;td&gt;$0.101 / image&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Nano Banana 2, 4K&lt;/td&gt;
&lt;td&gt;$0.151 / image&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Comparing OpenAI’s &lt;strong&gt;$30&lt;/strong&gt; with Google’s &lt;strong&gt;$60 per million image tokens&lt;/strong&gt; doesn’t establish that OpenAI costs half as much. The providers tokenize images differently, and OpenAI’s quality ladder changes consumption.&lt;/p&gt;

&lt;p&gt;Google’s &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;published pricing&lt;/a&gt; also includes Batch API image-output equivalents approximately 50% below standard pricing:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Resolution&lt;/th&gt;
&lt;th&gt;Standard&lt;/th&gt;
&lt;th&gt;Approximate Batch equivalent&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;0.5K&lt;/td&gt;
&lt;td&gt;$0.045&lt;/td&gt;
&lt;td&gt;$0.022&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1K&lt;/td&gt;
&lt;td&gt;$0.067&lt;/td&gt;
&lt;td&gt;$0.034&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;2K&lt;/td&gt;
&lt;td&gt;$0.101&lt;/td&gt;
&lt;td&gt;$0.050&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;4K&lt;/td&gt;
&lt;td&gt;$0.151&lt;/td&gt;
&lt;td&gt;$0.076&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For asynchronous asset production, that is a meaningful feature. For interactive generation, I’d evaluate the standard path separately.&lt;/p&gt;

&lt;p&gt;OpenAI’s quality controls support another useful pattern: inexpensive drafts followed by more expensive final renders. Both variants expose &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;medium&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;xhigh&lt;/code&gt;, &lt;code&gt;max&lt;/code&gt;, and &lt;code&gt;auto&lt;/code&gt;; the selected setting belongs in both cost and latency measurements.&lt;/p&gt;

&lt;h3&gt;
  
  
  A unified API can simplify routing
&lt;/h3&gt;

&lt;p&gt;If I wanted one integration for both families, CometAPI is one option. The source’s quoted rates for both Flare and Sunburst are $4 per million text-input tokens and $24 per million output tokens, versus the corresponding official $5/$30 rates; image-input pricing requires a current provider quote.&lt;/p&gt;

&lt;p&gt;The same quoted gateway pricing for Nano Banana 2 is &lt;strong&gt;$0.0360&lt;/strong&gt;, &lt;strong&gt;$0.0536&lt;/strong&gt;, &lt;strong&gt;$0.0808&lt;/strong&gt;, and &lt;strong&gt;$0.1208&lt;/strong&gt; per image at 0.5K, 1K, 2K, and 4K respectively. Batch availability needs checking with the provider.&lt;/p&gt;

&lt;p&gt;The engineering value is a common integration through which the application can select a model per request. I’d still keep provider-specific constraints visible in the routing logic.&lt;/p&gt;

&lt;h2&gt;
  
  
  The routing policy I’d put into an application
&lt;/h2&gt;

&lt;p&gt;I’d make model selection follow the request’s constraints and the cost of rework.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Request characteristic&lt;/th&gt;
&lt;th&gt;First model I’d evaluate&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Precise edits with minimal unrelated change&lt;/td&gt;
&lt;td&gt;Sunburst&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Final campaign or product asset&lt;/td&gt;
&lt;td&gt;Sunburst&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rapid OpenAI variants and everyday generation&lt;/td&gt;
&lt;td&gt;Flare&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transparent PNG or WebP asset&lt;/td&gt;
&lt;td&gt;GPT-Image-2.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Current web or visual context needed&lt;/td&gt;
&lt;td&gt;Nano Banana 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Extreme vertical or horizontal canvas&lt;/td&gt;
&lt;td&gt;Nano Banana 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Native resolution tiers through 4K&lt;/td&gt;
&lt;td&gt;Nano Banana 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multilingual text or translation inside a visual&lt;/td&gt;
&lt;td&gt;Nano Banana 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Asynchronous volume with published Batch pricing&lt;/td&gt;
&lt;td&gt;Nano Banana 2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Several recurring characters or objects&lt;/td&gt;
&lt;td&gt;Nano Banana 2 for its documented reference scale&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Photorealistic scenes without other constraints&lt;/td&gt;
&lt;td&gt;Evaluate both on representative prompts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For an editing-heavy product, I’d begin with Sunburst and test whether Flare meets the same acceptance criteria at lower latency. For a Gemini application generating grounded diagrams or localized assets in many shapes, I’d begin with Nano Banana 2.&lt;/p&gt;

&lt;p&gt;Mixed workloads justify using all three. One possible flow is grounded concept generation with Google followed by a controlled Sunburst revision. Another is Flare for initial variants, Sunburst for difficult final edits, and Google for canvas shapes outside OpenAI’s range.&lt;/p&gt;

&lt;p&gt;Those are architectures I’d evaluate, not measured guarantees that switching models improves a result. Any handoff belongs in the same reference-preservation tests as a single-model editing sequence.&lt;/p&gt;

&lt;p&gt;Before committing, I’d run identical prompt sets at comparable dimensions and concurrency, record actual latency and billed cost, and judge the final assets against explicit acceptance criteria. The preliminary leaderboard makes GPT-Image-2.5 a strong starting point for quality and editing; the production requirements determine whether it earns the request.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/gpt-image-2-5-vs-nano-banana-2/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=gpt-image-2-5-vs-nano-banana-2"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Free ChatGPT for Students: What I’d Check Before Paying for Plus</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Mon, 21 Sep 2026 09:32:44 +0000</pubDate>
      <link>https://dev.to/ryancole1/free-chatgpt-for-students-what-id-check-before-paying-for-plus-a5l</link>
      <guid>https://dev.to/ryancole1/free-chatgpt-for-students-what-id-check-before-paying-for-plus-a5l</guid>
      <description>&lt;p&gt;College students can use ChatGPT for free. A student email address, however, does not automatically unlock ChatGPT Plus.&lt;/p&gt;

&lt;p&gt;For the 2026 academic year, I’d separate access into three buckets: the public Free plan, a personal Plus subscription, and a university-managed ChatGPT Edu workspace. They have different limits, billing arrangements, and data policies. The two-month student promotion from 2025 is a fourth case, but its claim window has closed.&lt;/p&gt;

&lt;p&gt;My first step would be to check the university’s software portal. Paying for a personal subscription before checking campus access can mean buying capabilities the institution already provides.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 2025 student offer has expired
&lt;/h2&gt;

&lt;p&gt;OpenAI offered eligible college students in the &lt;strong&gt;United States and Canada two months of ChatGPT Plus at no cost&lt;/strong&gt;. The claim window ran from &lt;strong&gt;March 31 through May 31, 2025&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;That date matters. An article describing this as an offer available “right now” in 2026 is mixing a historical promotion with current access.&lt;/p&gt;

&lt;p&gt;The promotion worked as follows:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Students at qualifying institutions verified enrollment through &lt;strong&gt;SheerID&lt;/strong&gt; in the ChatGPT interface.&lt;/li&gt;
&lt;li&gt;The two free months started when the student claimed the offer.&lt;/li&gt;
&lt;li&gt;Existing paid Plus subscribers were also eligible, with the offer credited to their accounts.&lt;/li&gt;
&lt;li&gt;Normal subscription billing resumed after the promotional period unless the student canceled.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Students outside the eligible countries, or those who missed the May 31 deadline, could not claim this particular offer.&lt;/p&gt;

&lt;p&gt;There is no permanent, globally available free Plus entitlement for college students described here. Future seasonal offers are possible, but I wouldn’t budget around an unannounced promotion.&lt;/p&gt;

&lt;h3&gt;
  
  
  What those two months provided
&lt;/h3&gt;

&lt;p&gt;The offer gave students the normal Plus subscription benefits for the promotional period. It was a temporary upgrade to the paid product, with the same applicable usage limits.&lt;/p&gt;

&lt;p&gt;Those benefits included access to more capable models, priority availability during busy periods, higher messaging limits, and more generous file uploads. Expanded data analysis, image generation, multimodal interactions, and voice features also made Plus more useful for sustained academic work.&lt;/p&gt;

&lt;p&gt;Reporting around that period included research tools and the &lt;strong&gt;GPT-4.5 preview&lt;/strong&gt;. I’d treat those as historical feature references: a model available during the 2025 offer is not a promise about the 2026 model picker.&lt;/p&gt;

&lt;p&gt;For coursework, the practical gains were straightforward: longer sessions reviewing materials, more room to work with research documents, and fewer interruptions while generating practice problems or analyzing datasets.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check campus access before comparing subscription prices
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;ChatGPT Edu&lt;/strong&gt; is an institutional offering for universities. It supports students, faculty, researchers, and campus operations, with administrative and privacy controls suited to a managed workspace.&lt;/p&gt;

&lt;p&gt;The institution pays for access, so an eligible student may have no direct subscription charge.&lt;/p&gt;

&lt;p&gt;Arizona State University, the University of Oxford, and the Wharton School of the University of Pennsylvania have been cited as examples of university AI adoption. I would still check the actual campus entitlement. A university partnership alone does not establish that every student, department, or course receives identical access.&lt;/p&gt;

&lt;p&gt;The useful questions are concrete:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Is my account eligible for the university’s ChatGPT workspace?&lt;/li&gt;
&lt;li&gt;How do I activate access with university credentials?&lt;/li&gt;
&lt;li&gt;Which models and tools are enabled?&lt;/li&gt;
&lt;li&gt;What usage limits and research-data rules apply?&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Why Edu can matter more than a personal upgrade
&lt;/h3&gt;

&lt;p&gt;For me, the most consequential Edu feature is its data policy: &lt;strong&gt;workspace data is not used to train OpenAI’s models&lt;/strong&gt;. That is especially relevant when working with thesis drafts or unpublished research.&lt;/p&gt;

&lt;p&gt;It does not, by itself, authorize uploading every research dataset. The university’s rules still determine what belongs in that workspace.&lt;/p&gt;

&lt;p&gt;Edu also offers higher usage limits and tools for working with spreadsheets, CSVs, and PDFs. Those capabilities support statistical analysis, visualization, and document summarization. Universities can create and share custom GPTs inside their workspace, such as a course tutor or grant-writing assistant.&lt;/p&gt;

&lt;p&gt;Model access has included &lt;strong&gt;GPT-4o&lt;/strong&gt;, with access to newer models depending on the offering and rollout. I would avoid interpreting “higher limits” as “unlimited”: the workspace’s actual limits are what determine whether it supports a long analysis session.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where the public Free plan is enough
&lt;/h2&gt;

&lt;p&gt;The Free plan remains useful for bounded tasks: explaining a compiler error, reviewing a short function, brainstorming an essay outline, or summarizing a short reading.&lt;/p&gt;

&lt;p&gt;I’d start there if my workload mostly consists of short, independent requests. The friction becomes more visible when a task needs repeated file uploads, extended reasoning, or a long uninterrupted conversation.&lt;/p&gt;

&lt;p&gt;Free access can involve dynamic message caps and reduced access to more capable models after a limit is reached. Data analysis, uploads, image generation, and advanced reasoning features can also have tighter restrictions.&lt;/p&gt;

&lt;p&gt;Specific model names need care. References to &lt;strong&gt;GPT-4o mini&lt;/strong&gt;, &lt;strong&gt;GPT-4o&lt;/strong&gt;, or &lt;strong&gt;GPT-5.2&lt;/strong&gt; do not establish a stable Free-plan entitlement for all of 2026. The available models and their limits need to be checked against the current plan and account.&lt;/p&gt;

&lt;h3&gt;
  
  
  Compare the workload, then the price
&lt;/h3&gt;

&lt;p&gt;The personal Plus price discussed here is &lt;strong&gt;$20 per month&lt;/strong&gt;. Whether that makes sense depends on how often the Free plan interrupts useful work.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Decision point&lt;/th&gt;
&lt;th&gt;Public Free plan&lt;/th&gt;
&lt;th&gt;Personal Plus subscription&lt;/th&gt;
&lt;th&gt;University Edu workspace&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Who pays?&lt;/td&gt;
&lt;td&gt;No subscription charge&lt;/td&gt;
&lt;td&gt;Student or another individual payer&lt;/td&gt;
&lt;td&gt;Institution&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Access limits&lt;/td&gt;
&lt;td&gt;Lower, dynamic limits&lt;/td&gt;
&lt;td&gt;Higher limits, still subject to caps&lt;/td&gt;
&lt;td&gt;Higher institutional limits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Documents and datasets&lt;/td&gt;
&lt;td&gt;More restricted uploads and analysis&lt;/td&gt;
&lt;td&gt;More generous uploads and analysis&lt;/td&gt;
&lt;td&gt;Analysis tools within a managed workspace&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Training policy&lt;/td&gt;
&lt;td&gt;Data may be used, depending on controls&lt;/td&gt;
&lt;td&gt;Data may be used unless opted out&lt;/td&gt;
&lt;td&gt;Workspace data is not used for training&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Main access requirement&lt;/td&gt;
&lt;td&gt;Public account&lt;/td&gt;
&lt;td&gt;Paid subscription&lt;/td&gt;
&lt;td&gt;University eligibility and provisioning&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I wouldn’t use an old comparison table listing &lt;strong&gt;o1-preview&lt;/strong&gt;, “legacy models,” &lt;strong&gt;DALL-E 3&lt;/strong&gt;, or a blanket &lt;strong&gt;5×&lt;/strong&gt; limit increase as a current purchasing specification. Those details are tied to particular product periods, models, and tools.&lt;/p&gt;

&lt;h3&gt;
  
  
  Long documents introduce a separate constraint
&lt;/h3&gt;

&lt;p&gt;The context window limits how much material a model can consider at once. A long conversation can lose useful earlier detail, which becomes noticeable when discussing something like a &lt;strong&gt;20-page dissertation&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;A paid plan may provide access to models with larger context windows, but subscription tier alone is not a complete specification. The selected model and product behavior matter.&lt;/p&gt;

&lt;p&gt;I’d check that separately from message limits. More messages do not automatically mean the model can consider the entire document and conversation in every response.&lt;/p&gt;

&lt;h2&gt;
  
  
  Alternatives need the same eligibility check
&lt;/h2&gt;

&lt;p&gt;Google announced a student offer for &lt;strong&gt;Google One AI Premium&lt;/strong&gt;, which bundled Gemini Advanced and other tools, providing eligible U.S. college students free access through &lt;strong&gt;June 30, 2026&lt;/strong&gt;, subject to registering by the required deadline.&lt;/p&gt;

&lt;p&gt;The access end date and registration deadline are different things. Seeing “free through June 2026” does not establish that enrollment is still open.&lt;/p&gt;

&lt;p&gt;For model selection, Gemini is relevant to workflows involving images, video, and Google Search retrieval. Claude is another option for long-document reading and summarization. Neither observation establishes that the required features are free or unlimited; those depend on the applicable plan.&lt;/p&gt;

&lt;p&gt;For a student building software that needs multiple providers, CometAPI exposes models through a unified, OpenAI-style REST API, with model selection via a model string. That can reduce provider-specific integration work, but API access belongs in a separate cost comparison from a free ChatGPT account or a campus subscription.&lt;/p&gt;

&lt;p&gt;My purchase trigger would be a recurring constraint: hitting message caps during productive work, needing more document analysis, or requiring a capability unavailable through campus access. Until one of those shows up, I’d use the free account and check the university’s provisioning options first.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/is-chatgpt-free-for-college-students/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=is-chatgpt-free-for-college-students"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Building a Resilient Multi-Model AI Stack in 2026</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Mon, 21 Sep 2026 07:12:47 +0000</pubDate>
      <link>https://dev.to/ryancole1/building-a-resilient-multi-model-ai-stack-in-2026-1in</link>
      <guid>https://dev.to/ryancole1/building-a-resilient-multi-model-ai-stack-in-2026-1in</guid>
      <description>&lt;h2&gt;
  
  
  The Short Version
&lt;/h2&gt;

&lt;p&gt;For an OpenAI-compatible migration, change the SDK endpoint to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://api.cometapi.com/v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then replace the OpenAI key with a gateway token. The same integration exposes more than 500 frontier models, including GPT-5.5, Claude Opus 4.7, and Gemini 3.1 Pro. The advertised pricing reduction is 20% to 40% compared with official direct rates.&lt;/p&gt;

&lt;p&gt;That configuration change is small. The architectural change is more important: production AI features no longer need to depend on one provider, one region, or one model family.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why I Prefer a Multi-Model Architecture
&lt;/h2&gt;

&lt;p&gt;A single-provider dependency creates a predictable failure mode. If the provider develops elevated latency, throttles a tenant, or returns &lt;code&gt;429 Too Many Requests&lt;/code&gt;, every feature tied to that endpoint is affected.&lt;/p&gt;

&lt;p&gt;I have seen this happen even when an account is operating within its documented Tier 5 limits. If the application hardcodes one provider, the operational choices at 1:00 AM are limited to waiting for recovery or shipping an emergency patch.&lt;/p&gt;

&lt;p&gt;A unified gateway such as &lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt; provides another option: move a workload to a comparable model or region without changing the application’s message format. For example, reasoning traffic can be shifted to Claude Opus 4.7 from an operational dashboard while the rest of the application continues using the same client integration.&lt;/p&gt;

&lt;p&gt;This is useful even when OpenAI remains the default. Anthropic and Google models such as Claude Opus 4.7 and Gemini 3.1 Pro can outperform GPT-4o on particular coding and multimodal reasoning workloads. Model choice should be a workload decision, not an accidental consequence of whichever SDK was installed first.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pricing Snapshot
&lt;/h2&gt;

&lt;p&gt;Bulk token purchasing and routing can materially change the economics of an AI feature. The following prices were verified in May 2026 through the gateway’s pricing information.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Official Price (Input / 1M)&lt;/th&gt;
&lt;th&gt;Gateway Price (Input / 1M)&lt;/th&gt;
&lt;th&gt;Total Savings&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.cometapi.com/models/openai/gpt-5-5-pro/" rel="noopener noreferrer"&gt;GPT-5.5 Pro&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;$30.00&lt;/td&gt;
&lt;td&gt;$24.00&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.cometapi.com/models/openai/gpt-5-5/" rel="noopener noreferrer"&gt;GPT-5.5&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$4.00&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.cometapi.com/models/anthropic/claude-opus-4-7/" rel="noopener noreferrer"&gt;Claude Opus 4.7&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;td&gt;$3.00&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.cometapi.com/models/anthropic/claude-sonnet-4-6/" rel="noopener noreferrer"&gt;Claude Sonnet 4.6&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;$3.00&lt;/td&gt;
&lt;td&gt;$2.40&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.cometapi.com/models/google/gemini-3-1-pro-preview/" rel="noopener noreferrer"&gt;Gemini 3.1 Pro&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$1.60&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.cometapi.com/models/deepseek/deepseek-v4/" rel="noopener noreferrer"&gt;DeepSeek V4 Pro&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;$0.52&lt;/td&gt;
&lt;td&gt;$0.42&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;a href="https://www.cometapi.com/models/xai/grok-4-2/" rel="noopener noreferrer"&gt;Grok 4.20&lt;/a&gt;&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$1.60&lt;/td&gt;
&lt;td&gt;20%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At 100 million GPT-5.5 input tokens per month, direct billing is approximately $3,000. At the listed gateway rate, the same volume costs $2,400, leaving a $600 monthly difference.&lt;/p&gt;

&lt;p&gt;That is large enough to cover a staging environment or materially reduce the operating cost of a support agent. It also gives teams room to use different models for different request classes instead of forcing every request through the most expensive option.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model Selection by Workload
&lt;/h2&gt;

&lt;p&gt;The catalog covers text, image, video, and audio models. The practical mapping looks like this:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fresource.cometapi.com%2FUnified%2520Model%2520Catalog%2520Overview.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fresource.cometapi.com%2FUnified%2520Model%2520Catalog%2520Overview.webp" title="Unified Model Catalog Overview" alt="500+ AI models available via a unified API" width="799" height="382"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;Models&lt;/th&gt;
&lt;th&gt;Typical Use&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning&lt;/td&gt;
&lt;td&gt;GPT-5.5 Pro, Claude Opus 4.7&lt;/td&gt;
&lt;td&gt;Complex planning and autonomous agents&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agentic coding&lt;/td&gt;
&lt;td&gt;Kimi K2.6, Qwen3.6-Plus&lt;/td&gt;
&lt;td&gt;Repository-scale refactors and “Vibe Coding”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long context&lt;/td&gt;
&lt;td&gt;Grok 4.20, with 2M tokens&lt;/td&gt;
&lt;td&gt;Large logs and document analysis&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multimodal&lt;/td&gt;
&lt;td&gt;Gemini 3.1 Pro, GPT Image 2&lt;/td&gt;
&lt;td&gt;Video analysis and production design&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fast response&lt;/td&gt;
&lt;td&gt;DeepSeek V4 Flash&lt;/td&gt;
&lt;td&gt;High-volume classification&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The service claims a 99.9% availability SLA and an average response time below 400ms. Its routing layer is intended to bypass high-latency nodes when a regional provider is degraded.&lt;/p&gt;

&lt;p&gt;Those figures should still be validated against your own traffic patterns. Latency distributions, streaming behavior, payload size, and model-specific limits matter more than a single average.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Migration Is a Configuration Change
&lt;/h2&gt;

&lt;p&gt;The OpenAI SDK remains the client. The endpoint and credential change; the request structure stays the same.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;APIError&lt;/span&gt;

&lt;span class="c1"&gt;# Step 1: Initialize using the gateway credentials
&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="c1"&gt;# Point to the unified endpoint
&lt;/span&gt;    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="c1"&gt;# Use the gateway token from environment variables
&lt;/span&gt;    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;run_ai_task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-5.5&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Step 2: Swap models without changing the request shape
&lt;/span&gt;        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
            &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;

    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="n"&gt;APIError&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="c1"&gt;# Unified handling for 401, 429, and 500 responses
&lt;/span&gt;        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;API Error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;status_code&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt; - &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Unexpected connection error: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="nf"&gt;str&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="c1"&gt;# Test the switch with a reasoning-heavy model
&lt;/span&gt;&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;run_ai_task&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Analyze our system&lt;/span&gt;&lt;span class="sh"&gt;'&lt;/span&gt;&lt;span class="s"&gt;s scalability benchmarks.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;claude-opus-4-7&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;))&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The important operational detail is that model selection is now data. A request router can choose &lt;code&gt;gpt-5.5&lt;/code&gt;, &lt;code&gt;claude-opus-4-7&lt;/code&gt;, or another compatible model based on latency, task type, cost, or current availability.&lt;/p&gt;

&lt;h3&gt;
  
  
  Avoid Building the Proxy Yourself
&lt;/h3&gt;

&lt;p&gt;One documented migration case involved a mid-sized team maintaining its own internal API proxy. One senior engineer handled SDK updates, provider billing changes, and custom failover routing. That system cost more than $8,000 per month to operate while saving only $300 in API charges.&lt;/p&gt;

&lt;p&gt;The arithmetic was negative before considering the opportunity cost. A managed gateway provides the routing and provider integration without a platform fee, which is generally a better trade when the team’s actual savings are modest.&lt;/p&gt;

&lt;h2&gt;
  
  
  Security and Retention
&lt;/h2&gt;

&lt;p&gt;Using a unified endpoint introduces a governance boundary, so its data policies need to be evaluated like any other third-party dependency. The stated controls are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Prompts and completions are not used to train future model iterations.&lt;/li&gt;
&lt;li&gt;Logs are retained for a maximum of 3 months for debugging, then permanently deleted.&lt;/li&gt;
&lt;li&gt;The platform is SOC 2 certified and uses end-to-end data encryption.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For proprietary code and regulated workloads, I would still verify the current agreement, retention configuration, regional processing behavior, and incident procedures before routing production data.&lt;/p&gt;

&lt;h2&gt;
  
  
  Initial Account Credits
&lt;/h2&gt;

&lt;p&gt;The onboarding offer described for testing is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create a free account with no credit card required.&lt;/li&gt;
&lt;li&gt;Click &lt;strong&gt;Add Token&lt;/strong&gt; in the dashboard to receive a &lt;code&gt;$0.5&lt;/code&gt; bonus.&lt;/li&gt;
&lt;li&gt;Make the first API call to receive an additional &lt;code&gt;$1&lt;/code&gt; credit.&lt;/li&gt;
&lt;li&gt;Deposit &lt;code&gt;$10&lt;/code&gt; for the first time and receive a &lt;code&gt;$3&lt;/code&gt; reward.&lt;/li&gt;
&lt;li&gt;Update &lt;code&gt;base_url&lt;/code&gt; and begin using the listed rates.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The total onboarding balance is approximately &lt;code&gt;$1.50&lt;/code&gt; before the first deposit, which is enough to exercise several models in the Playground.&lt;/p&gt;

&lt;h2&gt;
  
  
  Questions I Would Resolve Before Production
&lt;/h2&gt;

&lt;h3&gt;
  
  
  How long does migration take?
&lt;/h3&gt;

&lt;p&gt;Documented migration cases report approximately 8 minutes for enterprise projects because only &lt;code&gt;base_url&lt;/code&gt; and &lt;code&gt;api_key&lt;/code&gt; change. Messages, &lt;code&gt;temperature&lt;/code&gt;, and streaming logic remain the same. Teams with codebases larger than 150,000 lines have reported their unit tests passing immediately after the configuration change without refactoring.&lt;/p&gt;

&lt;p&gt;That is the compatibility claim; I would still run integration tests against every model and response mode used by the application.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens during an outage?
&lt;/h3&gt;

&lt;p&gt;The design uses multi-region routing. During a major provider fluctuation, requests can be redirected to another region or a comparable model family. This reduces the single-provider failure surface, although it does not eliminate the need for timeouts, retries, circuit breakers, and application-level fallback behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  Are the cheaper models downgraded?
&lt;/h3&gt;

&lt;p&gt;The listed model versions are intended to be identical to their official counterparts, with requests routed to the original providers such as Anthropic or OpenAI. The stated 20% to 40% reduction comes from purchasing hundreds of billions of tokens annually at wholesale rates. There are no monthly subscription fees or hidden platform costs in the described pricing model.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can unused prepaid credit be refunded?
&lt;/h3&gt;

&lt;p&gt;The prepaid system supports refunds for unused balances. Starting with the approximately &lt;code&gt;$1.50&lt;/code&gt; in onboarding credits is the lowest-risk way to test the Playground and model behavior before depositing funds.&lt;/p&gt;

&lt;h3&gt;
  
  
  Are prompts used for training?
&lt;/h3&gt;

&lt;p&gt;The stated enterprise privacy policy says that neither input nor output is used to train models. Logs are retained for 3 months for debugging and then automatically purged, with no possibility of recovery.&lt;/p&gt;

&lt;h3&gt;
  
  
  Are image and video models supported?
&lt;/h3&gt;

&lt;p&gt;Yes. The same API key provides access to multimodal models, including GPT Image 2 for image generation and ByteDance’s Seedance 2.0 for video generation. That makes it possible to build an application spanning text, images, and video without maintaining separate accounts for each provider.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does it support high-concurrency workloads?
&lt;/h3&gt;

&lt;p&gt;The infrastructure is designed for production workloads with thousands of requests per second. Its stated global average latency is under 400ms, and dynamic rate limits are intended to scale with business demand. Actual throughput should be measured using the concurrency, payload sizes, and models that match your workload.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/openai-alternative-cometapi/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=openai-alternative-cometapi"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How to Use Midjourney on Discord in 2026</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Mon, 21 Sep 2026 05:45:21 +0000</pubDate>
      <link>https://dev.to/ryancole1/how-to-use-midjourney-on-discord-in-2026-g37</link>
      <guid>https://dev.to/ryancole1/how-to-use-midjourney-on-discord-in-2026-g37</guid>
      <description>&lt;p&gt;Midjourney remains one of the most powerful and aesthetically acclaimed AI image generators in 2026. While a polished web interface exists at midjourney.com, many creators — especially those who value community feedback, rapid iteration in public channels, and the full suite of Discord-specific commands — continue to prefer the original Discord experience.&lt;/p&gt;

&lt;p&gt;This comprehensive guide walks you through everything: setup, latest V8.1 model features, advanced prompting, parameters, best practices, troubleshooting, comparisons, and &lt;strong&gt;practical recommendations&lt;/strong&gt; for scaling your workflow efficiently via &lt;strong&gt;CometAPI&lt;/strong&gt; on Cometapi.com.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How to Use Midjourney on Discord in 2026: The Ultimate Beginner-to-Pro Guide for Stunning AI Art Creation&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Midjourney remains one of the most powerful and aesthetically acclaimed AI image generators in 2026. While a polished web interface exists at midjourney.com, many creators — especially those who value community feedback, rapid iteration in public channels, and the full suite of Discord-specific commands — continue to prefer the original Discord experience.&lt;/p&gt;

&lt;p&gt;This comprehensive guide walks you through everything: setup, latest V8.1 model features, advanced prompting, parameters, best practices, troubleshooting, comparisons, and &lt;strong&gt;practical recommendations&lt;/strong&gt; for scaling your workflow efficiently via &lt;strong&gt;CometAPI&lt;/strong&gt; on Cometapi.com.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Midjourney on Discord Still Matters in 2026
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Midjourney on Discord offers real-time community interaction, easy upscale/variation buttons (U/V), remix capabilities, and access to the latest V8.1 model with improved sharpness, faster generation (4-5x in standard jobs), HD 2K output, and better prompt adherence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Benefits Over Web-Only:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Live feedback in newbie/general channels.&lt;/li&gt;
&lt;li&gt;Direct bot interaction with slash commands.&lt;/li&gt;
&lt;li&gt;Privacy options via your own server.&lt;/li&gt;
&lt;li&gt;Seamless image referencing and blending.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Latest News (as of May 2026):&lt;/strong&gt; V8.1 (April 30, 2026) brings enhanced aesthetics, better small-detail retention, Raw mode, and HD support across Discord and web. Video generation (image-to-5s clips, extendable to 21s) is maturing.&lt;/p&gt;




&lt;h2&gt;
  
  
  Getting Started: Setting Up Midjourney on Discord
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Create or Log Into a Discord Account
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Download the Discord app (desktop/mobile) or use the web version.&lt;/li&gt;
&lt;li&gt;Sign up with email or use an existing account.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Step 2: Join the Official Midjourney Server
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;In Discord, click the &lt;strong&gt;+&lt;/strong&gt; icon in the server list.&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;Join a Server&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Paste: &lt;code&gt;discord.gg/midjourney&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Join and verify if required.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Tip:&lt;/strong&gt; Start in &lt;strong&gt;#newbies&lt;/strong&gt; or &lt;strong&gt;#general&lt;/strong&gt; channels to observe prompts from thousands of users.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Subscribe to a Plan
&lt;/h3&gt;

&lt;p&gt;Midjourney is subscription-based (no free tier for heavy use in 2026).&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Plan&lt;/th&gt;
&lt;th&gt;Monthly Price&lt;/th&gt;
&lt;th&gt;Annual Price (effective)&lt;/th&gt;
&lt;th&gt;Fast GPU Hours&lt;/th&gt;
&lt;th&gt;Relax Mode&lt;/th&gt;
&lt;th&gt;Stealth Mode&lt;/th&gt;
&lt;th&gt;Max Concurrent Jobs&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Basic&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;td&gt;$8/mo ($96/yr)&lt;/td&gt;
&lt;td&gt;3.3 hrs (~200 images)&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;3 Fast&lt;/td&gt;
&lt;td&gt;Testing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;$30&lt;/td&gt;
&lt;td&gt;$24/mo ($288/yr)&lt;/td&gt;
&lt;td&gt;15 hrs&lt;/td&gt;
&lt;td&gt;Unlimited&lt;/td&gt;
&lt;td&gt;No&lt;/td&gt;
&lt;td&gt;3 Fast/Relax&lt;/td&gt;
&lt;td&gt;Most users&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pro&lt;/td&gt;
&lt;td&gt;$60&lt;/td&gt;
&lt;td&gt;$48/mo ($576/yr)&lt;/td&gt;
&lt;td&gt;30 hrs&lt;/td&gt;
&lt;td&gt;Unlimited&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;12 Fast/3 Relax&lt;/td&gt;
&lt;td&gt;Professionals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mega&lt;/td&gt;
&lt;td&gt;$120&lt;/td&gt;
&lt;td&gt;$96/mo ($1,152/yr)&lt;/td&gt;
&lt;td&gt;60 hrs&lt;/td&gt;
&lt;td&gt;Unlimited&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;12 Fast/3 Relax&lt;/td&gt;
&lt;td&gt;High-volume&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Data Point:&lt;/strong&gt; Standard plan users generate ~15 hours of fast GPU time monthly, sufficient for hundreds of images depending on complexity. Relax mode offers slower but unlimited generations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recommendation:&lt;/strong&gt; Start with Standard for balanced speed and cost. Upgrade for Stealth (private images) if building commercial work.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Your First Image – The /imagine Command
&lt;/h3&gt;

&lt;p&gt;In any channel with the Midjourney bot:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Type &lt;code&gt;/imagine&lt;/code&gt; and press Tab.&lt;/li&gt;
&lt;li&gt;Add your &lt;strong&gt;prompt&lt;/strong&gt; after &lt;code&gt;prompt:&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Press Enter.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Example:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/imagine prompt: a serene mountain lake at dawn, misty fog, photorealistic, cinematic lighting --v 8.1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Midjourney generates a 2x2 grid of 4 images. Use &lt;strong&gt;U1-U4&lt;/strong&gt; to upscale and &lt;strong&gt;V1-V4&lt;/strong&gt; for variations.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Accept Terms of Service&lt;/strong&gt; on first use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Mastering Midjourney V8.1 Features on Discord (2026 Updates)
&lt;/h2&gt;

&lt;p&gt;V8.1 emphasizes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Faster generation&lt;/strong&gt;: Standard jobs 4-5x quicker.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Improved sharpness &amp;amp; detail&lt;/strong&gt; retention.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Better prompt adherence&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;HD 2K support&lt;/strong&gt; (add &lt;code&gt;--hd&lt;/code&gt; or toggle in settings).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Raw mode&lt;/strong&gt; for less stylized, more literal outputs.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Version Parameter:&lt;/strong&gt; &lt;code&gt;--v 8.1&lt;/code&gt; (default in many cases, but specify for consistency).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Niji 7&lt;/strong&gt; for anime/stylized: &lt;code&gt;--niji 7&lt;/code&gt;.&lt;/p&gt;




&lt;h2&gt;
  
  
  Advanced Prompt Engineering: From Beginner to 3000-Word Pro
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Core Structure (Featured Snippet Style):&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Subject&lt;/strong&gt; (main focus).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Details&lt;/strong&gt; (environment, lighting, mood).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Style/Artistic References&lt;/strong&gt; (artist, medium, era).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Technical Parameters&lt;/strong&gt; (--ar, --v, --q, etc.).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Best Practices (Supported by Data):&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Keep it concise&lt;/strong&gt;: Short prompts often outperform long ones. Midjourney processes core concepts best.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Word order matters&lt;/strong&gt;: Important elements first.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use commas&lt;/strong&gt; to separate ideas.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Synonyms over repetition&lt;/strong&gt;: "Gigantic" &amp;gt; "big big".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Weights&lt;/strong&gt;: &lt;code&gt;::2&lt;/code&gt; for emphasis (e.g., &lt;code&gt;red dragon::2&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Negative prompts&lt;/strong&gt;: Use &lt;code&gt;--no&lt;/code&gt; (e.g., &lt;code&gt;--no blur, text&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Prompt Examples by Category
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Photorealistic Portrait:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/imagine prompt: middle-aged Asian woman, thoughtful expression, soft natural window light, detailed skin texture, 85mm lens, f/2.8, cinematic, photorealistic --ar 2:3 --v 8.1 --q 2
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Fantasy Scene:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/imagine prompt: ancient elven city in glowing forest, bioluminescent plants, epic scale, dramatic volumetric lighting, in the style of Studio Ghibli and Alphonse Mucha --ar 16:9 --v 8.1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Product Visualization:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/imagine prompt: sleek wireless earbuds on marble surface, minimalist, studio lighting, octane render, 8k --ar 1:1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Tips for 10x Better Results:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Add &lt;strong&gt;lighting descriptors&lt;/strong&gt;: "golden hour, dramatic chiaroscuro, backlit".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Camera/ Lens&lt;/strong&gt;: "shot on Canon EOS R5, 50mm, shallow depth of field".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Artists&lt;/strong&gt;: "in the style of Greg Rutkowski, Artgerm, WLOP".&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Quality boosters&lt;/strong&gt;: &lt;code&gt;--q 2&lt;/code&gt; (higher quality, more GPU), &lt;code&gt;--stylize 750&lt;/code&gt; (default ~100, higher = more artistic).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Data Insight:&lt;/strong&gt; Users who iterate with V-buttons and Remix achieve 40-60% higher satisfaction rates in community polls (anecdotal from Discord activity).&lt;/p&gt;

&lt;h3&gt;
  
  
  Image Prompts &amp;amp; References
&lt;/h3&gt;

&lt;p&gt;Upload an image, then use its URL in prompt: &lt;code&gt;image_url description --iw 1.5&lt;/code&gt; (image weight).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;--sref&lt;/strong&gt; (style reference) and &lt;strong&gt;--cref&lt;/strong&gt; (character reference) for consistency across generations.&lt;/p&gt;




&lt;h2&gt;
  
  
  Essential Parameters &amp;amp; Commands Table
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Parameter&lt;/th&gt;
&lt;th&gt;Example&lt;/th&gt;
&lt;th&gt;Effect&lt;/th&gt;
&lt;th&gt;Use Case&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;--ar&lt;/td&gt;
&lt;td&gt;--ar 16:9&lt;/td&gt;
&lt;td&gt;Aspect ratio&lt;/td&gt;
&lt;td&gt;Cinematic (16:9), Portrait (2:3)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;--v&lt;/td&gt;
&lt;td&gt;--v 8.1&lt;/td&gt;
&lt;td&gt;Model version&lt;/td&gt;
&lt;td&gt;Latest features&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;--q&lt;/td&gt;
&lt;td&gt;--q 2&lt;/td&gt;
&lt;td&gt;Quality&lt;/td&gt;
&lt;td&gt;Sharper details (higher GPU)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;--stylize&lt;/td&gt;
&lt;td&gt;--stylize 400-1000&lt;/td&gt;
&lt;td&gt;Artistic intensity&lt;/td&gt;
&lt;td&gt;Low = literal, High = creative&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;--chaos&lt;/td&gt;
&lt;td&gt;--chaos 50&lt;/td&gt;
&lt;td&gt;Variation&lt;/td&gt;
&lt;td&gt;More diverse grids&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;--hd&lt;/td&gt;
&lt;td&gt;--hd&lt;/td&gt;
&lt;td&gt;High definition&lt;/td&gt;
&lt;td&gt;2K output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;--no&lt;/td&gt;
&lt;td&gt;--no people&lt;/td&gt;
&lt;td&gt;Negative&lt;/td&gt;
&lt;td&gt;Exclude elements&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;--tile&lt;/td&gt;
&lt;td&gt;--tile&lt;/td&gt;
&lt;td&gt;Seamless patterns&lt;/td&gt;
&lt;td&gt;Textures, wallpapers&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Other Commands:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;/describe&lt;/code&gt; (upload image → prompt suggestions).&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/remix&lt;/code&gt; for prompt editing.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/settings&lt;/code&gt; for defaults.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;/info&lt;/code&gt; for your usage stats.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Best Practices &amp;amp; Workflow Optimization
&lt;/h2&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Create Your Own Server:&lt;/strong&gt; Add Midjourney Bot for private generation (less noise).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Organize Outputs:&lt;/strong&gt; Use folders in Discord or export regularly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Iterate Efficiently:&lt;/strong&gt; Use V1-V4 → Upscale → Remix.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Community Learning:&lt;/strong&gt; Observe top prompts in popular channels.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batch Testing:&lt;/strong&gt; Use &lt;code&gt;--repeat 4&lt;/code&gt; for variations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Commercial Use:&lt;/strong&gt; All plans allow general commercial terms (check TOS for specifics).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Pro Tip:&lt;/strong&gt; Combine with external tools for post-processing (Photoshop, Topaz AI upscaler).&lt;/p&gt;

&lt;h2&gt;
  
  
  Midjourney on Discord vs. Web vs. Competitors
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Discord Strengths:&lt;/strong&gt; Community, speed of interaction, discoverability.&lt;br&gt;
&lt;strong&gt;Web Strengths:&lt;/strong&gt; Cleaner UI, easier organization.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Vs. Competitors:&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Tool&lt;/th&gt;
&lt;th&gt;Artistic Quality&lt;/th&gt;
&lt;th&gt;Prompt Adherence&lt;/th&gt;
&lt;th&gt;Text in Images&lt;/th&gt;
&lt;th&gt;Pricing (Entry)&lt;/th&gt;
&lt;th&gt;Speed&lt;/th&gt;
&lt;th&gt;Best For&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Midjourney&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Fair&lt;/td&gt;
&lt;td&gt;$10/mo&lt;/td&gt;
&lt;td&gt;Fast (V8.1)&lt;/td&gt;
&lt;td&gt;Art, Concepts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flux 1.1 Pro&lt;/td&gt;
&lt;td&gt;Very Good&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;API ~$0.02/img&lt;/td&gt;
&lt;td&gt;Very Fast&lt;/td&gt;
&lt;td&gt;Photorealism&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DALL-E 3&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Excellent&lt;/td&gt;
&lt;td&gt;Via ChatGPT&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;td&gt;Precision/Text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ideogram&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Very Good&lt;/td&gt;
&lt;td&gt;Best&lt;/td&gt;
&lt;td&gt;Subscription&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;td&gt;Typography&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Midjourney leads in "wow factor" and stylized output.&lt;/p&gt;

&lt;h2&gt;
  
  
  Scaling Beyond Discord: CometAPI Recommendations for Developers &amp;amp; Power Users
&lt;/h2&gt;

&lt;p&gt;Discord is fantastic for exploration, but for &lt;strong&gt;production workflows&lt;/strong&gt;, websites, apps, or high-volume needs, manual prompting becomes a bottleneck.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Enter CometAPI (Cometapi.com):&lt;/strong&gt; A reliable unofficial Midjourney API provider offering unified access to Midjourney (and 500+ other models) through clean REST endpoints.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why Integrate via CometAPI?&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Programmatic Access:&lt;/strong&gt; Generate images from code (Python, Node.js, etc.) without Discord.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cost Efficiency:&lt;/strong&gt; Pay-per-use or subscription models often cheaper than high-tier Midjourney plans for bulk.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Automation:&lt;/strong&gt; Build apps, e-commerce mockups, marketing tools, or SaaS features.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reliability:&lt;/strong&gt; Handles queuing, retries, and multiple models in one dashboard.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ease:&lt;/strong&gt; Simple API keys, detailed docs, and Apifox testing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;How to Get Started with CometAPI Midjourney:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Visit &lt;a href="https://www.cometapi.com" rel="noopener noreferrer"&gt;Cometapi.com&lt;/a&gt; and sign up.&lt;/li&gt;
&lt;li&gt;Generate API key in console.&lt;/li&gt;
&lt;li&gt;Call the Midjourney endpoint with your prompt (supports parameters like --ar, --v 8.1).&lt;/li&gt;
&lt;li&gt;Receive image URLs instantly or via webhook.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Use Cases:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dynamic product imagery for e-commerce.&lt;/li&gt;
&lt;li&gt;Automated social media content.&lt;/li&gt;
&lt;li&gt;AI art pipelines in design tools.&lt;/li&gt;
&lt;li&gt;Enterprise bulk generation with consistent styling.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Pro Recommendation:&lt;/strong&gt; Use Discord for creative ideation and discovery, then pipe refined prompts into CometAPI for scalable production. This hybrid approach maximizes quality and efficiency.&lt;/p&gt;

&lt;h2&gt;
  
  
  Troubleshooting Common Issues
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Rate Limits:&lt;/strong&gt; Wait or upgrade plan.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Blurred Images:&lt;/strong&gt; Increase --q or use HD.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Censorship:&lt;/strong&gt; Avoid sensitive content (strict filters).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Slow Generation:&lt;/strong&gt; Switch to Relax mode or check server status.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Bot Not Responding:&lt;/strong&gt; Ensure you're in a channel with permissions; try DM with bot.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Future Outlook &amp;amp; Tips for 2026 Success
&lt;/h2&gt;

&lt;p&gt;With V8.1 and upcoming V9/ video enhancements, Midjourney continues evolving. Focus on prompt mastery, consistent character references, and ethical use.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Actionable Next Steps:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Join Discord server today.&lt;/li&gt;
&lt;li&gt;Experiment with 10 prompts using V8.1.&lt;/li&gt;
&lt;li&gt;Sign up at Cometapi.com for API access and scale your creations.&lt;/li&gt;
&lt;li&gt;Track updates on midjourney.com/updates.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;By combining Discord's creative power with CometAPI's programmatic capabilities, you'll unlock professional-grade AI imagery workflows that save time and elevate output.&lt;/p&gt;

&lt;p&gt;Ready to create? Head to Discord or &lt;a href="https://www.cometapi.com" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt; and start generating today!&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-use-midjourney-on-discord-in-2026/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-use-midjourney-on-discord-in-2026"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How to Update Gemini CLI to the Latest Version: Complete Guide, New Features &amp; Pro Tips</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Mon, 21 Sep 2026 04:35:33 +0000</pubDate>
      <link>https://dev.to/ryancole1/how-to-update-gemini-cli-to-the-latest-version-complete-guide-new-features-pro-tips-2fa2</link>
      <guid>https://dev.to/ryancole1/how-to-update-gemini-cli-to-the-latest-version-complete-guide-new-features-pro-tips-2fa2</guid>
      <description>&lt;p&gt;Gemini CLI has rapidly evolved into one of the most powerful open-source AI agents for developers. By bringing Google's Gemini models directly into your terminal, it enables coding, debugging, deployment, data analysis, and complex agentic workflows without leaving your command line.&lt;/p&gt;

&lt;p&gt;As of May 2026, the latest &lt;strong&gt;stable release is v0.40.0&lt;/strong&gt; (April 28, 2026), with preview and nightly channels delivering even more experimental features. Regular updates bring critical improvements in offline capabilities, agent skills, resource management, themes, interactivity, and integration with the latest Gemini models like Gemini 3.x series.&lt;/p&gt;

&lt;p&gt;Failing to update can mean missing out on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Enhanced security and stability fixes&lt;/li&gt;
&lt;li&gt;New sub-agents and parallel task handling&lt;/li&gt;
&lt;li&gt;Better context management and MCP (Model Context Protocol) support&lt;/li&gt;
&lt;li&gt;Improved performance and lower latency&lt;/li&gt;
&lt;li&gt;Accessibility features like colorblind themes&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Is Gemini CLI? A Quick Overview
&lt;/h2&gt;

&lt;p&gt;Gemini CLI is Google's open-source AI agent that turns your terminal into a powerful reasoning and acting (ReAct) environment powered by Gemini models. It supports built-in tools, local/remote MCP servers, interactive shell commands, custom slash commands, and agentic workflows.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Capabilities:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agent Mode&lt;/strong&gt;: Multi-step planning, tool use, and execution.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Interactive Shell&lt;/strong&gt;: Run vim, top, or other interactive programs seamlessly.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Context Management&lt;/strong&gt;: GEMINI.md files, codebase ingestion.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Extensibility&lt;/strong&gt;: Custom tools, sub-agents, IDE plugins.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model Access&lt;/strong&gt;: Automatic updates to latest Gemini models (including experimental ones).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;It is free and open-source, with optional paid Google AI subscriptions for higher quotas.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Gemini CLI updates matter more now
&lt;/h2&gt;

&lt;p&gt;Gemini CLI is not a tiny utility you install once and forget. Google describes it as an open-source AI agent that brings Gemini directly into the terminal, with support for code understanding, file operations, shell commands, web fetching, and MCP-based integrations. The GitHub project also highlights a free tier for personal Google accounts, Gemini 3 model support, and a 1M token context window, so new releases can affect both capabilities and usage limits.&lt;/p&gt;

&lt;p&gt;That matters because Gemini CLI has been moving quickly in 2026. The official release notes separate nightly, preview, and stable channels, and they explicitly recommend the stable release for most users. The latest release notes also show active feature work such as offline search support, GitHub-style themes, MCP resource tools, and a newer memory-management approach.&lt;/p&gt;

&lt;p&gt;The practical takeaway is simple: if you use Gemini CLI daily, updates are not just about bug fixes. They can change model routing, auth behavior, available tools, and how safely the CLI behaves in more automated environments.&lt;/p&gt;

&lt;h2&gt;
  
  
  The latest Gemini CLI news you should know before updating
&lt;/h2&gt;

&lt;p&gt;Gemini CLI updates frequently across three channels: &lt;strong&gt;Stable&lt;/strong&gt; (recommended), &lt;strong&gt;Preview&lt;/strong&gt;, and &lt;strong&gt;Nightly&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  1) Google changed Gemini CLI service behavior in March 2026
&lt;/h3&gt;

&lt;p&gt;On March 18, 2026, the Gemini CLI team announced service changes that would add more robust abuse detection and prioritize traffic differently based on license type and account standing. The same update said that, starting March 25, 2026, free-tier users would be limited to Gemini Flash models, while Gemini Pro models would require paid subscriptions. Google also reminded users that they can regain more direct control of quotas and billing by using their own paid API key through AI Studio or Vertex AI.&lt;/p&gt;

&lt;p&gt;For readers, that means “updating Gemini CLI” is now partly a product-policy issue, not just a software-version issue. Two users on the same version may still experience different behavior depending on account type, model choice, and traffic conditions.&lt;/p&gt;

&lt;h3&gt;
  
  
  2) The April 2026 release stream added more utility
&lt;/h3&gt;

&lt;p&gt;The April 28, 2026 release notes for v0.40.0 describe several visible improvements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Offline Search Support&lt;/strong&gt;: Bundled ripgrep for fast local codebase searching without internet.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GitHub-Style Colorblind Themes&lt;/strong&gt;: Improved accessibility and customization.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Advanced MCP Resource &amp;amp; Memory Management&lt;/strong&gt;: New resource tools for better handling of external contexts and tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Improved Narrative Flow &amp;amp; UI/UX&lt;/strong&gt;: Smoother interactions and storytelling in agent responses.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Streamlined Local Model Support&lt;/strong&gt;: Easier integration with on-device or self-hosted models.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Recent preview/nightly additions include sub-agents for parallel workflows (around v0.36+), enhanced plan mode with review steps, tab autocomplete, notifications, and better error handling for transient issues.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why These Matter:&lt;/strong&gt; Developers report 2-5x productivity gains in complex tasks like bug fixing, deployment, and data pipelines due to these features. Sub-agents help prevent context overload by delegating subtasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  3) Security reporting has made update hygiene more important
&lt;/h3&gt;

&lt;p&gt;A recent security report warned that malicious campaigns were impersonating the official Gemini CLI, including fake websites, cloned repositories, misleading social posts, and typosquatted npm packages. The safest response is to use only official sources and verify the package name before installing or updating.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to update Gemini CLI the official way
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The built-in update command
&lt;/h3&gt;

&lt;p&gt;The official Gemini CLI cheatsheet lists &lt;code&gt;gemini update&lt;/code&gt; as the command to update to the latest version. It also lists &lt;code&gt;--version&lt;/code&gt; / &lt;code&gt;-v&lt;/code&gt; as the flag for showing the current CLI version, which is the quickest way to confirm whether the update worked.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;gemini update
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;





&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;gemini &lt;span class="nt"&gt;--version&lt;/span&gt;&lt;span class="c"&gt;# orgemini -v&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The same cheatsheet also shows that Gemini CLI is designed as a command-line app with REPL mode, prompt mode, resume mode, and extension/MCP management built in, so version upgrades can affect both the core assistant and the surrounding workflow commands.&lt;/p&gt;

&lt;h3&gt;
  
  
  Install and reinstall paths you may also see
&lt;/h3&gt;

&lt;p&gt;Gemini CLI can be installed globally with npm using &lt;code&gt;npm install -g @google/gemini-cli&lt;/code&gt;. Quick-install options via &lt;code&gt;npx @google/gemini-cli&lt;/code&gt;, global npm installation, and Homebrew on macOS/Linux.&lt;/p&gt;

&lt;p&gt;That gives you a useful rule of thumb: if your install is healthy, use the built-in &lt;code&gt;gemini update&lt;/code&gt; path. If the install itself is damaged, inconsistent, or stuck between package managers, a clean reinstall from the official package source is often the safer recovery route. That second sentence is an operational recommendation, not a direct vendor claim.&lt;/p&gt;

&lt;h3&gt;
  
  
  Release-channel choice matters
&lt;/h3&gt;

&lt;p&gt;Gemini CLI’s official release notes define three channels: nightly, preview, and stable. Nightly contains the most recent changes, preview is for experimental features and early feedback, and stable is the recommended option for general use.&lt;/p&gt;

&lt;p&gt;If you want to minimize surprises, the stable channel is the right default. If you publish content, manage teams, or rely on Gemini CLI in automation, stable should be your baseline and preview should be treated as a test lane.&lt;/p&gt;

&lt;h2&gt;
  
  
  Comparison table: the best ways to update Gemini CLI
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Update path&lt;/th&gt;
&lt;th&gt;Best for&lt;/th&gt;
&lt;th&gt;What you do&lt;/th&gt;
&lt;th&gt;Notes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;gemini update&lt;/td&gt;
&lt;td&gt;Most users on a healthy install&lt;/td&gt;
&lt;td&gt;Run the built-in CLI update command, then confirm with gemini --version / gemini -v.&lt;/td&gt;
&lt;td&gt;This is the most direct, official update path.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fresh reinstall via official package&lt;/td&gt;
&lt;td&gt;Broken or inconsistent installs&lt;/td&gt;
&lt;td&gt;Use the official package route shown in the docs: npm install -g @google/gemini-cli. The README also lists npx and Homebrew options.&lt;/td&gt;
&lt;td&gt;Best when your install method is tangled or your package manager state is unreliable.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;ACP Agent Registry inside an IDE&lt;/td&gt;
&lt;td&gt;JetBrains, Zed, or other ACP-compatible IDE users&lt;/td&gt;
&lt;td&gt;Gemini CLI is officially available in the ACP Agent Registry, which lets supported IDEs install and update it directly.&lt;/td&gt;
&lt;td&gt;Great for teams that prefer updating from inside their editor.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Switch release channels&lt;/td&gt;
&lt;td&gt;Testers and early adopters&lt;/td&gt;
&lt;td&gt;Use the release notes to decide between nightly, preview, and stable. The docs recommend stable for general use.&lt;/td&gt;
&lt;td&gt;Nightly and preview move faster, but they can be noisier.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  How to Update Gemini CLI: Step-by-Step Guide
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Check Your Current Version
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Bash
gemini &lt;span class="nt"&gt;--version&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Bash
gemini &lt;span class="nt"&gt;-v&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This confirms your installation and version.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Update Methods
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Recommended for Most Users (Global Install):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Bash
npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @google/gemini-cli@latest
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Specific Update Command:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Bash
npm update &lt;span class="nt"&gt;-g&lt;/span&gt; @google/gemini-cli
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Preview Channel (Experimental Features):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Bash
npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @google/gemini-cli@preview
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Nightly Channel:&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Bash
npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @google/gemini-cli@nightly
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Without Installation (npx – Always Latest):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;Bash
npx https://github.com/google-gemini/gemini-cli
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is ideal for testing or one-off use.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Post-Update Verification and Restart
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Close and reopen your terminal.&lt;/li&gt;
&lt;li&gt;Run gemini --version again.&lt;/li&gt;
&lt;li&gt;Start a new session: gemini and sign in if prompted (Google account or API key).&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Pro Tip:&lt;/strong&gt; Use npm install -g @google/gemini-cli@latest --force if you encounter permission or cache issues.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Verify your account and release channel
&lt;/h3&gt;

&lt;p&gt;The current docs note that most individual users can sign in with a personal Google account, while organizations and some enterprise setups may need a Google Cloud project or different auth path. The same docs recommend starting Gemini CLI and logging in with a Google account for the simplest local workflow.&lt;/p&gt;

&lt;p&gt;That matters because a successful upgrade may still feel “wrong” if your account type changed, your quota changed, or the CLI now routes you to a different model set. The March 2026 service update made those distinctions more visible.&lt;/p&gt;

&lt;h2&gt;
  
  
  Troubleshooting when the Gemini CLI update fails
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Mixed package managers can cause messy installs
&lt;/h3&gt;

&lt;p&gt;Update loops and PATH conflicts when legacy installs, npm, and pnpm get mixed together. Community issue reports also describe auto-update confusion when pnpm installations are detected as npm installs. The lesson is not that Gemini CLI is broken; the lesson is that package-manager consistency matters.&lt;/p&gt;

&lt;p&gt;If &lt;code&gt;gemini update&lt;/code&gt; does not behave as expected, check how the tool was originally installed, and avoid mixing global installs across npm, pnpm, and other package managers on the same machine.&lt;/p&gt;

&lt;h3&gt;
  
  
  Some environments have hit installation or runtime compatibility issues
&lt;/h3&gt;

&lt;p&gt;There are also historical issue reports about Node.js version compatibility and PATH problems after upgrades. If you manage a team environment, it is smart to standardize Node.js versions and document the exact install method alongside the CLI version.&lt;/p&gt;

&lt;h3&gt;
  
  
  Security-first update hygiene is not optional anymore
&lt;/h3&gt;

&lt;p&gt;Because recent reporting has highlighted fake Gemini CLI download campaigns, always verify that you are using the official &lt;code&gt;google-gemini/gemini-cli&lt;/code&gt; repository or the official Gemini CLI documentation before you update. The official repo and docs are the safest anchor points; random “early access” installers are not.&lt;/p&gt;

&lt;h3&gt;
  
  
  Troubleshooting Common Update Issues
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Permission Errors&lt;/strong&gt;: Use sudo npm install -g ... (macOS/Linux) or run PowerShell as Administrator (Windows). Better: Use nvm or fix npm permissions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Update Not Applying&lt;/strong&gt;: Clear npm cache: npm cache clean --force, then reinstall.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Version Mismatch&lt;/strong&gt;: Ensure you're using a fresh terminal session.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Proxy/Firewall&lt;/strong&gt;: Configure npm proxy settings.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Node.js Version&lt;/strong&gt;: Require Node.js 18+ (recommend 20+).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Broken Install&lt;/strong&gt;: Uninstall first: npm uninstall -g @google/gemini-cli, then reinstall.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Monitor GitHub issues for platform-specific bugs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Advanced Configuration After Updating
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Authentication&lt;/strong&gt;: gemini → Sign in with Google for higher quotas or use Gemini API key.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GEMINI.md&lt;/strong&gt;: Create in project root for persistent context and custom instructions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Slash Commands &amp;amp; Custom Tools&lt;/strong&gt;: Extend functionality.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;MCP Servers&lt;/strong&gt;: Connect local/remote tools for enhanced capabilities.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Themes &amp;amp; Settings&lt;/strong&gt;: Customize via config for accessibility.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plan Mode&lt;/strong&gt;: Enable for safer multi-step executions with review.&lt;/li&gt;
&lt;/ul&gt;




&lt;h2&gt;
  
  
  Comparison Table: Gemini CLI Update Channels
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature/Aspect&lt;/th&gt;
&lt;th&gt;Stable (v0.40.0)&lt;/th&gt;
&lt;th&gt;Preview&lt;/th&gt;
&lt;th&gt;Nightly&lt;/th&gt;
&lt;th&gt;npx (No Install)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Stability&lt;/td&gt;
&lt;td&gt;High (Recommended)&lt;/td&gt;
&lt;td&gt;Medium-High&lt;/td&gt;
&lt;td&gt;Low (Experimental)&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latest Features&lt;/td&gt;
&lt;td&gt;Balanced&lt;/td&gt;
&lt;td&gt;Early Access&lt;/td&gt;
&lt;td&gt;Cutting-Edge&lt;/td&gt;
&lt;td&gt;Always Latest&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Update Frequency&lt;/td&gt;
&lt;td&gt;Weekly/Monthly&lt;/td&gt;
&lt;td&gt;Frequent&lt;/td&gt;
&lt;td&gt;Daily&lt;/td&gt;
&lt;td&gt;On-demand&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Use Case&lt;/td&gt;
&lt;td&gt;Production/ Daily Work&lt;/td&gt;
&lt;td&gt;Testing New Features&lt;/td&gt;
&lt;td&gt;Developers/Contributors&lt;/td&gt;
&lt;td&gt;Quick Tests&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quota/Performance&lt;/td&gt;
&lt;td&gt;Optimized&lt;/td&gt;
&lt;td&gt;May vary&lt;/td&gt;
&lt;td&gt;Variable&lt;/td&gt;
&lt;td&gt;Same as installed&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Risk&lt;/td&gt;
&lt;td&gt;Lowest&lt;/td&gt;
&lt;td&gt;Moderate&lt;/td&gt;
&lt;td&gt;Highest&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;




&lt;h2&gt;
  
  
  Real-World Use Cases and Productivity Gains
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Code Generation &amp;amp; Refactoring&lt;/strong&gt;: Ingest entire repos via context.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Debugging &amp;amp; Deployment&lt;/strong&gt;: Interactive shell + agents for Cloud Run, etc.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data Analysis&lt;/strong&gt;: Combine with local tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Content &amp;amp; Research&lt;/strong&gt;: Offline search + Gemini reasoning.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agentic Workflows&lt;/strong&gt;: Sub-agents handle parallel tasks like testing, docs, and deployment.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Users report significant time savings, with features like plan mode reducing errors.&lt;/p&gt;




&lt;h2&gt;
  
  
  Integrating with APIs: Why CometAPI Complements Gemini CLI Perfectly
&lt;/h2&gt;

&lt;p&gt;While Gemini CLI excels in terminal interactivity, pairing it with a unified API provider like &lt;strong&gt;CometAPI&lt;/strong&gt; unlocks production-scale advantages:&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Key Benefits of CometAPI for Gemini Workflows:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cost Savings&lt;/strong&gt;: Up to 20%+ lower prices than direct Google Gemini APIs (e.g., Gemini 2.5 Pro at competitive input/output rates).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Unified Access&lt;/strong&gt;: One API key for Gemini models + others (GPT, Claude, etc.) – seamless switching.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;High Reliability &amp;amp; Speed&lt;/strong&gt;: Optimized routing, reduced latency for CLI-augmented scripts.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Easy Integration&lt;/strong&gt;: Standard OpenAI-compatible format. Call Gemini models from scripts or custom MCP tools via CometAPI endpoint (&lt;code&gt;https://api.cometapi.com/&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Scalability&lt;/strong&gt;: Higher rate limits and enterprise features without managing multiple keys.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Example Integration in Your Workflow:&lt;/strong&gt; Use Gemini CLI for interactive sessions, but route heavy batch jobs or custom agents through CometAPI SDKs/scripts for cost efficiency and reliability.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Recommendation for CometAPI.com Readers:&lt;/strong&gt; Sign up at &lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;CometAPI.com&lt;/a&gt;, grab your key, and configure it in custom tools or external scripts. It's the smartest way to scale Gemini-powered development beyond the terminal while keeping costs low and performance high. Whether building AI apps, automating pipelines, or experimenting with agents, CometAPI ensures you maximize value from the latest Gemini models.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion: Stay Ahead with Updated Gemini CLI + CometAPI
&lt;/h2&gt;

&lt;p&gt;Updating Gemini CLI to v0.40.0+ is straightforward yet unlocks transformative terminal AI capabilities. With rapid releases focused on agents, context, and usability, it's a must-have for modern developers.&lt;/p&gt;

&lt;p&gt;For optimal results, combine the interactive power of Gemini CLI with the cost-effective, unified API access from &lt;strong&gt;CometAPI&lt;/strong&gt;. This hybrid approach delivers the best of both worlds: seamless local workflows and scalable, affordable cloud intelligence.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Action Steps Today:&lt;/strong&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Update now: npm install -g @google/gemini-cli@latest&lt;/li&gt;
&lt;li&gt;Explore new features in v0.40.0&lt;/li&gt;
&lt;li&gt;Visit &lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt; for Gemini API integration and savings&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Stay productive, innovate faster, and build the future from your terminal.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-update-gemini-cli/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-update-gemini-cli"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Claude Sonnet 5: What the Leaks Mean for Developers</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Mon, 21 Sep 2026 04:06:37 +0000</pubDate>
      <link>https://dev.to/ryancole1/claude-sonnet-5-what-the-leaks-mean-for-developers-l22</link>
      <guid>https://dev.to/ryancole1/claude-sonnet-5-what-the-leaks-mean-for-developers-l22</guid>
      <description>&lt;p&gt;As of June 23, 2026, Anthropic has not announced Claude Sonnet 5. Nevertheless, the model identifier &lt;code&gt;claude-sonnet-5&lt;/code&gt; has reportedly appeared in internal configurations, error logs, Claude tooling, and partner developer platforms.&lt;/p&gt;

&lt;p&gt;That is not launch confirmation, but it usually means release preparation is well underway. Current speculation points to late June 2026, potentially June 24 or the week beginning June 29.&lt;/p&gt;

&lt;p&gt;The timing matters. Anthropic recently launched Fable 5 and Mythos 5, models positioned above Opus for demanding reasoning and agentic work, before suspending access worldwide after a US export control directive cited national security concerns related to a reported jailbreak vulnerability. The suspension affected domestic users as well. Anthropic has not said that development stopped; the accurate statement is that access was suspended.&lt;/p&gt;

&lt;p&gt;That leaves the Sonnet tier in an especially important position: capable enough for serious development work, but generally fast and economical enough to deploy at scale.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Sonnet Fits in Anthropic's Lineup
&lt;/h2&gt;

&lt;p&gt;Anthropic's current model tiers remain fairly straightforward:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Haiku 4.5&lt;/strong&gt;: optimized for speed.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sonnet 4.6&lt;/strong&gt;: the general-purpose balance of intelligence, latency, and cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Opus 4.8&lt;/strong&gt;: the highest-end option for difficult reasoning and long-running work.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Fable 5 and Mythos 5 were introduced as a more capable class above Opus, but their availability is currently suspended. Sonnet 4.6 therefore remains the practical default for many coding and agent workloads, while Opus 4.8 handles tasks where failure is expensive or the reasoning horizon is unusually long.&lt;/p&gt;

&lt;p&gt;Sonnet models have historically been popular because they deliver much of the capability developers need without Opus-level pricing and latency.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Developer-Platform Sightings Suggest
&lt;/h2&gt;

&lt;p&gt;Recent reports from June 22-23 mention &lt;code&gt;claude-sonnet-5&lt;/code&gt; in Anthropic configurations, Claude applications and tools, and cloud partner environments such as Vertex AI or similar platforms.&lt;/p&gt;

&lt;p&gt;There was also an earlier reference to Sonnet 5, reportedly under the codename Fennec, with dated identifiers such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;claude-sonnet-5-20260203
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The evidence being discussed publicly includes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Model references in partner infrastructure error logs and model lists.&lt;/li&gt;
&lt;li&gt;Community posts and X discussions showing “Sonnet 5” labels.&lt;/li&gt;
&lt;li&gt;A release cadence consistent with Anthropic's recent pattern of shipping major models every few weeks or months.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;None of this establishes a public release date. It does indicate that the model may have progressed beyond an internal research artifact.&lt;/p&gt;

&lt;h2&gt;
  
  
  Expected Release and Availability
&lt;/h2&gt;

&lt;p&gt;The current rumor window is late June 2026, with June 24 often cited as the earliest possible date. Another version of the rumor points to the week beginning June 29. Anthropic has not confirmed either date.&lt;/p&gt;

&lt;p&gt;If the model ships, the likely distribution path is Claude.ai, the Anthropic API, and cloud partners such as AWS Bedrock and Vertex AI. The exact rollout will depend on Anthropic's internal testing and safety checks.&lt;/p&gt;

&lt;p&gt;The suspension of Fable 5 and Mythos 5 is a useful reminder that technical readiness and public availability are separate questions. A model can appear in infrastructure before Anthropic is prepared to expose it broadly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Likely Areas of Improvement
&lt;/h2&gt;

&lt;p&gt;Everything below is projection. Sonnet 5 has not been officially launched, so its capabilities, pricing, and benchmarks remain unconfirmed.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. More Reliable Coding Agents
&lt;/h3&gt;

&lt;p&gt;Sonnet 4.6 improved codebase understanding, bug fixing, long-session consistency, and instruction following. Anthropic reported that early Claude Code users preferred Sonnet 4.6 over Sonnet 4.5 approximately 70% of the time, and preferred it over Opus 4.5 approximately 59% of the time.&lt;/p&gt;

&lt;p&gt;For Sonnet 5, I would expect the emphasis to remain on agentic software development:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Reading and navigating large repositories.&lt;/li&gt;
&lt;li&gt;Planning changes across multiple files.&lt;/li&gt;
&lt;li&gt;Calling tools consistently.&lt;/li&gt;
&lt;li&gt;Avoiding unnecessary rewrites.&lt;/li&gt;
&lt;li&gt;Running verification steps before reporting completion.&lt;/li&gt;
&lt;li&gt;Maintaining context over long coding sessions.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This is where the Sonnet tier has commercial leverage. Opus may provide deeper reasoning, but Sonnet is often the model teams can afford to call repeatedly inside production workflows.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. A Smaller Gap With Opus
&lt;/h3&gt;

&lt;p&gt;Anthropic's official pricing currently lists:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Claude Opus 4.8&lt;/strong&gt;: $5 per million input tokens and $25 per million output tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Claude Sonnet 4.6&lt;/strong&gt;: $3 per million input tokens and $15 per million output tokens.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If Sonnet 5 keeps the existing Sonnet pricing tier while improving its reasoning quality, it could become a particularly strong option for coding agents, enterprise automation, internal tools, and research assistants.&lt;/p&gt;

&lt;p&gt;The useful metric will not be benchmark score alone. Teams should measure cost per successful task, including retries, human corrections, and failed tool calls.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Better Use of Long Context
&lt;/h3&gt;

&lt;p&gt;Claude Sonnet 4.6 introduced a 1M-token context window in beta. Anthropic's pricing documentation states that Fable 5, Mythos 5, Opus 4.8, Opus 4.7, Opus 4.6, and Sonnet 4.6 include the full 1M-token context window at standard pricing.&lt;/p&gt;

&lt;p&gt;Sonnet 5 will likely continue in that direction, potentially offering a 1M+ token context window. The important improvement would not simply be the maximum token count. Long-context systems are useful only when the model can retrieve and apply information buried deep in a repository, contract, research archive, or support history.&lt;/p&gt;

&lt;p&gt;Better attention across long inputs and fewer missed details would matter more than a larger headline number.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. More Capable Computer and Tool Use
&lt;/h3&gt;

&lt;p&gt;Sonnet 4.6 made substantial progress on computer use, web tasks, spreadsheet navigation, and multi-step workflows. Opus 4.8 extended the agentic direction with dynamic workflows in Claude Code, including planning large tasks and running many parallel subagents.&lt;/p&gt;

&lt;p&gt;Sonnet 5 may bring some of that behavior to a cheaper and faster model. The likely target workloads include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Browser-based agents.&lt;/li&gt;
&lt;li&gt;Internal administration tools.&lt;/li&gt;
&lt;li&gt;Structured function and tool calling.&lt;/li&gt;
&lt;li&gt;Spreadsheet and document workflows.&lt;/li&gt;
&lt;li&gt;Software navigation.&lt;/li&gt;
&lt;li&gt;Automations that must interact with real systems rather than generate text alone.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  5. More Honest Self-Checking
&lt;/h3&gt;

&lt;p&gt;Anthropic's Opus 4.8 launch highlighted a claimed improvement in self-verification: Opus 4.8 was around four times less likely than its predecessor to overlook flaws in code it had written.&lt;/p&gt;

&lt;p&gt;That behavior is valuable in production. A model that reports uncertainty, asks for missing context, or admits that tests failed is more useful than one that confidently produces an incorrect result.&lt;/p&gt;

&lt;p&gt;If Sonnet 5 inherits some of Opus 4.8's self-checking behavior while retaining Sonnet-level cost and speed, the practical improvement could be larger than a raw benchmark gain.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sonnet 5 Versus Opus 4.8
&lt;/h2&gt;

&lt;p&gt;The following comparison combines current data for Sonnet 4.6 and Opus 4.8 with projections for Sonnet 5:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;Claude Sonnet 5 (expected)&lt;/th&gt;
&lt;th&gt;Claude Opus 4.8&lt;/th&gt;
&lt;th&gt;Practical takeaway&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Intelligence tier&lt;/td&gt;
&lt;td&gt;Mid-tier, balanced&lt;/td&gt;
&lt;td&gt;Frontier, highest&lt;/td&gt;
&lt;td&gt;Opus for maximum complexity; Sonnet may narrow the gap&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;SWE-Bench Verified&lt;/td&gt;
&lt;td&gt;~82%+ projected&lt;/td&gt;
&lt;td&gt;~80-81% for related Opus results&lt;/td&gt;
&lt;td&gt;Sonnet could lead on value&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context window&lt;/td&gt;
&lt;td&gt;1M+ tokens expected&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;td&gt;Roughly equivalent, with a possible Sonnet edge&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input/output pricing per MTok&lt;/td&gt;
&lt;td&gt;~$3 / $15 expected&lt;/td&gt;
&lt;td&gt;$5 / $25&lt;/td&gt;
&lt;td&gt;Sonnet offers substantial savings&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Fast expected&lt;/td&gt;
&lt;td&gt;Moderate&lt;/td&gt;
&lt;td&gt;Sonnet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding and agents&lt;/td&gt;
&lt;td&gt;Excellent, potentially with parallel agents&lt;/td&gt;
&lt;td&gt;Superior for long-horizon work&lt;/td&gt;
&lt;td&gt;Task-dependent&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vision and multimodal work&lt;/td&gt;
&lt;td&gt;Improved diagrams rumored&lt;/td&gt;
&lt;td&gt;Strong&lt;/td&gt;
&lt;td&gt;Unconfirmed Sonnet advantage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best fit&lt;/td&gt;
&lt;td&gt;Daily development, agents, high-volume workloads&lt;/td&gt;
&lt;td&gt;Complex research and autonomous work&lt;/td&gt;
&lt;td&gt;Sonnet for ROI, Opus for depth&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For many coding, analysis, and agent workflows, Sonnet 5 could provide better price-performance. The source estimate is that it may cover 70-80% of common workflows, leaving Opus 4.8 for tasks that require maximum reasoning depth.&lt;/p&gt;

&lt;p&gt;That estimate should be validated against real workloads. Blind preferences and benchmark results do not always predict production behavior.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I Would Route Workloads
&lt;/h2&gt;

&lt;p&gt;Opus 4.8 is currently the safer choice for the hardest tasks because it is official, documented, and available. Anthropic recommends it for complex reasoning, long-horizon agentic coding, and high-autonomy work.&lt;/p&gt;

&lt;p&gt;If Sonnet 5 launches, I would treat it as the likely production default for high-volume tasks:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Code review assistance.&lt;/li&gt;
&lt;li&gt;Customer-support copilots.&lt;/li&gt;
&lt;li&gt;Data extraction.&lt;/li&gt;
&lt;li&gt;Research summarization.&lt;/li&gt;
&lt;li&gt;Routine code generation and explanation.&lt;/li&gt;
&lt;li&gt;Workflow automation.&lt;/li&gt;
&lt;li&gt;Agent orchestration.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The routing policy would look something like this:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Use Sonnet 4.6 or Sonnet 5 for routine, high-volume operations.&lt;/li&gt;
&lt;li&gt;Escalate architecture reviews, risky code changes, financial analysis, legal reasoning, and long-running autonomous work to Opus 4.8.&lt;/li&gt;
&lt;li&gt;Test Fable-class models only when access, policy, and safeguards permit.&lt;/li&gt;
&lt;li&gt;Run A/B evaluations before changing the default model.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A unified multi-model API such as CometAPI can reduce the infrastructure work involved in comparing Claude, GPT, Gemini, and other providers, but the evaluation layer still belongs in your application. Routing without measurement is just guesswork.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prepare Before the Model Ships
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Establish a Real Baseline
&lt;/h3&gt;

&lt;p&gt;Capture results from Claude Sonnet 4.6 and Claude Opus 4.8 before testing Sonnet 5. Use representative prompts and production-shaped inputs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Coding tasks.&lt;/li&gt;
&lt;li&gt;Customer tickets.&lt;/li&gt;
&lt;li&gt;Internal documents.&lt;/li&gt;
&lt;li&gt;Long-context retrieval.&lt;/li&gt;
&lt;li&gt;Tool calls.&lt;/li&gt;
&lt;li&gt;Strict JSON output.&lt;/li&gt;
&lt;li&gt;Known failure cases.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Track at least:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Success rate.&lt;/li&gt;
&lt;li&gt;Cost per successful task.&lt;/li&gt;
&lt;li&gt;End-to-end latency.&lt;/li&gt;
&lt;li&gt;Human correction time.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A model that costs 20% more but cuts manual review by 50% may be cheaper overall.&lt;/p&gt;

&lt;h3&gt;
  
  
  Keep Model Names Out of Core Business Logic
&lt;/h3&gt;

&lt;p&gt;The Fable 5 and Mythos 5 suspension demonstrates why model availability cannot be treated as permanent. Avoid coupling product behavior directly to one provider or model identifier.&lt;/p&gt;

&lt;p&gt;Put model selection behind configuration or a routing layer. That makes it possible to test a new release, shift traffic during an outage, or use different models for different task classes without rewriting the application.&lt;/p&gt;

&lt;h3&gt;
  
  
  Test Replacement Behavior, Not Just Output Quality
&lt;/h3&gt;

&lt;p&gt;When Sonnet 5 becomes available, compare it with Sonnet 4.6 and Opus 4.8 on the same workload. Include tool failures, malformed output, missing context, retries, and long-running sessions.&lt;/p&gt;

&lt;p&gt;The best replacement is not necessarily the model with the highest score. It is the one that produces the lowest cost for an acceptable success rate and correction burden.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Is Confirmed?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Has Claude Sonnet 5 launched?
&lt;/h3&gt;

&lt;p&gt;No. As of June 23, 2026, Anthropic has not officially released Claude Sonnet 5. The current official Sonnet model is Claude Sonnet 4.6.&lt;/p&gt;

&lt;h3&gt;
  
  
  When is it expected?
&lt;/h3&gt;

&lt;p&gt;Rumors point to late June 2026. June 24 has been cited as a possible early release date, while another report points to the week beginning June 29. Anthropic has confirmed neither.&lt;/p&gt;

&lt;h3&gt;
  
  
  What is the model ID?
&lt;/h3&gt;

&lt;p&gt;The reported identifier is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;claude-sonnet-5
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That identifier is not official. Anthropic's current documented Sonnet model ID is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;claude-sonnet-4-6
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Will Sonnet 5 outperform Opus 4.8?
&lt;/h3&gt;

&lt;p&gt;That is unknown. The likely positioning is stronger cost-performance for Sonnet 5, with Opus 4.8 retaining an advantage in deep reasoning and long-horizon autonomous work. The only reliable answer will come from side-by-side testing on real applications.&lt;/p&gt;

&lt;h3&gt;
  
  
  What should developers use today?
&lt;/h3&gt;

&lt;p&gt;Keep production workloads on official, available models such as Claude Sonnet 4.6 and Claude Opus 4.8. Build evaluation, routing, and cost tracking now so a future Sonnet 5 release can be tested without turning the launch into an infrastructure migration.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/claude-sonnet-5-spotted/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=claude-sonnet-5-spotted"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Integrating GPT-5.6: API Access, Real Costs, and Fallback Design</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Mon, 21 Sep 2026 03:24:36 +0000</pubDate>
      <link>https://dev.to/ryancole1/integrating-gpt-56-api-access-real-costs-and-fallback-design-4ggl</link>
      <guid>https://dev.to/ryancole1/integrating-gpt-56-api-access-real-costs-and-fallback-design-4ggl</guid>
      <description>&lt;p&gt;When I evaluate a model API, the first successful request is only the starting point. I want to know which route I’m calling, what a completed workflow costs, and what happens when that route stops responding.&lt;/p&gt;

&lt;p&gt;For GPT-5.6, I’d approach integration in that order: verify access, test representative workloads, then build the routing and failure behavior around the results.&lt;/p&gt;

&lt;p&gt;An OpenAI-compatible interface can reduce integration work. It still leaves decisions about model selection, budgets, and fallback behavior in the application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Verify the model route before writing the integration
&lt;/h2&gt;

&lt;p&gt;API access lets applications use a model in coding assistants, research agents, support workflows, internal knowledge tools, data analysis, and SaaS features.&lt;/p&gt;

&lt;p&gt;The access workflow is straightforward:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create an account with the API provider.&lt;/li&gt;
&lt;li&gt;Generate an API key.&lt;/li&gt;
&lt;li&gt;Configure the endpoint.&lt;/li&gt;
&lt;li&gt;Select an available model route.&lt;/li&gt;
&lt;li&gt;Send a request and consume the response.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The detail I’d verify first is the model identifier. The source guide describes GPT-5.6 options named Sol, Terra, and Luna, and gives &lt;code&gt;gpt-5.6-sol&lt;/code&gt; and &lt;code&gt;gpt-5.6-terra&lt;/code&gt; as possible route names. Those names need checking against the provider’s current catalog before deployment.&lt;/p&gt;

&lt;p&gt;I wouldn’t infer a variant’s performance from its name. The useful comparison is how each available route handles the application’s requirements: reasoning quality, cost, latency, and throughput.&lt;/p&gt;

&lt;h3&gt;
  
  
  A minimal request
&lt;/h3&gt;

&lt;p&gt;CometAPI is one option for accessing multiple models through a shared API key and an OpenAI-compatible endpoint. The source’s illustrative request uses the following URL and payload:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl https://api.cometapi.com/v1/chat/completions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$COMETAPI_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "gpt-5.6",
    "messages": [
      {
        "role": "user",
        "content": "Explain how a unified API layer helps production AI apps."
      }
    ]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set &lt;code&gt;COMETAPI_KEY&lt;/code&gt; to the credential issued by the provider, and replace &lt;code&gt;gpt-5.6&lt;/code&gt; if the dashboard lists a different active route.&lt;/p&gt;

&lt;p&gt;This demonstrates the request structure; it doesn’t establish that a particular model alias is currently available. I’d confirm that separately before making the route a production dependency.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decide where model access belongs
&lt;/h2&gt;

&lt;p&gt;A product might begin with one text model and later add separate models for reasoning, coding, fast chat, images, video, and speech. A backup model adds another dependency.&lt;/p&gt;

&lt;p&gt;With direct integrations, that can mean multiple credentials, billing dashboards, SDKs, rate limits, and error formats. I’d decide early whether the application should manage those differences itself or put a shared access layer in front of them.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Integration&lt;/th&gt;
&lt;th&gt;Operational consequence&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;One direct provider integration&lt;/td&gt;
&lt;td&gt;An outage or limit at that provider can interrupt the application.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unified API layer&lt;/td&gt;
&lt;td&gt;The application keeps a common interface while the underlying model route can change.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unified API with configured fallback&lt;/td&gt;
&lt;td&gt;A failed primary request can be sent to another suitable model or provider route.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A shared interface makes model changes easier to manage. Fallback still needs to be designed and configured; a common request format alone doesn’t define what happens after an error.&lt;/p&gt;

&lt;p&gt;This matters for agents, automation, SaaS features, and developer workflows involving tools such as Claude Code or Cursor. As the model stack changes, I want route selection to remain a manageable part of the integration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Price completed work, including failures
&lt;/h2&gt;

&lt;p&gt;The source does not provide GPT-5.6 token rates, so there isn’t enough information here to calculate a production bill. Current prices need to come from the active route’s pricing information.&lt;/p&gt;

&lt;p&gt;Even with those rates, token price is only one input. Long prompts, large outputs, repeated agent calls, retries, and failed requests all affect the cost of delivering a result.&lt;/p&gt;

&lt;p&gt;I’d evaluate pricing through three questions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;What does a successful user action cost?&lt;/strong&gt; Include the requests and retries needed to finish it.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Can the route handle the expected traffic?&lt;/strong&gt; Check latency, rate limits, errors, and availability.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;What happens to cost and behavior during fallback?&lt;/strong&gt; A backup route becomes part of the production workload when the primary route fails.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A low listed price is useful only if the route also meets the application’s quality and performance requirements.&lt;/p&gt;

&lt;h3&gt;
  
  
  Compare models on the same workload
&lt;/h3&gt;

&lt;p&gt;I’d compare GPT-5.6 with available alternatives such as Claude, Gemini, DeepSeek, Grok, and Qwen using representative application prompts.&lt;/p&gt;

&lt;p&gt;For a coding assistant, I care about whether it solves the coding task. For an agent, I care about tool instructions and repeated execution. For customer support, I care about whether it handles the actual support workflow consistently.&lt;/p&gt;

&lt;p&gt;Generic prompts can confirm that an endpoint responds. They give much less evidence about whether the model fits the product.&lt;/p&gt;

&lt;p&gt;The comparison should include:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measurement&lt;/th&gt;
&lt;th&gt;What it helps answer&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Task quality&lt;/td&gt;
&lt;td&gt;Does the output meet the product’s requirements?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input and output token usage&lt;/td&gt;
&lt;td&gt;How much usage does each request generate?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Does the route fit the user experience?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Error and retry behavior&lt;/td&gt;
&lt;td&gt;How often does the workflow need recovery?&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost per completed workflow&lt;/td&gt;
&lt;td&gt;What does useful output actually cost?&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Treat trial credit as an evaluation budget
&lt;/h2&gt;

&lt;p&gt;The source advertises &lt;strong&gt;$1 in free credit for new registrations&lt;/strong&gt; at the provider used in the example. I’d check the current offer and supported routes before relying on it.&lt;/p&gt;

&lt;p&gt;Trial credit can support initial request tests, output comparisons, token measurements, and checks of latency and error behavior. It should be treated as a limited evaluation allowance.&lt;/p&gt;

&lt;p&gt;“Free API” commonly refers to trial credits, a testing quota, a promotion, or temporary access. It does not establish unlimited production usage.&lt;/p&gt;

&lt;p&gt;My evaluation sequence would be:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Select a small set of real application prompts.&lt;/li&gt;
&lt;li&gt;Record input and output token usage.&lt;/li&gt;
&lt;li&gt;Run the same tasks through suitable alternative models.&lt;/li&gt;
&lt;li&gt;Measure latency and errors.&lt;/li&gt;
&lt;li&gt;Estimate monthly usage from the expected workload.&lt;/li&gt;
&lt;li&gt;Test fallback before launch.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This gives the trial a concrete purpose: gather enough evidence to choose a route and budget for it.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make fallback an explicit application behavior
&lt;/h2&gt;

&lt;p&gt;A provider outage is only one reason a request might need another route. Rate limits, timeouts, latency spikes, and temporary model unavailability can also interrupt a workflow.&lt;/p&gt;

&lt;p&gt;The basic fallback sequence is:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Send the request to the primary GPT-5.6 route.&lt;/li&gt;
&lt;li&gt;Detect a failure or timeout.&lt;/li&gt;
&lt;li&gt;Send the request to a suitable backup model.&lt;/li&gt;
&lt;li&gt;Return the backup response if it succeeds.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The word &lt;em&gt;suitable&lt;/em&gt; matters. Different models can produce different outputs, so availability alone doesn’t make a model a good replacement.&lt;/p&gt;

&lt;p&gt;For chat, a different response may be acceptable. For coding, support automation, or an agent following tool instructions, I’d evaluate the backup against those same task requirements before depending on it.&lt;/p&gt;

&lt;p&gt;Fallback can reduce user-facing errors, but it does not guarantee every request will succeed. The backup also needs to be available and capable of completing the task.&lt;/p&gt;

&lt;h3&gt;
  
  
  Test recovery before an outage forces the issue
&lt;/h3&gt;

&lt;p&gt;I’d include fallback in the initial production work. Waiting for a provider incident leaves both routing behavior and backup output quality untested at the moment they matter most.&lt;/p&gt;

&lt;p&gt;The evaluation should establish whether the backup can complete the workflow, maintain acceptable quality, and stay within the application’s latency and cost constraints.&lt;/p&gt;

&lt;p&gt;That applies equally to chatbots, coding tools, internal automation, customer support, and high-traffic SaaS features.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep production measurements close to the model decision
&lt;/h2&gt;

&lt;p&gt;Model selection shouldn’t end when the integration ships. Long contexts, document-heavy requests, and agent loops can change usage substantially as real traffic arrives.&lt;/p&gt;

&lt;p&gt;I’d track average input tokens, average output tokens, cost per user action, cost per workflow, and projected monthly usage from the beginning. Alongside those, I’d watch latency, errors, and fallback behavior.&lt;/p&gt;

&lt;p&gt;Those measurements make later changes easier to justify. A different route might reduce cost, improve response time, or handle a task more consistently. A flexible model layer gives the application room to make that change without turning every model release into another integration project.&lt;/p&gt;

&lt;p&gt;My launch criteria would be concrete: a verified route identifier, working authentication, acceptable results on real tasks, measured usage, and a tested backup path. That is the evidence I’d want before letting a model API become a dependency users rely on.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/gpt-5-6-guide-consideration-api-key-access/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=gpt-5-6-guide-consideration-api-key-access"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>One OpenAI Client, Multiple Models: What I Check Before Shipping</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Mon, 21 Sep 2026 02:25:06 +0000</pubDate>
      <link>https://dev.to/ryancole1/one-openai-client-multiple-models-what-i-check-before-shipping-497i</link>
      <guid>https://dev.to/ryancole1/one-openai-client-multiple-models-what-i-check-before-shipping-497i</guid>
      <description>&lt;p&gt;A shared API client is the easy part of a multi-model application. The harder part is making sure that switching &lt;code&gt;model&lt;/code&gt; does not silently change instruction handling, break tool calls, or turn a cheap request into an expensive retry loop.&lt;/p&gt;

&lt;p&gt;I use an OpenAI-compatible gateway to keep provider integration out of application logic. The application supplies a gateway URL, a gateway credential, and a model identifier; the gateway handles upstream routing and schema translation. That reduces SDK maintenance, but it does not make the models interchangeable.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start With the Client, Not the Routing Policy
&lt;/h2&gt;

&lt;p&gt;With the Python SDK, the relevant configuration is &lt;code&gt;base_url&lt;/code&gt; and &lt;code&gt;api_key&lt;/code&gt;. The JavaScript/TypeScript client uses &lt;code&gt;baseURL&lt;/code&gt;. The credential must belong to the endpoint receiving the request: pointing at a gateway while keeping an unrelated provider key is not enough.&lt;/p&gt;

&lt;p&gt;A unified multi-model service such as CometAPI is relevant here because it exposes multiple providers through one integration. I keep the endpoint and model IDs in configuration rather than embedding a provider catalog in application code.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AI_BASE_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;reasoning_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AI_REASONING_MODEL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a precise technical assistant.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Analyze this system architecture for latency bottlenecks.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;reasoning_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;document_response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AI_DOCUMENT_MODEL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Refine this technical documentation for clarity.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;max_tokens&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;document_response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This uses the official &lt;code&gt;openai&lt;/code&gt; Python package. Set &lt;code&gt;AI_BASE_URL&lt;/code&gt; to the gateway's documented OpenAI-compatible endpoint and both model variables to identifiers verified in its current catalog. The prompts are smoke tests; a useful architecture review or documentation edit also needs the actual material in the request.&lt;/p&gt;

&lt;p&gt;The same client handles both calls. That is the integration benefit, not a promise that every parameter works on every route. For example, the document request's &lt;code&gt;max_tokens=1000&lt;/code&gt; needs to be supported and translated correctly. I would test temperature controls separately rather than assume that &lt;code&gt;temperature=0.2&lt;/code&gt; is accepted by every reasoning model.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Gateway Actually Does
&lt;/h2&gt;

&lt;p&gt;For a chat completion, the gateway authenticates the incoming request and reads its &lt;code&gt;model&lt;/code&gt; field. It resolves that identifier to an upstream destination, translates the OpenAI-shaped payload into the provider's native schema, and sends the request using the appropriate upstream authentication.&lt;/p&gt;

&lt;p&gt;On the return path, it converts the response into the shape the client expects, including content, token usage, and finish reasons where supported. The application can then read &lt;code&gt;response.choices[0].message.content&lt;/code&gt; without maintaining a separate response parser for each provider.&lt;/p&gt;

&lt;p&gt;Credential ownership is a separate deployment decision. A gateway may supply upstream access, or it may support credentials you configure yourself. I would verify that contract instead of assuming a particular service has a bring-your-own-key vault. Where upstream keys are yours, scope them narrowly, test authentication per provider, and configure billing alerts and usage limits.&lt;/p&gt;

&lt;p&gt;This arrangement centralizes integration work. It also centralizes failure: a gateway outage can affect every model behind the endpoint, even when those providers are healthy.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat Model Comparisons as Evaluation Inputs
&lt;/h2&gt;

&lt;p&gt;I would not commit a routing policy based on a model launch announcement or an old pricing table. First verify the exact gateway model ID, availability, supported parameters, context limits, and current price. A provider's public model name and a gateway's routing identifier are not necessarily identical.&lt;/p&gt;

&lt;p&gt;The supplied comparison describes a &lt;strong&gt;July 2026 snapshot&lt;/strong&gt;: GPT-5.5, reportedly released in April 2026, and Claude Sonnet 5, reportedly released in June 2026. Those release, retirement, pricing, and performance claims need confirmation against current official documentation. They are not established by a successful request through an OpenAI-compatible endpoint.&lt;/p&gt;

&lt;p&gt;For reference, these are the snapshot's specific claims, preserved as claims rather than a verified production catalog:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;GPT-5.5 claim&lt;/th&gt;
&lt;th&gt;Claude Sonnet 5 claim&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Positioning&lt;/td&gt;
&lt;td&gt;Flagship reasoning and agentic model&lt;/td&gt;
&lt;td&gt;Most agentic Sonnet release, approaching Opus-class performance at lower cost&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context and output&lt;/td&gt;
&lt;td&gt;Roughly 1.05M-token input context; 128K maximum output&lt;/td&gt;
&lt;td&gt;1M-token input context, default and maximum; 128K maximum output&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Evaluations&lt;/td&gt;
&lt;td&gt;Terminal-Bench 2.0: 82.7%; Expert-SWE: 73.1%; GDPval: 84.9%; FrontierMath Tiers 1–3: 51.7%&lt;/td&gt;
&lt;td&gt;Largest gains over Sonnet 4.6 concentrated in coding and agentic tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Price per 1M tokens&lt;/td&gt;
&lt;td&gt;Approximately $5 input / $30 output, standard tier&lt;/td&gt;
&lt;td&gt;$2 input / $10 output introductory pricing through August 31, 2026; $3 / $15 thereafter&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The same snapshot describes GPT-5.5 as supporting reasoning, tool use, and computer use, with those evaluation scores improving on GPT-5.4. It positions Sonnet 5 for long-document synthesis, legal and financial analysis, instruction following, self-verification, and low hallucination and sycophancy rates. I would treat those descriptions as hypotheses to test on actual workloads, not routing guarantees.&lt;/p&gt;

&lt;p&gt;It also claims that earlier &lt;code&gt;gpt-5-chat-latest&lt;/code&gt; variants were lightweight, non-reasoning tiers and that the GPT-5.2 Instant/Thinking/Pro line was deprecated in June 2026 with traffic migrated to GPT-5.5. Model lifecycle claims are particularly important to verify before relying on an alias. A familiar name is not evidence of unchanged behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  Route by Successful Task Cost
&lt;/h3&gt;

&lt;p&gt;My routing criteria start with the work: prompt complexity, required context, output contract, latency budget, and acceptable cost. Simple classification deserves a cheaper-model evaluation. Multi-step execution needs tests that include tool results and recovery. Long-document analysis needs retrieval and synthesis checks across the document, not just a context-window number.&lt;/p&gt;

&lt;p&gt;A flagship reasoning model may be unnecessary for high-volume classification. Likewise, a large document-oriented model may be unnecessary for a constrained code-generation task that a smaller model handles reliably. Neither observation establishes a universal winner; the relevant metric is cost per successful task, including failures and retries.&lt;/p&gt;

&lt;p&gt;I would compare candidate routes on task quality, time to first token, total latency, token consumption, JSON/schema validity, context handling, and fallback behavior. Published benchmarks can narrow the candidate set. They cannot validate an application's output contract.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compatibility Needs Its Own Test Suite
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Parameter and Instruction Translation
&lt;/h3&gt;

&lt;p&gt;Anthropic's Messages API expects a top-level &lt;code&gt;system&lt;/code&gt; parameter, while an OpenAI-style request can carry system instructions in the &lt;code&gt;messages&lt;/code&gt; array. The gateway needs to extract and translate those instructions without losing their meaning. I would explicitly test system prompts rather than infer support from a successful user-only request.&lt;/p&gt;

&lt;p&gt;Output limits deserve similar attention. &lt;code&gt;max_completion_tokens&lt;/code&gt; and &lt;code&gt;max_tokens&lt;/code&gt; are not names to swap blindly across routes. Confirm what the gateway accepts, how it maps the value, and whether the upstream model enforces the requested limit. A parameter that disappears during translation is harder to detect than an explicit validation error.&lt;/p&gt;

&lt;p&gt;Unsupported parameters may be rejected, mapped, or stripped, depending on the gateway. None of those behaviors should be assumed safe without inspection. For temperature boundaries, system-message structures, and other edge cases, I want integration tests against every active route.&lt;/p&gt;

&lt;h3&gt;
  
  
  Tools and Structured Output
&lt;/h3&gt;

&lt;p&gt;A normalized chat response does not establish tool-calling compatibility. OpenAI tool definitions, Anthropic tool use, and Google's function-calling schema have differences, especially around tool-choice constraints and complex nested schemas. Test the definitions the application actually sends, then validate the resulting tool arguments.&lt;/p&gt;

&lt;p&gt;The same applies to JSON and schema-constrained output. Receiving syntactically valid JSON is not the same as satisfying the required schema. I would make schema validation part of the success criteria used to compare models and approve fallbacks.&lt;/p&gt;

&lt;p&gt;Provider-specific token-bias controls, moderation options, or other proprietary features may have no equivalent in the shared interface. Where those features matter, a documented pass-through mechanism or a direct provider integration may be necessary. One SDK is useful, but losing a required capability is not an acceptable trade.&lt;/p&gt;

&lt;h3&gt;
  
  
  Streaming and Latency
&lt;/h3&gt;

&lt;p&gt;A streaming gateway must normalize upstream events into OpenAI-compatible Server-Sent Events, such as &lt;code&gt;data: {...}&lt;/code&gt;, and forward them incrementally. Buffering the whole response before sending it defeats the point of streaming, even if the final body looks correct.&lt;/p&gt;

&lt;p&gt;The source proposes &lt;strong&gt;5–30 milliseconds&lt;/strong&gt; as a gateway processing-overhead target, excluding upstream transit time. I would treat that as a measurement target, not a service guarantee. An extra network hop also introduces geography and connection-management effects; deployment near the application and upstream connection pooling both deserve inspection.&lt;/p&gt;

&lt;p&gt;Model time to first token and completion time can range from hundreds of milliseconds to several seconds, so small proxy overhead may be acceptable. Still, measure gateway processing, network transit, first-token latency, and total completion time separately. A single end-to-end timer cannot explain whether slow responses come from the proxy or the model.&lt;/p&gt;

&lt;h2&gt;
  
  
  Design Failure Handling Before Enabling Fallbacks
&lt;/h2&gt;

&lt;p&gt;Error normalization is useful only when it preserves enough information to act. Providers can report different status codes and bodies for rejected requests. The source's examples include &lt;code&gt;400 Bad Request&lt;/code&gt; for a safety-related rejection and &lt;code&gt;422 Unprocessable Entity&lt;/code&gt; for a context violation; those are examples, not universal provider contracts.&lt;/p&gt;

&lt;p&gt;If the gateway collapses every upstream failure into &lt;code&gt;502 Bad Gateway&lt;/code&gt; or &lt;code&gt;500 Internal Server Error&lt;/code&gt;, the application loses the distinction between a bad payload, a rate limit, and an outage. I want the original upstream status and diagnostic message preserved in response metadata, with enough route information to debug the failure.&lt;/p&gt;

&lt;p&gt;For transient conditions such as &lt;code&gt;429 Too Many Requests&lt;/code&gt; or &lt;code&gt;503 Service Unavailable&lt;/code&gt;, define explicit retry or alternate-model behavior. Then simulate those conditions in staging. A configured fallback is not proven until the application handles it without unhandled errors and the replacement model still satisfies the task's output requirements.&lt;/p&gt;

&lt;p&gt;Multi-region gateway deployments and automated failover can reduce gateway availability risk where the deployment supports them. A secondary client that calls a provider directly is another option. That bypass must be tested independently because its authentication, feature support, and error behavior may differ from the gateway route.&lt;/p&gt;

&lt;h2&gt;
  
  
  My Production Gate
&lt;/h2&gt;

&lt;p&gt;Before increasing traffic, I check that each route authenticates, uses a current model ID, and has documented parameter support. For credentials I control, I check least-privilege access, rotation procedures, billing alerts, and usage limits. I also verify what upstream credential management the gateway actually provides.&lt;/p&gt;

&lt;p&gt;Observability must attribute latency and token usage to the selected route and credential. Where the gateway exposes diagnostic headers, the logging stack should parse them. I track proxy overhead separately from generation time and watch token-consumption drift, since a model switch can change cost even when request volume stays flat.&lt;/p&gt;

&lt;p&gt;Finally, I run daily integration tests against active routes, covering system prompts, tool schemas, output limits, temperature support, streaming, and error translation. I start with a small subset of non-critical traffic and compare successful-task cost and latency before expanding. A unified endpoint earns its place when it reduces integration work without hiding the differences the application depends on.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-call-multiple-ai-models-using-an-openai-compatible-base-url/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-call-multiple-ai-models-using-an-openai-compatible-base-url"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Gemini 4 Pro’s Leaked Scores: What I’d Check Before Switching Models</title>
      <dc:creator>Ryan Cole</dc:creator>
      <pubDate>Mon, 21 Sep 2026 01:29:54 +0000</pubDate>
      <link>https://dev.to/ryancole1/gemini-4-pros-leaked-scores-what-id-check-before-switching-models-5ee5</link>
      <guid>https://dev.to/ryancole1/gemini-4-pros-leaked-scores-what-id-check-before-switching-models-5ee5</guid>
      <description>&lt;p&gt;Gemini 4 Pro leads GPT-6 Astra and Claude Fable 5.1 in a circulating benchmark table. That makes it interesting to evaluate. It does not establish that Google has shipped a better model.&lt;/p&gt;

&lt;p&gt;The reporting dated September 20, 2026 describes an unreleased model without an official model card, public API endpoint, or confirmed pricing. The scores come from community posts, alleged internal checkpoints, and Arena.ai sightings. Some earlier sheets were reportedly flagged as predictions or partially copied data.&lt;/p&gt;

&lt;p&gt;My read: the claimed improvements in coding, computer use, and multimodal reasoning deserve attention. I would still keep production decisions tied to reproducible evaluations and actual access.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the provenance
&lt;/h2&gt;

&lt;p&gt;Several different kinds of evidence have become bundled into the Gemini 4 story.&lt;/p&gt;

&lt;p&gt;The strongest context cited in the reporting is Google’s July 2026 earnings discussion of investment in a “larger Gemini 4 base model.” That supports an account of increased training investment. It does not validate a particular checkpoint, benchmark result, context limit, or token price.&lt;/p&gt;

&lt;p&gt;The remaining claims fall into three groups:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;th&gt;What it claims&lt;/th&gt;
&lt;th&gt;What remains unresolved&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Posts attributed to leaker Pankaj Kumar and subsequent community reporting&lt;/td&gt;
&lt;td&gt;An internal Pro checkpoint, an October release, possibly an earlier Flash-Lite variant&lt;/td&gt;
&lt;td&gt;Checkpoint identity and launch timing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Arena.ai encounters under names such as &lt;code&gt;gemini-3.8-flash&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Strong SVG, Three.js, structured output, and agent behavior&lt;/td&gt;
&lt;td&gt;Whether the model was Gemini 4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;A circulating comparison sheet&lt;/td&gt;
&lt;td&gt;Benchmark scores, prices, and context limits across several model families&lt;/td&gt;
&lt;td&gt;Authenticity, evaluation settings, and reproducibility&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The comparison sheet includes Gemini 4 Pro, Flash, and Flash-Lite alongside GPT-6 Astra, Claude Fable 5.1, Claude Opus 5, and GPT-5.6 Sol.&lt;/p&gt;

&lt;p&gt;Google has not confirmed the table or the reported Arena identities. The reporting points to a history of stealth evaluations before launches, but that pattern cannot identify a particular anonymous model.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the benchmark sheet actually says
&lt;/h2&gt;

&lt;p&gt;I find the numbers easier to assess when the comparison and its limitations sit together. &lt;strong&gt;Every result below is an unverified claim from the circulating sheet.&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Gemini 4 Pro&lt;/th&gt;
&lt;th&gt;GPT-6 Astra&lt;/th&gt;
&lt;th&gt;Claude Fable 5.1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE v1.1&lt;/td&gt;
&lt;td&gt;88.7%&lt;/td&gt;
&lt;td&gt;86.9%&lt;/td&gt;
&lt;td&gt;69.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-bench 2.1&lt;/td&gt;
&lt;td&gt;95.3%&lt;/td&gt;
&lt;td&gt;94.1%&lt;/td&gt;
&lt;td&gt;92.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-bench 4.0&lt;/td&gt;
&lt;td&gt;69.7%&lt;/td&gt;
&lt;td&gt;66.4%&lt;/td&gt;
&lt;td&gt;57.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GDPval-AA v2&lt;/td&gt;
&lt;td&gt;2064 Elo&lt;/td&gt;
&lt;td&gt;1994 Elo&lt;/td&gt;
&lt;td&gt;1853 Elo&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OSWorld-2.0&lt;/td&gt;
&lt;td&gt;86.8%&lt;/td&gt;
&lt;td&gt;84.5%&lt;/td&gt;
&lt;td&gt;77.9%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HLE-Verified&lt;/td&gt;
&lt;td&gt;72.1%&lt;/td&gt;
&lt;td&gt;67.3%&lt;/td&gt;
&lt;td&gt;62.1%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LVBench&lt;/td&gt;
&lt;td&gt;95.8% agentic / 95.0% static&lt;/td&gt;
&lt;td&gt;93.8%&lt;/td&gt;
&lt;td&gt;90.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The OSWorld-2.0 number is described as a &lt;strong&gt;partial score with batch tools enabled&lt;/strong&gt;. LVBench gives separate agentic and static results for Gemini, while the comparison supplies only one figure for each competitor. Those details matter when deciding whether the rows represent equivalent evaluation conditions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Coding and sustained agent work
&lt;/h3&gt;

&lt;p&gt;DeepSWE v1.1 is presented as a long-horizon software engineering evaluation. The sheet puts Gemini 4 Pro at 88.7%, ahead of Astra’s 86.9%, Fable’s 69.1%, and Claude Opus 5’s 75.0%.&lt;/p&gt;

&lt;p&gt;The two Terminal-bench versions measure different workloads in the reporting: version 2.1 covers agentic terminal coding, while version 4.0 covers general agent capabilities. I would keep their results separate rather than treat them as interchangeable measures of coding quality.&lt;/p&gt;

&lt;p&gt;If reproduced, these results would make Gemini 4 Pro a serious candidate for repository work: refactoring across files, debugging, running tools, and recovering from failed steps. They would still leave the practical question of how reliably it finishes &lt;em&gt;my&lt;/em&gt; tasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Knowledge work and specialist evaluations
&lt;/h3&gt;

&lt;p&gt;The sheet’s 2064 GDPval-AA v2 rating would make Gemini 4 Pro the only listed model above 2000 Elo. It also claims:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Vals Finance Agent v2:&lt;/strong&gt; 74.2%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Harvey’s Legal Agent Benchmark:&lt;/strong&gt; 18.7% on complex legal workflows, measured by all-pass rate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;CharXiv Reasoning:&lt;/strong&gt; 94.7% for reasoning over complex charts without tools.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;BioMysteryBench and LABBench2:&lt;/strong&gt; leading results on human-solvable and difficult subsets, without exact scores supplied in the article.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I would resist flattening these into a single “professional work” score. An all-pass legal workflow result, a chart reasoning accuracy, and an Elo rating describe different properties.&lt;/p&gt;

&lt;p&gt;The same applies to HLE-Verified’s claimed 72.1%: it suggests strong multidisciplinary expert reasoning under that evaluation, subject to verification. It does not establish reliability for every specialist application.&lt;/p&gt;

&lt;h2&gt;
  
  
  Price and context contain their own uncertainty
&lt;/h2&gt;

&lt;p&gt;The leaked pricing is attractive enough to influence deployment plans—if it survives launch.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input per 1M tokens&lt;/th&gt;
&lt;th&gt;Output per 1M tokens&lt;/th&gt;
&lt;th&gt;Claimed maximum input&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Pro&lt;/td&gt;
&lt;td&gt;$2.25&lt;/td&gt;
&lt;td&gt;$11.25&lt;/td&gt;
&lt;td&gt;2M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Flash&lt;/td&gt;
&lt;td&gt;$0.75&lt;/td&gt;
&lt;td&gt;$3.75&lt;/td&gt;
&lt;td&gt;2M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Flash-Lite&lt;/td&gt;
&lt;td&gt;$0.35&lt;/td&gt;
&lt;td&gt;$1.75&lt;/td&gt;
&lt;td&gt;2M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;$12.00&lt;/td&gt;
&lt;td&gt;$60.00&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5.1&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$50.00&lt;/td&gt;
&lt;td&gt;200k tokens in the sheet&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The Gemini Pro figures are quoted without caching. They amount to roughly one-fifth of Astra’s listed token prices. Fable 5.1’s reported cache-read discounts complicate any comparison based solely on uncached rates.&lt;/p&gt;

&lt;p&gt;The reporting also gives Gemini 3.x Pro pricing as $2–$4 per million input tokens and $12–$18 per million output tokens, depending on context length. That makes the leaked prices look plausible, but plausibility is a weak substitute for a published rate card.&lt;/p&gt;

&lt;h3&gt;
  
  
  The context numbers disagree
&lt;/h3&gt;

&lt;p&gt;The sheet lists 2M input tokens for the Gemini 4 family, 1M for Astra, and 200k for the Claude Fable/Opus 5 line.&lt;/p&gt;

&lt;p&gt;However, the same reporting says official Anthropic documentation commonly gives Fable 5.1 a 1M context window. The 200k entry may describe a particular evaluation configuration or simply be wrong. I would not use it to claim a confirmed tenfold context advantage.&lt;/p&gt;

&lt;p&gt;Other community accounts mention targets around 1.5M tokens or higher, and separate “ghost-routing” claims suggest 10M ceilings. These are distinct claims, not a coherent specification.&lt;/p&gt;

&lt;p&gt;For repository agents, I care about retrieval fidelity and task completion as the context fills. A maximum token count alone cannot answer either question.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Arena demos can tell us
&lt;/h2&gt;

&lt;p&gt;Reported encounters with the anonymous model include SVG scenes such as a pelican riding a bicycle, mechanical butterflies built in Three.js, long JSON outputs without obvious repetition, and capable agent behavior.&lt;/p&gt;

&lt;p&gt;Community comparisons reportedly judged some of these outputs superior or competitive with Astra. I understand why those examples spread: complex visual output makes differences immediately visible.&lt;/p&gt;

&lt;p&gt;Their evidentiary limit is substantial, though. A compelling result from a placeholder model does not establish its identity, typical performance, or production configuration.&lt;/p&gt;

&lt;p&gt;Claims about multi-million-token context, 256k output limits, cross-session memory, and native internet or tool access also need separate confirmation. A demo may combine model behavior with capabilities supplied by the surrounding application.&lt;/p&gt;

&lt;p&gt;I would save the prompts as future evaluation cases. I would avoid using the screenshots as a deployment ranking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Architecture and reasoning: separate reports from inference
&lt;/h2&gt;

&lt;p&gt;The reporting attributes “most ambitious pre-training run yet” language to Google and describes a significantly larger base model. It then infers possible increases in parameter count, effective capacity, training compute, and multimodal coverage.&lt;/p&gt;

&lt;p&gt;Those are reasonable possibilities. They do not reveal whether the architecture uses denser scaling, mixture-of-experts, different sparsity, or another approach.&lt;/p&gt;

&lt;h3&gt;
  
  
  The alleged &lt;code&gt;argon&lt;/code&gt; checkpoint
&lt;/h3&gt;

&lt;p&gt;Descriptions of an early checkpoint called &lt;code&gt;argon&lt;/code&gt; mention a &lt;strong&gt;High thinking-effort&lt;/strong&gt; mode, roughly &lt;strong&gt;2.4 minutes for one generation&lt;/strong&gt;, and an output ceiling of &lt;strong&gt;256k tokens&lt;/strong&gt;, compared with a reported previous limit of approximately &lt;strong&gt;64k&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;If accurate, those details suggest room for more computation and longer generations. They do not establish typical latency, effective output quality, or production limits.&lt;/p&gt;

&lt;p&gt;The reporting expects reasoning controls comparable to current Gemini Flash low, medium, and high thinking levels. Until documented, I would treat that as an interface expectation.&lt;/p&gt;

&lt;h3&gt;
  
  
  The expected capability direction
&lt;/h3&gt;

&lt;p&gt;The proposed focus is consistent across the leaks: sustained software engineering, tool use, context reliability, and multimodal understanding.&lt;/p&gt;

&lt;p&gt;The article connects those expectations to agentic features attributed to Gemini 3.5/3.8 Flash and the September 2026 Gemini 3.8 Live models. Expected improvements include function calling, code execution, search grounding, structured JSON/XML outputs, video comprehension, spatial reasoning, and voice interactions.&lt;/p&gt;

&lt;p&gt;For application development, I would turn those expectations into test cases. “Better tool use” becomes successful function selection and argument construction. “Long-context reliability” becomes retrieving the right evidence from a large input. “Stronger agents” becomes completing a task despite intermediate failures.&lt;/p&gt;

&lt;h2&gt;
  
  
  The RSI claim needs different evidence
&lt;/h2&gt;

&lt;p&gt;Recursive self-improvement has become attached to the launch speculation, but the reporting provides no independently verified evidence that Gemini 4 autonomously improved its own model weights through a closed loop.&lt;/p&gt;

&lt;p&gt;It references executive comments connecting AI investment to longer-term self-improvement goals and research such as Dream-RSI, where agents improve exploration strategies without changing model weights.&lt;/p&gt;

&lt;p&gt;That distinction is central. Improving an agent’s strategy, using AI in model research, and autonomously producing better model generations are different claims. The leaked benchmark table cannot establish the last one.&lt;/p&gt;

&lt;h2&gt;
  
  
  How I’d prepare an evaluation
&lt;/h2&gt;

&lt;p&gt;The September 20 account describes Astra and Fable 5.1 as publicly available and independently evaluated, with different strengths and closely matched overall results on third-party trackers. Gemini 4 Pro remains an alleged upcoming option in that snapshot.&lt;/p&gt;

&lt;p&gt;I would keep building against available models and prepare a repeatable comparison:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Freeze representative tasks.&lt;/strong&gt; Include repository changes, terminal work, structured outputs, and the multimodal inputs the application actually receives.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Record execution conditions.&lt;/strong&gt; Capture tool access, thinking settings, context size, caching, retry policy, and agent configuration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Measure completed work.&lt;/strong&gt; Track correctness, latency, failures, and total spend across retries and tool loops.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Re-run when public access exists.&lt;/strong&gt; Verify the documented model ID, limits, and pricing before interpreting results.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Route by observed performance.&lt;/strong&gt; A cheaper Flash variant may suit one workload while a Pro model earns its cost on another.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A unified multi-model API such as CometAPI can simplify that comparison through an OpenAI-compatible interface, provided the required models and features are available. I would still check feature support rather than assume that changing a model parameter preserves every provider-specific behavior.&lt;/p&gt;

&lt;p&gt;October 2026 is the reported release expectation, with some late-September speculation and a possible earlier Flash-Lite release. None is an official commitment in the supplied reporting.&lt;/p&gt;

&lt;p&gt;My threshold for switching is straightforward: public access, documented operating limits, and better results on the workloads I need to ship. The leaked numbers give me reasons to run those tests. They do not supply the results.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/has-gemini-4-already-beated-gpt-6-and-claude-fable-5-1/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=has-gemini-4-already-beated-gpt-6-and-claude-fable-5-1"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
