<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Dylan Foster</title>
    <description>The latest articles on DEV Community by Dylan Foster (@dylanfoster1).</description>
    <link>https://dev.to/dylanfoster1</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F4113588%2F5d383886-7dac-4091-84e6-d2931beb7876.png</url>
      <title>DEV Community: Dylan Foster</title>
      <link>https://dev.to/dylanfoster1</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/dylanfoster1"/>
    <language>en</language>
    <item>
      <title>How I’d Route Work Between GLM-5.3 Flash and GLM-5.3</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Thu, 24 Sep 2026 08:45:55 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/how-id-route-work-between-glm-53-flash-and-glm-53-338e</link>
      <guid>https://dev.to/dylanfoster1/how-id-route-work-between-glm-53-flash-and-glm-53-338e</guid>
      <description>&lt;p&gt;I’d start most workloads on &lt;code&gt;glm-5.3-flash&lt;/code&gt; and reserve &lt;code&gt;glm-5.3&lt;/code&gt; for difficult text tasks where better reasoning can pay for the extra tokens. Flash adds native visual input, matches the flagship’s 1M-token context window, and costs substantially less. The flagship has stronger reported results on demanding repository work, tool-assisted reasoning, and cybersecurity.&lt;/p&gt;

&lt;p&gt;The naming is less useful than the workload split. Flash is a 320B-parameter model, and independent API measurements show the flagship generating output faster. Neither “small” nor “faster” is a safe assumption.&lt;/p&gt;

&lt;p&gt;My decision would come down to three things: whether the task needs vision, how often the output passes validation, and what each accepted result costs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with the input and acceptance criteria
&lt;/h2&gt;

&lt;p&gt;For screenshots, rendered interfaces, charts, document layouts, and browser state, I’d choose Flash. It accepts native visual input; the flagship accepts text only.&lt;/p&gt;

&lt;p&gt;For routine repository maintenance, debugging, refactoring, and test generation, I’d also evaluate Flash first. The flagship’s coding advantage matters most when the task is difficult enough for that advantage to change the outcome.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Workload&lt;/th&gt;
&lt;th&gt;My starting model&lt;/th&gt;
&lt;th&gt;Reason&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;UI implementation and visual debugging&lt;/td&gt;
&lt;td&gt;Flash&lt;/td&gt;
&lt;td&gt;Can inspect screenshots and rendered results&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Documents, charts, and Office workflows&lt;/td&gt;
&lt;td&gt;Flash&lt;/td&gt;
&lt;td&gt;Can use visual structure and layout&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;High-volume automation&lt;/td&gt;
&lt;td&gt;Flash&lt;/td&gt;
&lt;td&gt;Lower token cost and strong tool-use results&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Routine repository changes&lt;/td&gt;
&lt;td&gt;Flash&lt;/td&gt;
&lt;td&gt;Measure acceptance before paying for more capacity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Difficult repository engineering&lt;/td&gt;
&lt;td&gt;GLM-5.3&lt;/td&gt;
&lt;td&gt;Higher reported coding scores&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-horizon research with tools&lt;/td&gt;
&lt;td&gt;GLM-5.3&lt;/td&gt;
&lt;td&gt;Stronger difficult-reasoning results&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Authorized vulnerability discovery&lt;/td&gt;
&lt;td&gt;GLM-5.3&lt;/td&gt;
&lt;td&gt;Dedicated cybersecurity evaluations&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Mixed workloads&lt;/td&gt;
&lt;td&gt;Both&lt;/td&gt;
&lt;td&gt;Default to Flash; escalate selected text tasks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both support function calling, streaming tool calls, always-on reasoning, and open-weight deployment. Context length alone won’t settle the choice: each supports 1M input context and up to 128K output tokens.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the two checkpoints actually contain
&lt;/h2&gt;

&lt;p&gt;These models followed different development paths. &lt;a href="https://docs.z.ai/guides/vlm/glm-5.3-flash" rel="noopener noreferrer"&gt;Z.ai describes Flash&lt;/a&gt; as the first native multimodal model in the GLM-5 series, built from a newly trained 30T-token multimodal base. GLM-5.3 retains the GLM-5.2 base, with its gains coming from scaled post-training.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Property&lt;/th&gt;
&lt;th&gt;GLM-5.3 Flash&lt;/th&gt;
&lt;th&gt;GLM-5.3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Model ID&lt;/td&gt;
&lt;td&gt;&lt;code&gt;glm-5.3-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;glm-5.3&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Architecture&lt;/td&gt;
&lt;td&gt;Native multimodal MoE&lt;/td&gt;
&lt;td&gt;Text MoE&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total parameters&lt;/td&gt;
&lt;td&gt;320B&lt;/td&gt;
&lt;td&gt;744B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Active parameters&lt;/td&gt;
&lt;td&gt;18B&lt;/td&gt;
&lt;td&gt;40B&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Base&lt;/td&gt;
&lt;td&gt;New 30T-token multimodal pre-training&lt;/td&gt;
&lt;td&gt;GLM-5.2 base&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;Text and visual inputs&lt;/td&gt;
&lt;td&gt;Text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;Text&lt;/td&gt;
&lt;td&gt;Text&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context / maximum output&lt;/td&gt;
&lt;td&gt;1M / 128K tokens&lt;/td&gt;
&lt;td&gt;1M / 128K tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning effort&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;low&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;max&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;low&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, &lt;code&gt;max&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning disabled mode&lt;/td&gt;
&lt;td&gt;No; always enabled&lt;/td&gt;
&lt;td&gt;No; always enabled&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Flash activates 55% fewer parameters per token. Its 45-layer architecture combines sparse attention, linear attention, mHC, and MTP. The flagship activates more than twice as many parameters, but I’d use task evaluations—not parameter count—to decide whether that capacity helps.&lt;/p&gt;

&lt;p&gt;Flash’s visual inputs include images and supported video/file workflows. That makes a practical difference in an agent loop: the model can inspect a rendered page or interface state and use that evidence when choosing its next tool call.&lt;/p&gt;

&lt;h3&gt;
  
  
  Long context has different infrastructure costs
&lt;/h3&gt;

&lt;p&gt;Flash alternates linear-attention and sparse-attention blocks, with mHC around attention and MoE components. Its IndexPool mechanism compresses four indexer key vectors into one.&lt;/p&gt;

&lt;p&gt;At a 1M-token sequence length, Z.ai reports:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;3.01× lower per-layer attention compute than the flagship.&lt;/li&gt;
&lt;li&gt;A 4.44× smaller per-layer KV cache.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those are architecture measurements. I wouldn’t translate them into an API latency promise: hardware, quantization, batching, serving software, provider load, and reasoning length all affect the actual request.&lt;/p&gt;

&lt;p&gt;Both models expose &lt;code&gt;low&lt;/code&gt;, &lt;code&gt;high&lt;/code&gt;, and &lt;code&gt;max&lt;/code&gt; reasoning effort. Z.ai recommends &lt;code&gt;max&lt;/code&gt; for difficult coding and benchmark reproduction. I’d keep that setting fixed when comparing models, then evaluate effort levels separately.&lt;/p&gt;

&lt;h2&gt;
  
  
  The benchmark split supports selective escalation
&lt;/h2&gt;

&lt;p&gt;The vendor results favor the flagship on difficult engineering and reasoning, but Flash wins some tool-use and automation comparisons.&lt;/p&gt;

&lt;p&gt;These are the closest matched values from the &lt;a href="https://docs.z.ai/guides/vlm/glm-5.3-flash" rel="noopener noreferrer"&gt;Flash documentation&lt;/a&gt; and &lt;a href="https://docs.z.ai/guides/llm/glm-5.3" rel="noopener noreferrer"&gt;flagship evaluation&lt;/a&gt;. Matching benchmark names does not guarantee identical harnesses, tool configurations, or context management.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Flash&lt;/th&gt;
&lt;th&gt;GLM-5.3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench 2.1&lt;/td&gt;
&lt;td&gt;84.3&lt;/td&gt;
&lt;td&gt;88.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE v1.1&lt;/td&gt;
&lt;td&gt;63.4&lt;/td&gt;
&lt;td&gt;66.9&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;NL2Repo&lt;/td&gt;
&lt;td&gt;56.3&lt;/td&gt;
&lt;td&gt;58.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Toolathlon Verified&lt;/td&gt;
&lt;td&gt;78.4&lt;/td&gt;
&lt;td&gt;73.0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AutomationBench&lt;/td&gt;
&lt;td&gt;48.8&lt;/td&gt;
&lt;td&gt;48.2&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agents’ Last Exam&lt;/td&gt;
&lt;td&gt;26.3&lt;/td&gt;
&lt;td&gt;28.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;HLE with tools&lt;/td&gt;
&lt;td&gt;55.3&lt;/td&gt;
&lt;td&gt;62.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GDPval-AA v2&lt;/td&gt;
&lt;td&gt;1773 Elo&lt;/td&gt;
&lt;td&gt;1769 Elo&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;I’d treat the small AutomationBench and GDPval-AA v2 differences as near ties. Toolathlon Verified gives Flash a clearer lead; HLE with tools gives the flagship a substantial one.&lt;/p&gt;

&lt;h3&gt;
  
  
  Coding: the flagship earns an evaluation slot
&lt;/h3&gt;

&lt;p&gt;GLM-5.3 leads by 3.5 points on DeepSWE and 1.7 points on NL2Repo.&lt;/p&gt;

&lt;p&gt;There is another useful comparison in Z.ai Code Bench v1.0: with both models evaluated through Claude Code 2.1.207, the flagship scores 34.5% at &lt;code&gt;max&lt;/code&gt; effort against Flash’s 29.0%. That is a 5.5-point difference.&lt;/p&gt;

&lt;p&gt;Those results justify testing the flagship on difficult repository tasks. They don’t establish that every maintenance ticket needs it. My escalation signal would be failed validation or unusually demanding engineering work, with the final decision based on acceptance rate.&lt;/p&gt;

&lt;h3&gt;
  
  
  Security: stronger evidence for the flagship
&lt;/h3&gt;

&lt;p&gt;Z.ai reports GLM-5.3 scores of 84.5 on CyberGym and 54.4 on ExploitBench. Under normalized two-hour and six-hour budgets, it completed 105 and 130 ExploitGym tasks, respectively.&lt;/p&gt;

&lt;p&gt;Flash does not have an equivalent matched public cybersecurity suite in the cited material. That makes the flagship the better-supported starting point for authorized vulnerability discovery and demanding security analysis.&lt;/p&gt;

&lt;p&gt;I’d keep human approval, isolated tooling, audit logs, expert review, and disclosure controls in that workflow. A benchmark score doesn’t replace those controls.&lt;/p&gt;

&lt;h2&gt;
  
  
  “Flash” doesn’t tell you which endpoint is faster
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://artificialanalysis.ai/models/comparisons/glm-5-3-flash-vs-glm-5-3" rel="noopener noreferrer"&gt;Artificial Analysis comparison&lt;/a&gt; separates startup latency from output throughput:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Measurement&lt;/th&gt;
&lt;th&gt;Flash&lt;/th&gt;
&lt;th&gt;GLM-5.3 at max effort&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Intelligence Index&lt;/td&gt;
&lt;td&gt;57&lt;/td&gt;
&lt;td&gt;60&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output speed&lt;/td&gt;
&lt;td&gt;50.2 tokens/s&lt;/td&gt;
&lt;td&gt;76.6 tokens/s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Time to first token&lt;/td&gt;
&lt;td&gt;1.49 s&lt;/td&gt;
&lt;td&gt;1.61 s&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context window&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Blended price per 1M tokens&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$0.90&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Flash starts slightly sooner; the tested flagship endpoint emits tokens faster once generation begins.&lt;/p&gt;

&lt;p&gt;I’d treat these as snapshots of the tested APIs. They don’t establish a permanent speed ranking across providers, and the blended prices are a different comparison from Z.ai’s standard input/output list rates.&lt;/p&gt;

&lt;p&gt;For an agent, I’d also measure the full trajectory. Time to first token and output throughput each describe only part of a run that may include reasoning, tool execution, retries, and validation.&lt;/p&gt;

&lt;h2&gt;
  
  
  Token economics leave room for retries and escalation
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://docs.z.ai/guides/overview/pricing" rel="noopener noreferrer"&gt;Z.ai’s standard pricing&lt;/a&gt; puts Flash at $0.15 per million input tokens and $0.50 per million output tokens. The flagship lists at $1.40 and $4.40.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Z.ai rate&lt;/th&gt;
&lt;th&gt;Flash&lt;/th&gt;
&lt;th&gt;GLM-5.3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Standard input&lt;/td&gt;
&lt;td&gt;$0.15&lt;/td&gt;
&lt;td&gt;$1.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard cached input&lt;/td&gt;
&lt;td&gt;$0.03&lt;/td&gt;
&lt;td&gt;$0.26&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard output&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$4.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Promotional input&lt;/td&gt;
&lt;td&gt;$0.075&lt;/td&gt;
&lt;td&gt;$1.40&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Promotional cached input&lt;/td&gt;
&lt;td&gt;$0.015&lt;/td&gt;
&lt;td&gt;$0.26&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Promotional output&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;td&gt;$4.40&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;All prices are USD per 1M tokens. Promotions can change.&lt;/p&gt;

&lt;p&gt;For a unified multi-model API, CometAPI exposes both model IDs through the same OpenAI-compatible endpoint; the quoted route prices are $0.06 input / $0.20 output for Flash and $1.12 input / $3.528 output for GLM-5.3, with cached pricing requiring a live-route check.&lt;/p&gt;

&lt;p&gt;At those route rates, a workload consuming 100M input tokens and 20M output tokens comes out as follows:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Component&lt;/th&gt;
&lt;th&gt;Flash&lt;/th&gt;
&lt;th&gt;GLM-5.3&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;100 × $0.06 = $6.00&lt;/td&gt;
&lt;td&gt;100 × $1.12 = $112.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;20 × $0.20 = $4.00&lt;/td&gt;
&lt;td&gt;20 × $3.528 = $70.56&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$182.56&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;That is about a 94.5% reduction for Flash before cache effects or tool charges.&lt;/p&gt;

&lt;p&gt;I’d use that difference to fund a measured escalation policy. A cheap request that repeatedly fails can still be expensive, while an expensive request that resolves a difficult task immediately may be worthwhile. Cost per accepted result is the metric I’d optimize.&lt;/p&gt;

&lt;h2&gt;
  
  
  A minimal comparison through one endpoint
&lt;/h2&gt;

&lt;p&gt;This snippet runs the same text prompt against both model IDs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;models&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;glm-5.3-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;glm-5.3&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;models&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
            &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Review this migration plan and identify its three highest-risk assumptions.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a connectivity smoke test. The prompt doesn’t include an actual migration plan, and the request doesn’t explicitly set reasoning effort or output limits.&lt;/p&gt;

&lt;p&gt;For a useful evaluation, I’d supply real task material and hold the prompt, reasoning effort, tool definitions, maximum output, and acceptance rubric constant. I’d record acceptance, token use, latency, tool failures, and human intervention for each run.&lt;/p&gt;

&lt;p&gt;Visual evaluation needs its own setup. Check the live Flash route documentation for the supported image-content schema. A flagship comparison needs a text-only equivalent, which also means the inputs are no longer identical.&lt;/p&gt;

&lt;h2&gt;
  
  
  Open weights still mean substantial deployment work
&lt;/h2&gt;

&lt;p&gt;Both models have downloadable FP8 and BF16 checkpoints. Their weight licenses differ:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://huggingface.co/zai-org/GLM-5.3-Flash/blob/main/LICENSE" rel="noopener noreferrer"&gt;Flash weights use the MIT License&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE" rel="noopener noreferrer"&gt;GLM-5.3 weights use the separate GLM-5.3 License&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;The supporting &lt;a href="https://github.com/zai-org/GLM-5/blob/main/LICENSE" rel="noopener noreferrer"&gt;GLM-5 repository code uses Apache-2.0&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The flagship license requires security review before commercial use by Model-as-a-Service operators whose aggregate revenue exceeds US$10 billion over any consecutive 12 months. I’d read the checkpoint’s license directly rather than infer its terms from the supporting code repository.&lt;/p&gt;

&lt;p&gt;Flash has the smaller total and active footprint, but a 320B checkpoint is still a substantial deployment. Multimodal components, KV cache, and operation at 1M context add to the infrastructure requirements. Neither model becomes inexpensive to host simply because its weights are downloadable.&lt;/p&gt;

&lt;p&gt;I’d consider self-hosting when data control, custom serving, or sustained utilization justifies operating that infrastructure.&lt;/p&gt;

&lt;h2&gt;
  
  
  The routing policy I’d deploy first
&lt;/h2&gt;

&lt;p&gt;My initial policy would stay simple:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Send visual tasks, routine coding, and high-volume automation to Flash.&lt;/li&gt;
&lt;li&gt;Validate the result using the task’s acceptance criteria.&lt;/li&gt;
&lt;li&gt;Escalate difficult or failed text tasks to GLM-5.3.&lt;/li&gt;
&lt;li&gt;Route authorized security work through the flagship with the required approval and isolation controls.&lt;/li&gt;
&lt;li&gt;Adjust routing using measured cost per accepted result.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The outstanding uncertainty is how these results transfer to a particular workload. Several scores are vendor-reported; evaluation tools, prompts, inference settings, and context management can change the outcome. Pricing, availability, rate limits, and route capabilities also need a live check before deployment.&lt;/p&gt;

&lt;p&gt;I’d make Flash the default because its vision support and token economics cover a broad range of work. I’d keep GLM-5.3 available because its stronger difficult-coding and reasoning results give a concrete reason to escalate when validation says the default route isn’t enough.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/glm-5-3-flash-vs-glm-5-3/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=glm-5-3-flash-vs-glm-5-3"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How Much Does Cursor Composer Cost?</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Tue, 22 Sep 2026 05:24:17 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/how-much-does-cursor-composer-cost-1425</link>
      <guid>https://dev.to/dylanfoster1/how-much-does-cursor-composer-cost-1425</guid>
      <description>&lt;p&gt;Cursor Composer is a new, frontier-grade coding model released as part of Cursor 2.0 that delivers much faster, agentic code-generation for complex, multi-file workflows. Access to Composer is governed by Cursor’s existing tiered subscriptions plus token-based usage when you exhaust plan allowances or use Cursor’s “Auto” routing — meaning costs are a mix of a fixed subscription fee and metered token charges. Below you’ll find a full, practical breakdown (features, advantages, pricing mechanics, worked examples and competitor comparisons) so you can estimate real-world costs and decide whether Composer is worth it for your team.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Cursor Composer?
&lt;/h2&gt;

&lt;p&gt;Composer is Cursor’s new “frontier model” introduced as part of Cursor 2.0. It was built and tuned specifically for software engineering workflows and agentic (multi-step) coding tasks. According to Cursor’s announcement, Composer delivers frontier-level coding performance while being optimized for low latency and fast iteration — Cursor says most conversational turns complete in under 30 seconds in practice and claims generation throughput roughly four times that of similarly capable models in their internal benchmarks. Composer was trained with codebase-wide search and tool access so it can reason about, and perform edits across, large projects.&lt;/p&gt;

&lt;h3&gt;
  
  
  Where Composer sits inside Cursor’s product
&lt;/h3&gt;

&lt;p&gt;Composer is not a separate “app” you buy on its own; it’s offered as a model option inside the Cursor product (desktop &amp;amp; web) and is routable through Cursor’s model router (Auto). You get model-level access depending on which Cursor subscription you have and whether you pay metered usage fees beyond your plan’s allowance. Cursor’s model docs list Composer among the available models and the company provides both subscription tiers and token-metering for model usage.&lt;/p&gt;

&lt;p&gt;Cursor’s mid-2025 changes to usage pools and credit systems illustrate this trend: rather than truly unlimited use of premium models, Cursor provides plan allowances (and Auto choices), then bills extra usage at API/token rates.&lt;/p&gt;

&lt;h2&gt;
  
  
  Key features and advantages of Composer
&lt;/h2&gt;

&lt;p&gt;Composer is aimed at developer productivity for nontrivial engineering tasks. The main selling points:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agentic code reasoning:&lt;/strong&gt; Composer supports multi-step workflows (e.g., understanding a bug, searching a repo, editing multiple files, running tests and iterating). This makes it better suited than single-shot completions for complex engineering work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Speed / low latency:&lt;/strong&gt; Cursor reports Composer is significantly faster in generation throughput compared to comparable models and that typical interactive turns finish quickly, enabling faster iteration loops.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tight codebase integration:&lt;/strong&gt; Composer was trained with access to Cursor’s retrieval and editing toolset as well as codebase indexing, which improves its ability to work with large repositories and maintain context across files.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent modes &amp;amp; tools:&lt;/strong&gt; Composer is designed to work with Cursor’s agent modes and the Model Context Protocol (MCP), letting it call specialized tools, read indexed code, and avoid repeatedly re-explaining the project structure. That reduces repetitive token usage in many workflows.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;strong&gt;Why that matters:&lt;/strong&gt; for teams doing deep code edits and multi-file refactors, Composer can reduce manual iteration and context switching — but because it is agentic and can perform more compute work per request, per-request token usage tends to be higher than simple completion models (which drives the metered costs discussion below).&lt;/p&gt;

&lt;h2&gt;
  
  
  How was Composer built ?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Architecture and training approach
&lt;/h3&gt;

&lt;p&gt;Composer is described as an MoE model fine-tuned with reinforcement learning and a custom, large-scale training pipeline. Key elements highlighted by Cursor:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Mixture-of-experts (MoE)&lt;/strong&gt; design to scale capacity efficiently for long-context code tasks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Reinforcement learning (RL)&lt;/strong&gt; with reward signals tuned to agentic behaviors useful in software engineering: plan writing, using search, editing code, writing tests, and maximizing parallel tool use.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Tool-aware training&lt;/strong&gt;: during training Composer had access to a set of tools (file read/write, semantic search, terminal, grep) so it learned to call tools when appropriate and integrate the outputs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Custom infra&lt;/strong&gt;: Cursor built PyTorch + Ray based pipelines, MXFP8 MoE kernels, and large VM clusters to enable asynchronous, tool-enabled RL at scale. The infra choices (low-precision training, expert parallelism) are intended to reduce communication costs and keep inference latency low.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Why moE + RL matters for code
&lt;/h3&gt;

&lt;p&gt;Code editing requires precise, multi-step reasoning over large repositories. MoE gives the model episodic capacity (lots of parameters available selectively) while RL optimizes for behaviors (don’t hallucinate, run tests, propose minimal diffs). Training with the agent toolset means Composer is not being fine-tuned purely on next-token prediction — it learned to use the tooling available in Cursor’s product setting. That’s why Cursor positions Composer as an “agentic” model rather than just a completion model.&lt;/p&gt;

&lt;h2&gt;
  
  
  How are Cursor subscription plans priced for Composer?
&lt;/h2&gt;

&lt;p&gt;Cursor’s pricing combines &lt;strong&gt;subscription tiers&lt;/strong&gt; (monthly plans) with &lt;strong&gt;usage-based charges&lt;/strong&gt; (tokens, cache, and certain agent/tool fees). The subscription tiers give you base capabilities and included, prioritized usage; the heavy or premium-model usage is then billed on top. Below are the public list prices and the high-level meaning of each plan.&lt;/p&gt;

&lt;h3&gt;
  
  
  Individual (personal) tiers
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Hobby (Free):&lt;/strong&gt; entry-level, limited agent requests / tab completions; includes a short Pro trial. Good for light experimentation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pro — $20 / month (individual):&lt;/strong&gt; everything in Hobby plus extended agent usage, unlimited tab completions, background agents, and maximum context windows. This is the common starting point for individual developers who want Composer-level features.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pro+ — $60 / month (individual, recommended for power users):&lt;/strong&gt; more included usage on premium models . Cursor’s June 2025 pricing rollout clarified that Pro plans include a pool of model credits (for “frontier model” usage) and that additional usage can be purchased at cost-plus rates or via token billing.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Ultra — $200 / month:&lt;/strong&gt; for heavy individuals needing substantially larger included model usage and priority access.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Team / Enterprise
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Teams — $40 / user / month:&lt;/strong&gt; adds centralized billing, usage analytics, role-based controls and SSO. Larger teams can also buy Enterprise (custom pricing) that includes pooled usage, invoice/PO billing, SCIM, audit logs and priority support.&lt;/p&gt;

&lt;h2&gt;
  
  
  Token-Based Pricing for Cursor Composer
&lt;/h2&gt;

&lt;p&gt;Cursor mixes per-user plans with per-token billing for premium or agentic requests. There are two related but distinct billing contexts to understand:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Auto / Max mode token rates&lt;/strong&gt; (Cursor’s “Auto” dynamic selection or Max/Auto billing buckets).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model-list / direct model pricing&lt;/strong&gt; (if you select a model like Composer directly, the model list APIs have per-model token rates).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;These different modes change the effective input/output token rates you’ll see on your bill. Below are the canonical figures Cursor publishes in its documentation and model pages — these are the most load-bearing numbers for cost calculations.&lt;/p&gt;

&lt;h3&gt;
  
  
  Auto / Max
&lt;/h3&gt;

&lt;p&gt;When you go beyond plan allowances (or explicitly use Auto to route to premium models), Cursor charges for model usage on a &lt;strong&gt;per-token&lt;/strong&gt; basis. The most commonly referenced rates for Cursor’s &lt;strong&gt;Auto&lt;/strong&gt; router (which picks a premium model on demand) are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input + Cache Write:&lt;/strong&gt; &lt;strong&gt;$1.25 per 1,000,000 tokens&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output (generation):&lt;/strong&gt; &lt;strong&gt;$6.00 per 1,000,000 tokens&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Cache Read:&lt;/strong&gt; &lt;strong&gt;$0.25 per 1,000,000 tokens&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those rates were documented in Cursor’s account/pricing docs describing Auto billing and are the backbone of Composer’s operating cost when Composer usage is billed via Auto or when you directly select model usage charged at API rates.&lt;/p&gt;

&lt;h3&gt;
  
  
  Composer and model-list prices
&lt;/h3&gt;

&lt;p&gt;Cursor’s model list / model-pricing reference shows per-model pricing entries. For some premium models inside Cursor, Composer in model-list prices : &lt;strong&gt;Input $1.25 / 1M; Output $10.00 / 1M&lt;/strong&gt;. In practice this means if you explicitly choose Composer as the model rather than running Auto, the output token rate you incur could be higher than Auto’s $6 output rate&lt;/p&gt;

&lt;h3&gt;
  
  
  Why input vs output tokens differ
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input tokens&lt;/strong&gt; are the tokens you send (prompts, instructions, code snippets, file context). Cursor charges for writing those into the system (and occasionally caching them).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output tokens&lt;/strong&gt; are what the model generates (the code edits, suggestions, diffs, etc.). Output generation is more expensive because it consumes more compute. Cursor’s published numbers reflect those relative costs.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Comparing Cursor Composer with competitors
&lt;/h2&gt;

&lt;p&gt;When judging cost and value, it’s useful to compare Composer’s unit economics to other widely-used developer AI services. Note that model capabilities, latency, integration, and included plan allowances also matter — price alone isn’t the whole story.&lt;/p&gt;

&lt;h3&gt;
  
  
  GitHub Copilot (individual tiers)
&lt;/h3&gt;

&lt;p&gt;GitHub Copilot is primarily priced per user with tiers (Free, Pro at ~$10/month, Pro+ and Business tiers higher). Copilot provides a number of “premium” requests per month and charges for additional premium requests (published per-request add-ons). Copilot bundles models (including Google/Anthropic/OpenAI options in some plans) and is sold as a per-developer SaaS. For many individual devs, Copilot’s all-in per-seat price can be simpler and cheaper for routine completions; for heavy multi-step agentic tasks, a token-metered model may be more transparent.&lt;/p&gt;

&lt;h3&gt;
  
  
  OpenAI (API / advanced models)
&lt;/h3&gt;

&lt;p&gt;OpenAI’s higher-end models (GPT-5 series and premium variants) have different per-token economics that can be higher than Cursor’s Composer rate for certain pro models. OpenAI also provides many performance tiers (and batch or cached discounts) that affect effective costs. If comparing, consider latency, accuracy on coding tasks, and the value of Cursor’s editor integration (which may offset a per-token cost delta).&lt;/p&gt;

&lt;h3&gt;
  
  
  Which is cheaper in practice?
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Small, frequent completions / autocompletes:&lt;/strong&gt; A per-seat SaaS (Copilot) is often cheapest and simplest.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Large multi-file, agentic tasks:&lt;/strong&gt; Token-metered models (Composer via Cursor Auto or Anthropic/OpenAI directly) give flexibility/quality but cost more per heavy request; careful modeling of token use is essential.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion — Is Composer “expensive”?
&lt;/h2&gt;

&lt;p&gt;Composer is &lt;strong&gt;not&lt;/strong&gt; billed as a single flat line-item — it’s part of a hybrid system. For light-to-moderate interactive use, a &lt;strong&gt;$20/month Pro&lt;/strong&gt; plan plus Auto-mode usage may keep your costs low (tens of dollars a month). For heavy, parallel agent workloads with many long outputs, Composer can drive hundreds or thousands per month because output-token rates and concurrency multiply costs. Compared to subscription-first competitors (e.g., GitHub Copilot), Cursor’s Composer trades a higher marginal inference cost for much faster, agentic, repository-aware capabilities.&lt;/p&gt;

&lt;p&gt;If your goals are multi-agent automation, repo-wide refactors, or shorter iteration cycles that save engineering time, Composer’s speed and tooling can deliver strong ROI.&lt;/p&gt;

&lt;h2&gt;
  
  
  How do I use CometAPI inside Cursor? (step-by-step)
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;Short summary: CometAPI is a model-aggregation gateway (single endpoint that can proxy many model vendors). To use it in Cursor you register at CometAPI, get an API key and model identifier, then add that key + endpoint into Cursor’s Models settings as a custom provider (override base URL) and select the CometAPI model in Composer/Agent mode.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;CometAPI also designed a proprietary coding model based on Claude specifically for cursor: &lt;code&gt;cometapi-sonnet-4-5-20250929-thinking&lt;/code&gt; and &lt;code&gt;cometapi-opus-4-1-20250805-thinking&lt;/code&gt; etc.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step A — Get your CometAPI credentials
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Sign up at CometAPI and&amp;nbsp;&lt;a href="https://www.cometapi.com/console/token" rel="noopener noreferrer"&gt;create an API key&lt;/a&gt;&amp;nbsp;from their dashboard. Keep the key secret (treat it like any bearer token).&lt;/li&gt;
&lt;li&gt;Create / copy an API key and note the model name/ID you want to use (e.g.,&amp;nbsp;&lt;code&gt;claude-sonnet-4.5&lt;/code&gt;&amp;nbsp;or another vendor model available via CometAPI).&lt;a href="https://www.cometapi.com/pricing/" rel="noopener noreferrer"&gt;CometAPI docs/guides&lt;/a&gt;&amp;nbsp;describe the process and list supported model names.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Step B — Add CometAPI as a custom model/provider in Cursor
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Open Cursor →&amp;nbsp;&lt;strong&gt;Settings&lt;/strong&gt;&amp;nbsp;→&amp;nbsp;&lt;strong&gt;Models&lt;/strong&gt;&amp;nbsp;(or Settings → API Keys).&lt;/li&gt;
&lt;li&gt;If Cursor shows an&amp;nbsp;&lt;strong&gt;“Add Custom Model”&lt;/strong&gt;&amp;nbsp;or&amp;nbsp;&lt;strong&gt;“Override OpenAI Base URL”&lt;/strong&gt;&amp;nbsp;option, use it:&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Base URL / Endpoint&lt;/strong&gt;: paste the CometAPI OpenAI-compatible base URL (CometAPI will document whether they expose an&amp;nbsp;&lt;code&gt;openai/v1&lt;/code&gt;&amp;nbsp;style endpoint or a provider-specific endpoint). (Example:&amp;nbsp;&lt;code&gt;https://api.cometapi.com/v1&lt;/code&gt;&amp;nbsp;— use the actual URL from CometAPI docs.)&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;API Key&lt;/strong&gt;: paste your CometAPI key in the API key field.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model name&lt;/strong&gt;: add the model identifier exactly as CometAPI documents (e.g.,&amp;nbsp;&lt;code&gt;claude-sonnet-4.5&lt;/code&gt;&amp;nbsp;or&amp;nbsp;&lt;code&gt;composer-like-model&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Verify&lt;/strong&gt;&amp;nbsp;the connection if Cursor offers a “Verify” / “Test” button. Cursor’s custom model mechanism commonly requires the provider to be OpenAI-compatible (or for Cursor to accept a base URL + key). Community guides show the same pattern (override base URL → provide key → verify).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;If you want to know more tips, guides and news on AI follow us on&amp;nbsp;&lt;a href="https://vk.com/id1078176061" rel="noopener noreferrer"&gt;VK&lt;/a&gt;,&amp;nbsp;&lt;a href="https://x.com/cometapi2025" rel="noopener noreferrer"&gt;X&lt;/a&gt;&amp;nbsp;and&amp;nbsp;&lt;a href="https://discord.com/invite/HMpuV6FCrG" rel="noopener noreferrer"&gt;Discord&lt;/a&gt;!&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;See also &lt;a href="https://www.cometapi.com/cursor-2-0-what-changed-and-why-it-matters/" rel="noopener noreferrer"&gt;Cursor 2.0 and Composer: how a multi-agent rethink surprised AI coding&lt;/a&gt;&lt;/strong&gt;&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-much-does-cursor-composer-cost/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-much-does-cursor-composer-cost"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Connecting Raycast AI to an OpenAI-Compatible Gateway</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Tue, 22 Sep 2026 02:09:07 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/connecting-raycast-ai-to-an-openai-compatible-gateway-4ae7</link>
      <guid>https://dev.to/dylanfoster1/connecting-raycast-ai-to-an-openai-compatible-gateway-4ae7</guid>
      <description>&lt;p&gt;I want model selection to be part of my desktop workflow, not a reason to open another application. Raycast’s custom providers make that possible: configure an OpenAI-compatible endpoint, supply your own API key, and select a model inside Raycast AI.&lt;/p&gt;

&lt;p&gt;For this setup, I’m using CometAPI, a unified gateway exposing hundreds of models through an OpenAI-style REST API. The integration hinges on &lt;code&gt;https://api.cometapi.com/v1&lt;/code&gt;, Raycast’s &lt;code&gt;providers.yaml&lt;/code&gt;, and a valid token. The main thing to watch is version-specific configuration: use the template shipped with your installed Raycast version.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decide What You’re Connecting
&lt;/h2&gt;

&lt;p&gt;Raycast is a macOS productivity launcher with commands, scripts, and AI features. Its AI surfaces include Quick AI for launcher prompts, AI Chat for conversations with attachments and context, and AI Commands or extensions for reusable workflows. It also supports local models through Ollama and remote providers through Bring Your Own Key (BYOK) and custom providers.&lt;/p&gt;

&lt;p&gt;Raycast added BYOK in &lt;strong&gt;v1.100.0&lt;/strong&gt;, with BYOK and Custom Providers rolling out during &lt;strong&gt;2025&lt;/strong&gt;. I’d start with a recent release and check &lt;strong&gt;Settings → AI&lt;/strong&gt; for the controls available in that installation.&lt;/p&gt;

&lt;p&gt;The gateway’s catalog spans text, images, embeddings, audio, and video, but I would not treat that catalog as a list of capabilities automatically available in Raycast AI. Chat integration and endpoint-specific integrations are different pieces of work.&lt;/p&gt;

&lt;p&gt;For code explanation, refactor suggestions, unit tests, PR summaries, and README drafts, the chat path is the relevant starting point. Image generation needs an extension that calls the image endpoint. Semantic search needs an embedding index and a script or cloud function that Raycast can query. Audio support, including TTS and STT, depends on the underlying model; specialized video backends include Sora and Veo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check the Prerequisites First
&lt;/h2&gt;

&lt;p&gt;You need macOS, a recent Raycast installation with the relevant custom-provider controls, and an account with a valid gateway API key. Create a token in the provider’s console and keep it out of shared configuration.&lt;/p&gt;

&lt;p&gt;Raycast also needs HTTPS access to &lt;code&gt;api.cometapi.com&lt;/code&gt;. On a corporate network, check the proxy and firewall before debugging YAML. Terminal and cURL are enough for a basic connectivity check; Python, Node, and OpenAI SDKs are optional tools for more detailed testing. Standard OpenAI-style clients can use the gateway by overriding &lt;code&gt;base_url&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;I’d test authentication and model-list access before touching Raycast. With your token already available locally as &lt;code&gt;GATEWAY_API_KEY&lt;/code&gt;, this requests the OpenAI-style models endpoint:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;--fail-with-body&lt;/span&gt; &lt;span class="nt"&gt;--silent&lt;/span&gt; &lt;span class="nt"&gt;--show-error&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s1"&gt;'https://api.cometapi.com/v1/models'&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;--header&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;GATEWAY_API_KEY&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This checks that endpoint, not a complete chat workflow. Automatic model discovery still depends on the provider exposing a compatible models response and your Raycast version supporting discovery.&lt;/p&gt;

&lt;h2&gt;
  
  
  Configure From Raycast’s Own Template
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Locate the Configuration
&lt;/h3&gt;

&lt;p&gt;Open &lt;strong&gt;Preferences → AI&lt;/strong&gt;, find &lt;strong&gt;Custom Providers&lt;/strong&gt; or &lt;strong&gt;Custom OpenAI-compatible APIs&lt;/strong&gt;, and choose &lt;strong&gt;Reveal Providers Config&lt;/strong&gt;. Use the directory Raycast reveals rather than guessing a configuration path.&lt;/p&gt;

&lt;p&gt;Raycast provides a template, usually named &lt;code&gt;providers.template.yaml&lt;/code&gt;. Copy it to &lt;code&gt;providers.yaml&lt;/code&gt; in that directory, or edit the existing configuration if you already have custom providers.&lt;/p&gt;

&lt;p&gt;The exact schema can differ across releases. Common provider entries include &lt;code&gt;id&lt;/code&gt;, &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;base_url&lt;/code&gt;, and an optional &lt;code&gt;models&lt;/code&gt; block, but the installed template should determine their nesting and syntax. I would not paste a supposedly universal YAML example over that template.&lt;/p&gt;

&lt;h3&gt;
  
  
  Set the Endpoint and Credentials
&lt;/h3&gt;

&lt;p&gt;Add a provider entry using the template’s structure. Give it a distinct identifier and a recognizable display name, then set &lt;code&gt;base_url&lt;/code&gt; to &lt;code&gt;https://api.cometapi.com/v1&lt;/code&gt;. There is no trailing period in that URL.&lt;/p&gt;

&lt;p&gt;Add the token through Raycast’s custom API-key or secure credential fields where supported. Never commit a real token in a shared &lt;code&gt;providers.yaml&lt;/code&gt;. macOS Keychain is another option where the integration supports it; environment-variable injection is appropriate for a local proxy you control, not something to assume Raycast’s YAML supports automatically.&lt;/p&gt;

&lt;p&gt;You may not need to enumerate every model manually. Raycast can discover models through a properly implemented OpenAI-style &lt;code&gt;GET /v1/models&lt;/code&gt; endpoint when that discovery path is supported. Otherwise, follow the template’s model configuration and use exact model identifiers from the provider.&lt;/p&gt;

&lt;h3&gt;
  
  
  Reload and Run a Small Test
&lt;/h3&gt;

&lt;p&gt;Return to Raycast and refresh models if your version offers that action. If the provider or models do not appear, restart the app.&lt;/p&gt;

&lt;p&gt;Open Quick AI, explicitly choose a model from the new provider, and submit a short prompt. I’d keep this first request minimal: it should establish that credentials, model selection, and the response path work before adding attachments, long context, or tools.&lt;/p&gt;

&lt;h2&gt;
  
  
  Debug the Boundary That Failed
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;No models in the picker:&lt;/strong&gt; Check that &lt;code&gt;providers.yaml&lt;/code&gt; is in the exact directory opened by &lt;strong&gt;Reveal Providers Config&lt;/strong&gt;. Compare its structure with the installed template, then refresh or restart. If discovery is unavailable, check whether explicit model entries are required.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;401 or invalid-token responses:&lt;/strong&gt; Confirm that the token is valid and has not expired. Run the direct request above and verify that authentication uses &lt;code&gt;Authorization: Bearer …&lt;/code&gt;. A failed direct request is a reason to resolve credentials or endpoint access before changing Raycast settings.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Model-specific failures:&lt;/strong&gt; Verify the model ID first. OpenAI compatibility can still leave differences in response shapes or streaming behavior. If a request fails while streaming, test a non-streaming request directly to narrow the problem, then raise the incompatibility with the provider if needed.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Slow responses:&lt;/strong&gt; Measure the models you actually intend to use. Gateway routing can introduce variable latency, and a model that works well for a long reasoning task may be a poor choice for an interactive launcher command.&lt;/p&gt;

&lt;h2&gt;
  
  
  Make the Workflow Worth Keeping
&lt;/h2&gt;

&lt;p&gt;I’d use lightweight, fast models for selected-text summaries, action-item extraction, and short lookups. For deeper reasoning, choose a higher-capacity model; for larger inputs, evaluate context capacity as well. Model selection is a practical cost and latency control, not just a preference.&lt;/p&gt;

&lt;p&gt;Track usage in the gateway dashboard and configure budget alerts where available. Shorter system messages and deliberate context management can reduce token use without turning every prompt into an optimization exercise.&lt;/p&gt;

&lt;p&gt;For recurring tasks, duplicate a built-in Raycast AI Command and adjust its prompt. Utility commands benefit from predictable instructions and consistent output formats; ideation commands can be more open-ended. Documentation drafting and PR summaries are useful starting points because the output is easy to inspect.&lt;/p&gt;

&lt;p&gt;Finally, local credentials do not make remote inference local. Before sending sensitive code, notes, or attachments, read both Raycast’s and the provider’s privacy documentation. I’d keep the initial setup narrow: one tested model, one useful command, and a clear understanding of where the request goes.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-use-cometapi-in-raycast/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-use-cometapi-in-raycast"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How to Prompt GPT Image 2.5 Like a Pro: The Complete 2026 Guide</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Tue, 22 Sep 2026 01:29:19 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/how-to-prompt-gpt-image-25-like-a-pro-the-complete-2026-guide-1ll0</link>
      <guid>https://dev.to/dylanfoster1/how-to-prompt-gpt-image-25-like-a-pro-the-complete-2026-guide-1ll0</guid>
      <description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt; GPT Image 2.5 changes image prompting in one important way: a good prompt no longer describes only what you want to create. For editing and reference-based generation, it should also define what must not change.&lt;/p&gt;

&lt;p&gt;That matters because &lt;a href="https://openai.com/index/introducing-chatgpt-images-2-5/" rel="noopener noreferrer"&gt;GPT Image 2.5 is designed for more precise editing and stronger reference preservation&lt;/a&gt;. OpenAI says the model is better at changing only the requested element while keeping the subject, composition, and surrounding details intact.&lt;/p&gt;

&lt;p&gt;Good GPT Image 2.5 Prompt = Deliverable + Subject + Scene + Composition + Visual Direction + Exact Text + Constraints + Output&lt;/p&gt;

&lt;p&gt;Good Editing Prompt = Change Target + Preserve List + Integration Rules + Exclusions + Output&lt;/p&gt;

&lt;p&gt;The second formula is especially important. Instead of saying “Change the background to a beach,” write: “Replace only the background with a quiet Mediterranean beach at sunset. Keep the person, face, pose, clothing, camera position, framing, lighting direction, and foreground unchanged.”&lt;/p&gt;

&lt;h2&gt;
  
  
  Key Takeaways
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://openai.com/index/introducing-chatgpt-images-2-5/" rel="noopener noreferrer"&gt;GPT Image 2.5 improves fidelity, multi-turn consistency&lt;/a&gt;, text rendering, and speed; Flare is the recommended starting point for most users.&lt;/li&gt;
&lt;li&gt;Core generation formula: Deliverable + Subject + Composition + Style + Exact Text (in quotes) + Constraints/Exclusions.&lt;/li&gt;
&lt;li&gt;For edits the golden rule is “Change only X. Preserve [full list of identity, pose, lighting, geometry, text, style].”&lt;/li&gt;
&lt;li&gt;Multiple reference images succeed when each is given an explicit job (subject, clothing, background, palette, etc.).&lt;/li&gt;
&lt;li&gt;Transparent backgrounds, exact typography, and sequential editing are markedly stronger when prompted correctly and parameters are set properly.&lt;/li&gt;
&lt;li&gt;Flare prioritizes speed and volume; Sunburst prioritizes edit precision and final-asset fidelity. Both share the same token rates.&lt;/li&gt;
&lt;li&gt;CometAPI offers unified, cost-effective access to &lt;code&gt;gpt-image-2.5-flare&lt;/code&gt; and &lt;code&gt;gpt-image-2.5-sunburst&lt;/code&gt; via familiar OpenAI-compatible routes.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  What Is GPT Image 2.5
&lt;/h2&gt;

&lt;p&gt;OpenAI released ChatGPT Images 2.5 on September 8, 2026. The update focuses on sharper detail, more natural lighting and texture, reference-image fidelity, precise edits, and stronger consistency across multiple editing turns. OpenAI also reports up to 50% lower image-generation latency compared with Images 2.0.&lt;/p&gt;

&lt;p&gt;For developers, the family is split into two API models:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;a href="https://developers.openai.com/api/docs/models/gpt-image-2.5-flare" rel="noopener noreferrer"&gt;GPT-Image-2.5 Flare&lt;/a&gt; — optimized for fast, high-quality everyday generation and high-volume workflows.&lt;/li&gt;
&lt;li&gt;
&lt;a href="https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst" rel="noopener noreferrer"&gt;GPT-Image-2.5 Sunburst&lt;/a&gt; — optimized for demanding generation and editing workflows where precision matters more than latency.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Developers can also use the &lt;a href="https://www.cometapi.com/what-is-gpt-image-2-5/" rel="noopener noreferrer"&gt;GPT Image 2.5 API in CometAPI&lt;/a&gt;, including the &lt;code&gt;gpt-image-2.5-flare&lt;/code&gt; and &lt;code&gt;gpt-image-2.5-sunburst&lt;/code&gt; model identifiers.&lt;/p&gt;

&lt;h3&gt;
  
  
  Benchmark and capabilities
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Arena text-to-image benchmark&lt;/th&gt;
&lt;th&gt;GPT Image 2.5 Sunburst&lt;/th&gt;
&lt;th&gt;GPT Image 2.5 Flare&lt;/th&gt;
&lt;th&gt;GPT Image 2&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Text-to-image score&lt;/td&gt;
&lt;td&gt;1421 ± 13&lt;/td&gt;
&lt;td&gt;1399 ± 13&lt;/td&gt;
&lt;td&gt;1381 ± 4&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Text-to-image rank&lt;/td&gt;
&lt;td&gt;#1&lt;/td&gt;
&lt;td&gt;#2&lt;/td&gt;
&lt;td&gt;#3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single-image-edit score&lt;/td&gt;
&lt;td&gt;1520 ± 9&lt;/td&gt;
&lt;td&gt;1491 ± 9&lt;/td&gt;
&lt;td&gt;1461 ± 3&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single-image-edit rank&lt;/td&gt;
&lt;td&gt;#1&lt;/td&gt;
&lt;td&gt;#2&lt;/td&gt;
&lt;td&gt;#3&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Independent tests and early Arena rankings placed both models at the top of text-to-image and image-editing leaderboards shortly after launch, with clear gains in composition, detail, style fidelity, subject preservation, and speed (Flare often finishing in ~20–40 seconds vs. significantly longer prior generations in some benchmarks).&lt;/p&gt;

&lt;p&gt;Key capabilities include text + image inputs, image outputs up to roughly 4K (with constraints: edges multiples of 16, max ~3840 px per side, aspect ratio ≤ 3:1), transparent backgrounds (PNG/WebP), quality tiers (low → max), progressive previews, inpainting/area-specific editing, and strong support for multi-reference and sequential editing. OpenAI reports more than 3 billion images created weekly across ChatGPT Images and the GPT-Image API family.&lt;/p&gt;

&lt;h2&gt;
  
  
  General Formula for Generating and Editing GPT Image 2.5
&lt;/h2&gt;

&lt;h3&gt;
  
  
  The Most Important Rule: Change vs Preserve
&lt;/h3&gt;

&lt;p&gt;This distinction is the single highest-leverage habit for GPT Image 2.5.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Generation (new image from text or references)&lt;/strong&gt;&lt;br&gt;
Describe everything you want to &lt;em&gt;appear&lt;/em&gt;. Be concrete about the subject, framing, lighting, materials, colors, and what must &lt;em&gt;not&lt;/em&gt; appear.&lt;/p&gt;

&lt;h3&gt;
  
  
  Editing (image-to-image or multi-turn refinement)
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Name &lt;strong&gt;exactly one change&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Explicitly list &lt;strong&gt;everything that must remain unchanged&lt;/strong&gt; (identity, pose, camera angle, framing, lighting direction, shadows, background elements, existing text, colors, geometry, overall style).&lt;/li&gt;
&lt;li&gt;Describe how the new element should integrate with the existing light, perspective, texture, and contact shadows.&lt;/li&gt;
&lt;li&gt;Add clear exclusions (no extra text, no logos, no watermarks, no heavy retouching).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Vague multi-change prompts or missing preserve lists cause cumulative drift. Restate the preserve list on every turn. When a region must stay pixel-identical, composite the approved local edit back into the original master rather than trusting the model alone.&lt;/p&gt;

&lt;h3&gt;
  
  
  Minimal generation skeleton
&lt;/h3&gt;

&lt;p&gt;[Deliverable / intended use] + Subject + Composition + Style / lighting / materials + Exact text in quotation marks + Constraints and exclusions&lt;/p&gt;

&lt;h3&gt;
  
  
  Minimal edit skeleton
&lt;/h3&gt;

&lt;p&gt;Change only [specific element] to [new description]. Preserve [identity, pose, framing, lighting, shadows, background, text, colors, style]. Match the new element to existing light and geometry. Do not add [unwanted elements].&lt;/p&gt;

&lt;h2&gt;
  
  
  How GPT Image 2.5 Prompting Works
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://developers.openai.com/api/docs/guides/image-prompting" rel="noopener noreferrer"&gt;OpenAI’s official prompting guidance&lt;/a&gt; is straightforward: start with the image you need, then describe the subject, composition, style, and constraints. For edits, clearly separate the requested change from the details that must stay the same. Refine one thing at a time and inspect every result.Structure prompts as Subject → Composition → Style → Text (quoted) → Constraints. Assign clear roles to every reference image, iterate one change at a time, and set size/quality/background as API parameters rather than prompt text.&lt;/p&gt;

&lt;h3&gt;
  
  
  Subject
&lt;/h3&gt;

&lt;p&gt;Name the primary person, object, or scene with visible, testable details: appearance, clothing, pose, action, relative scale, and interaction with objects.&lt;br&gt;
Example: “A young archivist with a short auburn bob, round black glasses, and a navy work coat holding a sealed paper envelope with a red wax moon emblem.”&lt;/p&gt;

&lt;h3&gt;
  
  
  Composition
&lt;/h3&gt;

&lt;p&gt;Specify framing, camera height and angle, subject placement (centered, left third, etc.), negative space, depth of field, and what must be visible (full body with feet, medium close-up, etc.).&lt;br&gt;
Example: “Full body, feet visible, low eye-level camera. Character on the left third; endless archive shelves recede behind her. Clean dark space at upper right for later title placement.”&lt;/p&gt;

&lt;h3&gt;
  
  
  Style
&lt;/h3&gt;

&lt;p&gt;Describe the visual medium, lighting (source + direction), materials, color palette, texture, and realism level. Prefer concrete visual language over pure mood adjectives. Explicitly request “photorealistic,” “real photograph,” film grain, or a named artistic style when required.&lt;br&gt;
Example: “Detailed hand-painted anime background, restrained cel shading, natural proportions. One warm desk lamp against cool blue moonlight from high windows.”&lt;/p&gt;

&lt;h3&gt;
  
  
  Text
&lt;/h3&gt;

&lt;p&gt;Place every required string inside quotation marks. Specify location, hierarchy, approximate typography, and how many times the text should appear. Spell unusual or brand names letter-by-letter when critical. Always request “no other text.”&lt;br&gt;
Example: the headline "Open Late" in bold condensed type across the top, and "Thursday to Sunday" in smaller type at the bottom. No other text.&lt;/p&gt;

&lt;h3&gt;
  
  
  Constraints / Exclusions
&lt;/h3&gt;

&lt;p&gt;List what must not appear and what must remain unchanged. High-value exclusions include no extra text, no logos, no watermarks, no heavy retouching, and no unwanted objects.&lt;/p&gt;

&lt;p&gt;For longer or more complex requests, use labeled sections (SCENE / SUBJECT / STYLE / TEXT / CONSTRAINTS). This improves readability and makes later edits easier to maintain.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Critical technical note:&lt;/strong&gt; Set model, quality, size (or aspect), and background (auto / opaque / transparent) as API parameters. Do not rely on the prompt text to enforce resolution or transparency.&lt;/p&gt;

&lt;h2&gt;
  
  
  How Do You Use Multiple Reference Images With GPT Image 2.5?
&lt;/h2&gt;

&lt;p&gt;GPT Image 2.5 accepts multiple image inputs. Results improve dramatically when every reference is given an explicit role.&lt;/p&gt;

&lt;p&gt;Recommended pattern:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Number or clearly name each reference image.&lt;/li&gt;
&lt;li&gt;State its precise job: primary subject identity, clothing/style reference, background/setting, color palette only, lighting reference, etc.&lt;/li&gt;
&lt;li&gt;Describe how the elements should combine and which parts move where.&lt;/li&gt;
&lt;li&gt;Restate preservation constraints for the primary subject and any protected details.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Example: “Use image 1 as the product (matte black water bottle with exact label text). Use image 2 only for the pale concrete ledge and natural outdoor lighting direction. Place the bottle from image 1 on the ledge from image 2, matching the existing soft shadows and light angle. Preserve the bottle’s exact shape, label text, colors, and reflections. Do not change the concrete texture or add any extra objects.”&lt;/p&gt;

&lt;p&gt;Clear role assignment prevents unwanted hybridization of subjects or styles.&lt;/p&gt;

&lt;h2&gt;
  
  
  Advanced Prompting Techniques
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Exact Text &amp;amp; Typography Tutorial
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;Always enclose required strings in quotation marks.&lt;/li&gt;
&lt;li&gt;Specify position, hierarchy, approximate weight and style, and the number of times the text should appear.&lt;/li&gt;
&lt;li&gt;Request “no other text,” “no logos,” and “no watermarks.”&lt;/li&gt;
&lt;li&gt;For dense layouts or multiple fonts, test medium or high quality and visually verify spelling and legibility in the output.&lt;/li&gt;
&lt;li&gt;Critical brand names can be spelled letter-by-letter for extra reliability.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Transparent Backgrounds
&lt;/h3&gt;

&lt;p&gt;Set background="transparent" in the API call (or the equivalent UI control) and request PNG or WebP output. Explicitly ask for clean edges and nothing behind the subject. Inspect the alpha channel carefully around hair, glass, fine details, and contact areas. Add exclusions such as “no background elements, no drop shadows unless requested.” Transparent generation is significantly more reliable in 2.5 when the parameter and prompt constraints are both correct.&lt;/p&gt;

&lt;h3&gt;
  
  
  Multi-turn Editing Workflow
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Start from a strong base generation or approved reference.&lt;/li&gt;
&lt;li&gt;Feed the previous output back as the new input image.&lt;/li&gt;
&lt;li&gt;Request exactly one change.&lt;/li&gt;
&lt;li&gt;Restate the complete preserve list on every turn.&lt;/li&gt;
&lt;li&gt;Inspect the result before continuing.&lt;/li&gt;
&lt;li&gt;For regions that must remain pixel-perfect, composite the approved local edit into the original master image.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This disciplined loop takes full advantage of GPT Image 2.5’s improved multi-turn consistency while minimizing cumulative drift.&lt;/p&gt;

&lt;h2&gt;
  
  
  Flare vs. Sunburst: Which Model Fits Which Prompt?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Criterion&lt;/th&gt;
&lt;th&gt;GPT-Image-2.5 Flare&lt;/th&gt;
&lt;th&gt;GPT-Image-2.5 Sunburst&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Primary strength&lt;/td&gt;
&lt;td&gt;Speed + high everyday quality&lt;/td&gt;
&lt;td&gt;Edit precision &amp;amp; final-asset fidelity&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Up to ~50% lower than GPT Image 2&lt;/td&gt;
&lt;td&gt;Longer generation times&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Best use cases&lt;/td&gt;
&lt;td&gt;Social content, prototyping, high-volume, rapid iteration&lt;/td&gt;
&lt;td&gt;Production campaign creative, polished product shots, complex multi-turn edits&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Quality vs GPT Image 2&lt;/td&gt;
&lt;td&gt;Higher&lt;/td&gt;
&lt;td&gt;Highest control / tightest preservation&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Recommended starting point&lt;/td&gt;
&lt;td&gt;Most workflows&lt;/td&gt;
&lt;td&gt;When quality or edit precision is the bottleneck&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token pricing&lt;/td&gt;
&lt;td&gt;Same rates for both&lt;/td&gt;
&lt;td&gt;Same rates for both&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Transparent backgrounds &amp;amp; quality tiers&lt;/td&gt;
&lt;td&gt;Fully supported (low → max)&lt;/td&gt;
&lt;td&gt;Fully supported&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Practical decision rule:&lt;/strong&gt; Explore ideas, generate variations, and iterate quickly on Flare. When you need the final polished version or when a complex edit sequence begins to drift, re-run the identical prompt + references + settings on Sunburst and compare. Many production teams keep both models available and route traffic accordingly.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prompt Templates (Ready to Adapt)
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Clean Product Hero
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Purpose: Ecommerce or landing-page hero, square or 4:5.
Subject: [exact product description including materials and any label text in quotes].
Composition: Centered or slight three-quarter view, soft contact shadow, generous negative space at top for headline.
Style: Clean studio lighting from the left, photorealistic, high texture detail.
Constraints: No extra text, no logos, no people reflections, seamless background.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  2. Exact-Text Poster / Ad
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;A vertical poster. Headline “YOUR EXACT HEADLINE” in bold condensed sans-serif across the upper third. Subhead “Secondary line here” in smaller weight directly below. [Full scene and subject description]. [Style]. No other text or logos.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. Precise Single-Element Edit
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Change only the [specific element] to [new description]. Preserve the subject’s exact face/identity, pose, clothing details, camera angle, lighting direction, shadows, background, and all existing text. Match the new element’s lighting, perspective, and contact shadows to the original scene. Do not add any new objects, text, or watermarks.
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4. Multi-Reference Character Consistency
&lt;/h3&gt;

&lt;p&gt;Use image 1 for the character’s face, body proportions, and identity. Use image 2 only for the outfit and fabric texture. Place the character from image 1 wearing the outfit from image 2 inside [scene description]. Match the lighting direction of image 1. Preserve identity and pose exactly. No extra accessories or text.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Transparent Cutout / Asset
&lt;/h3&gt;

&lt;p&gt;[Detailed subject description] isolated on a pure transparent background. Clean edges, no drop shadow, no background elements, no floor contact shadow unless explicitly requested. Photorealistic. Output ready for compositing.&lt;/p&gt;

&lt;p&gt;Replace the bracketed sections with your specifics and always restate constraints on subsequent editing turns.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why Prompts Fail
&lt;/h2&gt;

&lt;p&gt;Common failure patterns and how to fix them:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Relying on vague mood words (“cozy,” “epic,” “premium”) without attached visual detail → add concrete scale, light direction, materials, and color.&lt;/li&gt;
&lt;li&gt;Packing multiple changes into one prompt → limit to one change per turn.&lt;/li&gt;
&lt;li&gt;Omitting or shortening the preserve list on later edits → details drift; restate the full list every time.&lt;/li&gt;
&lt;li&gt;Exact text not placed in quotation marks or without location/typography guidance → wrong words or extra text appear.&lt;/li&gt;
&lt;li&gt;Reference images uploaded without assigned roles → the model blends them unpredictably.&lt;/li&gt;
&lt;li&gt;Expecting a higher quality tier to rescue a weak description → the prompt carries the result; quality mainly affects refinement, cost, and latency.&lt;/li&gt;
&lt;li&gt;Using special syntax, weights, or “masterpiece 8K ultra detailed” keyword stacking → clear, ordered sentences consistently outperform keyword stuffing.&lt;/li&gt;
&lt;li&gt;Forgetting integration instructions on edits → new elements look composited rather than natural.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Higher quality settings (xhigh or max) help with dense text, fine detail, or complex layouts, but they cannot compensate for an underspecified prompt.&lt;/p&gt;

&lt;h2&gt;
  
  
  PT Image 2.5 Prompting Mistakes Should You Avoid?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Mistake&lt;/th&gt;
&lt;th&gt;Weak approach&lt;/th&gt;
&lt;th&gt;Better approach&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;No deliverable&lt;/td&gt;
&lt;td&gt;“Make something cinematic”&lt;/td&gt;
&lt;td&gt;“Create a 4:5 paid-social product photograph”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vague edit&lt;/td&gt;
&lt;td&gt;“Improve the background”&lt;/td&gt;
&lt;td&gt;“Replace only the background with…”&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No preservation rules&lt;/td&gt;
&lt;td&gt;“Change the shirt”&lt;/td&gt;
&lt;td&gt;Lock face, pose, camera, and background&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Unassigned references&lt;/td&gt;
&lt;td&gt;“Use these images”&lt;/td&gt;
&lt;td&gt;Give each reference one role&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Too many changes&lt;/td&gt;
&lt;td&gt;Five major edits in one request&lt;/td&gt;
&lt;td&gt;One controlled change per pass&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Assuming prompt = compliance&lt;/td&gt;
&lt;td&gt;Trust the generation&lt;/td&gt;
&lt;td&gt;Perform visual QA&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Accessing GPT Image 2.5 Through CometAPI
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt; provides OpenAI-compatible access to both &lt;a href="https://www.cometapi.com/models/openai/gpt-image-2-5-flare/" rel="noopener noreferrer"&gt;gpt-image-2.5-flare&lt;/a&gt; and &lt;a href="https://www.cometapi.com/models/openai/gpt-image-2-5/" rel="noopener noreferrer"&gt;gpt-image-2.5-sunburst&lt;/a&gt; under a single API key and the familiar /v1/images/generations and edits endpoints. Benefits include centralized authentication and billing, competitive token rates, easy switching between Flare and Sunburst, usage visibility, and the ability to combine image generation with the broader catalog of 500+ models without managing multiple vendor accounts.&lt;/p&gt;

&lt;p&gt;A typical migration requires only changing the base URL to &lt;code&gt;https://api.cometapi.com/v1&lt;/code&gt;, supplying your CometAPI key, and selecting the desired model ID. This setup is especially valuable for production pipelines that need to A/B the two variants, control costs, or route different quality tiers efficiently.&lt;/p&gt;

&lt;p&gt;For the latest parameter lists, pricing, and code examples, consult CometAPI’s image generation documentation alongside OpenAI’s official image prompting guide.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;GPT Image 2.5 rewards clarity, structure, and discipline. Define the deliverable, describe the visible subject and composition with concrete detail, name the style through lighting/materials/color, put exact text in quotes, and (for every edit) ruthlessly separate the single change from the full preserve list. Give every reference image a clear job, iterate one decision at a time, and let API parameters handle size, quality, and transparency. Start exploration and high-volume work on Flare; move to Sunburst when final precision or complex edit sequences demand it.&lt;/p&gt;

&lt;p&gt;Master the Change-vs-Preserve rule and the Subject–Composition–Style–Text–Constraints structure, and GPT Image 2.5 becomes a reliable production tool rather than a source of endless regeneration. Pair it with a unified gateway such as CometAPI and you gain both creative power and operational simplicity.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQs
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Should GPT Image 2.5 prompts be long?
&lt;/h3&gt;

&lt;p&gt;Not necessarily. A long prompt full of vague adjectives can be weaker than a short prompt with precise composition and preservation instructions. Include detail where failure would matter.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I stop GPT Image 2.5 from changing the whole image?
&lt;/h3&gt;

&lt;p&gt;Use a narrow instruction such as “Change only the jacket,” then explicitly list what must remain unchanged: identity, pose, framing, camera position, lighting, background, text, and other approved elements.&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I preserve a person’s face?
&lt;/h3&gt;

&lt;p&gt;Use a clear reference photo, assign it as the authoritative identity reference, and state which characteristics must remain stable. Do not rely entirely on “same person.”&lt;/p&gt;

&lt;h3&gt;
  
  
  How do I make GPT Image 2.5 generate exact text?
&lt;/h3&gt;

&lt;p&gt;Provide the final copy verbatim, state where it should appear, specify its hierarchy, and instruct the model not to add other words. Always verify the final text visually.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can GPT Image 2.5 create transparent PNG images?
&lt;/h3&gt;

&lt;p&gt;Yes. Set the background to transparent and use PNG or WebP output.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can GPT Image 2.5 generate 4K images?
&lt;/h3&gt;

&lt;p&gt;Yes. The documented maximum includes 3840 × 2160 landscape and 2160 × 3840 portrait output. Resolutions above 2560 × 1440 are experimental, and custom dimensions must satisfy the API’s divisibility and aspect-ratio constraints.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is GPT Image 2.5 better for editing than GPT Image 2?
&lt;/h3&gt;

&lt;p&gt;OpenAI highlights improved precision editing, reference preservation, and multi-turn consistency. Arena’s early preference results also favored both 2.5 variants at the dated snapshot, but these preliminary rankings can change.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-prompt-gpt-image-2-5-like-a-pro/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-prompt-gpt-image-2-5-like-a-pro"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How to Integrat Make to 500+ AI Models with CometAPI</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Mon, 21 Sep 2026 07:26:19 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/how-to-integrat-make-to-500-ai-models-with-cometapi-200m</link>
      <guid>https://dev.to/dylanfoster1/how-to-integrat-make-to-500-ai-models-with-cometapi-200m</guid>
      <description>&lt;p&gt;In 2026, AI is no longer a standalone tool—it's the engine driving automated business processes. Platforms like &lt;strong&gt;Make.com&lt;/strong&gt; (formerly Integromat) empower users to build complex visual workflows connecting thousands of apps, while AI models handle intelligent decision-making, content generation, data analysis, and more.&lt;/p&gt;

&lt;p&gt;However, integrating dozens of AI providers (OpenAI, Anthropic, Google, xAI, etc.) means managing multiple API keys, billing accounts, rate limits, and inconsistent SDKs. This creates friction, vendor lock-in, and higher costs.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CometAPI&lt;/strong&gt; solves this by offering unified access to &lt;strong&gt;500+ cutting-edge AI models&lt;/strong&gt; through a single OpenAI-compatible API endpoint. Users get one key, one dashboard for billing and analytics, real-time model access, and typical savings of 20-40% compared to direct provider pricing.&lt;/p&gt;

&lt;p&gt;Pairing &lt;strong&gt;Make&lt;/strong&gt; with &lt;strong&gt;CometAPI&lt;/strong&gt; creates a powerful no-code/low-code solution for AI-powered automations. Whether you're generating content, classifying support tickets, building AI agents, or creating multimodal workflows (text, image, video), this integration delivers speed, flexibility, and scalability.&lt;/p&gt;

&lt;p&gt;Make's CometAPI integration includes dedicated modules: &lt;strong&gt;Create a Chat&lt;/strong&gt; (with fallback models), &lt;strong&gt;Create an API Call&lt;/strong&gt; (arbitrary authorized requests), and &lt;strong&gt;List Models&lt;/strong&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What is Make? Why It's Ideal for AI Automations
&lt;/h2&gt;

&lt;p&gt;Make.com is a visual workflow automation platform supporting 3,000+ pre-built app integrations. It excels at:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Drag-and-drop scenario builder with routers, iterators, filters, and error handlers.&lt;/li&gt;
&lt;li&gt;Native support for webhooks, scheduling, data parsing, and JSON mapping.&lt;/li&gt;
&lt;li&gt;Built-in AI tools and agents (next-gen agents with multimodal support announced in 2026).&lt;/li&gt;
&lt;li&gt;Enterprise features: SSO, audit logs, team collaboration.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why Use CometAPI with Make
&lt;/h2&gt;

&lt;blockquote&gt;
&lt;p&gt;Users consolidate traffic (LLM + images) for savings. Developers praise support and pricing transparency. Integration is verified and maintained by CometAPI on Make.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;For no-code developers, the traditional method of building AI workflows involves installing separate modules for OpenAI, Anthropic, and Google. This creates "vendor sprawl," where you must monitor multiple billing dashboards and manage separate API credits. Using CometAPI with Make simplifies this architecture by providing one single connection that grants access to over 500 models. Instead of switching modules when you want to move from GPT to Claude, you simply change a text field in your configuration.&lt;/p&gt;

&lt;p&gt;Cost efficiency is another primary driver for this integration. CometAPI utilizes institutional bulk purchasing power to set prices permanently 20% to 40% below official vendor rates. In high-volume production environments—such as a Make scenario that processes thousands of customer emails daily—these savings can translate into hundreds of dollars of reclaimed margin every month. Furthermore, CometAPI provides a 99.9% Service Availability SLA, ensuring that if a specific provider like OpenAI experiences a regional outage, your Make scenario remains operational through intelligent multi-region routing.&lt;/p&gt;

&lt;h2&gt;
  
  
  Prerequisites
&lt;/h2&gt;

&lt;p&gt;To follow this guide, you will need the following:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;Make account&lt;/strong&gt; (Works on all plans, including Free and Pro).&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;CometAPI account&lt;/strong&gt; (Registration includes free trial credits with no credit card required).&lt;/li&gt;
&lt;li&gt;An active &lt;strong&gt;CometAPI API Key&lt;/strong&gt; from your personal dashboard.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  &lt;a href="https://apidoc.cometapi.com/integrations/make" rel="noopener noreferrer"&gt;Step-by-Step Setup Guide&lt;/a&gt;
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Step 1: Get Your CometAPI API Key
&lt;/h3&gt;

&lt;p&gt;First, log in to your CometAPI dashboard. Navigate to the &lt;strong&gt;API&lt;/strong&gt; &lt;strong&gt;Token&lt;/strong&gt; section and click the &lt;strong&gt;Add API Key&lt;/strong&gt; button. This will generate a unique key (formatted as &lt;code&gt;sk-xxxx&lt;/code&gt;) that acts as your "master key" for all 500+ models. Copy this key and keep it secure. Note the unified Base URL provided in the documentation: &lt;code&gt;https://api.cometapi.com/v1.&lt;/code&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4ylu9l75wk2im6pdy9dv.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F4ylu9l75wk2im6pdy9dv.webp" alt="How to Integrat Make to 500+ AI Models with CometAPI" width="800" height="396"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Create a New Scenario in Make
&lt;/h3&gt;

&lt;p&gt;Log in to your Make account and click &lt;strong&gt;Create a new scenario&lt;/strong&gt;. In the scenario editor, click the large plus icon to add your first module. Search for &lt;strong&gt;CometAPI&lt;/strong&gt; in the search bar. You will see the official CometAPI module listed; select it to see the available actions. For most workflows, you will use the &lt;strong&gt;Make an API Call&lt;/strong&gt; action.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fllhvro84dywk3vgiiwyq.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fllhvro84dywk3vgiiwyq.webp" alt="How to Integrat Make to 500+ AI Models with CometAPI" width="799" height="404"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxlftryvdpa9i97r2ai7o.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fxlftryvdpa9i97r2ai7o.webp" alt="How to Integrat Make to 500+ AI Models with CometAPI" width="800" height="395"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Connect Your CometAPI Account
&lt;/h3&gt;

&lt;p&gt;After selecting the action, a configuration window will appear. Click the &lt;strong&gt;Add&lt;/strong&gt; button next to the Connection field. In the "API Key" field, paste the secret key you copied from the CometAPI dashboard in Step 1. Give your connection a descriptive name, such as "My Production CometAPI," and click &lt;strong&gt;Save&lt;/strong&gt;. This connection is now authorized to call any model in the catalog.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbtwv8qmpz6uggu6ol20m.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbtwv8qmpz6uggu6ol20m.webp" alt="How to Integrat Make to 500+ AI Models with CometAPI" width="799" height="425"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 4: Configure the API Call
&lt;/h3&gt;

&lt;p&gt;Example using &lt;strong&gt;Create a Chat&lt;/strong&gt;:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Choose model (e.g., &lt;code&gt;claude-opus-4-7&lt;/code&gt; or &lt;code&gt;gpt-5-4-pro&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;Set messages, temperature, max_tokens, etc.&lt;/li&gt;
&lt;li&gt;Add fallback models for resilience.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Now you must define which model you want to use and what data you want to send.&lt;/p&gt;

&lt;p&gt;For text tasks, set the &lt;strong&gt;URL&lt;/strong&gt; to &lt;code&gt;/v1/chat/completions&lt;/code&gt; and the &lt;strong&gt;Method&lt;/strong&gt; to &lt;code&gt;POST&lt;/code&gt;. In the &lt;strong&gt;Body&lt;/strong&gt; field, use the following JSON structure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{  "model": "gpt-5.5",  "messages": [ &amp;nbsp;  { &amp;nbsp; &amp;nbsp;  "role": "user", &amp;nbsp; &amp;nbsp;  "content": "{{1.text}}" &amp;nbsp;  }  ],  "stream": false}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;code&gt;{{1.text}}&lt;/code&gt; syntax is standard Make mapping. You can replace this by clicking into the field and selecting a variable from a previous module (like a Gmail message or a Google Sheet cell). If you want to generate images, change the URL to &lt;code&gt;/v1/images/generations&lt;/code&gt; and use the image-specific body format.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fat8bnt9ypgp7x3qypbcp.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fat8bnt9ypgp7x3qypbcp.webp" alt="How to Integrat Make to 500+ AI Models with CometAPI" width="800" height="415"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: Test and Publish
&lt;/h3&gt;

&lt;p&gt;Click the &lt;strong&gt;Run once&lt;/strong&gt; button at the bottom of the scenario editor. Make will execute the scenario and send your request to CometAPI. Once finished, click the bubble above the CometAPI module to inspect the output. You should see a successful &lt;code&gt;200 OK&lt;/code&gt; response with the AI-generated text or image URL. If everything looks correct, toggle the &lt;strong&gt;Scheduling&lt;/strong&gt; switch to "On" to activate your automation.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsjtaz8guxzqnc84vh6ls.webp" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fsjtaz8guxzqnc84vh6ls.webp" alt="How to Integrat Make to 500+ AI Models with CometAPI" width="800" height="427"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Which Models Can You Use
&lt;/h2&gt;

&lt;p&gt;The versatility of a unified API means you can use the best tool for every specific task in your Make no-code AI workflow.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model Category&lt;/th&gt;
&lt;th&gt;Example Model ID&lt;/th&gt;
&lt;th&gt;Best Make Scenario Use Case&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Logic &amp;amp; Reasoning&lt;/td&gt;
&lt;td&gt;claude-opus-4-7&lt;/td&gt;
&lt;td&gt;Analyzing complex legal contracts or multi-step support tickets.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Coding &amp;amp; Data&lt;/td&gt;
&lt;td&gt;deepseek-v4-pro&lt;/td&gt;
&lt;td&gt;Writing SQL queries or refactoring code snippets from Airtable.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Efficient Chat&lt;/td&gt;
&lt;td&gt;gpt-5.5&lt;/td&gt;
&lt;td&gt;Daily conversational assistants and email drafting.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Image Creation&lt;/td&gt;
&lt;td&gt;flux-2-max&lt;/td&gt;
&lt;td&gt;Creating high-fidelity marketing banners and product mockups.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video Automation&lt;/td&gt;
&lt;td&gt;sora-2&lt;/td&gt;
&lt;td&gt;Turning social media posts into short cinematic clips with audio.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Ready-to-Use Make Scenario Templates
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Template 1: Customer Support Auto-Reply
&lt;/h3&gt;

&lt;p&gt;This workflow reduces human response time for common inquiries while escalating complex issues.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Trigger&lt;/strong&gt;: A &lt;strong&gt;Gmail&lt;/strong&gt; or &lt;strong&gt;Typeform&lt;/strong&gt; module detects a new customer message.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Processing&lt;/strong&gt;: Use &lt;strong&gt;Claude Opus 4.7&lt;/strong&gt; via the CometAPI module to analyze the message. This model is chosen for its superior context window and low hallucination rate.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Router&lt;/strong&gt;: Use a &lt;strong&gt;Router&lt;/strong&gt; module to check the AI's "Sentiment" or "Urgency" output.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Branch A&lt;/strong&gt;: If the issue is simple, the scenario sends an &lt;strong&gt;Automated Reply&lt;/strong&gt; via Gmail.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Branch B&lt;/strong&gt;: If the issue is a high-priority bug, the scenario sends a &lt;strong&gt;Slack notification&lt;/strong&gt; to the engineering team.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Parameters&lt;/strong&gt;: Set the body to request a JSON response containing &lt;code&gt;{"category": "bug", "urgency": 10}&lt;/code&gt; for easy filtering.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Template 2: Content Repurposing Pipeline
&lt;/h3&gt;

&lt;p&gt;This template allows you to scale your social media presence across multiple languages with extreme cost efficiency.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Trigger&lt;/strong&gt;: A new row is added to &lt;strong&gt;Google Sheets&lt;/strong&gt; containing a blog post URL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Action 1&lt;/strong&gt;: An &lt;strong&gt;HTTP&lt;/strong&gt; module scrapes the text from the URL.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Processing 1&lt;/strong&gt;: Use &lt;strong&gt;GPT&lt;/strong&gt; &lt;strong&gt;5.5&lt;/strong&gt; to generate a high-quality 200-word summary in English.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Processing 2&lt;/strong&gt;: Send that summary to &lt;strong&gt;DeepSeek V3&lt;/strong&gt; to translate it into Chinese and generate SEO keywords.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Why Two Models?&lt;/strong&gt;: DeepSeek is used for the translation step because it is significantly cheaper ($0.216/M tokens vs $4/M for GPT 5.5), allowing you to optimize your per-run costs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output&lt;/strong&gt;: Post the results to a &lt;strong&gt;Buffer&lt;/strong&gt; or &lt;strong&gt;Notion&lt;/strong&gt; module.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Template 3: Image Generation Automation
&lt;/h3&gt;

&lt;p&gt;Automate your e-commerce design process by turning product descriptions into visual assets.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Trigger&lt;/strong&gt;: A new record is created in &lt;strong&gt;Airtable&lt;/strong&gt; with a product name and "Design Brief."&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Action&lt;/strong&gt;: Use the CometAPI module with the &lt;code&gt;/v1/images/generations&lt;/code&gt; endpoint and the &lt;strong&gt;Flux 2 Max&lt;/strong&gt; model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;JSON&lt;/strong&gt; &lt;strong&gt;Body&lt;/strong&gt;:
&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{  "model": "flux-2-max",  "prompt": "E-commerce product shot of {{1.Product_Name}}, {{1.Brief}}, photorealistic, 4k",  "n": 1,  "size": "1024x1024"}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Storage&lt;/strong&gt;: Use an &lt;strong&gt;Airtable Update Record&lt;/strong&gt; module to save the resulting image URL back into a checkbox field for review.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Comparison Table: CometAPI + Make vs. Alternatives
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Aspect&lt;/th&gt;
&lt;th&gt;CometAPI + Make&lt;/th&gt;
&lt;th&gt;Direct Provider + Make&lt;/th&gt;
&lt;th&gt;Other Aggregators (e.g., OpenRouter)&lt;/th&gt;
&lt;th&gt;Zapier + Single Provider&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;# Models&lt;/td&gt;
&lt;td&gt;500+ unified&lt;/td&gt;
&lt;td&gt;Limited per provider&lt;/td&gt;
&lt;td&gt;Many, but variable&lt;/td&gt;
&lt;td&gt;Fewer&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Setup Complexity&lt;/td&gt;
&lt;td&gt;Low (pre-built modules)&lt;/td&gt;
&lt;td&gt;Medium (multiple connections)&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost Savings&lt;/td&gt;
&lt;td&gt;20-40% + unified billing&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;Varies&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fallbacks &amp;amp; Routing&lt;/td&gt;
&lt;td&gt;Native in modules&lt;/td&gt;
&lt;td&gt;Manual&lt;/td&gt;
&lt;td&gt;Some&lt;/td&gt;
&lt;td&gt;Limited&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Observability&lt;/td&gt;
&lt;td&gt;Excellent (unified dashboard)&lt;/td&gt;
&lt;td&gt;Fragmented&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Basic&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Multimodal&lt;/td&gt;
&lt;td&gt;Full support&lt;/td&gt;
&lt;td&gt;Per provider&lt;/td&gt;
&lt;td&gt;Good&lt;/td&gt;
&lt;td&gt;Varies&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;No-Code Ease&lt;/td&gt;
&lt;td&gt;Highest&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vendor Lock-in&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Cost Optimization Tips for Make + CometAPI
&lt;/h2&gt;

&lt;p&gt;To get the most out of your automation budget, implement these three strategies:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;DeepSeek Routing&lt;/strong&gt;: For classification or simple data extraction tasks, route your traffic to &lt;strong&gt;DeepSeek V4 Flash&lt;/strong&gt;. This model offers a 1M-token context window but costs 90% less than flagship models. By using DeepSeek for the "dirty work" of your scenario and reserving GPT or Claude for the final "polished" output, you can reduce your total scenario cost by over 60%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Make Filter Modules&lt;/strong&gt;: Always use a &lt;strong&gt;Filter&lt;/strong&gt; module before your CometAPI call. If a field is empty or does not meet specific criteria, the filter will stop the scenario, preventing unnecessary API calls and saving you "Operations" in Make as well as tokens in CometAPI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Aggregator Batching&lt;/strong&gt;: If your scenario processes many records at once, use the &lt;strong&gt;Array Aggregator&lt;/strong&gt; module to combine them into a single list, then send them to CometAPI in one large prompt. This reduces the number of separate API requests, which helps you stay within rate limits and simplifies your usage logs in the dashboard.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Pricing Insights and ROI Calculation
&lt;/h3&gt;

&lt;p&gt;CometAPI: Pay-as-you-go, credits roll over, volume discounts. Examples show significant savings vs. official rates.&lt;/p&gt;

&lt;p&gt;Make: Starts low (e.g., ~$9/mo for operations). Combined, ideal for high-ROI automations (time saved &amp;gt;&amp;gt; subscription).&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Example ROI&lt;/strong&gt;: Automate content for 10x output at fraction of manual cost; support triage reduces tickets by 50%+.&lt;/p&gt;

&lt;h2&gt;
  
  
  Troubleshooting Common Issues
&lt;/h2&gt;

&lt;h3&gt;
  
  
  401 Unauthorized Error
&lt;/h3&gt;

&lt;p&gt;This error nearly always points to an issue with your API Key. Double-check that you have not added an extra space at the beginning or end of the key when pasting it into Make. Also, ensure your CometAPI account has a positive credit balance; although signup is free, you must have active credits to make calls beyond the trial.&lt;/p&gt;

&lt;h3&gt;
  
  
  422 Unprocessable Entity
&lt;/h3&gt;

&lt;p&gt;If you receive a 422 error, check your JSON formatting in the body field. Ensure every opening brace &lt;code&gt;{&lt;/code&gt; has a corresponding closing brace &lt;code&gt;}&lt;/code&gt; and that you are using straight quotes &lt;code&gt;"&lt;/code&gt; rather than "curly" quotes. Additionally, verify that the &lt;code&gt;model&lt;/code&gt; name you entered matches the official identifier in the CometAPI model catalog exactly (e.g., &lt;code&gt;gpt-5.5&lt;/code&gt; instead of &lt;code&gt;GPT 5.5&lt;/code&gt;).&lt;/p&gt;

&lt;h3&gt;
  
  
  Scenario Timeouts
&lt;/h3&gt;

&lt;p&gt;Some advanced reasoning models take longer to generate a response. If your Make scenario times out, first ensure that &lt;code&gt;"stream": false&lt;/code&gt; is set in your JSON body, as Make does not support raw stream processing in its standard API call module. If the error persists, consider switching to a "Flash" tier model like &lt;strong&gt;Gemini 3.1 Flash-Lite&lt;/strong&gt; or &lt;strong&gt;DeepSeek V4 Flash&lt;/strong&gt;, which are optimized for sub-second responses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Future-Proofing Your AI Stack with CometAPI on Make
&lt;/h2&gt;

&lt;p&gt;As AI evolves (new models weekly in 2026), this integration lets you adopt instantly without refactoring. Combine with Make Grid, AI Agents, and CometAPI's continuous updates for a robust, scalable system.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;CometAPI Recommendation&lt;/strong&gt;: Start with free credits on CometAPI. Use the playground to test models, then refer to &lt;a href="https://apidoc.cometapi.com/integrations/make" rel="noopener noreferrer"&gt;guide&lt;/a&gt; and build your first Make scenario. For high volume, explore enterprise options for custom SLAs and dedicated support.&lt;/p&gt;

&lt;h2&gt;
  
  
  Conclusion
&lt;/h2&gt;

&lt;p&gt;&lt;a href="https://apidoc.cometapi.com/integrations/make" rel="noopener noreferrer"&gt;Integrating Make with CometAPI&lt;/a&gt; unlocks the full potential of no-code AI automations with unparalleled model choice, cost efficiency, and simplicity. One integration gives access to the entire AI ecosystem—saving time, money, and engineering effort while delivering production-grade reliability.&lt;/p&gt;

&lt;p&gt;Ready to get started?&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Sign up for CometAPI (free credits) → &lt;a href="https://www.cometapi.com/console/login" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt;
&lt;/li&gt;
&lt;li&gt;Build your first scenario on Make.com&lt;/li&gt;
&lt;li&gt;Explore more templates and guides on both platforms.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This powerful combo positions your workflows for 2026 and beyond. Experiment, iterate, and scale confidently.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Q: Is there an official CometAPI module in Make?
&lt;/h3&gt;

&lt;p&gt;A: Yes. You can find it by searching for "CometAPI" in the module selector. It provides a standardized way to call any model in the catalog without writing custom code.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: Can I use multiple different models in a single Make scenario?
&lt;/h3&gt;

&lt;p&gt;A: Absolutely. You can add as many CometAPI modules as you need to a single workflow. For example, you can use one module for text analysis and another for image generation within the same automation path.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: Is the CometAPI integration compatible with the Make Free plan?
&lt;/h3&gt;

&lt;p&gt;A: Yes. As long as you have your API key and are using the "Make an API Call" action, it will function perfectly on the Free tier.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: How does this integration compare to the native &lt;strong&gt;OpenAI&lt;/strong&gt; &lt;strong&gt;module in Make?&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;A: While the native OpenAI module is restricted to OpenAI models, CometAPI gives you access to 500+ models from all providers (OpenAI, Google, Anthropic, etc.) using the exact same connection, plus a 20-40% cost saving.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: Does the integration support image generation?
&lt;/h3&gt;

&lt;p&gt;A: Yes. You can call the &lt;code&gt;/v1/images/generations&lt;/code&gt; endpoint to access models like GPT Image 2, Flux 2 Max, and Nano Banana 2 directly from Make.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-connect-make-with-cometapi/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-connect-make-with-cometapi"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>The 2026 LLM API Pricing Comparison: GPT-5.5, Claude Sonnet 4.6, Gemini 3.5 Flash and DeepSeek V4</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Mon, 21 Sep 2026 06:57:17 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/the-2026-llm-api-pricing-comparison-gpt-55-claude-sonnet-46-gemini-35-flash-and-deepseek-v4-4g4</link>
      <guid>https://dev.to/dylanfoster1/the-2026-llm-api-pricing-comparison-gpt-55-claude-sonnet-46-gemini-35-flash-and-deepseek-v4-4g4</guid>
      <description>&lt;p&gt;Pricing is the single most consequential decision in choosing a frontier LLM, and it is also the dimension where most published comparisons are out of date within a quarter. This article cuts through that. Below is a current, sourced view of input and output token pricing across the four models that account for the majority of production frontier-model traffic in 2026 (OpenAI’s GPT-5.5, Anthropic’s Claude Sonnet 4.6, Google’s Gemini 3.5 Flash, and DeepSeek’s V4), together with the levers that meaningfully change your bill at scale: prompt caching, batch processing, and long-context surcharges.&lt;/p&gt;

&lt;p&gt;The piece is built around two questions. First: at list price, what does each model cost per million tokens, and how do the quoted rates compare on the inputs and outputs that actually drive a production bill? Second: when you apply a representative workload (100 million tokens a month, 80% input and 20% output, with realistic cache hit rates), what is the monthly bill in dollars on each model? The first answer establishes the rate card; the second tells you what that rate card becomes once it touches a real production pattern.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;Quick read:&lt;/strong&gt; Across the four frontier models, list pricing spans roughly two orders of magnitude. DeepSeek V4 is the cheapest at $0.435 per million input tokens; Claude Opus 4.7 is the most expensive at $5.00. The shape of your workload, particularly your cache hit rate and your input-to-output ratio, changes which model is cheapest in practice, often by more than the rate card suggests.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Why a like-for-like pricing comparison is harder than it looks
&lt;/h2&gt;

&lt;p&gt;Provider pricing pages are written for that provider's own customers, not for someone evaluating four options side by side. The result is that comparing them produces three persistent traps:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Tokens are not the same across providers.&lt;/strong&gt; Claude Opus 4.7 ships with a new tokenizer that can produce up to 35% more tokens for the same input text than Opus 4.6. Gemini's tokenizer differs from OpenAI's. The rate card is per million tokens, but the token count for the identical prompt varies between providers, meaning the headline rate is only a first approximation of relative cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Long-context pricing tiers create cost cliffs.&lt;/strong&gt; OpenAI's GPT-5.5 family has separate short-context and long-context rates that kick in around 270,000 tokens. Anthropic, conversely, holds the same per-token rate across its full 1M context window. Workloads that sit near these thresholds are priced very differently to workloads that sit comfortably inside them.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Discounts are stacked, not separate.&lt;/strong&gt; Prompt caching, batch processing, and provider-specific volume tiers can each cut effective cost dramatically, and they stack. A cached batch request on Anthropic can cost as little as 5% of a standard non-cached request. A pricing comparison that ignores these levers overstates list cost, sometimes by an order of magnitude.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The comparison below normalises for these traps where it can, and flags them explicitly where it cannot.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 2026 frontier LLM pricing comparison
&lt;/h2&gt;

&lt;p&gt;All figures in US dollars per million tokens. Sourced from each provider's official pricing documentation as of May 2026.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Cached input&lt;/th&gt;
&lt;th&gt;Batch (50% off)&lt;/th&gt;
&lt;th&gt;Context window&lt;/th&gt;
&lt;th&gt;Long-context surcharge&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.5&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$30.00&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$2.50 / $15.00&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;Yes (~270K)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;$3.00&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;td&gt;$0.30&lt;/td&gt;
&lt;td&gt;$1.50 / $7.50&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.7&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$25.00&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$2.50 / $12.50&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.5 Flash&lt;/td&gt;
&lt;td&gt;$1.50&lt;/td&gt;
&lt;td&gt;$9.00&lt;/td&gt;
&lt;td&gt;$0.15&lt;/td&gt;
&lt;td&gt;$1.00 / $6.00&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;Yes (200K)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSeek V4&lt;/td&gt;
&lt;td&gt;$0.435&lt;/td&gt;
&lt;td&gt;$0.87&lt;/td&gt;
&lt;td&gt;$0.0028&lt;/td&gt;
&lt;td&gt;Not offered&lt;/td&gt;
&lt;td&gt;384K&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;Reading the table:&lt;/em&gt; Cached input is the rate paid on tokens served from prompt cache (typically system prompts, few-shot examples, or document prefixes that recur across requests). Batch is the rate paid for asynchronous workloads with up to 24-hour latency. Long-context surcharge denotes whether the provider raises rates above a context-length threshold; for those that do, the threshold is given in parentheses.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where each model wins
&lt;/h2&gt;

&lt;h3&gt;
  
  
  GPT-5.5: the highest-capability default for hard reasoning and agentic work
&lt;/h3&gt;

&lt;p&gt;GPT-5.5 is OpenAI's frontier model for complex professional workloads: coding agents, multi-step planning, long-running tool use, and document analysis where reasoning depth is the dominant requirement. It is also the most expensive of the major US frontier models on input ($5.00 per million) and the highest on output ($30.00 per million), which means it earns its position on workloads where the alternative is paying a flagship rate to a different model that solves the problem less reliably. GPT-5.5 supports caching at a 90% discount, batch processing at 50% off, and long-context pricing kicks in around the 270K-token mark, which is relevant for very long codebases or full-repository contexts but not for typical RAG workloads.&lt;/p&gt;

&lt;h3&gt;
  
  
  Claude Sonnet 4.6: the recommended default for most production traffic
&lt;/h3&gt;

&lt;p&gt;Sonnet 4.6 is Anthropic's recommended model for the majority of production workloads, and the price-to-capability ratio is the reason. At $3 input and $15 output per million tokens, it sits below GPT-5.5 on both rates while delivering near-Opus quality on the workloads that dominate most production systems: coding, analysis, RAG pipelines, customer-facing chat, and structured output generation. Sonnet's distinguishing pricing feature is that the full 1M token context window is available at standard rates (there is no long-context surcharge), which makes it the cheapest credible option for workloads that occasionally need to ingest very long documents or full repositories. Prompt caching cuts cached input to 10% of standard, which is decisive for any workload with a stable system prompt.&lt;/p&gt;

&lt;h3&gt;
  
  
  Gemini 3.5 Flash: the most aggressively-priced flagship for short-context work
&lt;/h3&gt;

&lt;p&gt;Gemini 3.5 Flash is the cheapest flagship-class model from a major US provider on raw API pricing, at $1.50 input and $9.00 output per million tokens. For most production traffic, that is the relevant pricing tier, and it materially undercuts both GPT-5.5 and Claude Opus 4.7. Higher price than prior Flash models leads to increased overall costs in token-heavy agentic scenarios (5.5x Intelligence Index cost vs. Gemini 3 Flash due to pricing + usage).. Gemini's other distinguishing feature is the genuinely free tier in Google AI Studio, which is useful for prototyping but not relevant to production cost models.&lt;/p&gt;

&lt;h3&gt;
  
  
  DeepSeek V4: dramatically cheaper, with caveats worth understanding
&lt;/h3&gt;

&lt;p&gt;DeepSeek V4 lists at $0.435 per million input tokens and $0.87 per million output tokens, which is between five and seventy times cheaper than the US frontier models depending on which one you compare against. The model itself is competitive on many benchmarks, particularly reasoning and code. The caveats are worth being explicit about: data is processed in China, which is a non-starter for some regulated workloads; English-language quality is strong but the model is optimised differently to the US frontier models, and head-to-head testing on your specific workload is essential rather than optional. For workloads where these caveats are acceptable, DeepSeek genuinely changes the cost equation.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;A note on Claude Opus 4.7 vs Sonnet 4.6.&lt;/strong&gt; Opus is included in the table for completeness, but for the great majority of production traffic, Sonnet 4.6 is the better economic choice. Opus costs 1.67x Sonnet on both input and output, and for workloads where Sonnet is sufficient (which is most of them), that premium has no offsetting benefit. Reach for Opus when evaluations show Sonnet is failing on a specific class of task: highly autonomous coding agents, long-horizon professional workflows, and tasks where instruction-following at the margin is decisive.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Worked example: what 100 million tokens a month actually costs
&lt;/h2&gt;

&lt;p&gt;Headline pricing per million tokens means little until it touches a representative workload. The example below uses a profile that approximates a non-trivial production system: 100 million total tokens per month, split 80% input (80M) and 20% output (20M), with a 30% cache hit rate on the input portion. This pattern is broadly representative of a customer-facing chat or RAG workload with a stable system prompt and document context.&lt;/p&gt;

&lt;p&gt;The math for each model: cached input cost + uncached input cost + output cost. Cached input is billed at 10% of standard for the providers that offer caching.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Cached input (24M)&lt;/th&gt;
&lt;th&gt;Uncached input (56M)&lt;/th&gt;
&lt;th&gt;Output (20M)&lt;/th&gt;
&lt;th&gt;Total monthly bill&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.5&lt;/td&gt;
&lt;td&gt;$12.00&lt;/td&gt;
&lt;td&gt;$280.00&lt;/td&gt;
&lt;td&gt;$600.00&lt;/td&gt;
&lt;td&gt;$892.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 4.6&lt;/td&gt;
&lt;td&gt;$7.20&lt;/td&gt;
&lt;td&gt;$168.00&lt;/td&gt;
&lt;td&gt;$300.00&lt;/td&gt;
&lt;td&gt;$475.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.7&lt;/td&gt;
&lt;td&gt;$12.00&lt;/td&gt;
&lt;td&gt;$280.00&lt;/td&gt;
&lt;td&gt;$500.00&lt;/td&gt;
&lt;td&gt;$792.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;em&gt;What this tells you.&lt;/em&gt; On a representative workload,Sonnet 4.6 in turn is roughly half the cost of GPT-5.5. DeepSeek is in a different cost universe entirely. These are list-price numbers; applying batch processing where eligible cuts each total by a further 50% on the inputs and outputs (though not the cache hits).&lt;/p&gt;

&lt;p&gt;Two observations worth carrying forward. First: caching is the single most impactful lever you control. The example above assumes a 30% cache hit rate; raise it to 60% (entirely achievable for workloads with a stable system prompt), and total cost drops by roughly another 25%. Second: the input-to-output ratio matters a lot. Workloads that are output-heavy (summarisation, long-form writing) bias toward providers with cheaper output rates, while input-heavy workloads (long-context analysis, large RAG retrievals) bias toward providers with cheaper input rates and no long-context surcharge.&lt;/p&gt;

&lt;h2&gt;
  
  
  The hidden costs not on the pricing page
&lt;/h2&gt;

&lt;p&gt;List pricing is the floor, not the ceiling. Five additional costs are worth budgeting for explicitly, because they routinely surprise teams scaling from prototype to production:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Reasoning tokens.&lt;/strong&gt; Models with extended reasoning modes (GPT-5.5 Thinking, DeepSeek V4 thinking mode) generate internal reasoning content that counts as output tokens. A single high-effort reasoning call on a long prompt can run 20,000 reasoning tokens, which is $0.60 of output cost on GPT-5.5 before the visible response is produced. Budget per workload, not per request.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Long-context surcharges.&lt;/strong&gt; Both Gemini 3.5 Flash and GPT-5.5 raise rates above a context-length threshold. RAG pipelines that include large documents can silently push every request into the higher bracket without anyone noticing until the bill arrives. Measure your actual prompt lengths in production and check whether you are crossing the threshold.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Data residency multipliers.&lt;/strong&gt; Anthropic charges a 10% premium for US-only inference on Opus 4.7 and Sonnet 4.6. OpenAI applies a 10% uplift on data residency endpoints for the GPT-5.4 family. For regulated workloads where this matters, factor it into the rate card from day one.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output verbosity drift.&lt;/strong&gt; When a new model version is more thorough by default (as Opus 4.7 reportedly is compared to Opus 4.6), output tokens per response can creep up even if input length is constant. Output is priced 5x higher than input on the Anthropic line, so a 20% creep in output verbosity is a 20% increase in the dominant cost driver.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Failed and retried requests.&lt;/strong&gt; Most providers do not bill for 4xx and 5xx errors, but they do bill for partial generations and retries that succeed on the second attempt. In production systems with active retry logic, this can add a few percent to the bill. Worth knowing about when reconciling provider invoices against expected cost.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  How CometAPI fits in
&lt;/h2&gt;

&lt;p&gt;All four of these models, plus 500+ others, are available through &lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt; on a single OpenAI-compatible endpoint, with one credential, unified billing, and no per-provider account setup. Pricing on CometAPI is metered per token at the same per-model rates published by the underlying providers, with credits purchased upfront and applied across any model in the catalogue. The value of routing through CometAPI is operational rather than per-token: one credential to manage, one invoice to reconcile, and the ability to swap from GPT-5.5 to Claude Sonnet 4.6 to Gemini 3.5 Flash by changing a single string in your code.&lt;/p&gt;

&lt;p&gt;There are workloads where direct provider access is the right call. If you run a single-model workload at very high volume on one provider, with a negotiated enterprise contract, the unit economics of going direct are better. If your compliance posture requires a specific vendor-of-record relationship, an aggregator complicates rather than simplifies that conversation. For the majority of teams running multi-model production workloads, however, the operational friction of managing three or four direct provider relationships is itself a meaningful cost, one that the rate card does not capture.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Try the comparison on your workload.&lt;/strong&gt; The free tier on CometAPI lets you run the same prompt against GPT-5.5, Sonnet 4.6, Gemini 3.5 Flash, and DeepSeek V4 from a single endpoint, with no separate signups. For a workload-specific cost decision, that one-hour exercise is worth more than any pricing comparison ever published.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to use this comparison
&lt;/h2&gt;

&lt;p&gt;The right model for your workload depends on which dimension of the rate card matters most for your traffic shape. A practical decision framework:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;If reasoning depth is the bottleneck (&lt;/strong&gt;agentic workflows, complex multi-step planning, the hardest coding tasks), start with GPT-5.5 or Claude Opus 4.7. The premium is real but earned on these workloads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you want the best price-to-capability ratio for general production traffic,&lt;/strong&gt; Claude Sonnet 4.6 is the recommended default. Near-frontier capability, full 1M context at standard rates, and strong caching support.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If you are cost-sensitive and your workload sits below 200K context,&lt;/strong&gt; Gemini 3.5 Flash is the cheapest credible flagship-class option from a major US provider.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;If your workload is high-volume and price-dominated, and DeepSeek’s data-residency posture is acceptable,&lt;/strong&gt; V4 changes the cost equation enough to be worth a serious evaluation, particularly for batch-shaped workloads.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;em&gt;Want to go further on cost optimization?&lt;/em&gt; The pricing data above is the foundation for routing: the practice of sending different queries to different models based on which one can handle them at the lowest cost. The companion piece, &lt;em&gt;Cutting LLM API Costs in Half: A Model Routing Guide for Production Workloads in 2026&lt;/em&gt;, walks through the routing patterns that turn this rate card into actual savings on your monthly bill.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/2026-llm-api-pricing-comparison-gpt-5-5-claude-gemini/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=2026-llm-api-pricing-comparison-gpt-5-5-claude-gemini"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>How to Add AI Image Generation to a Web App</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Mon, 21 Sep 2026 04:38:58 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/how-to-add-ai-image-generation-to-a-web-app-f25</link>
      <guid>https://dev.to/dylanfoster1/how-to-add-ai-image-generation-to-a-web-app-f25</guid>
      <description>&lt;p&gt;In 2026, AI image generation has transformed from a novelty into a core feature for modern web applications. Whether you're building an e-commerce platform with personalized product visuals, a content creation tool, a social media app, or an educational platform, embedding AI-powered image generation can dramatically enhance user experience, boost engagement, and create new revenue streams.&lt;/p&gt;

&lt;p&gt;The global AI image generator market was valued at approximately USD 412-484 million in 2025/early 2026 and is projected to reach USD 1.7 billion by 2034, growing at a CAGR of around 17.4%. Other analyses show even faster expansion in the broader generative AI segment, with daily image creation exceeding tens of millions. Over 150 million people use these tools monthly, producing massive volumes of content.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Why integrate now?&lt;/strong&gt; Users expect dynamic, personalized visuals. Static images lead to higher bounce rates; AI-generated ones increase time-on-site by enabling customization (e.g., "generate a beach scene with my dog"). Leading models in 2026—such as OpenAI's GPT Image series, Google's Nano Banana / Imagen variants, Black Forest Labs' Flux 2 Pro, and Midjourney—deliver photorealism, accurate text rendering, 4K output, real-time grounding, and conversational editing.&lt;/p&gt;

&lt;p&gt;This comprehensive guide covers everything: market context, technical implementation with code, best practices, comparisons, security/ethics, optimization, and tailored recommendations for &lt;strong&gt;CometAPI&lt;/strong&gt; (a unified gateway to 500+ models including image generation like Midjourney, GPT Image, and more). By the end, you'll have actionable knowledge to ship production-ready features.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why AI Image Generation Matters for Web Apps in 2026
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Quick Answer:&lt;/strong&gt; Adding AI image generation involves choosing an API (e.g., CometAPI for multi-model access), handling frontend prompts and backend calls securely, displaying results with error handling, and optimizing for cost/latency. Key benefits include personalization, faster content creation, and competitive edge.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Supporting Data:&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;82% of large enterprises use generative AI in at least one function.&lt;/li&gt;
&lt;li&gt;Photorealism and text-in-image capabilities have improved dramatically; models like Flux 2 Pro and GPT Image 1.5/2 lead benchmarks.&lt;/li&gt;
&lt;li&gt;Cost per image ranges from $0.005 (budget models) to $0.06+ for premium, making high-volume apps viable.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Long-tail keywords covered: "integrate Flux AI image API web app", "Midjourney API React tutorial 2026", "cost-effective AI image generation for SaaS".&lt;/p&gt;

&lt;h2&gt;
  
  
  Understanding the 2026 AI Image Generation Landscape
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Latest Trends and Models
&lt;/h3&gt;

&lt;p&gt;2026 is the year of the "AI image arms race." Key advancements:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;4K output and real-time grounding&lt;/strong&gt;: Models incorporate live data for context-aware images.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Conversational editing&lt;/strong&gt;: Iterative refinement via chat (strong in GPT Image and Gemini-based models).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Specialized strengths&lt;/strong&gt;: Flux for photorealism/product shots; Ideogram for text; Midjourney for artistic/consistent characters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Top models (per LM Arena and comparisons):&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPT Image 1.5/2 (OpenAI): High quality, strong prompting.&lt;/li&gt;
&lt;li&gt;Flux 2 Pro (Black Forest Labs): Excellent fidelity.&lt;/li&gt;
&lt;li&gt;Imagen 4 / Nano Banana (Google): Speed and integration.&lt;/li&gt;
&lt;li&gt;Midjourney: Creative excellence via API.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Market Impact on Web Devs
&lt;/h3&gt;

&lt;p&gt;Integrating these reduces dependency on stock libraries (costly licensing) and enables features like user-generated mockups or dynamic avatars, driving metrics like conversion rates up 20-30% in e-commerce tests (industry benchmarks).&lt;/p&gt;

&lt;h3&gt;
  
  
  Choosing the Right AI Image Generation API: Comparison Table
&lt;/h3&gt;

&lt;p&gt;Selecting an API is critical. Direct provider APIs work but lead to vendor lock-in and multiple keys. Unified services like &lt;strong&gt;CometAPI&lt;/strong&gt; excel here.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Comparison Table (2026 Data):&lt;/strong&gt;&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model/Provider&lt;/th&gt;
&lt;th&gt;Quality (Elo/Score)&lt;/th&gt;
&lt;th&gt;Speed&lt;/th&gt;
&lt;th&gt;Price/Image (approx.)&lt;/th&gt;
&lt;th&gt;Strengths&lt;/th&gt;
&lt;th&gt;Best For Web Apps&lt;/th&gt;
&lt;th&gt;CometAPI Access?&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT Image 1.5/2 (OpenAI)&lt;/td&gt;
&lt;td&gt;Top (1264+)&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;td&gt;$0.04-$0.06&lt;/td&gt;
&lt;td&gt;Prompt adherence, editing&lt;/td&gt;
&lt;td&gt;General, conversational&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Flux 2 Pro&lt;/td&gt;
&lt;td&gt;1265+&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;$0.03-$0.055&lt;/td&gt;
&lt;td&gt;Photorealism, detail&lt;/td&gt;
&lt;td&gt;E-commerce, products&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Imagen 4 / Nano Banana&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Very Fast&lt;/td&gt;
&lt;td&gt;$0.02-$0.04&lt;/td&gt;
&lt;td&gt;Speed, text, multimodal&lt;/td&gt;
&lt;td&gt;Real-time apps&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Midjourney&lt;/td&gt;
&lt;td&gt;Artistic leader&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Varies&lt;/td&gt;
&lt;td&gt;Creativity, consistency&lt;/td&gt;
&lt;td&gt;Design, social&lt;/td&gt;
&lt;td&gt;Yes (via CometAPI)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Ideogram v3&lt;/td&gt;
&lt;td&gt;Strong text&lt;/td&gt;
&lt;td&gt;Fast&lt;/td&gt;
&lt;td&gt;Competitive&lt;/td&gt;
&lt;td&gt;Typography in images&lt;/td&gt;
&lt;td&gt;Marketing banners&lt;/td&gt;
&lt;td&gt;Available&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;strong&gt;Recommendation:&lt;/strong&gt; Start with &lt;strong&gt;CometAPI&lt;/strong&gt; for one OpenAI-compatible endpoint, access to 500+ models (LLMs + images + video), pay-as-you-go, free tier credits, and no lock-in. It simplifies switching models based on task (e.g., cheap for prototypes, premium for production).&lt;/p&gt;

&lt;h2&gt;
  
  
  Step-by-Step: How to Integrate AI Image Generation into a Web App
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Planning and Architecture
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Frontend&lt;/strong&gt;: React/Vue/Svelte for prompt input, preview, gallery.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Backend&lt;/strong&gt;: Node.js/Express, Python/FastAPI, or Next.js API routes for security (hide API keys).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Flow&lt;/strong&gt;: User prompt → Backend validation/rate limiting → API call → Store/return URL → Display with lazy loading.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Additional&lt;/strong&gt;: Async queues (e.g., BullMQ) for high traffic; caching (Redis) for repeats.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Setting Up with CometAPI (Recommended)
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Sign up at CometAPI.com and get your API key (free credits available).&lt;/li&gt;
&lt;li&gt;Use OpenAI-compatible endpoint: &lt;code&gt;https://api.cometapi.com/v1/images/generations&lt;/code&gt; (or specific model endpoints).&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;strong&gt;Example Node.js Backend (Express):&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;express&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;express&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;axios&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;axios&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;express&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;use&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;express&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;COMETAPI_KEY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// Never expose client-side&lt;/span&gt;

&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/generate-image&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;gpt-image-2&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// Or flux, midjourney etc. via CometAPI&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;prompt&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="mi"&gt;4000&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Invalid prompt&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;axios&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;https://api.cometapi.com/v1/images/generations&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;n&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;size&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;1024x1024&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="c1"&gt;// or higher for 2026 models&lt;/span&gt;
      &lt;span class="c1"&gt;// quality, style params as supported&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Authorization&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="s2"&gt;`Bearer &lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;imageUrl&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;url&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
    &lt;span class="c1"&gt;// Optional: Save to S3/Cloudinary, log usage&lt;/span&gt;
    &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="nx"&gt;imageUrl&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="na"&gt;revised_prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nx"&gt;revised_prompt&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;error&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;?.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nx"&gt;error&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;500&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;error&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Generation failed. Try again.&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;app&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;listen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;3000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Server running&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Security Best Practices&lt;/strong&gt;: Use environment variables, rate limiting (express-rate-limit), input sanitization, and monitor for prompt injection (OWASP GenAI guidelines).&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Frontend Implementation (React Example)
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight jsx"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;React&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;useState&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;react&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;axios&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;axios&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;ImageGenerator&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;setPrompt&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;''&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;imageUrl&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;setImageUrl&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;loading&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;setLoading&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;useState&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;generate&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;async &lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="nf"&gt;setLoading&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;axios&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;/generate-image&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;prompt&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
      &lt;span class="nf"&gt;setImageUrl&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;data&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;imageUrl&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nf"&gt;alert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Error generating image&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="nf"&gt;setLoading&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;};&lt;/span&gt;

  &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;

       &lt;span class="nf"&gt;setPrompt&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;e&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;target&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;value&lt;/span&gt;&lt;span class="p"&gt;)}&lt;/span&gt; &lt;span class="nx"&gt;placeholder&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;A futuristic city at sunset...&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;/&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;lt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="nx"&gt;button&lt;/span&gt; &lt;span class="nx"&gt;onClick&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="nx"&gt;disabled&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;loading&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;gt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;loading&lt;/span&gt; &lt;span class="p"&gt;?&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Generating...&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt; &lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;Generate Image&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
      &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;lt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sr"&gt;/button&amp;amp;gt&lt;/span&gt;&lt;span class="err"&gt;;
&lt;/span&gt;      &lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;imageUrl&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;amp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;amp&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;lt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="nx"&gt;img&lt;/span&gt; &lt;span class="nx"&gt;src&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="nx"&gt;imageUrl&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="nx"&gt;alt&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;AI Generated&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="nx"&gt;style&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{{&lt;/span&gt;&lt;span class="na"&gt;maxWidth&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="s1"&gt;100%&lt;/span&gt;&lt;span class="dl"&gt;'&lt;/span&gt;&lt;span class="p"&gt;}}&lt;/span&gt; &lt;span class="sr"&gt;/&amp;amp;gt;&lt;/span&gt;&lt;span class="err"&gt;}
&lt;/span&gt;    &lt;span class="o"&gt;&amp;amp;&lt;/span&gt;&lt;span class="nx"&gt;lt&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="sr"&gt;/div&amp;amp;gt&lt;/span&gt;&lt;span class="err"&gt;;
&lt;/span&gt;  &lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Enhance with galleries, history (localStorage or DB), and variations (call API with &lt;code&gt;variation&lt;/code&gt; params where supported).&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Python/FastAPI Alternative (for Data-Heavy Apps)
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;fastapi&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;FastAPI&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;httpx&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;

&lt;span class="n"&gt;app&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;FastAPI&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="n"&gt;COMETAPI_KEY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nd"&gt;@app.post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;/generate&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;generate&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nb"&gt;str&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;flux-2-pro&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;async&lt;/span&gt; &lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;httpx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;AsyncClient&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
        &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1/images/generations&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;model&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;prompt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Authorization&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Bearer &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;COMETAPI_KEY&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;
        &lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Deploy with Uvicorn + Docker for scalability.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Advanced Features
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Image Editing/Inpainting&lt;/strong&gt;: Use edit endpoints (mask + prompt).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Batch Generation&lt;/strong&gt;: Loop with async/await for multiple variants.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Upscaling &amp;amp; Post-processing&lt;/strong&gt;: Chain with dedicated upscaler models via CometAPI.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Real-time&lt;/strong&gt;: WebSockets for progress updates on longer generations.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Mobile Optimization&lt;/strong&gt;: Responsive design + PWA for on-device previews.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Best Practices, Optimization, and Scaling
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Cost Management&lt;/strong&gt;: Route cheap models for testing, premium for final output. Monitor with CometAPI dashboards. Implement user quotas.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Performance&lt;/strong&gt;: CDN for images, lazy loading, progressive enhancement. Aim for &amp;lt;5s response (many 2026 models achieve 2-5s).&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;UX/UI&lt;/strong&gt;: Prompt suggestions (AI-powered), negative prompts, style selectors, history gallery, download/share buttons.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Error Handling &amp;amp; Fallbacks&lt;/strong&gt;: Graceful degradation, retry logic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Accessibility&lt;/strong&gt;: Alt text generation (pair with vision LLM via same API), color contrast checks.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Legal/Ethics&lt;/strong&gt;: Disclose AI-generated content, respect copyrights (use models with commercial licenses), comply with data privacy (GDPR). Avoid harmful content filters.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At 10k users/day with moderate usage, expect $100s-$1000s/month—optimize via model routing and caching.&lt;/p&gt;

&lt;h2&gt;
  
  
  Case Studies and Real-World Examples
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;E-commerce&lt;/strong&gt;: Dynamic product visualizations (e.g., "red sneakers in mountain setting") increase conversions.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;SaaS Design Tools&lt;/strong&gt;: Instant mockups.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Content Platforms&lt;/strong&gt;: Auto-generate thumbnails or illustrations.
Many apps using unified APIs like CometAPI report 40-60% reduction in integration time vs. multiple providers.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Common Challenges and Troubleshooting
&lt;/h2&gt;

&lt;ul&gt;
&lt;li&gt;Latency: Use faster models or edge caching.&lt;/li&gt;
&lt;li&gt;Quality Inconsistency: Refine prompts with examples; use system prompts for style consistency.&lt;/li&gt;
&lt;li&gt;Costs Overruns: Set budgets/alerts.&lt;/li&gt;
&lt;li&gt;API Changes: Unified services like CometAPI abstract this.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Conclusion: Get Started with CometAPI Today
&lt;/h2&gt;

&lt;p&gt;Integrating AI image generation is no longer optional—it's a superpower for web apps. With robust models, straightforward APIs, and services like &lt;strong&gt;CometAPI&lt;/strong&gt; providing one-key access to Midjourney, GPT Image, Flux, and hundreds more, developers can focus on innovation rather than infrastructure.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Call to Action:&lt;/strong&gt; Visit &lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt;, grab your free credits, and implement the code above. Experiment with different models to find the perfect fit for your app. Your users (and metrics) will thank you.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQs
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Q: Can I use DALL-E 3 to generate multiple images in one API call?
&lt;/h3&gt;

&lt;p&gt;No. DALL-E 3 only supports &lt;code&gt;n=1&lt;/code&gt; — one image per request. If you need multiple variations, you'll need to make separate requests, either sequentially or in parallel. DALL-E 2 is the model that supports batch generation (up to &lt;code&gt;n=10&lt;/code&gt; per request).&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: How long does a DALL-E image URL stay valid?
&lt;/h3&gt;

&lt;p&gt;About 1 hour. OpenAI's image URLs are temporary — don't store the URL and expect it to work the next day. Download the image immediately after generation and save it to your own storage (S3, Cloudflare R2, etc.). Alternatively, use &lt;code&gt;response_format: "b64_json"&lt;/code&gt; to get the image data directly in the response, bypassing the URL expiration issue entirely.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: What's the difference between GPT Image 2 and DALL-E 3?
&lt;/h3&gt;

&lt;p&gt;GPT Image 2 is better at rendering text inside images, supports quality tiers (&lt;code&gt;low&lt;/code&gt;/&lt;code&gt;medium&lt;/code&gt;/&lt;code&gt;high&lt;/code&gt;), and generates faster. DALL-E 3 returns a URL by default (easier to handle), supports batch-friendly workflows via &lt;code&gt;response_format&lt;/code&gt;, and is the safer default for general creative use. The two models also use different parameter sets — &lt;code&gt;response_format&lt;/code&gt; works on DALL-E 3 but not GPT Image 2.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: Why does my Qwen Image request fail when I set n=2?
&lt;/h3&gt;

&lt;p&gt;Qwen Image only supports &lt;code&gt;n=1&lt;/code&gt;. Passing any higher value will return a 400 error. If you need multiple images, make separate requests.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: Do I need a separate API key for each model?
&lt;/h3&gt;

&lt;p&gt;No. CometAPI uses a single API key across all models — &lt;a href="https://www.cometapi.com/models/openai/dall-e-3/" rel="noopener noreferrer"&gt;DALL-E 3&lt;/a&gt;, GPT Image 2, Qwen Image, and everything else in their catalog. You switch models by changing the &lt;code&gt;model&lt;/code&gt; field in your request, not by managing multiple keys.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: What sizes does GPT Image 2 support?
&lt;/h3&gt;

&lt;p&gt;&lt;a href="https://www.cometapi.com/models/openai/gpt-image-2/" rel="noopener noreferrer"&gt;GPT Image 2&lt;/a&gt; supports &lt;code&gt;1024x1024&lt;/code&gt; (square), &lt;code&gt;1536x1024&lt;/code&gt; (landscape), &lt;code&gt;1024x1536&lt;/code&gt; (portrait), and &lt;code&gt;auto&lt;/code&gt; (model picks based on the prompt). It does not support arbitrary custom resolutions.&lt;/p&gt;

&lt;h3&gt;
  
  
  Q: My prompt keeps getting filtered. How do I debug it?
&lt;/h3&gt;

&lt;p&gt;Two things to check: first, look at the &lt;code&gt;revised_prompt&lt;/code&gt; field in the response — providers sometimes rewrite your prompt, and seeing what they changed tells you what triggered the filter. Second, check if the &lt;code&gt;data&lt;/code&gt; array in the response is empty — that's the signal that generation was blocked rather than a network or auth error. Rephrase the prompt to be more neutral and avoid specific names, brands, or sensitive subjects.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/how-to-add-ai-image-generation-to-a-web-app/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=how-to-add-ai-image-generation-to-a-web-app"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Why Your AI Spend Looks Nothing Like Your Actual Usage Patterns</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Mon, 21 Sep 2026 04:16:51 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/why-your-ai-spend-looks-nothing-like-your-actual-usage-patterns-1j2i</link>
      <guid>https://dev.to/dylanfoster1/why-your-ai-spend-looks-nothing-like-your-actual-usage-patterns-1j2i</guid>
      <description>&lt;p&gt;&lt;em&gt;Your monthly AI invoice is a single line that traces nowhere — not to specific features, not to specific teams, not to the workloads that drove the cost. For AI-native startups, the gap between what the bill says and what the product actually does is the reason next quarter's AI forecast is mostly guesswork.&lt;/em&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The mismatch
&lt;/h2&gt;

&lt;p&gt;Open the most recent monthly invoice from any of the major AI providers. The format is consistent: a top-line dollar figure, a breakdown by model, possibly a breakdown by API key if you have set that up deliberately. What you will not find is any meaningful mapping to your actual product. Which feature drove most of the cost? Which team's experiments accounted for which slice? How much was production traffic versus internal R&amp;amp;D? Was the spike on the 14th a one-off or a new baseline? The invoice does not answer any of these questions, because the invoice was not designed to.&lt;/p&gt;

&lt;p&gt;This is a structural mismatch between how AI providers bill and how AI-native startups actually run. Provider billing is organised around the unit of inference — tokens consumed, requests made, seconds of video generated. Startups are organised around the unit of product — features shipped, experiments run, teams that own things, customers being served. The two shapes do not align, and the cost of that misalignment compounds every time someone asks a question the invoice cannot answer.&lt;/p&gt;

&lt;p&gt;This article is the version of that conversation that takes the problem seriously. The argument is not that providers should change their billing — they will not, and frankly do not need to. The argument is that the gap between provider billing and product reality is bridgeable by the team running the product, and the bridge unlocks decisions that are otherwise impossible to make. Most AI-native startups in 2026 are flying without instruments on this; the ones that have instrumented properly are making better calls about pricing, prioritisation, and forecasting than the ones that haven't.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The headline finding:&lt;/strong&gt; AI spend is bursty, multi-model, and feature-driven. AI billing is monthly, single-line, and provider-organised. The mismatch makes forecasting unreliable, makes feature-level pricing impossible, and makes the AI line item the one your CFO trusts least. The fix is not provider-side — it is at the metering layer, and most teams can build it in a week.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  Three patterns that don't fit subscription thinking
&lt;/h2&gt;

&lt;p&gt;To understand why standard billing infrastructure fails AI workloads, it helps to name the three workload patterns that make AI spend behave differently from the SaaS spend that came before it. Each pattern individually creates a forecasting challenge; together, they explain why AI line items are systematically the least predictable category on most startup budgets.&lt;/p&gt;

&lt;h3&gt;
  
  
  Bursty usage on feature launches
&lt;/h3&gt;

&lt;p&gt;AI workloads do not have a steady-state baseline in the way that SaaS workloads do. A typical AI-native startup's monthly token consumption can spike 5–10x in the week following a feature launch, then drop back to baseline as the launch traffic subsides. The spike is real — it represents actual customers using a new feature — but it is not the new baseline. Anyone forecasting from the spike will overstate next quarter's AI budget; anyone forecasting from the baseline will underestimate the cost of the next launch.&lt;/p&gt;

&lt;p&gt;The conventional response — "average it out across the quarter" — is the wrong answer. Averaged numbers hide both the launch behaviour and the steady-state, which means they cannot inform decisions about either. The right framing is to forecast launches and baseline separately, but doing that requires usage data tagged in a way that lets you separate them after the fact. Standard provider invoices do not have that data.&lt;/p&gt;

&lt;h3&gt;
  
  
  Multi-model workflows where one request touches &lt;strong&gt;several providers&lt;/strong&gt;
&lt;/h3&gt;

&lt;p&gt;A single product feature in 2026 routinely calls more than one model. A document analysis pipeline might use GPT-5.5 for synthesis, Claude Sonnet 4.6 for re-ranking, and Gemini 3.1 Pro for structured extraction — three providers, three&amp;nbsp; &lt;a href="https://www.cometapi.com/pricing/" rel="noopener noreferrer"&gt;rate cards&lt;/a&gt;, three contributions to the cost of a single user interaction. From the user's perspective, this is one feature. From the provider invoices' perspective, it is three independent line items distributed across three monthly bills.&lt;/p&gt;

&lt;p&gt;The result is that feature-level cost analysis becomes a manual reconciliation problem. Which slice of the OpenAI invoice belongs to the document analysis feature versus the chat feature versus the agent feature? Without explicit tagging at the request level, the answer is unknowable. Most teams either give up on the question or produce rough estimates that can move 50% in either direction depending on how the math is done. Neither is good enough for a product decision.&lt;/p&gt;

&lt;h3&gt;
  
  
  Internal R&amp;amp;D usage indistinguishable from production
&lt;/h3&gt;

&lt;p&gt;Engineers running prompt experiments, evaluation suites, or new-model comparisons generate genuine API traffic that lands on the same monthly invoice as production usage. When the invoice arrives, there is no native way to separate "production traffic our customers generated" from "R&amp;amp;D our team consumed." For early-stage startups, the R&amp;amp;D fraction can be 30–50% of total spend; for mature ones, it is smaller but still meaningful. Without separation, you cannot answer simple questions like "is our per-customer AI cost going up or are we just experimenting more this month?"&lt;/p&gt;

&lt;p&gt;This is the failure mode that hits hardest at series A / series B fundraising. Investors who see flat per-customer AI cost (because experiments and production are being counted together) cannot distinguish efficient products from inefficient ones; the wrong framing can hurt the conversation. Teams that have instrumented R&amp;amp;D versus production separately walk into those conversations with a much sharper story about their unit economics.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why this matters for forecasting
&lt;/h2&gt;

&lt;p&gt;Forecasting is the activity where the cost of unattributed AI spend shows up most painfully. A finance team trying to model next quarter's AI line needs to answer questions like:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;What does our AI cost look like at the current customer count versus 2x it?&lt;/li&gt;
&lt;li&gt;How much of last quarter's spend was production traffic versus internal experiments?&lt;/li&gt;
&lt;li&gt;If we launch the new agent feature in October, what does that do to the November and December bills?&lt;/li&gt;
&lt;li&gt;Which features have the highest AI cost per active user, and are we charging enough to cover them?&lt;/li&gt;
&lt;li&gt;What's the marginal AI cost of adding a new enterprise customer of size X?&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Each of these questions is answerable with properly attributed data. None of them is answerable from a standard provider invoice. The result is that AI forecasts produced from invoice data are typically either wildly optimistic (smoothing over launch spikes that will recur) or wildly pessimistic (anchoring on a single high-usage month). Both are wrong in different directions, and the finance team learns over time that the AI line is the one they cannot trust — which means it becomes the line they pad most conservatively, which means the budget conversation becomes more contentious than it needs to be.&lt;/p&gt;

&lt;p&gt;The shift that fixes this is moving from invoice-level data to request-level data, with each request tagged for the dimensions that matter for forecasting: which feature it served, which team owns it, whether it was production traffic or R&amp;amp;D, which customer or customer-tier triggered it, and which workflow path it took. Once the metering captures these dimensions at the request layer, every forecasting question above becomes a query against that data, not a guess against the invoice.&lt;/p&gt;

&lt;h2&gt;
  
  
  What proper cost attribution unlocks
&lt;/h2&gt;

&lt;p&gt;The case for instrumenting cost attribution is not just better forecasting. Once the per-request data exists, four downstream decisions become possible that are otherwise either guesswork or impossible to make defensibly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pricing the product accurately
&lt;/h3&gt;

&lt;p&gt;AI-native products that charge per seat, per usage, or per outcome all need to know what their underlying inference cost looks like by user, by usage tier, or by outcome category. A product priced at $99/month per user that turns out to cost $112 in AI inference per active user is in trouble; the same product priced at $99/month with $34 of AI cost per user is healthy. The difference between these two situations is invisible from the invoice and obvious from per-feature attribution data. Teams that have this data price their products with confidence; teams that don't are guessing — and the guess goes wrong in both directions often enough that it matters.&lt;/p&gt;

&lt;h3&gt;
  
  
  Prioritising engineering work
&lt;/h3&gt;

&lt;p&gt;Product roadmap decisions are routinely shaped by cost considerations: "can we afford to ship this feature given the AI bill it will add?" Without attribution, this question is unanswerable in advance. With attribution — specifically, the ability to look at similar existing features and estimate the AI cost of the proposed one — the question becomes a 20-minute analysis. Teams that prioritise this way ship more confidently, sequence work better, and avoid the awkward conversation six months later when a beloved feature turns out to be financially unsustainable.&lt;/p&gt;

&lt;h3&gt;
  
  
  Defending the AI budget line in CFO conversations
&lt;/h3&gt;

&lt;p&gt;Every AI-native startup's CFO at some point asks the same question: "why is the AI line so volatile, and what are we getting for it?" The teams that can answer in detail — here is the cost broken down by feature, here is the R&amp;amp;D fraction, here are the customer cohorts that consume most, here is the trend over the last six months — have a different conversation than the teams whose only answer is "because of the OpenAI invoice." The CFO's confidence in the budget directly determines how much friction the line item generates each quarter. Detailed attribution buys that confidence cheaply.&lt;/p&gt;

&lt;h3&gt;
  
  
  Identifying optimisation opportunities surgically
&lt;/h3&gt;

&lt;p&gt;When the AI bill jumps unexpectedly, the question is always "why?" — and the speed of answering that question determines whether the team gets to a fix in a day or a week. With attribution, you can isolate the spike to a specific feature, a specific user cohort, or a specific code path. Without attribution, you have to do detective work across multiple provider dashboards to figure out what changed. Most teams that have done both consistently report that proper attribution turns multi-hour or multi-day investigations into 15-minute queries.&lt;/p&gt;

&lt;h2&gt;
  
  
  The metering that makes this possible
&lt;/h2&gt;

&lt;p&gt;The shift from invoice-level to request-level cost data depends on metering infrastructure that captures the right dimensions at the moment each request happens. Most teams in 2026 build this on top of one of three patterns, listed in order of increasing investment and capability.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pattern 1: Per-key segmentation
&lt;/h3&gt;

&lt;p&gt;The simplest pattern, and the one most teams start with. You issue separate API keys for each major dimension you want to attribute on — one key per feature, one per team, one for R&amp;amp;D, one for production. The aggregator's billing dashboard (or, with significantly more effort, the underlying provider dashboards) shows usage broken down by key. At month-end, you have an attribution view that maps cleanly to the dimensions you cared about.&lt;/p&gt;

&lt;p&gt;Per-key segmentation is enough for many teams. It handles the production-vs-R&amp;amp;D split, the per-feature attribution for products with a handful of features, and the per-team attribution for small engineering organisations. Where it falls down is when you need finer-grained slicing — per-customer, per-workflow, per-user-tier — because the number of keys becomes unmanageable. For teams that hit that ceiling, the next pattern is the answer.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pattern 2: Request-level tagging at the application layer
&lt;/h3&gt;

&lt;p&gt;Instead of (or in addition to) per-key segmentation, you instrument your application to tag every AI request with the dimensions that matter: feature, customer ID, workflow step, environment, experiment cohort. The tags are logged to your own observability system alongside the request metadata; cost attribution becomes a query against that data, not a query against the provider invoice.&lt;/p&gt;

&lt;p&gt;This pattern is meaningfully more flexible than per-key segmentation because the dimensions are independent — you can slice by customer and feature simultaneously, or by workflow path and team simultaneously, in ways that key-based attribution cannot. The cost is the engineering investment in the metering layer (typically 3–10 days of work for a team that does not already have observability infrastructure) and the discipline of consistently tagging requests in application code.&lt;/p&gt;

&lt;h3&gt;
  
  
  Pattern 3: Integrated observability platforms
&lt;/h3&gt;

&lt;p&gt;For teams whose AI spend is large enough that the engineering investment in attribution pays back quickly, dedicated &lt;a href="https://www.cometapi.com/enterprise/" rel="noopener noreferrer"&gt;AI observability platforms&lt;/a&gt; (Helicone, Langfuse, Phoenix, and others in the 2026 landscape) provide request-level tracking out of the box. These platforms sit in the request path, capture all the dimensions you would otherwise build into your own metering layer, and produce dashboards and queries against the data. The trade-off is the vendor relationship and the routing change to put requests through the platform; the benefit is faster time-to-attribution and richer analysis capabilities than most teams would build internally.&lt;/p&gt;

&lt;p&gt;Most well-instrumented AI-native startups in 2026 use a combination — per-key segmentation for the coarse dimensions (production vs R&amp;amp;D, team boundaries) and either application-layer tagging or an observability platform for the finer dimensions. The combination scales well as the organisation grows; starting with per-key segmentation gives you immediate value while you decide whether to invest in deeper instrumentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  A worked example: a 12-person AI-native startup
&lt;/h2&gt;

&lt;p&gt;Concrete numbers help. Below, the per-feature attribution view for a representative 12-person AI-native startup running three core product features, with an additional row for internal R&amp;amp;D and one for shared infrastructure (embeddings, evals). All figures are illustrative but proportionally representative of what teams at this scale typically see.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cost dimension&lt;/th&gt;
&lt;th&gt;Monthly spend&lt;/th&gt;
&lt;th&gt;% of total&lt;/th&gt;
&lt;th&gt;Per active user&lt;/th&gt;
&lt;th&gt;Models used&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Feature A: AI chat&lt;/td&gt;
&lt;td&gt;$8,200&lt;/td&gt;
&lt;td&gt;32%&lt;/td&gt;
&lt;td&gt;$0.41&lt;/td&gt;
&lt;td&gt;GPT-5.5, Sonnet&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feature B: Document analysis&lt;/td&gt;
&lt;td&gt;$6,800&lt;/td&gt;
&lt;td&gt;26%&lt;/td&gt;
&lt;td&gt;$1.36&lt;/td&gt;
&lt;td&gt;Sonnet, Gemini&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Feature C: Agent workflows&lt;/td&gt;
&lt;td&gt;$4,500&lt;/td&gt;
&lt;td&gt;17%&lt;/td&gt;
&lt;td&gt;$3.21&lt;/td&gt;
&lt;td&gt;Opus, GPT-5.5&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Shared infra (embeddings, evals)&lt;/td&gt;
&lt;td&gt;$3,200&lt;/td&gt;
&lt;td&gt;12%&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Multiple&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Internal R&amp;amp;D and experiments&lt;/td&gt;
&lt;td&gt;$3,300&lt;/td&gt;
&lt;td&gt;13%&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;Multiple&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total&lt;/td&gt;
&lt;td&gt;$26,000&lt;/td&gt;
&lt;td&gt;100%&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;td&gt;—&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The conversation this table enables, that an invoice never would, is the per-active-user cost column. Feature A serves 20,000 active users; Feature B serves 5,000; Feature C serves 1,400. The per-user cost variation (41 cents, $1.36, $3.21) is genuinely useful information for the product team: it tells them that Feature C is the most expensive per user to run, and forces an honest conversation about whether the pricing or the underlying architecture needs to change. None of this is visible from a $26,000 monthly invoice without breakdown.&lt;/p&gt;

&lt;p&gt;The internal R&amp;amp;D fraction (13%) tells another important story: a healthy investment in experimentation, neither too low (suggesting the team is not exploring new models or prompt strategies) nor too high (suggesting R&amp;amp;D may be eating into production budget). Investors who see this fraction broken out separately are seeing the team's R&amp;amp;D investment explicitly, which is what they need to evaluate the company's engineering culture and unit economics independently.&lt;/p&gt;

&lt;h2&gt;
  
  
  The forecasting model that emerges
&lt;/h2&gt;

&lt;p&gt;Once the attribution data exists, forecasting next quarter's AI spend becomes a structured calculation rather than a guess. The model has three components — and once it is set up, the team can update it in 15 minutes whenever assumptions change.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Production baseline.&lt;/strong&gt; For each feature, take the trailing 90 days of per-active-user cost, multiplied by the forecast of active users in the period. This produces a baseline that grows linearly with customer count, which is the right shape for most production AI traffic.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Launch and event spikes.&lt;/strong&gt; For each planned product launch or major marketing moment, estimate the spike duration (typically 1–3 weeks) and the multiplier (typically 3–10x baseline traffic). Multiply into a one-time addition. This component captures the bursty pattern that breaks naive forecasting.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;R&amp;amp;D allocation.&lt;/strong&gt; Set the R&amp;amp;D budget as a percentage of total (10–20% is typical for AI-native startups in steady state) or as an absolute monthly cap. This component is a planning decision, not a forecast — but it should be set explicitly rather than absorbed silently into the production budget.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The sum of these three is the forecast. When something changes — a new launch added to the roadmap, a customer cohort growing faster than expected, a new model coming online that changes the per-user cost — the forecast updates immediately because the inputs are all explicit. Compare this to the current state in most AI-native startups, where the forecast is "last quarter's total times a growth factor we made up" — and the difference in forecasting accuracy is substantial.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What this means in practice:&lt;/strong&gt; Teams that move to attribution-based forecasting consistently report two changes. First, the variance between forecast and actual drops from typical ranges of 30–50% down to 5–15%. Second, the conversations between engineering and finance get easier — both sides are looking at the same data, the same assumptions are explicit, and disagreements about the AI line are about real questions ("should we cap R&amp;amp;D this quarter?") rather than about whose number is right.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;h2&gt;
  
  
  How to get started this week
&lt;/h2&gt;

&lt;p&gt;If your team is currently flying without instruments on AI cost attribution, the path from invoice-only to properly attributed is shorter than it looks. A practical sequence:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Define the dimensions you actually need to attribute on.&lt;/strong&gt; For most teams, the starting list is: feature (3–6 categories), environment (production vs R&amp;amp;D), and team (if you have multiple teams using AI). Customer-level attribution is the next layer up but can wait until the first three are working. Resist the urge to track every dimension you might want — start with what answers the questions your CFO is actually asking.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Issue one API key per dimension you want to track coarsely.&lt;/strong&gt; If your aggregator supports per-key billing dashboards, this is the fastest path to immediate value. One key per feature, one key for R&amp;amp;D, one key for shared infrastructure. The attribution shows up in the dashboard automatically. Time investment: an hour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Run for one month before drawing conclusions.&lt;/strong&gt; A single month of data is enough to see the per-feature shape but not enough to identify seasonal patterns or trend lines. Don't make big decisions from the first month; do start a habit of looking at the data weekly so the patterns become familiar.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Decide whether the coarse view is enough.&lt;/strong&gt; After 30 days, you will know whether per-key segmentation answers the questions you actually need answered. For many teams, it does. For teams that need finer-grained slicing (per-customer, per-workflow), now is the time to add application-layer tagging or evaluate an observability platform — informed by 30 days of real data about what you need.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Build the forecasting model.&lt;/strong&gt; Once you have three months of attributed data, the three-component forecast (production baseline + launch spikes + R&amp;amp;D allocation) can be built in an afternoon. This is the deliverable that changes the conversation with your CFO. Most teams report it as the single highest-leverage piece of finance instrumentation they ship in their first year.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Where this leaves you
&lt;/h2&gt;

&lt;p&gt;Your monthly AI invoice does not look like your product, and that mismatch is the reason AI forecasting feels harder than it should. The fix is not on the provider side. It is at the metering layer — making sure each request is tagged for the dimensions you actually care about, so that attribution becomes a query against your data rather than a guess against the invoice. Once that infrastructure exists, four things become possible that are otherwise impossible: accurate pricing, defensible prioritisation, credible CFO conversations, and surgical optimisation when things go wrong.&lt;/p&gt;

&lt;blockquote&gt;
&lt;p&gt;Provider billing is organised around tokens. Your product is organised around features. The mismatch is bridgeable, the bridge is cheap to build, and it unlocks decisions you cannot otherwise make. The teams that have instrumented attribution properly forecast AI cost within 5–15% accuracy; the teams that haven't run 30–50% off. The instrumentation is the difference.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;Ready to integrate reliably? Head to&amp;nbsp;&lt;a href="https://cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt;&amp;nbsp;and&amp;nbsp;&lt;a href="https://apidoc.cometapi.com/" rel="noopener noreferrer"&gt;API doc&lt;/a&gt;&amp;nbsp;for seamless Claude Fable 5 access alongside other frontier models, unified billing, and enterprise-grade reliability. Sign up today and get started with generous credits for new users—your next breakthrough project awaits.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/why-your-ai-spend-looks-nothing-like-your-actual-usage-patterns/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=why-your-ai-spend-looks-nothing-like-your-actual-usage-patterns"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>AGI vs ASI: The Distinction I Use When Evaluating AI Agents</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Mon, 21 Sep 2026 03:34:22 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/agi-vs-asi-the-distinction-i-use-when-evaluating-ai-agents-2n4b</link>
      <guid>https://dev.to/dylanfoster1/agi-vs-asi-the-distinction-i-use-when-evaluating-ai-agents-2n4b</guid>
      <description>&lt;p&gt;I find the AGI-versus-ASI debate useful only when it changes what I measure or how I build. A model producing a convincing answer is one thing. A system reliably completing unfamiliar work, across domains, without supervision is another.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Artificial General Intelligence (AGI)&lt;/strong&gt; means broadly human-level cognitive capability across tasks. &lt;strong&gt;Artificial Superintelligence (ASI)&lt;/strong&gt; means capability substantially beyond the best humans across virtually every cognitive domain. Generality is the central requirement for AGI; broad superiority is the additional requirement for ASI. Neither label follows from a strong coding demo.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start With Three Separate Questions
&lt;/h2&gt;

&lt;p&gt;When evaluating a system, I separate breadth, performance, and autonomy. Can it handle unfamiliar task categories? How well does it perform against competent humans? How long can it operate before someone must intervene? Collapsing those questions into one “intelligence” score hides the failures that matter in production.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Dimension&lt;/th&gt;
&lt;th&gt;AGI&lt;/th&gt;
&lt;th&gt;ASI&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Breadth&lt;/td&gt;
&lt;td&gt;General competence across cognitive tasks&lt;/td&gt;
&lt;td&gt;General competence with broad superhuman performance&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Performance&lt;/td&gt;
&lt;td&gt;Human-comparable, with thresholds varying by definition&lt;/td&gt;
&lt;td&gt;Substantially beyond the best human experts&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Learning&lt;/td&gt;
&lt;td&gt;Transfers knowledge and adapts to unfamiliar tasks&lt;/td&gt;
&lt;td&gt;Could develop better learning methods and improve AI itself&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Scientific work&lt;/td&gt;
&lt;td&gt;Performs or assists with human-level research&lt;/td&gt;
&lt;td&gt;Could accelerate scientific and technological progress&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Status&lt;/td&gt;
&lt;td&gt;No broad consensus that current systems qualify&lt;/td&gt;
&lt;td&gt;Theoretical&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Principal concerns&lt;/td&gt;
&lt;td&gt;Reliability, displacement, bias, misuse, control&lt;/td&gt;
&lt;td&gt;Those concerns plus potentially severe loss-of-control risks&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A narrow system can outperform humans at chess, protein structure prediction, or image classification without being AGI. Conversely, a broadly capable system does not become ASI merely because it runs faster than a person. Digital speed, parallel execution, and scalable deployment matter, but they are not substitutes for demonstrated competence.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Would Count as AGI?
&lt;/h2&gt;

&lt;p&gt;The idea reaches back to researchers including Alan Turing and John McCarthy. The practical target is a system that can learn, reason, plan, and apply knowledge across domains without needing a separate task-specific training process for each new assignment.&lt;/p&gt;

&lt;p&gt;Definitions differ on the human reference point. Some use roughly median human performance; others require expert-level competence across a broad task distribution. DeepMind’s levels framework distinguishes performance levels, including expert, virtuoso, and superhuman capability. That makes “human-level” an incomplete specification unless the task distribution and comparison population are also stated.&lt;/p&gt;

&lt;p&gt;For engineering purposes, I would look for transferable knowledge, long-horizon planning, effective tool use, adaptation to unfamiliar conditions, and reliable self-correction. None of these is established by an isolated benchmark result. A system that solves an Olympiad problem but misses a simple instruction has demonstrated an impressive capability, not consistent general competence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Why Current Models Remain Hard to Classify
&lt;/h3&gt;

&lt;p&gt;Frontier systems can code, analyze documents, use browsers and other tools, process multiple modalities, and assist with research. The source’s 2026 discussion names GPT-5 variants, Claude Opus, Gemini, and OpenAI’s o-series as examples of progress. It also identifies persistent weaknesses in factual grounding, causal reasoning, memory, embodied learning, long-term planning, and self-correction.&lt;/p&gt;

&lt;p&gt;The source reports approximately &lt;strong&gt;35–65% on Humanity’s Last Exam&lt;/strong&gt;, compared with approximately &lt;strong&gt;90% for human experts&lt;/strong&gt;, and leaders &lt;strong&gt;near or at 85% on ARC-AGI-2&lt;/strong&gt;. Those figures need dated evaluation references, model configurations, and scoring conditions before I would use them in a technical comparison. They are not interchangeable measurements of “percentage of AGI achieved.”&lt;/p&gt;

&lt;p&gt;The more revealing contrast is with deployed work. The source cites approximately &lt;strong&gt;2.5% automation of freelance work for top models in some assessments&lt;/strong&gt;. That is an assessment-specific result, not a universal automation rate, but it illustrates the benchmark-to-workflow gap. GAIA, BIG-bench, and ARC evaluations probe different capabilities; the Legg-Hutter intelligence measure is a theoretical formalization, not an equivalent operational leaderboard.&lt;/p&gt;

&lt;p&gt;My working description is &lt;strong&gt;advanced general-purpose AI with increasingly agentic capabilities&lt;/strong&gt;. It communicates what these systems can do without treating a disputed milestone as settled.&lt;/p&gt;

&lt;h2&gt;
  
  
  Autonomy and Reliability Matter More Than the Label
&lt;/h2&gt;

&lt;p&gt;An answer-generating model and an agent that browses, writes code, runs tests, purchases services, and messages people expose very different operational risks. Extending the time horizon also increases the opportunity for a small mistake to propagate through subsequent actions.&lt;/p&gt;

&lt;p&gt;METR’s 2025 time-horizon study provides a more concrete signal than impressions of intelligence. It measured task difficulty using the time a human expert would need, then evaluated agent success. The study found that the &lt;strong&gt;50% task-completion time horizon had doubled approximately every seven months over six years&lt;/strong&gt;. That is an observed trend, not proof that it will continue or that AGI arrives on a particular date.&lt;/p&gt;

&lt;p&gt;I would also resist treating 50% completion as a deployment target. An agent can improve rapidly on that metric while still needing extensive supervision for consequential work. Code that passes an initial check but introduces a security vulnerability remains a failed outcome.&lt;/p&gt;

&lt;p&gt;The source attributes continuing weaknesses in multi-step planning, financial analysis, coherent realistic video, and some expert academic exams to the Stanford AI Index 2026. It also cites the International AI Safety Report 2026 on fabricated information, flawed code, and misleading advice. These are central evaluation concerns, not cosmetic defects around otherwise complete general intelligence.&lt;/p&gt;

&lt;h2&gt;
  
  
  ASI Would Be a System, Not Just a Better Model
&lt;/h2&gt;

&lt;p&gt;ASI is usually defined as intelligence vastly exceeding the smartest humans across scientific discovery, strategic planning, creativity, social reasoning, and other cognitive domains. It remains hypothetical. Claims that it would instantly solve global problems, discover new physics, or make its reasoning incomprehensible are possibilities or speculation, not established properties.&lt;/p&gt;

&lt;p&gt;I find the system-level framing more useful: models, memory, tools, compute, data access, deployment infrastructure, and feedback loops working together. A single conversational interface could conceal that infrastructure, but the interface would tell us little about the system’s actual capability.&lt;/p&gt;

&lt;p&gt;Recursive self-improvement is one possible mechanism. An AI might improve training software, evaluations, synthetic data generation, inference efficiency, architectures, or hardware design. If those improvements substantially accelerate further improvements, capability growth could speed up. An “intelligence explosion” is a proposed outcome of that feedback loop, not something guaranteed by the definition of AGI.&lt;/p&gt;

&lt;p&gt;The upside could include faster progress in medicine, materials, education, climate science, and productivity. The downside includes amplified cyberattacks, deception, biological risks, economic disruption, and loss of human control. If a system outperforms human experts, checking every answer directly becomes harder. Oversight would need multiple approaches: scalable supervision, interpretability, adversarial testing, formal verification where applicable, model-to-model critique, audits, and institutional controls.&lt;/p&gt;

&lt;h2&gt;
  
  
  Four Possible Routes Beyond Human-Level AI
&lt;/h2&gt;

&lt;p&gt;The source attributes four pathways to a Google DeepMind report titled &lt;em&gt;From AGI to ASI&lt;/em&gt;, dated June 2026. I would treat that report attribution and publication date as requiring verification before citing them as established evidence. The pathways themselves are useful hypotheses to distinguish.&lt;/p&gt;

&lt;h3&gt;
  
  
  Scaling and Algorithmic Change
&lt;/h3&gt;

&lt;p&gt;Continued scaling means more compute, better data, larger or more efficient models, stronger inference-time reasoning, and improved tool use. The open question is which bottlenecks yield to those investments. Data quality, energy, memory, reasoning, embodiment, and alignment do not necessarily improve at the same rate.&lt;/p&gt;

&lt;p&gt;Algorithmic change is a different route: new architectures, world models, stronger memory, neuro-symbolic reasoning, active learning, self-play, or planning systems. Future progress need not come exclusively from larger versions of current transformers. Neither pathway establishes a predictable date for superintelligence.&lt;/p&gt;

&lt;h3&gt;
  
  
  Recursive Improvement and Agent Collectives
&lt;/h3&gt;

&lt;p&gt;Recursive improvement starts with AI contributing to AI development. The important distinction is between useful engineering assistance and a feedback loop powerful enough to accelerate the underlying research process dramatically.&lt;/p&gt;

&lt;p&gt;Multi-agent collectives offer another possibility. Specialized systems could search, simulate, debate, verify, and coordinate, potentially producing capability beyond any individual component. The source describes this as an AI organization or virtual agent economy. I would still evaluate the collective as a whole: adding agents is an architectural choice, not evidence that the resulting system is more reliable or superintelligent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Forecasts Are Scenarios, Not Delivery Dates
&lt;/h2&gt;

&lt;p&gt;The source’s forecasts span &lt;strong&gt;2027–2035&lt;/strong&gt; for AGI, while also citing leader expectations around &lt;strong&gt;2026–2029&lt;/strong&gt;, a reference by Dario Amodei to powerful systems in &lt;strong&gt;late 2026 or early 2027&lt;/strong&gt;, and Ray Kurzweil’s &lt;strong&gt;2029&lt;/strong&gt; prediction. It attributes approximately &lt;strong&gt;25% probability by 2029&lt;/strong&gt; and &lt;strong&gt;50% by 2033&lt;/strong&gt; to forecasting communities. Such figures require the exact question, resolution criteria, and forecast date to be meaningfully compared.&lt;/p&gt;

&lt;p&gt;For ASI, the source gives &lt;strong&gt;1–10 years after AGI&lt;/strong&gt;, speculation around &lt;strong&gt;2030–2040&lt;/strong&gt;, and an optimistic scenario of &lt;strong&gt;AGI in 2027 followed by ASI in 2028–2030&lt;/strong&gt;. Skeptical scenarios extend for decades. These are competing expectations, not a consolidated expert schedule.&lt;/p&gt;

&lt;p&gt;Compute, data, energy, regulation, geopolitical constraints, embodiment, and alignment can all affect progress. Synthetic data and efficiency improvements might relieve some bottlenecks without removing others. I would not make an infrastructure commitment depend on one forecast winning.&lt;/p&gt;

&lt;p&gt;Economic activity is easier to measure than arrival dates. The source attributes &lt;strong&gt;$581.7 billion in global corporate AI investment in 2025&lt;/strong&gt;, &lt;strong&gt;$344.7 billion in private AI investment&lt;/strong&gt;, &lt;strong&gt;88% organizational AI adoption&lt;/strong&gt;, and &lt;strong&gt;70% of organizations using generative AI in at least one business function&lt;/strong&gt; to Stanford’s 2026 AI Index. These are investment and adoption indicators, not demonstrations of AGI.&lt;/p&gt;

&lt;h2&gt;
  
  
  What I Would Build Around Today
&lt;/h2&gt;

&lt;p&gt;For cross-model experiments, a unified API such as CometAPI can simplify access to multiple providers through an OpenAI-compatible interface. That is useful infrastructure for comparing agents and evaluators; it does not resolve differences in capability, reliability, or oversight requirements.&lt;/p&gt;

&lt;p&gt;My priorities would be representative task evaluations, explicit success criteria, cost and failure tracking, and human review for consequential actions. I would test long-running workflows as workflows, rather than infer their reliability from single-turn answers. Open benchmarks and safety research are useful complements to those application-specific evaluations.&lt;/p&gt;

&lt;p&gt;The distinction I keep is straightforward: &lt;strong&gt;AGI concerns broad human-level competence; ASI concerns broad capability beyond human experts.&lt;/strong&gt; For a system I might deploy now, the decisive question is still what work it completes reliably, under which conditions, and how I detect when it fails.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/agi-vs-asi-whate28099s-the-difference-updated-2026/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=agi-vs-asi-whate28099s-the-difference-updated-2026"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>A Practical Multi-Model Gemini Integration with One API Surface</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Mon, 21 Sep 2026 02:36:27 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/a-practical-multi-model-gemini-integration-with-one-api-surface-9d2</link>
      <guid>https://dev.to/dylanfoster1/a-practical-multi-model-gemini-integration-with-one-api-surface-9d2</guid>
      <description>&lt;p&gt;By July 2026, running a production AI feature usually means choosing between several models rather than standardizing on one. Gemini 3.1 Pro may be the right choice for long-context reasoning or multimodal analysis, while another provider may be preferable for latency, price, or a particular capability.&lt;/p&gt;

&lt;p&gt;The integration cost is the part I dislike. Native SDKs bring separate authentication flows, payload formats, error models, quotas, dashboards, and billing systems. They also make model fallback harder: application code ends up coupled to provider-specific helpers instead of expressing the actual operation—send messages, receive a completion—in a provider-neutral way.&lt;/p&gt;

&lt;p&gt;For this kind of architecture, I use a unified gateway such as &lt;a href="https://www.cometapi.com/" rel="noopener noreferrer"&gt;CometAPI&lt;/a&gt;. It exposes more than 500 models through one endpoint, supports the OpenAI SDK as well as the native Gemini request format, and claims up to 20% savings on input and output tokens compared with official native pricing.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Gemini lineup brings
&lt;/h2&gt;

&lt;p&gt;Google’s 2026 model family covers more than text generation:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 3.1 Pro&lt;/strong&gt;: the flagship reasoning and long-context model for agentic workflows, document analysis, and code generation. The &lt;a href="https://www.cometapi.com/how-to-use-gemini-3-1-pro-api/" rel="noopener noreferrer"&gt;Gemini 3.1 Pro API guide&lt;/a&gt; covers its integration.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 3.5 Flash&lt;/strong&gt;: the speed- and cost-optimized option for high-volume, latency-sensitive workloads.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Nano Banana 2 (Gemini 3 Pro Image)&lt;/strong&gt;: an image generation and editing model focused on high-fidelity, prompt-accurate visuals. See the &lt;a href="https://www.cometapi.com/gemini-3-pro-image-nano-banana-2-api/" rel="noopener noreferrer"&gt;Nano Banana 2 API guide&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Veo 3.1&lt;/strong&gt;: a text-to-video and image-to-video model that generates video clips with synchronized audio. See the &lt;a href="https://www.cometapi.com/how-to-use-veo-3-1-api/" rel="noopener noreferrer"&gt;Veo 3.1 API guide&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini Omni&lt;/strong&gt;: a unified multimodal model that reasons across text, images, audio, and video in one request. See &lt;a href="https://www.cometapi.com/what-is-gemini-omni-googlee28099s-new-multimodal-video-model-explained/" rel="noopener noreferrer"&gt;What Is Gemini Omni?&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The appeal of a single access layer is operational as much as technical. Instead of separately provisioning Google Cloud IAM, quotas, and billing for every integration, I can use one API key and base URL. Switching between Gemini 3.1 Pro, Nano Banana 2, Veo 3.1, or models from other providers becomes a model-selection change rather than a client-library migration.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why native multi-SDK integrations become expensive
&lt;/h2&gt;

&lt;p&gt;Each provider makes different assumptions about:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Authentication and credential management&lt;/li&gt;
&lt;li&gt;System instructions&lt;/li&gt;
&lt;li&gt;Multimodal content schemas&lt;/li&gt;
&lt;li&gt;Error and retry behavior&lt;/li&gt;
&lt;li&gt;Rate-limit headers and quota accounting&lt;/li&gt;
&lt;li&gt;Response and token-usage formats&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Those differences are manageable in a proof of concept. They become expensive when the same product supports multiple models and dynamic routing. Middleware has to normalize requests, translate provider-specific failures, and keep downstream parsing stable. Every new model adds another compatibility surface.&lt;/p&gt;

&lt;p&gt;There is also a structural vendor-lock-in problem. If business logic depends on native SDK helpers, moving traffic from one provider to another—or adding a fallback when latency or quotas deteriorate—can require a substantial refactor. Fragmented consoles and invoices make cost attribution harder too, especially when several teams share multiple models.&lt;/p&gt;

&lt;p&gt;A gateway does not remove all model differences, but it puts the provider-specific translation in one place. The application talks to a standardized interface; the gateway translates requests for the selected backend and normalizes the response.&lt;/p&gt;

&lt;h2&gt;
  
  
  Two ways to call Gemini
&lt;/h2&gt;

&lt;p&gt;The unified endpoint is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://api.cometapi.com/v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is an API base URL for an SDK or HTTP client, not a browser page. The client appends a route such as &lt;code&gt;/chat/completions&lt;/code&gt;; opening the base URL directly returns a 404, which is expected and indicates that the server is reachable.&lt;/p&gt;

&lt;p&gt;There are two supported calling styles:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;OpenAI-compatible&lt;/strong&gt;: use the OpenAI SDK and set &lt;code&gt;model&lt;/code&gt; to a Gemini model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Native Gemini format&lt;/strong&gt;: call the &lt;code&gt;generateContent&lt;/code&gt; endpoint directly using Google’s request schema.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The native endpoint has this form:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://api.cometapi.com/v1beta/models/{model}:generateContent
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Native requests use the &lt;code&gt;x-goog-api-key&lt;/code&gt; header. The &lt;a href="https://apidoc.cometapi.com/quickstarts/text/gemini-api" rel="noopener noreferrer"&gt;native Gemini API quickstart&lt;/a&gt; documents that format.&lt;/p&gt;

&lt;h2&gt;
  
  
  Drop-in usage with the OpenAI Python client
&lt;/h2&gt;

&lt;p&gt;For an existing OpenAI integration, the practical change is usually the base URL, API key, and model name:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;completion&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3.1-pro&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;system&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;You are a helpful technical assistant.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;How does a unified API endpoint simplify multi-model routing?&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;completion&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;choices&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="n"&gt;message&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;content&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The useful property here is not merely that the request succeeds. The response follows the OpenAI JSON schema, so existing parsing, token accounting, and error-handling wrappers can remain in place.&lt;/p&gt;

&lt;p&gt;The same abstraction makes routing logic straightforward:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3.1-pro&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;}],&lt;/span&gt;
        &lt;span class="n"&gt;temperature&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mf"&gt;0.7&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="k"&gt;try&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gpt-5.4&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="k"&gt;except&lt;/span&gt; &lt;span class="nb"&gt;Exception&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;complete&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;prompt&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3.1-pro&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In a real service I would replace the broad exception with provider-aware handling and add bounded retries, but the important routing decision is just the &lt;code&gt;model&lt;/code&gt; parameter.&lt;/p&gt;

&lt;h2&gt;
  
  
  Multimodal requests use the familiar content shape
&lt;/h2&gt;

&lt;p&gt;Gemini 3.1 Pro can process visual and auditory inputs. Through the OpenAI-compatible interface, media can be supplied either as public URLs or as base64-encoded data embedded in the request.&lt;/p&gt;

&lt;p&gt;An image request looks like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="n"&gt;response&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;chat&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;completions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3.1-pro&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;messages&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;role&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;user&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
                &lt;span class="p"&gt;{&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Analyze the trends shown in this chart and &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;summarize the key takeaways.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                    &lt;span class="p"&gt;),&lt;/span&gt;
                &lt;span class="p"&gt;},&lt;/span&gt;
                &lt;span class="p"&gt;{&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;image_url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;url&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://example.com/charts/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;performance-summary.png&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                        &lt;span class="p"&gt;),&lt;/span&gt;
                    &lt;span class="p"&gt;},&lt;/span&gt;
                &lt;span class="p"&gt;},&lt;/span&gt;
            &lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The gateway translates the &lt;code&gt;image_url&lt;/code&gt; structure into the backend-specific format. It does not enhance, compress, or otherwise change Gemini’s multimodal capabilities. Accuracy, latency, and processing limits still come from Gemini 3.1 Pro.&lt;/p&gt;

&lt;p&gt;The benefit is downstream consistency: generated text, usage data, and finish reasons can be handled using the same parsing path whether the request is served by Gemini 3.1 Pro or another multimodal model.&lt;/p&gt;

&lt;h2&gt;
  
  
  The trade-off against direct Google integration
&lt;/h2&gt;

&lt;p&gt;A direct integration with Vertex AI or Google AI Studio has one obvious advantage: fewer network hops and direct access to Google’s infrastructure. If absolute minimum latency is the only metric that matters, native access may be the better choice.&lt;/p&gt;

&lt;p&gt;A unified endpoint adds an intermediary hop, although optimized routing is intended to keep the additional latency negligible for most applications. In return, I get several operational advantages:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;One invoice for usage across Gemini 3.1 Pro, GPT-5.4, and more than 500 supported models&lt;/li&gt;
&lt;li&gt;Centralized usage analytics for tokens, latency, and model-level cost distribution&lt;/li&gt;
&lt;li&gt;Fewer production credentials to secure&lt;/li&gt;
&lt;li&gt;Easier A/B testing and fallback routing&lt;/li&gt;
&lt;li&gt;Up to 20% savings on Gemini input and output tokens compared with official native pricing&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The cost and management benefits matter most for high-volume workloads such as large-scale document analysis and continuous agentic workflows. They matter less for a low-volume service where native integration is already simple and the gateway’s flexibility is unnecessary.&lt;/p&gt;

&lt;h2&gt;
  
  
  Limitations worth designing for
&lt;/h2&gt;

&lt;p&gt;A compatibility layer is not the same thing as identical model behavior.&lt;/p&gt;

&lt;h3&gt;
  
  
  New features may arrive later
&lt;/h3&gt;

&lt;p&gt;Google can release experimental parameters and provider-specific capabilities before they are represented by a standardized gateway schema. There may be a short propagation delay before those features are available through the translation layer.&lt;/p&gt;

&lt;p&gt;If day-one access to Google-specific experimental features is critical, I would retain a native integration for the relevant sandboxed workloads rather than forcing every request through the unified path.&lt;/p&gt;

&lt;h3&gt;
  
  
  Quotas move to the gateway
&lt;/h3&gt;

&lt;p&gt;When traffic is routed through the unified endpoint, rate limits and quotas are managed there rather than directly in Google AI Studio or Vertex AI. The application should inspect gateway rate-limit headers and implement appropriate backoff and retry behavior.&lt;/p&gt;

&lt;p&gt;The centralized quota is simpler to administer, but total token consumption across all active models must be coordinated within that quota.&lt;/p&gt;

&lt;h3&gt;
  
  
  Schemas are standardized, not identical
&lt;/h3&gt;

&lt;p&gt;Different models still interpret prompts differently. System instructions, temperature bounds, safety thresholds, and other behavior can vary between GPT models and Gemini 3.1 Pro even when the request schema is the same.&lt;/p&gt;

&lt;p&gt;For dynamic routing, I recommend:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Validating system prompts against every target model&lt;/li&gt;
&lt;li&gt;Avoiding assumptions about parameter ranges&lt;/li&gt;
&lt;li&gt;Handling model-specific API errors explicitly&lt;/li&gt;
&lt;li&gt;Testing both text and multimodal payloads during failover&lt;/li&gt;
&lt;li&gt;Treating output quality as a routing metric, not just latency and price&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Migration plan
&lt;/h2&gt;

&lt;p&gt;I would migrate an existing native Gemini integration in three passes.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Find provider-specific code
&lt;/h3&gt;

&lt;p&gt;Search for imports and calls from the native Google SDKs, including:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;@google/generative-ai
google-generativeai
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then inventory every Gemini call and record parameters such as &lt;code&gt;temperature&lt;/code&gt;, &lt;code&gt;top-p&lt;/code&gt;, system instructions, safety settings, and media handling. These details are where compatibility issues tend to surface.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Move configuration out of code
&lt;/h3&gt;

&lt;p&gt;Store the key in an environment variable such as &lt;code&gt;API_KEY&lt;/code&gt; rather than hardcoding it. Keep the base URL configurable too, so routing can change without another application release.&lt;/p&gt;

&lt;p&gt;For the OpenAI-compatible client, the relevant configuration is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;openai&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;OpenAI&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;base_url&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;getenv&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;AI_BASE_URL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://api.cometapi.com/v1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;api_key&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. Add and test routing before removing the native path
&lt;/h3&gt;

&lt;p&gt;Wrap model calls behind an application-level interface. Select the &lt;code&gt;model&lt;/code&gt; based on latency, cost, quota, or capability, and test simulated rate limits and API failures.&lt;/p&gt;

&lt;p&gt;A migration test suite should verify that the application can fail over from GPT-5.4 to Gemini 3.1 Pro—or in the other direction—without exposing an unhandled exception to the user. It should also validate response parsing, token accounting, and both image and audio workflows across target models.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://apidoc.cometapi.com/overview/quick-start" rel="noopener noreferrer"&gt;quick-start documentation&lt;/a&gt; provides the setup sequence, but the important architectural step is keeping the gateway configuration and model choice outside feature code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Final architecture notes
&lt;/h2&gt;

&lt;p&gt;For a production multi-model application in July 2026, I would separate three concerns:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Application behavior&lt;/strong&gt;: prompts, tools, business logic, and output validation&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Model policy&lt;/strong&gt;: selection based on capability, cost, latency, and availability&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Provider transport&lt;/strong&gt;: authentication, endpoint formatting, retries, and response normalization&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;A unified Gemini endpoint is useful when the goal is to keep those layers separate. It lets an existing OpenAI SDK call Gemini 3.1 Pro, exposes the broader Gemini family—including Nano Banana 2, Veo 3.1, and Gemini Omni—and supports native Gemini requests when that schema is preferable.&lt;/p&gt;

&lt;p&gt;Native SDKs remain the right escape hatch for immediate access to experimental provider-specific features or for systems where an extra network hop is unacceptable. For most applications that need model switching, centralized billing, multimodal support, and simpler operational management, standardizing the transport layer is the more maintainable choice.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/connecting-to-the-gemini-api-via-a-single-access/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=connecting-to-the-gemini-api-via-a-single-access"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
      <category>opensource</category>
    </item>
    <item>
      <title>Choosing a Coding Agent Model: Claude Opus 5, GPT-5.6 Sol, or Gemini 3.7 Flash</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Mon, 21 Sep 2026 01:37:39 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/choosing-a-coding-agent-model-claude-opus-5-gpt-56-sol-or-gemini-37-flash-4127</link>
      <guid>https://dev.to/dylanfoster1/choosing-a-coding-agent-model-claude-opus-5-gpt-56-sol-or-gemini-37-flash-4127</guid>
      <description>&lt;p&gt;There is no benchmark result that settles this choice for every coding agent.&lt;/p&gt;

&lt;p&gt;The published &lt;strong&gt;Terminal-Bench 2.1&lt;/strong&gt; numbers are close: &lt;strong&gt;Claude Opus 5 scores 89% at Max effort&lt;/strong&gt;, &lt;strong&gt;GPT-5.6 Sol scores 88.8% in the reported single-agent run and 91.9% in Ultra&lt;/strong&gt;, and &lt;strong&gt;Gemini 3.7 Flash scores 85.8%&lt;/strong&gt;. Those results are useful for narrowing the field, but they are not a clean ranking. The reasoning effort, harness configuration, and Ultra's multi-agent setup differ.&lt;/p&gt;

&lt;p&gt;For a fixed request containing &lt;strong&gt;20,000 input tokens and 2,000 output tokens&lt;/strong&gt;, the listed prices are &lt;strong&gt;USD 0.018 for Gemini 3.7 Flash&lt;/strong&gt;, &lt;strong&gt;USD 0.120 for Claude Opus 5&lt;/strong&gt;, and &lt;strong&gt;USD 0.128 for GPT-5.6 Sol&lt;/strong&gt; at its short-context rate.&lt;/p&gt;

&lt;p&gt;My production metric would be cost per accepted patch, measured against the same repository tasks, tools, tests, retry policy, and sandbox.&lt;/p&gt;

&lt;h2&gt;
  
  
  The Short Version
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Published coding result&lt;/th&gt;
&lt;th&gt;Listed price&lt;/th&gt;
&lt;th&gt;Model ID&lt;/th&gt;
&lt;th&gt;Main integration routes&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;Terminal-Bench 2.1: 89% at Max effort&lt;/td&gt;
&lt;td&gt;$4/M input, $20/M output; $0.120 for the fixed scenario&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-opus-5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Chat Completions, Anthropic-compatible Messages&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;Terminal-Bench 2.1: 88.8%; 91.9% Ultra&lt;/td&gt;
&lt;td&gt;$4/M input and $24/M output up to 272K input tokens; $8/M input and $36/M output above that; $0.128 for the fixed scenario&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gpt-5.6-sol&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Chat Completions, OpenAI Responses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;Terminal-Bench 2.1: 85.8%&lt;/td&gt;
&lt;td&gt;$0.60/M input, $3/M output; $0.018 for the fixed scenario&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemini-3.7-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Chat Completions, Gemini-native generation&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;As of &lt;strong&gt;August 24, 2026&lt;/strong&gt;, these are the listed rates. Prices and supported routes can change, so I would recheck the individual model pages and text API documentation before deployment.&lt;/p&gt;

&lt;p&gt;The fixed-request calculation follows the provider pricing guide:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;20,000 input tokens + 2,000 output tokens

Claude Opus 5:
20,000 * $4/M + 2,000 * $20/M = $0.120

GPT-5.6 Sol:
20,000 * $4/M + 2,000 * $24/M = $0.128

Gemini 3.7 Flash:
20,000 * $0.60/M + 2,000 * $3/M = $0.018
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;That is token cost only. Retries, tool calls, rejected patches, and long-context pricing can change the economics substantially.&lt;/p&gt;

&lt;h2&gt;
  
  
  Model-Specific Constraints
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Claude Opus 5
&lt;/h3&gt;

&lt;p&gt;Claude Opus 5 is available as &lt;code&gt;claude-opus-5&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Input: text, image, and PDF&lt;/li&gt;
&lt;li&gt;Context window: 1M tokens&lt;/li&gt;
&lt;li&gt;Routes: Chat Completions and Anthropic-compatible Messages&lt;/li&gt;
&lt;li&gt;Price: $4/M input and $20/M output&lt;/li&gt;
&lt;li&gt;Fixed scenario: $0.120&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;I would use Messages when the agent depends on Claude-specific controls. Chat Completions is more useful for a portable harness that switches between providers.&lt;/p&gt;

&lt;p&gt;The 89% Terminal-Bench 2.1 result was measured at Max effort. That makes it a strong signal, but the higher output-token price needs to translate into a higher accepted-patch rate.&lt;/p&gt;

&lt;h3&gt;
  
  
  GPT-5.6 Sol
&lt;/h3&gt;

&lt;p&gt;GPT-5.6 Sol uses the model ID &lt;code&gt;gpt-5.6-sol&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Input: text and image&lt;/li&gt;
&lt;li&gt;Routes: Chat Completions and OpenAI Responses&lt;/li&gt;
&lt;li&gt;Price up to 272K input tokens: $4/M input and $24/M output&lt;/li&gt;
&lt;li&gt;Price above 272K input tokens: $8/M input and $36/M output&lt;/li&gt;
&lt;li&gt;Fixed scenario at the short-context rate: $0.128&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Responses is the natural route for OpenAI-native agent features. Chat Completions is the better fit for a common comparison harness.&lt;/p&gt;

&lt;p&gt;The reported Terminal-Bench results need careful interpretation. The single-agent result is &lt;strong&gt;88.8%&lt;/strong&gt;. The &lt;strong&gt;91.9% Ultra&lt;/strong&gt; result uses a multi-agent setup, so I would not compare it directly with a single-agent run.&lt;/p&gt;

&lt;h3&gt;
  
  
  Gemini 3.7 Flash
&lt;/h3&gt;

&lt;p&gt;Gemini 3.7 Flash uses the model ID &lt;code&gt;gemini-3.7-flash&lt;/code&gt;.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Input: text, image, video, audio, and PDF&lt;/li&gt;
&lt;li&gt;Context window: 1,048,576 tokens&lt;/li&gt;
&lt;li&gt;Routes: Chat Completions and Gemini-native generation&lt;/li&gt;
&lt;li&gt;Price: $0.60/M input and $3/M output&lt;/li&gt;
&lt;li&gt;Fixed scenario: $0.018&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Gemini-native generation makes sense when Google-specific features matter. Chat Completions is more convenient for a shared harness.&lt;/p&gt;

&lt;p&gt;The large context window and multimodal input expand the possible test surface, especially for design-to-code workflows. They do not prove that repository navigation or stale-context handling will be good.&lt;/p&gt;

&lt;h2&gt;
  
  
  What the Other Benchmark Says
&lt;/h2&gt;

&lt;p&gt;Terminal-Bench is not enough to characterize a coding agent. The published DeepSWE v1.1 figures provide another signal:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Evidence&lt;/th&gt;
&lt;th&gt;Claude Opus 5&lt;/th&gt;
&lt;th&gt;GPT-5.6 Sol&lt;/th&gt;
&lt;th&gt;Gemini 3.7 Flash&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench 2.1&lt;/td&gt;
&lt;td&gt;89% at Max effort&lt;/td&gt;
&lt;td&gt;88.8%; 91.9% Ultra&lt;/td&gt;
&lt;td&gt;85.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE v1.1&lt;/td&gt;
&lt;td&gt;73.7%&lt;/td&gt;
&lt;td&gt;72.7%&lt;/td&gt;
&lt;td&gt;65.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context window&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;td&gt;1,050,000 tokens&lt;/td&gt;
&lt;td&gt;1,048,576 tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;20K input + 2K output&lt;/td&gt;
&lt;td&gt;USD 0.120&lt;/td&gt;
&lt;td&gt;USD 0.128 at short-context rate&lt;/td&gt;
&lt;td&gt;USD 0.018&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;DeepSWE v1.1 measures long-horizon software engineering in real codebases. I would only compare those numbers when the scaffold and evaluation setup match.&lt;/p&gt;

&lt;p&gt;Context capacity is also not retrieval quality. A model can accept a million-token repository and still fail symbol discovery, cross-file consistency, or stale-context handling.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build a Repository-Level Evaluation
&lt;/h2&gt;

&lt;p&gt;A useful comparison requires more than sending the same prompt three times. Pin the repository, expose the same tools, use the same dependency versions and test commands, and enforce the same executable acceptance checks.&lt;/p&gt;

&lt;p&gt;I use four task classes because each reveals a different failure mode:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;What the agent must do&lt;/th&gt;
&lt;th&gt;Acceptance condition&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Repair a failing CI build&lt;/td&gt;
&lt;td&gt;Read the failure log, trace a dependency or type error through the manifest, source, and tests, then run the affected suite.&lt;/td&gt;
&lt;td&gt;CI is green without disabling checks or introducing regressions.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Complete an SDK migration&lt;/td&gt;
&lt;td&gt;Update imports, configuration, types, and tests across multiple packages while preserving the public interface.&lt;/td&gt;
&lt;td&gt;The full suite passes with no deprecated calls or interface breaks.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Patch a security finding&lt;/td&gt;
&lt;td&gt;Follow the vulnerable call path, make the smallest safe change, add a regression test, and explain the risk boundary.&lt;/td&gt;
&lt;td&gt;The exploit test fails, the regression test passes, and unrelated behavior is unchanged.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Implement a UI from a reference&lt;/td&gt;
&lt;td&gt;Use a screenshot or design-system input, reuse existing components, and update visual or interaction tests.&lt;/td&gt;
&lt;td&gt;Functional tests and visual thresholds pass without adding a duplicated component layer.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Keep these variables fixed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Dependency versions&lt;/li&gt;
&lt;li&gt;Tool schemas&lt;/li&gt;
&lt;li&gt;Token limits&lt;/li&gt;
&lt;li&gt;Retry policy&lt;/li&gt;
&lt;li&gt;Execution sandbox&lt;/li&gt;
&lt;li&gt;Test commands&lt;/li&gt;
&lt;li&gt;Fallback behavior&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Disable fallback during the model comparison. Otherwise, a successful patch may be attributed to the wrong model.&lt;/p&gt;

&lt;p&gt;The primary result should be accepted-task rate. After that, report:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Tool-call success&lt;/li&gt;
&lt;li&gt;Regression count&lt;/li&gt;
&lt;li&gt;p50 and p95 end-to-end latency&lt;/li&gt;
&lt;li&gt;Total token spend&lt;/li&gt;
&lt;li&gt;Cost per accepted patch&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Store the repository commit hash, raw responses, tool traces, usage records, and environment details. Without those artifacts, the result will be difficult to reproduce.&lt;/p&gt;

&lt;h2&gt;
  
  
  Choosing the Primary Model
&lt;/h2&gt;

&lt;p&gt;I would choose &lt;strong&gt;Claude Opus 5&lt;/strong&gt; when the agent needs Messages-native controls or when its Max-effort performance remains strong on multi-file repository tasks. Its published benchmark result is slightly ahead of GPT-5.6 Sol's single-agent result, but the output price needs to be justified by fewer retries or more accepted patches.&lt;/p&gt;

&lt;p&gt;I would choose &lt;strong&gt;GPT-5.6 Sol&lt;/strong&gt; when the agent is built around Responses or when difficult terminal and long-horizon tasks justify the premium. Track the &lt;strong&gt;272K input-token threshold&lt;/strong&gt; in cost estimates, and keep the Ultra result separate from single-agent measurements.&lt;/p&gt;

&lt;p&gt;I would choose &lt;strong&gt;Gemini 3.7 Flash&lt;/strong&gt; when raw token cost, multimodal input, or a &lt;strong&gt;1,048,576-token context window&lt;/strong&gt; matters. It is the cheapest in the fixed request scenario, but rejected patches and retries can remove that advantage quickly.&lt;/p&gt;

&lt;h2&gt;
  
  
  One Harness, Multiple Providers
&lt;/h2&gt;

&lt;p&gt;The common path for these three models can use one OpenAI-compatible base URL, one key, and Chat Completions, switching only the model ID:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://api.cometapi.com/v1

claude-opus-5
gpt-5.6-sol
gemini-3.7-flash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;A unified multi-model API such as CometAPI is useful here when the goal is to keep authentication, request plumbing, retries, and usage accounting in one harness.&lt;/p&gt;

&lt;p&gt;That does not eliminate provider-specific work. Messages, Responses, and Gemini-native generation expose different semantics and features. If production uses one of those native routes, benchmark that exact route rather than assuming the common path is equivalent.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Common route&lt;/th&gt;
&lt;th&gt;Native route&lt;/th&gt;
&lt;th&gt;Integration decision&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5&lt;/td&gt;
&lt;td&gt;Chat Completions&lt;/td&gt;
&lt;td&gt;Anthropic-compatible Messages&lt;/td&gt;
&lt;td&gt;Use Chat for common tests; Messages for Claude-specific behavior.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;Chat Completions&lt;/td&gt;
&lt;td&gt;OpenAI Responses&lt;/td&gt;
&lt;td&gt;Use the endpoint that production orchestration actually calls.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.7 Flash&lt;/td&gt;
&lt;td&gt;Chat Completions&lt;/td&gt;
&lt;td&gt;Gemini-native generation&lt;/td&gt;
&lt;td&gt;Use Chat for common tests; native generation for Google-specific behavior.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;h2&gt;
  
  
  Cost and Fallback Policy
&lt;/h2&gt;

&lt;p&gt;Raw token price is a shortlisting metric. The production metric is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cost per accepted patch =
total spend across attempts, tool calls, retries, and rejected patches
/
tasks that pass the acceptance suite
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;After selecting a primary model, evaluate fallback separately. Retryable failures include rate limits and HTTP &lt;code&gt;500&lt;/code&gt;, &lt;code&gt;503&lt;/code&gt;, &lt;code&gt;504&lt;/code&gt;, and &lt;code&gt;524&lt;/code&gt; responses. Log which model ultimately served every request.&lt;/p&gt;

&lt;p&gt;Invalid requests and authentication failures, including HTTP &lt;code&gt;400&lt;/code&gt; and &lt;code&gt;401&lt;/code&gt;, should be fixed rather than rerouted.&lt;/p&gt;

&lt;h2&gt;
  
  
  Practical Answers
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Claude Opus 5 better than GPT-5.6 Sol?
&lt;/h3&gt;

&lt;p&gt;Their published Terminal-Bench 2.1 results are close: &lt;strong&gt;89% for Claude Opus 5 at Max effort&lt;/strong&gt;, versus &lt;strong&gt;88.8% for GPT-5.6 Sol in the cited single-agent run&lt;/strong&gt; and &lt;strong&gt;91.9% in Ultra's parallel-agent setup&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Use Claude when Messages-native controls matter. Use GPT when Responses-native orchestration matters. For quality, rerun the same repository tasks because the published settings are not identical.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is Gemini 3.7 Flash production-ready for coding agents?
&lt;/h3&gt;

&lt;p&gt;It can be, if it passes the repository's acceptance gate. Its published results are &lt;strong&gt;85.8% on Terminal-Bench 2.1&lt;/strong&gt; and &lt;strong&gt;65.3% on DeepSWE v1.1&lt;/strong&gt;, with the lowest raw token price in this comparison.&lt;/p&gt;

&lt;p&gt;Measure accepted-patch rate, regression count, tool-call validity, latency, and retry rate before deploying it.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which model has the best price-to-performance ratio?
&lt;/h3&gt;

&lt;p&gt;There is no universal answer. Start with Gemini 3.7 Flash if raw token cost is the dominant constraint, then calculate total spend per accepted patch.&lt;/p&gt;

&lt;p&gt;Claude Opus 5 or GPT-5.6 Sol can still be cheaper per successful task if they require fewer retries or produce fewer rejected changes.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which is better for large repositories?
&lt;/h3&gt;

&lt;p&gt;Claude Opus 5 lists a &lt;strong&gt;1M-token&lt;/strong&gt; context window, while Gemini 3.7 Flash lists &lt;strong&gt;1,048,576 tokens&lt;/strong&gt;. GPT-5.6 Sol lists &lt;strong&gt;1,050,000 tokens&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;Those capacities do not establish repository performance. Test symbol discovery, cross-file consistency, stale-context handling, and full-suite pass rate on the same pinned codebase.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is GPT-5.6 Sol more expensive than Gemini 3.7 Flash?
&lt;/h3&gt;

&lt;p&gt;For the fixed &lt;strong&gt;20K input and 2K output&lt;/strong&gt; scenario, GPT-5.6 Sol costs &lt;strong&gt;USD 0.128&lt;/strong&gt; at its short-context rate, compared with &lt;strong&gt;USD 0.018&lt;/strong&gt; for Gemini 3.7 Flash. GPT-5.6 Sol also moves to a higher rate above &lt;strong&gt;272K input tokens&lt;/strong&gt;.&lt;/p&gt;

&lt;p&gt;The relevant comparison remains cost per accepted patch, not token price alone.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can one coding-agent infrastructure switch among all three?
&lt;/h3&gt;

&lt;p&gt;Yes, for the common Chat Completions path. The harness can retain one client and key while changing the model ID. Claude Messages, OpenAI Responses, and Gemini-native features require route-specific request handling.&lt;/p&gt;

&lt;p&gt;That distinction should be reflected in the evaluation: compare the common route for portability, and benchmark native routes separately when they are part of the production design.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/claude-vs-gpt-vs-gemini-for-coding-agents/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=claude-vs-gpt-vs-gemini-for-coding-agents"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
    <item>
      <title>GPT-5.6 vs Claude: I’d Benchmark the Whole Coding Route, Not the Token Price</title>
      <dc:creator>Dylan Foster</dc:creator>
      <pubDate>Fri, 18 Sep 2026 17:07:36 +0000</pubDate>
      <link>https://dev.to/dylanfoster1/gpt-56-vs-claude-id-benchmark-the-whole-coding-route-not-the-token-price-3pj3</link>
      <guid>https://dev.to/dylanfoster1/gpt-56-vs-claude-id-benchmark-the-whole-coding-route-not-the-token-price-3pj3</guid>
      <description>&lt;p&gt;The coding model I want in production is the one that gets an accepted patch through validation at the lowest total cost. That includes failed attempts, escalation, tool execution, and reviewer time. A cheap response that creates another debugging session is not a cheap result.&lt;/p&gt;

&lt;p&gt;GPT-5.6 and Claude both offer enough tiers to make routing worth evaluating. Luna and Haiku cover lightweight work; Terra and Sonnet are general coding candidates; Sol, Opus, and Fable belong in harder-task evaluations. Those are starting points, not equivalent capability classes. The specific model names, benchmark results, and prices below are the source article’s reported figures; verify current provider documentation before using them in a budget or deployment.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start With an Accepted Task
&lt;/h2&gt;

&lt;p&gt;I would track this metric before optimizing token spend:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Cost per successful task =
  (primary model cost + retry cost + fallback cost
   + tool cost + human review cost) / successful tasks
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The denominator matters. Expected API spend per submitted task is not automatically cost per successful task: some tasks can still fail after escalation. Record first-pass success and final success separately, then account for everything spent on unsuccessful attempts too.&lt;/p&gt;

&lt;p&gt;For each task, I would log model and effort level, input and output usage, cache reads and writes, tool calls, retries, escalation, end-to-end latency, and human review minutes. Model-and-effort combinations are the unit of comparison. Higher reasoning effort only earns its additional spend when it improves acceptance or reduces downstream correction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Build a Shortlist, Not a Benchmark Leaderboard
&lt;/h2&gt;

&lt;p&gt;The comparison attributed to &lt;a href="https://openai.com/index/gpt-5-6/" rel="noopener noreferrer"&gt;OpenAI’s GPT-5.6 evaluation&lt;/a&gt; reports these results:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Artificial Analysis Coding Agent Index v1.1&lt;/th&gt;
&lt;th&gt;SWE-Bench Pro&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;80&lt;/td&gt;
&lt;td&gt;64.60%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Terra&lt;/td&gt;
&lt;td&gt;77.4&lt;/td&gt;
&lt;td&gt;63.40%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Luna&lt;/td&gt;
&lt;td&gt;74.6&lt;/td&gt;
&lt;td&gt;62.70%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5&lt;/td&gt;
&lt;td&gt;77.2&lt;/td&gt;
&lt;td&gt;80.00%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 4.8&lt;/td&gt;
&lt;td&gt;72.5&lt;/td&gt;
&lt;td&gt;69.20%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sol leads the Coding Agent Index in this table; Fable leads SWE-Bench Pro. The source also reports different results across DeepSWE and Terminal-Bench 2.1. I would use that disagreement to broaden an eval, not declare an overall winner. Harnesses, available tools, reasoning settings, and execution environments all affect coding-agent performance.&lt;/p&gt;

&lt;p&gt;My initial lightweight candidates would be Luna and Haiku 4.5 for classification, routing, and simple explanations. For repository Q&amp;amp;A, tests, reviews, and scoped fixes, I would compare Terra and Sonnet 5, escalating to Sol or Opus 4.8. Multi-file refactors justify testing Sol or higher-effort Sonnet first, with Opus or Fable as additional candidates. For architecture migrations and security-sensitive changes, stronger models do not remove the need for human review.&lt;/p&gt;

&lt;h2&gt;
  
  
  Put the Rates in One Place
&lt;/h2&gt;

&lt;p&gt;These are the source’s reported prices in dollars per million tokens. GPT-5.6 figures are for Standard short-context requests; long-context, Batch, Flex, and Priority processing have separate rates. Check the applicable service tier in the &lt;a href="https://developers.openai.com/api/docs/pricing" rel="noopener noreferrer"&gt;OpenAI pricing documentation&lt;/a&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model / period&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Cached read&lt;/th&gt;
&lt;th&gt;Cache write&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Luna&lt;/td&gt;
&lt;td&gt;$1.00&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$1.25&lt;/td&gt;
&lt;td&gt;$6.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Terra&lt;/td&gt;
&lt;td&gt;$2.50&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;td&gt;$3.13&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-5.6 Sol&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$6.25&lt;/td&gt;
&lt;td&gt;$30.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Haiku 4.5&lt;/td&gt;
&lt;td&gt;$1.00&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$1.25 / $2.00&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet 5, through Aug. 31, 2026&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$2.50 / $4.00&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Sonnet 5, from Sept. 1, 2026&lt;/td&gt;
&lt;td&gt;$3.00&lt;/td&gt;
&lt;td&gt;$0.30&lt;/td&gt;
&lt;td&gt;$3.75 / $6.00&lt;/td&gt;
&lt;td&gt;$15.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 4.8&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$6.25 / $10.00&lt;/td&gt;
&lt;td&gt;$25.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fable 5&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$1.00&lt;/td&gt;
&lt;td&gt;$12.50 / $20.00&lt;/td&gt;
&lt;td&gt;$50.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Claude’s two write prices correspond to five-minute and one-hour caches. Sonnet’s introductory $2 input / $10 output pricing runs through August 31, 2026; the reported standard $3 / $15 rates begin September 1.&lt;/p&gt;

&lt;p&gt;On headline rates, introductory Sonnet is cheaper than Terra. Haiku matches Luna’s input price and charges $5 rather than $6 for output. Opus matches Sol’s $5 input price and charges $25 rather than $30 for output. After Sonnet’s pricing transition, Terra has cheaper input at $2.50 versus $3, with both charging $15 for output. None of those comparisons includes reliability or actual token consumption.&lt;/p&gt;

&lt;h3&gt;
  
  
  Token Counts Are Part of the Price
&lt;/h3&gt;

&lt;p&gt;I would not apply one provider’s token estimate to another provider’s bill. The source attributes to Anthropic a newer tokenizer in Sonnet 5, Fable 5, and newer Opus models that can produce approximately 30% more tokens for the same text, depending on workload. Log returned usage rather than assuming identical token counts.&lt;/p&gt;

&lt;p&gt;There is also a model-selection trap: the source states that the generic &lt;code&gt;gpt-5.6&lt;/code&gt; alias maps to Sol. Where Terra or Luna passes the eval, select that tier explicitly instead of allowing an alias to determine flagship usage.&lt;/p&gt;

&lt;h2&gt;
  
  
  Work Through the Fallback Arithmetic
&lt;/h2&gt;

&lt;p&gt;Assume each attempt consumes 80,000 input tokens and 10,000 output tokens, without caching. There is one primary attempt and a stronger fallback when it fails. This is a pricing illustration, not measured model performance; tokenization, tool use, effort, and actual success rates can change the result.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Route&lt;/th&gt;
&lt;th&gt;Primary attempt&lt;/th&gt;
&lt;th&gt;Fallback attempt&lt;/th&gt;
&lt;th&gt;Expected API spend with 25% fallback&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Terra → Sol&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.08 × $2.50 + 0.01 × $15 = $0.35&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.08 × $5 + 0.01 × $30 = $0.70&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;$0.35 + 0.25 × $0.70 = $0.525&lt;/code&gt;, about $0.53&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Introductory Sonnet → Opus&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.08 × $2 + 0.01 × $10 = $0.26&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.08 × $5 + 0.01 × $25 = $0.65&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;$0.26 + 0.25 × $0.65 = $0.4225&lt;/code&gt;, about $0.42&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard Sonnet → Opus&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.08 × $3 + 0.01 × $15 = $0.39&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;$0.65&lt;/td&gt;
&lt;td&gt;&lt;code&gt;$0.39 + 0.25 × $0.65 = $0.5525&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;At equal fallback rates, introductory Sonnet wins this example; Terra becomes slightly cheaper after the reported September 1 change. But reduce Terra’s fallback rate to 10% and its expected spend becomes &lt;code&gt;$0.35 + 0.10 × $0.70 = $0.42&lt;/code&gt;. That is slightly below introductory Sonnet’s unrounded $0.4225 and below standard Sonnet’s $0.5525. Rounding both introductory results to cents hides the first difference.&lt;/p&gt;

&lt;p&gt;I would not call any row a production winner yet. These figures exclude tools and review, assume identical usage on primary and fallback attempts, and do not establish whether the fallback succeeds.&lt;/p&gt;

&lt;h2&gt;
  
  
  Measure Cache Reuse Alongside Escalation
&lt;/h2&gt;

&lt;p&gt;The source describes OpenAI matching prompt prefixes through implicit caching, with GPT-5.6 additionally supporting explicit breakpoints and &lt;code&gt;prompt_cache_key&lt;/code&gt;. It gives cache writes a 1.25× normal-input rate and reads the discounted cached-input rate.&lt;/p&gt;

&lt;p&gt;Claude caching is enabled through &lt;code&gt;cache_control&lt;/code&gt;, using a request-level automatic breakpoint or explicit content-block breakpoints. Its default lifetime is five minutes, with an optional, more expensive one-hour write. Reads cost 0.1× the base input rate. These implementation differences matter when repository instructions, tool definitions, coding standards, or project context recur across calls.&lt;/p&gt;

&lt;p&gt;For the Terra example, serving 40,000 of the 80,000 input tokens from cache changes a subsequent request to &lt;code&gt;$0.10&lt;/code&gt; regular input + &lt;code&gt;$0.01&lt;/code&gt; cached input + &lt;code&gt;$0.15&lt;/code&gt; output = &lt;strong&gt;$0.26&lt;/strong&gt;, down from $0.35. Writing that 40,000-token prefix instead gives &lt;strong&gt;$0.375&lt;/strong&gt; for the request using the stated 1.25× multiplier. The displayed $3.13/MTok write price is rounded; the multiplier corresponds to $3.125/MTok.&lt;/p&gt;

&lt;p&gt;Caching pays through reuse, not necessarily on the first request. I would measure writes, reads, retries, and fallback together. A lower input bill cannot rescue a route that repeatedly produces unusable patches, and context-window capacity alone says little about effective task cost.&lt;/p&gt;

&lt;h2&gt;
  
  
  Run a Small Repository Eval Before Building a Clever Router
&lt;/h2&gt;

&lt;p&gt;Around 30 representative tasks is a practical first pass: 10 bug fixes, 10 implementation or test-generation tasks, 5 refactors, and 5 reviews. Compare Terra, Sol, Sonnet 5, and Opus 4.8 where they fit the workload. Add Luna and Haiku for lightweight subtasks, and Fable as a higher-capability reference for difficult work.&lt;/p&gt;

&lt;p&gt;Keep acceptance criteria identical: passing tests, successful builds, lint and type checks, resolution of the requested issue, and the amount of human correction required. Parser or AST checks, &lt;code&gt;pytest&lt;/code&gt;, &lt;code&gt;npm test&lt;/code&gt;, and isolated patch execution provide useful automatic validation. They make cheaper-first routing testable; they do not prove every change correct.&lt;/p&gt;

&lt;p&gt;Segment results by task class. A model that is economical for reviews may be a poor default for bug fixes. Track first-pass and final success, total API spend, retries, fallback frequency, cache-hit rate, latency, and review time. Otherwise, an aggregate average can conceal exactly the workload distinction the router needs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Keep the First Router Rule-Based
&lt;/h2&gt;

&lt;p&gt;My starting policy would be: classify the task, choose the lowest-cost route that passes the relevant eval, validate, then escalate on failure. A candidate sequence is Luna or Haiku → Terra or Sonnet → Sol or Opus → Fable or human review. That is an evaluation hypothesis, not a requirement to traverse every tier.&lt;/p&gt;

&lt;p&gt;High fallback rates suggest strengthening the initial route. Premium calls that rarely improve acceptance suggest reducing escalation. Extra effort without better outcomes suggests lowering effort. Repeated context dominating spend suggests improving cache reuse. For security-sensitive or architectural work, I would require human review rather than treating automatic checks as sufficient.&lt;/p&gt;

&lt;p&gt;A unified interface can reduce integration work when comparing providers: CometAPI exposes supported models through an OpenAI-compatible Chat Completions interface, allowing model selection through the &lt;code&gt;model&lt;/code&gt; parameter. That simplifies switching, but it does not replace provider-specific usage accounting or quality evaluation.&lt;/p&gt;

&lt;p&gt;The deployment decision I care about is specific: which route delivers accepted changes for this task class, at an acceptable latency, with the lowest combined API, execution, and review cost? Public benchmarks identify candidates. Production telemetry decides which ones stay.&lt;/p&gt;




&lt;p&gt;&lt;em&gt;Originally published at &lt;a href="https://www.cometapi.com/gpt-5-6-vs-claude-api-coding-cost-per-task/?utm_source=dev.to&amp;amp;utm_medium=social&amp;amp;utm_campaign=content&amp;amp;utm_content=gpt-5-6-vs-claude-api-coding-cost-per-task"&gt;cometapi.com&lt;/a&gt;&lt;/em&gt;&lt;/p&gt;

</description>
      <category>ai</category>
      <category>agents</category>
    </item>
  </channel>
</rss>
