<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>DEV Community: Hassann</title>
    <description>The latest articles on DEV Community by Hassann (@hassann).</description>
    <link>https://dev.to/hassann</link>
    <image>
      <url>https://media2.dev.to/dynamic/image/width=90,height=90,fit=cover,gravity=auto,format=auto/https:%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Fuser%2Fprofile_image%2F3890506%2F89a141f2-4995-48b3-b5f2-e00ba5055afb.png</url>
      <title>DEV Community: Hassann</title>
      <link>https://dev.to/hassann</link>
    </image>
    <atom:link rel="self" type="application/rss+xml" href="https://dev.to/feed/hassann"/>
    <language>en</language>
    <item>
      <title>Gemini 4 Argon Benchmarks: All 19 Rows, How Google Ran Them, and the 5 It Loses</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:28:44 +0000</pubDate>
      <link>https://dev.to/hassann/gemini-4-argon-benchmarks-all-19-rows-how-google-ran-them-and-the-5-it-loses-596d</link>
      <guid>https://dev.to/hassann/gemini-4-argon-benchmarks-all-19-rows-how-google-ran-them-and-the-5-it-loses-596d</guid>
      <description>&lt;p&gt;Gemini 4 Argon leads 13 of the 19 rows in Google’s launch table outright, ties one row (CWE-bench v1, with GPT-6 Astra), and trails on five: FrontierSWE v2, Terminal-bench 4.0, PostTrainBench, Terminal-Bench Science 0.1, and OSWorld-2.0. The access limitation matters: Argon is currently available only to Fairwind Program defenders, so developers outside Google’s partner list and a few benchmark organizations cannot rerun these numbers yet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;Below is the full table, including who measured each row, where Argon loses, methodology caveats, cyber scores, third-party scoreboards, and a practical way to prepare your own evaluation in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; before access opens. For the model overview, start with &lt;a href="http://apidog.com/blog/what-is-gemini-4-argon?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what Gemini 4 Argon is&lt;/a&gt;. For a buying comparison, see &lt;a href="http://apidog.com/blog/gemini-4-argon-vs-gpt-6-astra-vs-claude-opus-5-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon vs GPT-6 Astra vs Claude Opus 5.5&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The full table, with who measured each row
&lt;/h2&gt;

&lt;p&gt;These are Google-reported results from the &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/" rel="noopener noreferrer"&gt;launch post&lt;/a&gt; and the &lt;a href="https://deepmind.google/models/evals-methodology/gemini-4-argon" rel="noopener noreferrer"&gt;evals methodology PDF&lt;/a&gt;, which is headed “results as of October, 2026.”&lt;/p&gt;

&lt;p&gt;Google computed 10 of the 19 Argon scores itself. The other nine come from public leaderboards. Rival scores come from leaderboards, Google runs, vendor system cards, or vendor blog posts, depending on the row. Harnesses are not uniform across every benchmark.&lt;/p&gt;

&lt;p&gt;Bold indicates the row leader.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Gemini 4 Argon&lt;/th&gt;
&lt;th&gt;GPT-6 Astra&lt;/th&gt;
&lt;th&gt;Claude Fable 5.1&lt;/th&gt;
&lt;th&gt;Claude Opus 5.5&lt;/th&gt;
&lt;th&gt;Who measured it&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Vals Index&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;63.1%&lt;/td&gt;
&lt;td&gt;65.8%&lt;/td&gt;
&lt;td&gt;67.0%&lt;/td&gt;
&lt;td&gt;Vals AI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;AutomationBench&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;51.3%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;41.4%&lt;/td&gt;
&lt;td&gt;31.4%&lt;/td&gt;
&lt;td&gt;42.5%&lt;/td&gt;
&lt;td&gt;Zapier leaderboard (private set)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vals Finance Agent v2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;65.4%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;53.5%&lt;/td&gt;
&lt;td&gt;58.9%&lt;/td&gt;
&lt;td&gt;58.6%&lt;/td&gt;
&lt;td&gt;Vals AI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Harvey Legal Agent&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;19.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5.4%&lt;/td&gt;
&lt;td&gt;6.7%&lt;/td&gt;
&lt;td&gt;3.8%&lt;/td&gt;
&lt;td&gt;Vals AI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE v1.1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;77.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;74.1%&lt;/td&gt;
&lt;td&gt;67.4%&lt;/td&gt;
&lt;td&gt;74.2%&lt;/td&gt;
&lt;td&gt;Argon self-computed; Astra leaderboard; Anthropic system cards&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FrontierSWE v2&lt;/td&gt;
&lt;td&gt;55.0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;65.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;56.3%&lt;/td&gt;
&lt;td&gt;62.3%&lt;/td&gt;
&lt;td&gt;Proximal leaderboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vibe Code Bench&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;91.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;89.6%&lt;/td&gt;
&lt;td&gt;90.3%&lt;/td&gt;
&lt;td&gt;90.3%&lt;/td&gt;
&lt;td&gt;Vals AI&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-bench 4.0&lt;/td&gt;
&lt;td&gt;57.4%&lt;/td&gt;
&lt;td&gt;58.2%&lt;/td&gt;
&lt;td&gt;57.9%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;66.4%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Argon self-computed; rivals leaderboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PostTrainBench&lt;/td&gt;
&lt;td&gt;45.3%&lt;/td&gt;
&lt;td&gt;44.3%&lt;/td&gt;
&lt;td&gt;40.2%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;49.3%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Self-computed for all models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-Bench Science 0.1&lt;/td&gt;
&lt;td&gt;57.6%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.1%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;52.6%&lt;/td&gt;
&lt;td&gt;63.3%&lt;/td&gt;
&lt;td&gt;Argon self-computed; rivals leaderboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LABBench 2&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;88.8%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;85.4%&lt;/td&gt;
&lt;td&gt;68.6%&lt;/td&gt;
&lt;td&gt;73.1%&lt;/td&gt;
&lt;td&gt;Self-computed for all models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;RiemannBench&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;76.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;72.0%&lt;/td&gt;
&lt;td&gt;65.6%&lt;/td&gt;
&lt;td&gt;69.6%&lt;/td&gt;
&lt;td&gt;Surge leaderboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GraphWalks up to 128K&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;99.7%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;98.7%&lt;/td&gt;
&lt;td&gt;91.4%&lt;/td&gt;
&lt;td&gt;90.6%&lt;/td&gt;
&lt;td&gt;Self-computed for all models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GraphWalks 256K to 1M&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;84.2%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;71.8%&lt;/td&gt;
&lt;td&gt;65.0%&lt;/td&gt;
&lt;td&gt;66.8%&lt;/td&gt;
&lt;td&gt;Self-computed for all models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agent’s Last Exam&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;39.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;34.2%&lt;/td&gt;
&lt;td&gt;n/r&lt;/td&gt;
&lt;td&gt;38.2%&lt;/td&gt;
&lt;td&gt;Argon self-computed; rivals leaderboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;OSWorld-2.0 (offline)&lt;/td&gt;
&lt;td&gt;69.2%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;72.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;n/r&lt;/td&gt;
&lt;td&gt;n/r&lt;/td&gt;
&lt;td&gt;Argon self-computed; Astra from OpenAI’s blog post&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Chartography&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;71.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;71.0%&lt;/td&gt;
&lt;td&gt;46.2%&lt;/td&gt;
&lt;td&gt;66.3%&lt;/td&gt;
&lt;td&gt;Surge leaderboard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;LVBench&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;91.7%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;87.5%&lt;/td&gt;
&lt;td&gt;79.7%&lt;/td&gt;
&lt;td&gt;83.7%&lt;/td&gt;
&lt;td&gt;Self-computed for all models&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CWE-bench v1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;58.0%&lt;/td&gt;
&lt;td&gt;67.0%&lt;/td&gt;
&lt;td&gt;Public leaderboard&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;&lt;code&gt;n/r&lt;/code&gt; means not reported.&lt;/p&gt;

&lt;p&gt;Grok 4.7 is not in Google’s table, but it also scores 68% on the CWE-bench leaderboard. That makes CWE-bench a three-way tie.&lt;/p&gt;

&lt;h3&gt;
  
  
  Read the source distribution before using the headline count
&lt;/h3&gt;

&lt;p&gt;The source mix is important when you use these results to select a model:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Nine rows&lt;/strong&gt; come from public leaderboards for every model: Vals AI, Zapier, Proximal, Surge, and CWE-bench.

&lt;ul&gt;
&lt;li&gt;Argon leads seven.&lt;/li&gt;
&lt;li&gt;Argon ties one.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Five rows&lt;/strong&gt; were run by Google for all four models.

&lt;ul&gt;
&lt;li&gt;Argon leads four.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Five rows&lt;/strong&gt; combine a Google-run Argon score with rival scores from another source.

&lt;ul&gt;
&lt;li&gt;Three of Argon’s five losses occur in these mixed-source rows.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;This does not eliminate methodology concerns, but it is relevant context: if Google self-computing systematically favored Argon, the losses would be expected to cluster in Google-run rows rather than mixed rows.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Argon trails, and why it matters for agentic coding
&lt;/h2&gt;

&lt;p&gt;Argon’s five losses are concentrated in benchmarks that matter for terminal-based and repository-scale agents:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;FrontierSWE v2:&lt;/strong&gt; 55.0% versus Astra’s 65.5%. Argon is last among the four models, behind Fable 5.1 at 56.3% and Opus 5.5 at 62.3%. This is a public leaderboard result.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Terminal-bench 4.0:&lt;/strong&gt; 57.4% versus Opus 5.5’s 66.4%. Argon is last again, although Astra and Fable 5.1 are within one point.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;PostTrainBench:&lt;/strong&gt; 45.3% versus Opus 5.5’s 49.3%. Argon finishes second in a row Google ran for all models.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Terminal-Bench Science 0.1:&lt;/strong&gt; 57.6% versus Astra’s 68.1%. Argon is third, ahead of Fable 5.1 only. Google ran Argon’s attempt with a 6x verifier timeout.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OSWorld-2.0 (offline subset):&lt;/strong&gt; 69.2% versus Astra’s 72.6%. Anthropic reports only a combined online and offline score, so neither Claude model appears.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For agentic coding specifically, the results split two to two:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark type&lt;/th&gt;
&lt;th&gt;Result&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE v1.1&lt;/td&gt;
&lt;td&gt;Argon wins&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Vibe Code Bench&lt;/td&gt;
&lt;td&gt;Argon wins&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FrontierSWE v2&lt;/td&gt;
&lt;td&gt;Argon finishes last&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-bench 4.0&lt;/td&gt;
&lt;td&gt;Argon finishes last&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The DeepSWE result mixes harnesses: Argon used a mini-swe agent harness, Astra’s score comes from the public leaderboard, and Claude scores come from system cards. Vibe Code Bench is cleaner because Vals AI scored all four models, but Argon’s winning margin is only 1.6 points.&lt;/p&gt;

&lt;p&gt;If your application depends on terminal automation, long-running repository tasks, or shell-based debugging, prioritize FrontierSWE v2 and Terminal-bench 4.0 over the aggregate win count. Astra holds the FrontierSWE lead; see the &lt;a href="http://apidog.com/blog/gpt-6-astra-hands-on?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GPT-6 Astra hands-on&lt;/a&gt; for current usage observations. For Anthropic’s results, see the &lt;a href="http://apidog.com/blog/claude-fable-5-1-benchmarks?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Fable 5.1 benchmarks breakdown&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Bloomberg reports that anonymous Google employees with access say Argon underwhelms on some coding and front-end design work relative to its benchmark scores. Google called that characterization inaccurate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Methodology caveats that change how you read the table
&lt;/h2&gt;

&lt;p&gt;Before treating a row as a production forecast, check the evaluation configuration.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Maximum effort settings:&lt;/strong&gt; Argon ran “with the Gemini API with the highest thinking settings.” Google used the “maximum thinking/reasoning settings available” for Astra, Fable 5.1, and Opus 5.5. These results compare ceilings, not default latency, default quality, or cost.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;LVBench input differences:&lt;/strong&gt; Google ran LVBench for all four models without tools. Gemini sampled video at 1 FPS. Astra received 800 frames, Fable 5.1 received 300, and Opus 5.5 received 600 due to API limitations. The models did not receive identical inputs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;OSWorld best-of-three scoring:&lt;/strong&gt; Argon’s score was “maxed over 3 runs with a single attempt per run.” This is the best run, not an average across runs.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Agent’s Last Exam execution window:&lt;/strong&gt; Argon ran on the ALE-Claw harness with a five-hour window.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Safety filters remained enabled:&lt;/strong&gt; Agent’s Last Exam and OSWorld used safety filters. Flagged responses returned as empty strings. This could reduce Argon’s result rather than inflate it.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  The cyber rows from DeepMind’s cyber page
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://deepmind.google/models/gemini/cyber/" rel="noopener noreferrer"&gt;DeepMind cyber page&lt;/a&gt; adds four scores beyond the launch table:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cyber eval&lt;/th&gt;
&lt;th&gt;Gemini 4 Argon&lt;/th&gt;
&lt;th&gt;Comparison&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Real-world vulnerability discovery&lt;/td&gt;
&lt;td&gt;85.8%&lt;/td&gt;
&lt;td&gt;Gemini 3.8 Flash Cyber: 71.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wiz Penetration Test Benchmark&lt;/td&gt;
&lt;td&gt;70.9%&lt;/td&gt;
&lt;td&gt;Gemini 3.8 Flash Cyber: 58.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gray Swan IPI attack success (&lt;code&gt;k=15&lt;/code&gt;, lower is better)&lt;/td&gt;
&lt;td&gt;0.7%&lt;/td&gt;
&lt;td&gt;Lowest on the chart; Kimi K3 highest at 52.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;CWE-bench v1&lt;/td&gt;
&lt;td&gt;68%&lt;/td&gt;
&lt;td&gt;Three-way tie with Grok 4.7 and GPT-6 Astra; Opus 5.5 at 67%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The vulnerability evaluation used an internal Antigravity harness that Google describes as “not cyber specialized,” with source access.&lt;/p&gt;

&lt;p&gt;The Wiz test gives the model access only to the running website and its public behavior, without application source code.&lt;/p&gt;

&lt;p&gt;Both headline comparisons are against Google’s prior cyber model rather than direct competitors. The &lt;a href="http://apidog.com/blog/gemini-4-argon-cyber-defense?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon cyber defense explainer&lt;/a&gt; covers what these scores mean for APIs you operate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Third-party scoreboards
&lt;/h2&gt;

&lt;p&gt;Several organizations published scores within minutes of launch, which indicates pre-release access.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Artificial Analysis:&lt;/strong&gt; Lists Gemini 4 Argon (High) at an Intelligence Index of 53, ranked #8 of 223. The ranking counts each reasoning setting separately. By distinct model, Argon ties Fable 5.1 and GPT-6 Astra, while trailing Opus 5.5 at 58 maximum and Sonnet 5.5 at 56 maximum. Artificial Analysis reports a 15% hallucination rate on AA-Omniscience, compared with Astra at maximum effort at 51%, with accuracy of 50% versus 63%. Running the index cost $1.99 per task, with approximately 62K output tokens per task versus approximately 27K for Astra at maximum settings. &lt;a href="https://artificialanalysis.ai/models/gemini-4-argon" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt; shows no speed or latency data.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Vals AI:&lt;/strong&gt; Places Argon at 68.90% on the Vals Index, #1 of 41 and the first Gemini model to top the index, at $15.68 per test. &lt;a href="https://www.vals.ai/models/google_gemini-4-argon" rel="noopener noreferrer"&gt;Vals AI&lt;/a&gt; lists a 262K maximum output for the tested configuration, below Google’s stated 1M.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Arena:&lt;/strong&gt; Ranks Argon #1 on &lt;a href="https://arena.ai/leaderboard/text" rel="noopener noreferrer"&gt;Text&lt;/a&gt; at 1525, marked Preliminary with 4,942 votes, and #8 on WebDev.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Why outlets publish 12, 13, and 14 wins
&lt;/h2&gt;

&lt;p&gt;Different headlines use different comparison rules. Here is the recount from Google’s 19 rows:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Counted against&lt;/th&gt;
&lt;th&gt;Argon wins&lt;/th&gt;
&lt;th&gt;Ties&lt;/th&gt;
&lt;th&gt;Argon loses&lt;/th&gt;
&lt;th&gt;Not reported&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Best of all three rivals&lt;/td&gt;
&lt;td&gt;13&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;5&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra only&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5.5 only&lt;/td&gt;
&lt;td&gt;14&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;4&lt;/td&gt;
&lt;td&gt;1&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5.1 only&lt;/td&gt;
&lt;td&gt;15&lt;/td&gt;
&lt;td&gt;0&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;td&gt;2&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Use the count that matches your decision:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;13 wins&lt;/strong&gt; is correct when Argon must beat the entire field.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;14 wins&lt;/strong&gt; is correct in a head-to-head comparison with Astra or Opus 5.5.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;12 wins&lt;/strong&gt; does not match these full-table framings, so inspect which rows the source excluded or counted differently.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Also check the margins. Harvey Legal Agent is a wide gap at 19.6% versus 6.7%, while Chartography is separated by only 0.6 points.&lt;/p&gt;

&lt;h2&gt;
  
  
  How to run your own eval the day access opens
&lt;/h2&gt;

&lt;p&gt;Public benchmarks measure someone else’s harness. Your evaluation should measure your prompts, tools, schemas, latency requirements, and cost limits.&lt;/p&gt;

&lt;p&gt;Build the scenario now using a model you can call. When Argon becomes available, change one environment variable and rerun the same tests.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Select production-like test cases
&lt;/h3&gt;

&lt;p&gt;Create prompts that map to the work you actually need:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A repository bug fix with expected tests or a patch.&lt;/li&gt;
&lt;li&gt;A terminal task requiring command sequencing.&lt;/li&gt;
&lt;li&gt;A long-document question with a verifiable answer.&lt;/li&gt;
&lt;li&gt;A structured API response that must match a JSON schema.&lt;/li&gt;
&lt;li&gt;A tool-calling workflow with error handling.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use real sanitized incidents when possible. Include both successful and failure cases.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Create environment variables in Apidog
&lt;/h3&gt;

&lt;p&gt;In &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;, create an environment with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GEMINI_API_KEY=&amp;lt;your-key&amp;gt;
GEMINI_MODEL=gemini-3.8-flash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Google has not published Argon’s model ID, so &lt;code&gt;GEMINI_MODEL&lt;/code&gt; should be the only value you need to change later.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Parameterize each request
&lt;/h3&gt;

&lt;p&gt;Save each prompt as a request and reference the model variable rather than hardcoding the model ID.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"{{GEMINI_MODEL}}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"parts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Fix the failing test. Return the patch and a short explanation."&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Group related requests into a scenario, such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;agentic-coding/
  repository-fix
  terminal-debugging
  code-review
  long-context-question
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4. Add assertions for quality and cost controls
&lt;/h3&gt;

&lt;p&gt;Do not evaluate only HTTP status. Assert the properties your integration requires:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Response status is successful.&lt;/li&gt;
&lt;li&gt;Required JSON fields exist.&lt;/li&gt;
&lt;li&gt;The output includes expected terms, files, commands, or citations.&lt;/li&gt;
&lt;li&gt;The response is valid JSON when your application expects JSON.&lt;/li&gt;
&lt;li&gt;Token usage remains below your cost ceiling.&lt;/li&gt;
&lt;li&gt;Tool calls, if used, match expected names and arguments.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, track &lt;code&gt;usageMetadata&lt;/code&gt; and cap thought-token usage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;returns a successful response&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;have&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;status&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usageMetadata&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;have&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;property&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;thoughtsTokenCount&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;usageMetadata&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;thoughtsTokenCount&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;be&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;below&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;50000&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;have&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;property&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;candidates&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Size the &lt;code&gt;thoughtsTokenCount&lt;/code&gt; ceiling to your own budget and latency target rather than treating maximum reasoning as the default.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Capture a baseline before Argon is public
&lt;/h3&gt;

&lt;p&gt;Run the scenario with Gemini 3.8 Flash and save the report.&lt;/p&gt;

&lt;p&gt;Record at least:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Metric&lt;/th&gt;
&lt;th&gt;Why it matters&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Pass rate&lt;/td&gt;
&lt;td&gt;Measures task success against your assertions&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output validity&lt;/td&gt;
&lt;td&gt;Detects schema and parsing failures&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Token usage&lt;/td&gt;
&lt;td&gt;Tracks model cost behavior&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Latency&lt;/td&gt;
&lt;td&gt;Shows whether a higher-quality result is usable in your workflow&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Failure examples&lt;/td&gt;
&lt;td&gt;Helps diagnose regressions after a model swap&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;When Argon’s model ID is released, update:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GEMINI_MODEL=&amp;lt;Argon model ID&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then rerun the exact same scenario and compare reports.&lt;/p&gt;

&lt;p&gt;Google says that “all new models” will launch on the Interactions API, so save an Interactions API version of each request as well. For expected price and limit differences, see &lt;a href="http://apidog.com/blog/gemini-4-argon-vs-gemini-3-8-flash?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon vs Gemini 3.8 Flash&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Are Gemini 4 Argon’s benchmarks independently verified?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Partly. Nine of the 19 rows come from public leaderboards for every model, and Artificial Analysis, Vals AI, and Arena posted their own results. The remaining rows are Google-run for Argon.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is Gemini 4 Argon’s DeepSWE score?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Argon scores 77.9% on DeepSWE v1.1, ahead of Opus 5.5 at 74.2% and Astra at 74.1%. Argon used a mini-swe agent harness, while rival scores come from a leaderboard and system cards.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Argon beat GPT-6 Astra on benchmarks?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Head to head in Google’s table, Argon wins 14 rows, ties one, and loses four. On Artificial Analysis, they tie at 53.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I rerun these benchmarks myself?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not yet. Argon is limited to Fairwind partners, and Google has not provided a public release date. The &lt;a href="http://apidog.com/blog/gemini-4-argon-release-date?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon release date tracker&lt;/a&gt; follows the rollout.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next step
&lt;/h2&gt;

&lt;p&gt;Build the scenario now, run it on Gemini 3.8 Flash, and keep the report. When Argon becomes available, swap the model variable and you will have comparable numbers from your own workload within minutes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; to set up the evaluation.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Gemini 4 Argon vs Gemini 3.8 Flash: Should You Wait for Argon or Ship on Flash Now?</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:26:54 +0000</pubDate>
      <link>https://dev.to/hassann/gemini-4-argon-vs-gemini-38-flash-should-you-wait-for-argon-or-ship-on-flash-now-39in</link>
      <guid>https://dev.to/hassann/gemini-4-argon-vs-gemini-38-flash-should-you-wait-for-argon-or-ship-on-flash-now-39in</guid>
      <description>&lt;p&gt;Gemini 3.8 Flash is available today: a stable model with a free tier, priced at $0.75 input and $3.75 output per 1M tokens through 2026-12-31, then $1.50 and $7.50, with a 64K output cap. Gemini 4 Argon is not generally available. Google’s &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/" rel="noopener noreferrer"&gt;launch post&lt;/a&gt; lists Argon at $2 input and $10 output during an intro period of unstated length, then $4 and $20. Google states a 1M-token output limit, but availability is currently limited to Fairwind Program defenders. If you need to ship this quarter, build on Flash and keep Argon behind a configuration switch.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This guide compares the models’ published specs, shared benchmark rows, cyber scores, and the cost of an identical call. It also shows how to make the eventual model switch a single environment-variable change. For background, see &lt;a href="http://apidog.com/blog/what-is-gemini-4-argon?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what Gemini 4 Argon is&lt;/a&gt; and &lt;a href="http://apidog.com/blog/what-is-gemini-3-8-flash?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what Gemini 3.8 Flash is&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Side by side: what you can call today
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Gemini 3.8 Flash&lt;/th&gt;
&lt;th&gt;Gemini 4 Argon&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Availability&lt;/td&gt;
&lt;td&gt;Stable on the Gemini API, AI Studio, Antigravity, and Gemini Enterprise&lt;/td&gt;
&lt;td&gt;A subset of Fairwind Program partners, as a managed model on Gemini Enterprise&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model ID&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gemini-3.8-flash&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Price per 1M tokens (input / output)&lt;/td&gt;
&lt;td&gt;$0.75 / $3.75 through 2026-12-31, then $1.50 / $7.50&lt;/td&gt;
&lt;td&gt;$2 / $10 intro pricing, then $4 / $20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cached input per 1M&lt;/td&gt;
&lt;td&gt;$0.075, then $0.15&lt;/td&gt;
&lt;td&gt;$0.10 intro, then $0.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output limit&lt;/td&gt;
&lt;td&gt;65,536 tokens&lt;/td&gt;
&lt;td&gt;1M tokens, according to Google&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input window&lt;/td&gt;
&lt;td&gt;1,048,576 tokens&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thinking levels&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;low&lt;/code&gt;, &lt;code&gt;medium&lt;/code&gt; (default), &lt;code&gt;high&lt;/code&gt;; &lt;code&gt;minimal&lt;/code&gt; returns HTTP 400&lt;/td&gt;
&lt;td&gt;Not documented&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Free API tier&lt;/td&gt;
&lt;td&gt;Yes, rate-limited&lt;/td&gt;
&lt;td&gt;None announced&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Consumer access&lt;/td&gt;
&lt;td&gt;Gemini app on AI Pro and Ultra; AI Mode in Search&lt;/td&gt;
&lt;td&gt;Not yet; AI Ultra subscribers are in the first wave, with no date announced&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A few details need qualification:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Google has not published Argon’s input window, although its long-context evaluation used prompts up to 1M tokens.&lt;/li&gt;
&lt;li&gt;Google states a 1M-token Argon output limit. Third-party evaluator Vals AI lists a 262K maximum output for the configuration it tested.&lt;/li&gt;
&lt;li&gt;Argon evaluations used “the highest thinking settings,” which suggests a high setting exists, but Google has not documented the level names or default.&lt;/li&gt;
&lt;li&gt;Flash thinking levels are documented in Google’s &lt;a href="https://ai.google.dev/gemini-api/docs/thinking" rel="noopener noreferrer"&gt;thinking documentation&lt;/a&gt; and in this &lt;a href="http://apidog.com/blog/gemini-3-8-flash-thinking-levels?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.8 Flash thinking-levels guide&lt;/a&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Flash’s free tier is rate-limited, and Google says free-tier data is “used to improve our products.” Google has not announced a free way to use Argon. See the &lt;a href="http://apidog.com/blog/gemini-4-argon-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon free-access explainer&lt;/a&gt; for currently available alternatives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Shared benchmark rows
&lt;/h2&gt;

&lt;p&gt;Google’s launch tables share two benchmark rows. These are Google-reported results, and Argon’s methodology attributes both rows to Vals AI.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Gemini 3.8 Flash&lt;/th&gt;
&lt;th&gt;Gemini 4 Argon&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Vals Finance Agent v2&lt;/td&gt;
&lt;td&gt;61.4%&lt;/td&gt;
&lt;td&gt;65.4%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Harvey Legal Agent Benchmark&lt;/td&gt;
&lt;td&gt;10.0%&lt;/td&gt;
&lt;td&gt;19.6%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Treat these results as directional, not as a controlled head-to-head test: Google published the tables a month apart and compared them against different model sets.&lt;/p&gt;

&lt;p&gt;Still, the published gap is notable:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Finance:&lt;/strong&gt; Argon leads by 4.0 percentage points.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Legal:&lt;/strong&gt; Argon scores 19.6% versus Flash’s 10.0%, nearly doubling the result.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Other rows do not have a direct Flash counterpart in the launch-post text. For example, Flash reported HLE-Verified while Argon did not, and Flash’s DeepSWE score appeared only in model-card images. On Arena’s Text leaderboard, a third-party ranking, Argon (High) is #1 on preliminary votes while Gemini 3.8 Flash (High) is #11.&lt;/p&gt;

&lt;p&gt;For the full table and measurement sources, see this &lt;a href="http://apidog.com/blog/gemini-4-argon-benchmarks?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon benchmarks breakdown&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cyber: Argon vs. 3.8 Flash Cyber
&lt;/h2&gt;

&lt;p&gt;The &lt;a href="https://deepmind.google/models/gemini/cyber/" rel="noopener noreferrer"&gt;DeepMind cyber page&lt;/a&gt; compares Argon with Gemini 3.8 Flash Cyber, a Fairwind-gated cyber-focused Flash variant.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Cyber evaluation&lt;/th&gt;
&lt;th&gt;Gemini 4 Argon&lt;/th&gt;
&lt;th&gt;Gemini 3.8 Flash Cyber&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Real-world vulnerability discovery&lt;/td&gt;
&lt;td&gt;85.8%&lt;/td&gt;
&lt;td&gt;71.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Wiz Penetration Test Benchmark&lt;/td&gt;
&lt;td&gt;70.9%&lt;/td&gt;
&lt;td&gt;58.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Both models are gated:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 3.8 Flash Cyber&lt;/strong&gt; is generally available behind an allowlist on Gemini Enterprise Agent Platform.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Gemini 4 Argon&lt;/strong&gt; is Fairwind-only.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If you are not a Fairwind partner, these scores do not change the model you can use today. The generally available Gemini 3.8 Flash model is the practical implementation choice.&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost of the same call on both
&lt;/h2&gt;

&lt;p&gt;Use the same workload to compare pricing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input:&lt;/strong&gt; 100,000 tokens&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output:&lt;/strong&gt; 8,000 tokens&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Assumption:&lt;/strong&gt; For current Gemini models, including Flash, thinking tokens are billed as output tokens. Google has not documented Argon’s thinking-token billing, so the Argon calculation assumes it follows current Gemini billing.&lt;/li&gt;
&lt;/ul&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model and period&lt;/th&gt;
&lt;th&gt;Input cost&lt;/th&gt;
&lt;th&gt;Output cost&lt;/th&gt;
&lt;th&gt;Total per call&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash, intro&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.1M × $0.75 = $0.075&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.008M × $3.75 = $0.030&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.105&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash, from 2027-01-01&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.1M × $1.50 = $0.150&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.008M × $7.50 = $0.060&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.210&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Argon, intro&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.1M × $2 = $0.200&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.008M × $10 = $0.080&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.280&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Argon, standard&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.1M × $4 = $0.400&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;0.008M × $20 = $0.160&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$0.560&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Argon’s per-token rates are &lt;code&gt;8 / 3&lt;/code&gt;, or about &lt;strong&gt;2.67× Flash pricing&lt;/strong&gt;, during both intro and standard pricing periods.&lt;/p&gt;

&lt;p&gt;At 1,000 calls per day using this workload:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Gemini 3.8 Flash intro pricing: 1,000 × $0.105 = $105/day
Gemini 4 Argon intro pricing: 1,000 × $0.280 = $280/day
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Three implementation details can change the real ratio:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Token counts will differ.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The same prompt can produce different answer lengths and thinking-token counts. Argon’s published evaluations used its highest thinking settings, while Google warns that higher Flash effort levels can consume more tokens.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Long outputs may require request splitting.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
A 200,000-token answer exceeds Flash’s 65,536-token limit, so Flash would require multiple calls. Under Argon’s stated limit, it fits in one response.&lt;br&gt;
&lt;/p&gt;&lt;/li&gt;
&lt;/ol&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;   Argon, 200K output tokens:
   0.2M × $10 = $2.00 at intro pricing
   0.2M × $20 = $4.00 at standard pricing

   Argon, 1M output tokens:
   1M × $10 = $10.00 at intro pricing
   1M × $20 = $20.00 at standard pricing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Only Flash has a published intro-pricing end date.&lt;/strong&gt;
Flash intro pricing ends on 2026-12-31. Google has not stated when Argon intro pricing ends.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Flash also supports Batch at a 50% discount: $0.375 input and $1.875 output per 1M tokens through December 31, according to the &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Gemini API pricing page&lt;/a&gt;. Google has not published Batch pricing for Argon.&lt;/p&gt;

&lt;p&gt;For more detail, see the &lt;a href="http://apidog.com/blog/gemini-4-argon-pricing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon pricing guide&lt;/a&gt; and &lt;a href="http://apidog.com/blog/gemini-3-8-flash-pricing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.8 Flash pricing guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Decision rule: ship on Flash or plan for Argon?
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Ship on Gemini 3.8 Flash now when
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;You are cost-bound.&lt;/strong&gt; Flash costs 3/8 of Argon per token, and Batch can halve Flash costs again.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You run high-volume workloads.&lt;/strong&gt; Classification, extraction, routing, and chat costs compound quickly at $10 per 1M output tokens.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You need a model today.&lt;/strong&gt; Flash has a stable model ID, free-tier prototyping access, and a published price schedule.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Your responses fit within 65,536 tokens.&lt;/strong&gt; Most API workloads do.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Plan for Gemini 4 Argon when
&lt;/h3&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Your workflows are long-horizon.&lt;/strong&gt; Google describes Argon as built for deep reasoning across complex, long-running workflows.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You need outputs beyond 64K tokens.&lt;/strong&gt; Large migrations, full-module rewrites, and long reports are the clearest use cases for the stated 1M-token output limit.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You build legal or finance agents.&lt;/strong&gt; These are the two benchmark rows Google published for both models, and Argon leads on both.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;You are a Fairwind-eligible defender.&lt;/strong&gt; Argon is the model currently led by that program.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The practical approach is not to choose permanently. Route normal production traffic to Flash, then test and route only the hard tail to Argon when it becomes available.&lt;/p&gt;

&lt;h2&gt;
  
  
  One saved request, two models
&lt;/h2&gt;

&lt;p&gt;Keep the model ID in an environment variable. Run production-like requests against Flash now, then update the variable when Google publishes Argon’s model ID.&lt;/p&gt;

&lt;p&gt;Do not guess an Argon model ID.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;GEMINI_MODEL&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;gemini&lt;/span&gt;&lt;span class="p"&gt;-3.8-flash&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="s2"&gt;"https://generativelanguage.googleapis.com/v1beta/models/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;MODEL&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:generateContent"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-goog-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$GEMINI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "contents": [
      {
        "parts": [
          {
            "text": "Summarize this contract clause in 3 bullets."
          }
        ]
      }
    ],
    "generationConfig": {
      "thinkingConfig": {
        "thinkingLevel": "medium"
      },
      "maxOutputTokens": 8192
    }
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set the environment values:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GEMINI_API_KEY&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"your-api-key"&lt;/span&gt;
&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GEMINI_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"gemini-3.8-flash"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;When Argon receives a published model ID, the intended switch is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GEMINI_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"published-argon-model-id"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Add cost and response checks
&lt;/h3&gt;

&lt;p&gt;Every &lt;code&gt;generateContent&lt;/code&gt; response includes &lt;code&gt;usageMetadata&lt;/code&gt;. Use it to enforce a cost ceiling and compare model behavior.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create an environment with &lt;code&gt;GEMINI_API_KEY&lt;/code&gt; and &lt;code&gt;GEMINI_MODEL&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Save the request above.&lt;/li&gt;
&lt;li&gt;Add assertions for:

&lt;ul&gt;
&lt;li&gt;HTTP status &lt;code&gt;200&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;The response fields your application consumes&lt;/li&gt;
&lt;li&gt;A maximum &lt;code&gt;usageMetadata.thoughtsTokenCount&lt;/code&gt; based on your token budget&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Keep &lt;code&gt;maxOutputTokens&lt;/code&gt; set so an unexpectedly long response cannot exceed your output budget.&lt;/li&gt;
&lt;li&gt;Add the request to a test scenario and run it against Flash.&lt;/li&gt;
&lt;li&gt;Save the report as a baseline for the Argon comparison.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For example, set an explicit output cap per request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"generationConfig"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"thinkingConfig"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"thinkingLevel"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"medium"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"maxOutputTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Prepare for the Interactions API
&lt;/h3&gt;

&lt;p&gt;There are two launch-day caveats:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Google says “all new models” will launch on the Interactions API, but has not confirmed whether Argon supports &lt;code&gt;generateContent&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Argon’s thinking-level names are not documented, so Flash’s &lt;code&gt;medium&lt;/code&gt; setting may not transfer directly.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep an Interactions API version of your request beside the &lt;code&gt;generateContent&lt;/code&gt; version:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;POST https://generativelanguage.googleapis.com/v1beta/interactions
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use &lt;code&gt;generation_config.thinking_level&lt;/code&gt; for the Interactions API request.&lt;/p&gt;

&lt;p&gt;When Argon is available:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Update only &lt;code&gt;GEMINI_MODEL&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Run the same prompt suite.&lt;/li&gt;
&lt;li&gt;Compare output quality, latency, and &lt;code&gt;usageMetadata&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Calculate cost from actual input, output, and thinking-token usage.&lt;/li&gt;
&lt;li&gt;Route only workloads where Argon’s result justifies its higher cost.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Gemini 4 Argon better than Gemini 3.8 Flash?
&lt;/h3&gt;

&lt;p&gt;On the two benchmark rows shared by both launch tables, yes:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Vals Finance Agent v2: &lt;strong&gt;65.4% vs. 61.4%&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Harvey Legal Agent Benchmark: &lt;strong&gt;19.6% vs. 10.0%&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;However, Argon costs about 2.67× more per token and is not generally callable yet.&lt;/p&gt;

&lt;h3&gt;
  
  
  Should I wait for Gemini 4 Argon?
&lt;/h3&gt;

&lt;p&gt;Not if you need to ship. Google has not provided a release date, only “as soon as possible,” starting with paid API customers and AI Ultra subscribers.&lt;/p&gt;

&lt;h3&gt;
  
  
  Gemini 3.8 Flash or Argon for a new project?
&lt;/h3&gt;

&lt;p&gt;Start with Gemini 3.8 Flash and place the model ID in an environment variable. Move requests requiring long outputs or stronger legal and finance performance to Argon only after testing it against your own prompts.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is there a free way to use Argon?
&lt;/h3&gt;

&lt;p&gt;No. Google has not announced a free tier. See the &lt;a href="http://apidog.com/blog/gemini-4-argon-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon free-access explainer&lt;/a&gt; for free Gemini models available today.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does Argon replace Gemini 3.8 Flash?
&lt;/h3&gt;

&lt;p&gt;Google has not said so. Google calls Argon its new frontier model, while Reuters reports it is larger than the previous Pro line. It appears to be a different tier from Flash.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next step
&lt;/h2&gt;

&lt;p&gt;Set up one saved request today, run it on Gemini 3.8 Flash, and store the output and usage report. When Argon ships, you can evaluate its quality and cost against your real traffic within minutes.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; to build and save the comparison.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 13:24:39 +0000</pubDate>
      <link>https://dev.to/hassann/gemini-4-argon-vs-gpt-6-astra-vs-claude-opus-55-1dli</link>
      <guid>https://dev.to/hassann/gemini-4-argon-vs-gpt-6-astra-vs-claude-opus-55-1dli</guid>
      <description>&lt;p&gt;Gemini 4 Argon leads Google’s benchmark table on 13 of 19 rows against GPT-6 Astra, Claude Opus 5.5, and Claude Fable 5.1. But availability matters more than any benchmark: you can buy Astra and Opus 5.5 today, while Argon is available only to Fairwind Program defenders. Its introductory price is $2/$10 per million input/output tokens, increasing to $4/$20 afterward—the same list price as Opus 5.5.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This comparison breaks down pricing and limits, groups Google’s benchmark results by winner, highlights evaluation caveats, and provides a practical routing guide. It also shows how to compare the available models with your own prompts in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;. If you are new to Argon, start with &lt;a href="http://apidog.com/blog/what-is-gemini-4-argon?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what is Gemini 4 Argon&lt;/a&gt;. For availability details, see the &lt;a href="http://apidog.com/blog/gemini-4-argon-release-date?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon release date guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Price and limits side by side
&lt;/h2&gt;

&lt;p&gt;Google’s comparison includes Claude Fable 5.1, so it belongs in the table. Prices below are per 1M tokens and come from Google’s &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/" rel="noopener noreferrer"&gt;launch post&lt;/a&gt;, OpenAI’s &lt;a href="https://developers.openai.com/api/docs/models/gpt-6-astra" rel="noopener noreferrer"&gt;GPT-6 Astra model page&lt;/a&gt;, and Anthropic’s &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Gemini 4 Argon&lt;/th&gt;
&lt;th&gt;GPT-6 Astra&lt;/th&gt;
&lt;th&gt;Claude Opus 5.5&lt;/th&gt;
&lt;th&gt;Claude Fable 5.1&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Can you call it today?&lt;/td&gt;
&lt;td&gt;No, Fairwind only&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;td&gt;Yes&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API model ID&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gpt-6-astra&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-opus-5-5&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;claude-fable-5-1&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;$2 intro, then $4&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;td&gt;$4&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cached input&lt;/td&gt;
&lt;td&gt;$0.10 intro, then $0.20&lt;/td&gt;
&lt;td&gt;$1&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;$10 intro, then $20&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max output&lt;/td&gt;
&lt;td&gt;1M (Google’s stated limit)&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;td&gt;128K (300K on Batch, beta)&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Context window&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;1,050,000 (922K max input)&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-prompt surcharge&lt;/td&gt;
&lt;td&gt;Not stated&lt;/td&gt;
&lt;td&gt;Over 272K input: 2x input and cache, 1.5x output, on the full request&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;td&gt;None&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Three implementation implications stand out:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Argon and Opus 5.5 have the same standard rates.&lt;/strong&gt; Argon’s cost advantage over Opus lasts only during Google’s introductory period, which has no published end date.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Astra and Fable 5.1 cost more per token.&lt;/strong&gt; Their listed input and output rates are 2.5x Argon’s standard rates and 5x Argon’s introductory rates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Treat Argon’s 1M output limit carefully.&lt;/strong&gt; Google states a 1M-token output limit, but Vals AI lists a 262K maximum output for the Argon configuration it tested.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For per-request calculations, see &lt;a href="http://apidog.com/blog/gemini-4-argon-pricing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon pricing&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where each model leads in Google’s table
&lt;/h2&gt;

&lt;p&gt;Google published this table, so interpret it accordingly. Argon’s scores are Google-reported, using a mix of self-computed results and leaderboard results. Competitor scores are mostly vendor-reported figures or public leaderboard results, often at maximum reasoning settings. Different harnesses mean small score gaps may not be meaningful.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Who leads&lt;/th&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Argon&lt;/th&gt;
&lt;th&gt;Astra&lt;/th&gt;
&lt;th&gt;Fable 5.1&lt;/th&gt;
&lt;th&gt;Opus 5.5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Argon: knowledge work&lt;/td&gt;
&lt;td&gt;Vals Index&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;63.1%&lt;/td&gt;
&lt;td&gt;65.8%&lt;/td&gt;
&lt;td&gt;67.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Argon: knowledge work&lt;/td&gt;
&lt;td&gt;AutomationBench&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;51.3%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;41.4%&lt;/td&gt;
&lt;td&gt;31.4%&lt;/td&gt;
&lt;td&gt;42.5%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Argon: knowledge work&lt;/td&gt;
&lt;td&gt;Harvey’s Legal Agent Benchmark&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;19.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;5.4%&lt;/td&gt;
&lt;td&gt;6.7%&lt;/td&gt;
&lt;td&gt;3.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Argon: coding&lt;/td&gt;
&lt;td&gt;DeepSWE v1.1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;77.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;74.1%&lt;/td&gt;
&lt;td&gt;67.4%&lt;/td&gt;
&lt;td&gt;74.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Argon: long context&lt;/td&gt;
&lt;td&gt;GraphWalks 256K to 1M (F1)&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;84.2%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;71.8%&lt;/td&gt;
&lt;td&gt;65.0%&lt;/td&gt;
&lt;td&gt;66.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Argon: multimodal&lt;/td&gt;
&lt;td&gt;LVBench&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;91.7%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;87.5%&lt;/td&gt;
&lt;td&gt;79.7%&lt;/td&gt;
&lt;td&gt;83.7%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Argon: multimodal&lt;/td&gt;
&lt;td&gt;Chartography&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;71.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;71.0%&lt;/td&gt;
&lt;td&gt;46.2%&lt;/td&gt;
&lt;td&gt;66.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Astra&lt;/td&gt;
&lt;td&gt;FrontierSWE v2&lt;/td&gt;
&lt;td&gt;55.0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;65.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;56.3%&lt;/td&gt;
&lt;td&gt;62.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Astra&lt;/td&gt;
&lt;td&gt;Terminal-Bench Science 0.1&lt;/td&gt;
&lt;td&gt;57.6%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.1%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;52.6%&lt;/td&gt;
&lt;td&gt;63.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Astra&lt;/td&gt;
&lt;td&gt;OSWorld-2.0 (offline, partial)&lt;/td&gt;
&lt;td&gt;69.2%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;72.6%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;not reported&lt;/td&gt;
&lt;td&gt;not reported&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 5.5&lt;/td&gt;
&lt;td&gt;Terminal-bench 4.0&lt;/td&gt;
&lt;td&gt;57.4%&lt;/td&gt;
&lt;td&gt;58.2%&lt;/td&gt;
&lt;td&gt;57.9%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;66.4%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Opus 5.5&lt;/td&gt;
&lt;td&gt;PostTrainBench&lt;/td&gt;
&lt;td&gt;45.3%&lt;/td&gt;
&lt;td&gt;44.3%&lt;/td&gt;
&lt;td&gt;40.2%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;49.3%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Tie&lt;/td&gt;
&lt;td&gt;CWE-bench v1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.0%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;58.0%&lt;/td&gt;
&lt;td&gt;67.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Argon’s other outright wins are Vals Finance Agent v2, Vibe Code Bench, LABBench 2, RiemannBench, GraphWalks up to 128K, and Agent’s Last Exam.&lt;/p&gt;

&lt;p&gt;Fable 5.1 does not lead a row in Google’s table. On CWE-bench v1, DeepMind’s &lt;a href="https://deepmind.google/models/gemini/cyber/" rel="noopener noreferrer"&gt;cyber leaderboard&lt;/a&gt; shows a three-way 68% tie between Argon, Astra, and Grok 4.7, with Opus 5.5 at 67%.&lt;/p&gt;

&lt;p&gt;For Argon versus Fable 5.1, Argon scores higher on 15 of the 17 rows where both report a result. Fable exceeds Argon only on:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;FrontierSWE v2: 56.3% vs. 55.0%&lt;/li&gt;
&lt;li&gt;Terminal-bench 4.0: 57.9% vs. 57.4%&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;See &lt;a href="http://apidog.com/blog/claude-fable-5-1-benchmarks?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Fable 5.1 benchmarks&lt;/a&gt; for Anthropic’s reported numbers and &lt;a href="http://apidog.com/blog/gemini-4-argon-benchmarks?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon benchmarks&lt;/a&gt; for the full 19-row table.&lt;/p&gt;

&lt;h2&gt;
  
  
  Three caveats before you trust the gaps
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Competitor scores are not Google reruns
&lt;/h3&gt;

&lt;p&gt;Google’s &lt;a href="https://deepmind.google/models/evals-methodology/gemini-4-argon" rel="noopener noreferrer"&gt;methodology&lt;/a&gt; says non-Gemini results are “sourced from providers’ self reported numbers unless otherwise mentioned.” Several rows also come from public leaderboards operated by Vals AI, Proximal, and Surge.&lt;/p&gt;

&lt;p&gt;Use these numbers for model selection hypotheses, not as final production evidence.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Benchmark harnesses differ
&lt;/h3&gt;

&lt;p&gt;On DeepSWE v1.1, Google calculated Argon’s 77.9% using a mini-swe-agent harness. Astra’s number comes from a public leaderboard, while Anthropic’s scores come from system cards.&lt;/p&gt;

&lt;p&gt;On LVBench, Gemini sampled video at one frame per second. Astra received 800 frames, Opus 5.5 received 600, and Fable 5.1 received 300, reportedly because of API limits.&lt;/p&gt;

&lt;p&gt;Do not treat these as strictly equivalent runs.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. A model can score differently across charts
&lt;/h3&gt;

&lt;p&gt;Opus 5.5’s Terminal-bench 4.0 score changes between charts based on the harness and effort setting. &lt;a href="https://www.anthropic.com/news/claude-opus-5-5" rel="noopener noreferrer"&gt;Anthropic reports&lt;/a&gt; its 66.4% result at &lt;code&gt;xhigh&lt;/code&gt; effort.&lt;/p&gt;

&lt;p&gt;When creating an internal evaluation spreadsheet, keep the following fields alongside each score:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;model
model_version
reasoning_or_effort_setting
benchmark_version
agent_harness
tool_configuration
date_tested
source
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Never combine benchmark results from different harnesses into a single ranking without labeling the differences.&lt;/p&gt;

&lt;h2&gt;
  
  
  What third-party evaluators say
&lt;/h2&gt;

&lt;p&gt;Independent leaderboards narrow the gap.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://artificialanalysis.ai/leaderboards/models" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt; lists Argon at #8 out of 223 entries, although that ranking counts each reasoning setting separately. Grouped by distinct model, Argon scores 53, tying GPT-6 Astra and Claude Fable 5.1. It trails Claude Opus 5.5 at 58 at maximum settings and Claude Sonnet 5.5 at 56.&lt;/p&gt;

&lt;p&gt;Artificial Analysis also lists Argon’s hallucination rate on AA-Omniscience at 15%, compared with 51% for Astra at its maximum setting.&lt;/p&gt;

&lt;p&gt;On &lt;a href="https://arena.ai/leaderboard/text" rel="noopener noreferrer"&gt;Arena&lt;/a&gt;, Argon ranks first in Text at 1525, marked Preliminary with 4,942 votes. Opus 5.5 ranks fourth.&lt;/p&gt;

&lt;p&gt;Vals AI ranks Argon first among 41 models on the Vals Index, making it the first Gemini model to top that ranking.&lt;/p&gt;

&lt;p&gt;There are also reports of internal skepticism. &lt;a href="https://www.bloomberg.com/news/articles/2026-09-30/google-grapples-with-employee-skepticism-about-new-gemini-model" rel="noopener noreferrer"&gt;Bloomberg reported&lt;/a&gt; that some Google employees with access found Argon less impressive on certain coding and front-end design tasks than its benchmark results suggest. Google called that characterization inaccurate.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which model should you route to?
&lt;/h2&gt;

&lt;p&gt;There is no universal best frontier model. Route requests by workload, quality requirements, latency, output limits, and cost.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Task&lt;/th&gt;
&lt;th&gt;Route today&lt;/th&gt;
&lt;th&gt;When Argon ships&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Legal, finance, and office automation agents&lt;/td&gt;
&lt;td&gt;Opus 5.5 (67.0% Vals Index)&lt;/td&gt;
&lt;td&gt;Test Argon: it leads all four knowledge-work rows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-heavy coding agents&lt;/td&gt;
&lt;td&gt;Opus 5.5 (66.4% Terminal-bench 4.0)&lt;/td&gt;
&lt;td&gt;Opus 5.5 still leads&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Agentic software engineering&lt;/td&gt;
&lt;td&gt;Astra (65.5% FrontierSWE v2) or Opus 5.5&lt;/td&gt;
&lt;td&gt;Argon leads DeepSWE but trails FrontierSWE; test both&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Computer use and GUI agents&lt;/td&gt;
&lt;td&gt;Astra (72.6% OSWorld-2.0)&lt;/td&gt;
&lt;td&gt;Astra leads OSWorld; Argon leads Agent’s Last Exam&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reasoning over 256K-token prompts&lt;/td&gt;
&lt;td&gt;Opus 5.5 (no surcharge) or Astra (surcharge over 272K)&lt;/td&gt;
&lt;td&gt;Argon (84.2% GraphWalks 256K to 1M)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Video and chart understanding&lt;/td&gt;
&lt;td&gt;Astra (87.5% LVBench)&lt;/td&gt;
&lt;td&gt;Argon (91.7%), with the frame-count caveat&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Science and ML engineering&lt;/td&gt;
&lt;td&gt;Astra (Terminal-Bench Science), Opus 5.5 (PostTrainBench)&lt;/td&gt;
&lt;td&gt;Argon leads LABBench 2 and RiemannBench; test it&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Single responses over 128K tokens&lt;/td&gt;
&lt;td&gt;Opus 5.5 on Batch (300K, beta)&lt;/td&gt;
&lt;td&gt;Argon, up to Google’s stated 1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;High-volume, cost-sensitive work&lt;/td&gt;
&lt;td&gt;Opus 5.5 ($4/$20)&lt;/td&gt;
&lt;td&gt;Argon at $2/$10 while the intro lasts&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If price matters more than peak benchmark results, compare OpenAI’s lower-cost line with Opus using &lt;a href="http://apidog.com/blog/gpt-6-sol-vs-claude-opus-5-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GPT-6 Sol vs Claude Opus 5.5&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compare the models on your own prompts
&lt;/h2&gt;

&lt;p&gt;Vendor benchmarks cannot tell you how models handle your repository, documents, tools, response formats, or failure modes. Build a small repeatable evaluation instead.&lt;/p&gt;

&lt;p&gt;You can set this up today in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Create three requests
&lt;/h3&gt;

&lt;p&gt;Create one project with three requests:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;GPT-6 Astra&lt;/li&gt;
&lt;li&gt;Claude Opus 5.5&lt;/li&gt;
&lt;li&gt;Gemini with a configurable model ID&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For the Gemini request, use an environment variable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"{{GEMINI_MODEL}}"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"contents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"role"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"user"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"parts"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
          &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"{{PROMPT}}"&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set &lt;code&gt;GEMINI_MODEL&lt;/code&gt; to &lt;code&gt;gemini-3.8-flash&lt;/code&gt; until Google publishes Argon’s API model ID.&lt;/p&gt;

&lt;p&gt;Use the &lt;a href="http://apidog.com/blog/gpt-6-astra-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GPT-6 Astra API guide&lt;/a&gt; and &lt;a href="http://apidog.com/blog/what-is-claude-opus-5-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what is Claude Opus 5.5&lt;/a&gt; for the vendor-specific request formats.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Store credentials and shared inputs as variables
&lt;/h3&gt;

&lt;p&gt;Create environment variables for each vendor key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;OPENAI_API_KEY
ANTHROPIC_API_KEY
GOOGLE_API_KEY
GEMINI_MODEL
PROMPT
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Put the test prompt in &lt;code&gt;PROMPT&lt;/code&gt; so every request receives exactly the same input.&lt;/p&gt;

&lt;p&gt;For multi-case testing, add a dataset with fields such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"case_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"bugfix-001"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"prompt"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Fix the failing test without changing the public API."&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"max_output_tokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"max_cost_usd"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;0.25&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 3: Add assertions
&lt;/h3&gt;

&lt;p&gt;At minimum, validate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;HTTP status is &lt;code&gt;200&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;The response contains non-empty output&lt;/li&gt;
&lt;li&gt;Output tokens remain below your configured ceiling&lt;/li&gt;
&lt;li&gt;Estimated cost stays within budget&lt;/li&gt;
&lt;li&gt;Required structured fields are present, if using JSON output&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, your assertions should enforce the equivalent of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="nf"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;status&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nf"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;answer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nf"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;outputTokens&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;maxOutputTokens&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="nf"&gt;assert&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;estimatedCostUsd&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;maxCostUsd&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Step 4: Track cost per request
&lt;/h3&gt;

&lt;p&gt;The basic cost formula is:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cost =
  (input_tokens / 1_000_000 × input_rate) +
  (output_tokens / 1_000_000 × output_rate)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For a request with 20,000 input tokens and 4,000 output tokens:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Astra:
20,000 × $10 / 1M + 4,000 × $50 / 1M
= $0.20 + $0.20
= $0.40

Opus 5.5:
20,000 × $4 / 1M + 4,000 × $20 / 1M
= $0.08 + $0.08
= $0.16

Argon at intro rates:
20,000 × $2 / 1M + 4,000 × $10 / 1M
= $0.04 + $0.04
= $0.08

Argon at standard rates:
20,000 × $4 / 1M + 4,000 × $20 / 1M
= $0.08 + $0.08
= $0.16
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Per-token pricing is not the same as per-task cost. Artificial Analysis measured Argon at roughly 62K output tokens per task versus about 27K for Astra at maximum effort. Measure actual token use for your own tasks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: Run one scenario and compare results
&lt;/h3&gt;

&lt;p&gt;Run all three requests as one test scenario. Compare:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Task completion rate&lt;/li&gt;
&lt;li&gt;Output validity&lt;/li&gt;
&lt;li&gt;Tool-call correctness&lt;/li&gt;
&lt;li&gt;Latency&lt;/li&gt;
&lt;li&gt;Input and output token usage&lt;/li&gt;
&lt;li&gt;Estimated request cost&lt;/li&gt;
&lt;li&gt;Human review score for a sample of outputs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When Argon becomes generally available, update only this variable:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GEMINI_MODEL=&amp;lt;published-argon-model-id&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then rerun the same scenario. You will have a same-prompt comparison against Astra and Opus 5.5 without rebuilding your test setup.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is Gemini 4 Argon better than GPT-6 Astra?
&lt;/h3&gt;

&lt;p&gt;They tie at 53 on Artificial Analysis. In Google’s table, Argon scores higher on most rows, but Astra wins FrontierSWE v2, Terminal-Bench Science 0.1, and OSWorld-2.0. Astra is also available today.&lt;/p&gt;

&lt;h3&gt;
  
  
  Gemini 4 Argon vs Claude Opus 5.5: which is better?
&lt;/h3&gt;

&lt;p&gt;Google’s table favors Argon on most rows. Opus 5.5 wins Terminal-bench 4.0 and PostTrainBench, and it leads Artificial Analysis 58 to 53. At Argon’s standard rates, the two models cost the same.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is Gemini 4 Argon cheaper than Claude Opus 5.5?
&lt;/h3&gt;

&lt;p&gt;During the introductory period, Argon is half the price: $2/$10 versus $4/$20 per million input/output tokens. After the introductory period, their list prices are identical. Google has not announced when the introductory pricing ends.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which model has the largest output limit?
&lt;/h3&gt;

&lt;p&gt;Argon has Google’s stated 1M-token output limit, although Vals AI lists 262K for the configuration it tested. The other models cap synchronous output at 128K. See the &lt;a href="http://apidog.com/blog/gemini-4-argon-1m-output-tokens?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;1M output tokens guide&lt;/a&gt; for client-side implications.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I use Gemini 4 Argon today?
&lt;/h3&gt;

&lt;p&gt;Only through Google’s Fairwind Program, which is available to a set of vetted cyber defense partners. Paid API customers and Google AI Ultra subscribers are expected later, but no date has been published.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pick by workload, then test
&lt;/h2&gt;

&lt;p&gt;Argon looks strongest in Google’s results for knowledge work, long context, and multimodal tasks. Astra leads on FrontierSWE, Terminal-Bench Science, and OSWorld. Opus 5.5 leads on terminal work and ML engineering while matching Argon’s eventual standard price.&lt;/p&gt;

&lt;p&gt;Build production workflows on the models you can call now. Keep the Gemini model ID in an environment variable, record task-level quality and cost, and rerun the same scenario when Argon opens access.&lt;/p&gt;

&lt;p&gt;To build the three-request comparison in one project, &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;download Apidog&lt;/a&gt;.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Gemini 4 Argon Pricing: $2/$10 Intro, $4/$20 After, and What a 1M-Token Answer Costs</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:55:44 +0000</pubDate>
      <link>https://dev.to/hassann/gemini-4-argon-pricing-210-intro-420-after-and-what-a-1m-token-answer-costs-ko9</link>
      <guid>https://dev.to/hassann/gemini-4-argon-pricing-210-intro-420-after-and-what-a-1m-token-answer-costs-ko9</guid>
      <description>&lt;p&gt;Gemini 4 Argon costs $2 per 1M input tokens and $10 per 1M output tokens during an introductory period, then $4 and $20. Cached input is 95% off the input price: $0.10 per 1M tokens at the intro rate and $0.20 afterward. Google has not said how long the intro period lasts, and you cannot buy Argon yet: it is announced but not generally available, and is currently rolling out only to Fairwind Program cyber defenders.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This post turns the rate card into implementation-ready budget numbers: comparisons with current models, four pricing scenarios, third-party per-task costs, and cost controls you can configure now. For the model overview, start with &lt;a href="http://apidog.com/blog/what-is-gemini-4-argon?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what Gemini 4 Argon is&lt;/a&gt;; for availability, see &lt;a href="http://apidog.com/blog/gemini-4-argon-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;is Gemini 4 Argon free&lt;/a&gt;. You can build cost assertions in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; against a model you can call today, then swap the model name when Argon becomes available.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj7e77ememx3nfiqsd9tv.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fj7e77ememx3nfiqsd9tv.png" alt="fe" width="800" height="802"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  The price sheet
&lt;/h2&gt;

&lt;p&gt;Google published these prices in the &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/" rel="noopener noreferrer"&gt;Gemini 4 Argon launch post&lt;/a&gt;. Footnote 1 specifies the post-intro rate: “After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.”&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Gemini 4 Argon, per 1M tokens&lt;/th&gt;
&lt;th&gt;Intro&lt;/th&gt;
&lt;th&gt;Standard (after intro)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$4.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cached input (95% off input)&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$20.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max output per response (Google)&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cost of one full 1M-token output&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$20.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The cached-input rates are derived from Google’s 95% discount rule:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;intro cached input    = $2 × 0.05 = $0.10 per 1M tokens
standard cached input = $4 × 0.05 = $0.20 per 1M tokens
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Google has not published an input context window or a model ID. It says the output limit is 1M tokens, “up from the previous 64K tokens.” At least one third-party evaluator, Vals AI, lists a 262K max output for the configuration it tested.&lt;/p&gt;

&lt;h2&gt;
  
  
  Argon vs. current frontier and Gemini models
&lt;/h2&gt;

&lt;p&gt;All prices are per 1M tokens.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Cached input&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Max output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Argon (intro)&lt;/td&gt;
&lt;td&gt;$2&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Argon (standard)&lt;/td&gt;
&lt;td&gt;$4&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;td&gt;$1&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;128,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5.5&lt;/td&gt;
&lt;td&gt;$4&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$20&lt;/td&gt;
&lt;td&gt;128K (300K on Batch, beta)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Sonnet 5.5&lt;/td&gt;
&lt;td&gt;$2&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;td&gt;128K (300K on Batch, beta)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6.1 Sol&lt;/td&gt;
&lt;td&gt;$2&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;td&gt;128,000&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5.1&lt;/td&gt;
&lt;td&gt;$10&lt;/td&gt;
&lt;td&gt;$0.25&lt;/td&gt;
&lt;td&gt;$50&lt;/td&gt;
&lt;td&gt;128K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemini-3.1-pro-preview (prompt &amp;lt;=200K / &amp;gt;200K)&lt;/td&gt;
&lt;td&gt;$2 / $4&lt;/td&gt;
&lt;td&gt;$0.20 / $0.40&lt;/td&gt;
&lt;td&gt;$12 / $18&lt;/td&gt;
&lt;td&gt;65,536&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;gemini-3.8-flash (intro / from 2027-01-01)&lt;/td&gt;
&lt;td&gt;$0.75 / $1.50&lt;/td&gt;
&lt;td&gt;$0.075 / $0.15&lt;/td&gt;
&lt;td&gt;$3.75 / $7.50&lt;/td&gt;
&lt;td&gt;65,536&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Competitor prices come from vendor documentation: &lt;a href="https://developers.openai.com/api/docs/models/gpt-6-astra" rel="noopener noreferrer"&gt;OpenAI’s GPT-6 Astra page&lt;/a&gt;, &lt;a href="https://platform.claude.com/docs/en/about-claude/pricing" rel="noopener noreferrer"&gt;Anthropic’s pricing page&lt;/a&gt;, and &lt;a href="https://ai.google.dev/gemini-api/docs/pricing" rel="noopener noreferrer"&gt;Google’s Gemini API pricing page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Key comparisons:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Standard Argon matches Claude Opus 5.5.&lt;/strong&gt; Both are $4 input, $20 output, and $0.20 cached input.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Intro Argon matches Claude Sonnet 5.5 and GPT-6.1 Sol&lt;/strong&gt; at $2 input and $10 output. Its $0.10 cached-input rate also matches GPT-6.1 Sol.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;GPT-6 Astra and Claude Fable 5.1 cost 5x intro Argon&lt;/strong&gt; and 2.5x standard Argon on a per-token basis. See &lt;a href="http://apidog.com/blog/claude-fable-5-1-pricing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Fable 5.1 pricing&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Argon’s intro output price ($10) is below gemini-3.1-pro-preview’s ($12)&lt;/strong&gt;, Google’s previous top Pro model.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Long prompts can change the calculation.&lt;/strong&gt; OpenAI bills prompts over 272K input tokens at 2x input and cache prices and 1.5x output price for the full request. Gemini 3.1 Pro costs more above 200K input tokens; Anthropic does not use that tier. Google has not said whether Argon will have a long-prompt tier.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For capability comparisons alongside price, see &lt;a href="http://apidog.com/blog/gemini-4-argon-vs-gpt-6-astra-vs-claude-opus-5-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Treat thinking tokens as output cost
&lt;/h2&gt;

&lt;p&gt;On current Gemini models, thinking tokens are billed as output. Google’s &lt;a href="https://ai.google.dev/gemini-api/docs/thinking" rel="noopener noreferrer"&gt;thinking documentation&lt;/a&gt; states that “response pricing is the sum of output tokens and thinking tokens.”&lt;/p&gt;

&lt;p&gt;Google has not documented Argon-specific billing behavior, but budget as though the same rule applies. Google ran Argon evaluations with its highest thinking settings, so the $10 per 1M output-token intro price should be treated as the cost of both visible answer tokens and reasoning tokens.&lt;/p&gt;

&lt;p&gt;Use this cost model:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cost =
  (uncached_input_tokens / 1,000,000 × input_rate)
  + (cached_input_tokens / 1,000,000 × cached_input_rate)
  + ((output_tokens + thinking_tokens) / 1,000,000 × output_rate)
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Four worked scenarios
&lt;/h2&gt;

&lt;p&gt;Every example below uses:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;tokens / 1,000,000 × per-1M rate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Scenarios 1–3 assume no caching.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Typical API call: 20K input, 5K output
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Intro pricing&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;input  = 20,000 / 1M × $2  = $0.04
output = 5,000 / 1M × $10  = $0.05
total  = $0.09
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Standard pricing&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;input  = 20,000 / 1M × $4  = $0.08
output = 5,000 / 1M × $20 = $0.10
total  = $0.18
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At 1,000 calls per day:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pricing period&lt;/th&gt;
&lt;th&gt;Daily cost&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Intro&lt;/td&gt;
&lt;td&gt;$90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;$180&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;For the same 20K-input, 5K-output call:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Opus 5.5: $0.18&lt;/li&gt;
&lt;li&gt;GPT-6 Astra: $0.45 ($0.20 input + $0.25 output)&lt;/li&gt;
&lt;li&gt;gemini-3.1-pro-preview: $0.10 ($0.04 input + $0.06 output)&lt;/li&gt;
&lt;li&gt;gemini-3.8-flash at its intro rate: about $0.034 ($0.015 input + $0.019 output)&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Long agent turn: 200K input, 50K output
&lt;/h3&gt;

&lt;p&gt;&lt;strong&gt;Intro pricing&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;input  = 200,000 / 1M × $2  = $0.40
output = 50,000 / 1M × $10  = $0.50
total  = $0.90
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;strong&gt;Standard pricing&lt;/strong&gt;&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;input  = 200,000 / 1M × $4  = $0.80
output = 50,000 / 1M × $20 = $1.00
total  = $1.80
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For comparison:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;GPT-6 Astra: $4.50 ($2.00 input + $2.50 output)&lt;/li&gt;
&lt;li&gt;gemini-3.1-pro-preview: $1.00 ($0.40 input + $0.60 output)&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The 200K input size is Gemini 3.1 Pro’s pricing threshold. Prompts beyond that threshold use $4 input and $18 output pricing. If Argon introduces a similar tier, longer agent turns will cost more than the flat-rate examples above.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Maximum-size response: 1M output tokens
&lt;/h3&gt;

&lt;p&gt;Output alone costs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;1,000,000 / 1M × $10 = $10.00 intro
1,000,000 / 1M × $20 = $20.00 standard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Add a 100K-token prompt:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Pricing period&lt;/th&gt;
&lt;th&gt;Output&lt;/th&gt;
&lt;th&gt;Input&lt;/th&gt;
&lt;th&gt;Total&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Intro&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;td&gt;$10.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Standard&lt;/td&gt;
&lt;td&gt;$20.00&lt;/td&gt;
&lt;td&gt;$0.40&lt;/td&gt;
&lt;td&gt;$20.40&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;No other model in the comparison table produces 1M tokens in one synchronous response. Competitors cap output at 128K, and the previous Gemini Pro model caps output at 65,536 tokens.&lt;/p&gt;

&lt;p&gt;At Vals AI’s listed 262K max output, a full response would cost approximately:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;262,000 / 1M × $10 = $2.62 intro
262,000 / 1M × $20 = $5.24 standard
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For engineering considerations such as streaming, timeouts, and gateway limits, see &lt;a href="http://apidog.com/blog/gemini-4-argon-1m-output-tokens?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon’s 1M output tokens&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Cached 500K-token prompt reused 10 times
&lt;/h3&gt;

&lt;p&gt;Assume you send the same 500K-token codebase or contract set 10 times and receive 5K output tokens per call. The first call pays the full input price; the next nine calls use cached input.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;10 calls: 500K prompt, 5K output each&lt;/th&gt;
&lt;th&gt;Intro&lt;/th&gt;
&lt;th&gt;Standard&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input without caching (10 × 500K)&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$20.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input with caching (1 full + 9 cached)&lt;/td&gt;
&lt;td&gt;$1.00 + 9 × $0.05 = $1.45&lt;/td&gt;
&lt;td&gt;$2.00 + 9 × $0.10 = $2.90&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output (10 × 5K)&lt;/td&gt;
&lt;td&gt;$0.50&lt;/td&gt;
&lt;td&gt;$1.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Total with caching&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$1.95&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;$3.90&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Total without caching&lt;/td&gt;
&lt;td&gt;$10.50&lt;/td&gt;
&lt;td&gt;$21.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Caching cuts input cost by 85.5% in this example.&lt;/p&gt;

&lt;p&gt;However, account for two unknowns:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Google has not said whether Argon charges cache-storage fees. Gemini 3.1 Pro charges $4.50 per 1M tokens per hour, which would add $2.25 to hold 500K tokens for one hour.&lt;/li&gt;
&lt;li&gt;A 500K-token prompt is above Gemini 3.1 Pro’s 200K threshold. Argon may introduce a similar tier.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Argon’s standard cached-input rate matches Claude Opus 5.5’s $0.20 rate, so the break-even approach in &lt;a href="http://apidog.com/blog/claude-opus-5-5-prompt-caching-cost-math?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Opus 5.5 prompt caching cost math&lt;/a&gt; also applies here.&lt;/p&gt;

&lt;h2&gt;
  
  
  Per-token price is not per-task cost
&lt;/h2&gt;

&lt;p&gt;Do not budget from rate cards alone. Measure how many tokens your workload actually uses.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://artificialanalysis.ai/models/gemini-4-argon" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt; lists $1.99 per task to run its Intelligence Index on Gemini 4 Argon (High) at intro prices. &lt;a href="https://the-decoder.com/google-gemini-4-argon-closes-the-gap-with-openai-and-anthropic-but-doesnt-take-a-clear-lead/" rel="noopener noreferrer"&gt;The Decoder&lt;/a&gt; puts the same run at $3.98 at standard prices. Artificial Analysis scores Argon at 53, the same as GPT-6 Astra (max), which it lists at $3.26 per task.&lt;/p&gt;

&lt;p&gt;The important variable is output volume:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Artificial Analysis counts roughly 62K output tokens per Argon task at High.&lt;/li&gt;
&lt;li&gt;It counts roughly 27K output tokens per GPT-6 Astra task at its maximum setting.&lt;/li&gt;
&lt;li&gt;That is about 2.3x more output tokens for Argon.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;At intro pricing:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Argon output cost = 62,000 / 1M × $10 = $0.62
Astra output cost = 27,000 / 1M × $50 = $1.35
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At standard Argon prices, The Decoder’s full-task estimate for Argon ($3.98) exceeds Artificial Analysis’s listed $3.26 per task for GPT-6 Astra (max), despite Argon’s standard per-token rates being 60% lower than Astra’s.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://www.vals.ai/models/google_gemini-4-argon" rel="noopener noreferrer"&gt;Vals AI&lt;/a&gt; shows the same pattern. It lists $15.68 per test for Argon on the Vals Index, using the $4/$20 standard rate, compared with:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Claude Opus 5.5: $32.14 per test&lt;/li&gt;
&lt;li&gt;Claude Sonnet 5.5: $21.34 per test&lt;/li&gt;
&lt;li&gt;GPT-6.1 Sol: $3.24 per test&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Sonnet 5.5 and GPT-6.1 Sol share Argon’s intro price on paper, but their per-test costs differ by more than 6x.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Use your own production token counts to budget per-task cost.&lt;/strong&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Cost controls to implement before Argon ships
&lt;/h2&gt;

&lt;h3&gt;
  
  
  1. Cap output tokens
&lt;/h3&gt;

&lt;p&gt;On &lt;code&gt;generateContent&lt;/code&gt;, set &lt;code&gt;generationConfig.maxOutputTokens&lt;/code&gt; to limit output cost.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"generationConfig"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"maxOutputTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8192&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Google says new models launch on the Interactions API, so configure the equivalent cap there once Argon-specific documentation is available.&lt;/p&gt;

&lt;p&gt;A 1M-token output limit is also a maximum output charge of:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;$10 per request at intro pricing
$20 per request at standard pricing
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set a lower limit for normal production paths. Reserve large outputs for endpoints that explicitly require them.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Cache stable prompt content
&lt;/h3&gt;

&lt;p&gt;Cache content that repeats across requests:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;System instructions&lt;/li&gt;
&lt;li&gt;Tool schemas&lt;/li&gt;
&lt;li&gt;API specifications&lt;/li&gt;
&lt;li&gt;Reference documents&lt;/li&gt;
&lt;li&gt;Codebase context&lt;/li&gt;
&lt;li&gt;Contract sets&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Avoid caching user-specific or frequently changing content. The 95% cached-input discount is most useful when a large, stable prompt is reused across multiple calls.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Select thinking levels intentionally
&lt;/h3&gt;

&lt;p&gt;Google has not published Argon’s thinking levels. Current defaults include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;gemini-3.1-pro-preview&lt;/code&gt;: &lt;code&gt;high&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;gemini-3.8-flash&lt;/code&gt;: &lt;code&gt;medium&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;When Argon’s settings are documented, benchmark lower thinking levels against your task-quality requirements. Since thinking tokens are expected to be billed as output, reducing unnecessary reasoning can reduce cost.&lt;/p&gt;

&lt;h3&gt;
  
  
  4. Add usage and cost assertions
&lt;/h3&gt;

&lt;p&gt;Every &lt;code&gt;generateContent&lt;/code&gt; response ends with &lt;code&gt;usageMetadata&lt;/code&gt;. Save the model name in a &lt;code&gt;GEMINI_MODEL&lt;/code&gt; environment variable, then calculate cost from the returned token counts.&lt;/p&gt;

&lt;p&gt;In &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;, you can:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Save a Gemini request in a collection.&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;{{GEMINI_MODEL}}&lt;/code&gt; in the request URL or body.&lt;/li&gt;
&lt;li&gt;Run the request against &lt;code&gt;gemini-3.8-flash&lt;/code&gt; today.&lt;/li&gt;
&lt;li&gt;Add a post-response script that reads &lt;code&gt;usageMetadata&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Fail the request when calculated cost exceeds your per-call budget.&lt;/li&gt;
&lt;li&gt;Replace only &lt;code&gt;GEMINI_MODEL&lt;/code&gt; when Argon’s model ID ships.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Use this calculation:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;cost = uncached_input / 1,000,000 × input_rate
     + cached_input   / 1,000,000 × cached_rate
     + (output + thinking) / 1,000,000 × output_rate
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  What Google has not priced yet
&lt;/h2&gt;

&lt;p&gt;Before committing a production budget, track these unresolved items:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;The intro period length and the date that $4/$20 pricing starts&lt;/li&gt;
&lt;li&gt;Whether prompts above 200K tokens have a higher price tier&lt;/li&gt;
&lt;li&gt;Batch or Flex pricing&lt;/li&gt;
&lt;li&gt;Cache-storage fees&lt;/li&gt;
&lt;li&gt;Rate limits&lt;/li&gt;
&lt;li&gt;Whether a free tier exists&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For reference, Gemini 3.1 Pro has Batch and Flex pricing at $1/$2 input and $6/$9 output.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;How much does Gemini 4 Argon cost per million tokens?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;$2 input and $10 output during the intro period, then $4 input and $20 output. Cached input is 95% off: $0.10 during intro pricing and $0.20 afterward.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How long does the intro price last?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Google has not said. The launch post gives no duration or end date.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Are thinking tokens billed as output?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;For current Gemini models, yes. Google has not documented Argon’s rules specifically.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How much does a 1M-token output response cost?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;$10 at intro pricing and $20 at standard pricing, before input-token charges.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Argon cheaper than Claude Opus 5.5 or GPT-6 Astra?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Per token, standard Argon matches Claude Opus 5.5 and is 60% cheaper than GPT-6 Astra. Per task, the result depends on token volume: Artificial Analysis measured Argon using about 2.3x Astra’s output tokens. For a cheaper Gemini model available today, see &lt;a href="http://apidog.com/blog/gemini-3-8-flash-pricing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.8 Flash pricing&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I pay for Argon today?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. It is rolling out only to Fairwind Program partners. Paid API customers and Google AI Ultra subscribers are next, but Google has not provided a date.&lt;/p&gt;

&lt;h2&gt;
  
  
  Budget it now, measure it later
&lt;/h2&gt;

&lt;p&gt;Use your current traffic data to estimate an Argon range:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;intro estimate    = input_tokens × $2/1M + output_tokens × $10/1M
standard estimate = input_tokens × $4/1M + output_tokens × $20/1M
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then implement measurement now:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Save a request using &lt;code&gt;gemini-3.8-flash&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Store the model name in &lt;code&gt;GEMINI_MODEL&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Add a cost assertion based on &lt;code&gt;usageMetadata&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Set an output cap and per-request cost ceiling.&lt;/li&gt;
&lt;li&gt;When Google publishes Argon’s model ID, update one variable and rerun the suite.&lt;/li&gt;
&lt;/ol&gt;

</description>
    </item>
    <item>
      <title>Gemini 4 Argon's 1M Output Tokens: What a Million-Token Response Does to Your API Stack</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:54:56 +0000</pubDate>
      <link>https://dev.to/hassann/gemini-4-argons-1m-output-tokens-what-a-million-token-response-does-to-your-api-stack-111l</link>
      <guid>https://dev.to/hassann/gemini-4-argons-1m-output-tokens-what-a-million-token-response-does-to-your-api-stack-111l</guid>
      <description>&lt;p&gt;Gemini 4 Argon’s headline number—1M tokens—is its &lt;strong&gt;output limit&lt;/strong&gt;, not its context window. Google says a single Argon response can run to 1 million tokens, roughly 16× the previous 64K cap, but it has not published Argon’s input window. There is also nothing to call yet: Argon is currently available only to Fairwind Program defenders, with paid API customers next when Google opens access. See the &lt;a href="http://apidog.com/blog/gemini-4-argon-release-date?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;release date and access guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;A response this large breaks common API assumptions: requests do not finish in seconds, response bodies do not safely fit in memory, and request cost is no longer small or predictable. This guide shows how to plan for cost, streaming, timeouts, output caps, storage, and testing in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; before access arrives. For the model overview, see &lt;a href="http://apidog.com/blog/what-is-gemini-4-argon?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what Gemini 4 Argon is&lt;/a&gt;. For request shapes, see the &lt;a href="http://apidog.com/blog/gemini-4-argon-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon API guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  1M output tokens is not a 1M context window
&lt;/h2&gt;

&lt;p&gt;Do not treat Argon’s 1M figure as a context-window specification. Google’s &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/" rel="noopener noreferrer"&gt;launch post&lt;/a&gt; says it expanded “the model’s output token limit to an industry-leading 1M tokens.”&lt;/p&gt;

&lt;p&gt;That 64K comparison matches the 65,536-token output cap of Gemini 3.1 Pro Preview, Google’s previous top Pro model.&lt;/p&gt;

&lt;p&gt;Argon’s input window is a separate, unpublished value. Google’s long-context evaluation, described in its &lt;a href="https://deepmind.google/models/evals-methodology/gemini-4-argon" rel="noopener noreferrer"&gt;evals methodology&lt;/a&gt;, used prompts between 256K and 1M tokens. That is a benchmark range, not an API specification.&lt;/p&gt;

&lt;p&gt;There is also an output caveat: &lt;a href="https://www.vals.ai/models/google_gemini-4-argon" rel="noopener noreferrer"&gt;Vals AI&lt;/a&gt; lists a 262K maximum output for the Argon configuration it tested. Google states a 1M-token model limit, while at least one third-party evaluator observed a lower cap on the endpoint it used.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Max output per response&lt;/th&gt;
&lt;th&gt;Input or context window&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Argon&lt;/td&gt;
&lt;td&gt;1M (Google’s stated limit)&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.1 Pro Preview&lt;/td&gt;
&lt;td&gt;65,536&lt;/td&gt;
&lt;td&gt;1,048,576&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 3.8 Flash&lt;/td&gt;
&lt;td&gt;65,536&lt;/td&gt;
&lt;td&gt;1,048,576&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;128,000&lt;/td&gt;
&lt;td&gt;1,050,000 (922K max input)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5.5&lt;/td&gt;
&lt;td&gt;128K (300K on Batch with a beta header)&lt;/td&gt;
&lt;td&gt;1M&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Every competitor in this table caps synchronous output at 128K, making Argon’s stated 1M limit about 8× larger. Google’s stated reason is reasoning depth: the model can “generate hundreds of thousands of tokens in a single trajectory” to solve hard problems in one pass.&lt;/p&gt;

&lt;p&gt;For other long-running API patterns, see &lt;a href="http://apidog.com/blog/claude-opus-5-5-18-hour-tasks-api-design?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Claude Opus 5.5’s 18-hour tasks&lt;/a&gt; and the &lt;a href="http://apidog.com/blog/gpt-6-astra-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GPT-6 Astra API guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnad3coux68pvy6y5vbid.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fnad3coux68pvy6y5vbid.png" alt="Gemini 4 Argon output token illustration" width="800" height="802"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Calculate the worst-case cost first
&lt;/h2&gt;

&lt;p&gt;At Argon’s stated rates, output costs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Intro period: &lt;strong&gt;$10 per 1M output tokens&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;Standard pricing: &lt;strong&gt;$20 per 1M output tokens&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A full-length response therefore costs:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Intro: &lt;code&gt;1,000,000 × $10 / 1,000,000 = $10.00&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Standard: &lt;code&gt;1,000,000 × $20 / 1,000,000 = $20.00&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Input is additional. For example, a 200,000-token prompt adds:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Intro: &lt;code&gt;200,000 × $2 / 1,000,000 = $0.40&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Standard: &lt;code&gt;$0.80&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;That makes a maxed-out call:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;$10.40&lt;/strong&gt; at intro rates&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;$20.80&lt;/strong&gt; at standard rates&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A nightly job that runs 100 such requests would cost &lt;strong&gt;$1,040&lt;/strong&gt; at intro pricing.&lt;/p&gt;

&lt;p&gt;Thinking usage can make this less obvious. On current Gemini models, thinking tokens bill as output. Google has not stated whether Argon follows the same rule or whether thinking counts toward its 1M ceiling. Either way, a short visible answer can still incur substantial output usage.&lt;/p&gt;

&lt;p&gt;For additional scenarios, including cached input at 95% off, see the &lt;a href="http://apidog.com/blog/gemini-4-argon-pricing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon pricing guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Stream every long response
&lt;/h2&gt;

&lt;p&gt;A non-streaming request returns nothing until the entire response is complete. For outputs that may reach hundreds of thousands of tokens, that creates a long silent connection. Any client, proxy, gateway, load balancer, or serverless timeout can terminate it before the first byte arrives.&lt;/p&gt;

&lt;p&gt;Use server-sent events (SSE) instead.&lt;/p&gt;

&lt;p&gt;For &lt;code&gt;generateContent&lt;/code&gt;, use:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;:streamGenerateContent?alt=sse
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Read each event as it arrives and write each text chunk directly to storage. Do not accumulate the full response body in memory.&lt;/p&gt;

&lt;p&gt;This example runs against Gemini 3.8 Flash today. Keep the model name in an environment variable because Google has not published Argon’s model ID. See the &lt;a href="http://apidog.com/blog/how-to-use-gemini-3-8-flash-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.8 Flash API guide&lt;/a&gt; for setup details.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;

&lt;span class="n"&gt;MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GEMINI_MODEL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3.8-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;URL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://generativelanguage.googleapis.com/v1beta/models/&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
    &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="s"&gt;:streamGenerateContent?alt=sse&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;contents&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
        &lt;span class="p"&gt;{&lt;/span&gt;
            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;parts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
                &lt;span class="p"&gt;{&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Write a test plan for every endpoint in a payments API.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;
                &lt;span class="p"&gt;}&lt;/span&gt;
            &lt;span class="p"&gt;]&lt;/span&gt;
        &lt;span class="p"&gt;}&lt;/span&gt;
    &lt;span class="p"&gt;],&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;generationConfig&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;maxOutputTokens&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;60000&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="n"&gt;usage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="bp"&gt;None&lt;/span&gt;

&lt;span class="k"&gt;with&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;post&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;URL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;body&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;stream&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;120&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x-goog-api-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GEMINI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nf"&gt;open&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;response.txt&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;a&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;encoding&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;utf-8&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;as&lt;/span&gt; &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;iter_lines&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;decode_unicode&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="bp"&gt;True&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt; &lt;span class="ow"&gt;or&lt;/span&gt; &lt;span class="ow"&gt;not&lt;/span&gt; &lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;startswith&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="k"&gt;continue&lt;/span&gt;

        &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;loads&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;line&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="mi"&gt;5&lt;/span&gt;&lt;span class="p"&gt;:])&lt;/span&gt;

        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;candidate&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;candidates&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]):&lt;/span&gt;
            &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;part&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;candidate&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{}).&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;parts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[]):&lt;/span&gt;
                &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;""&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;

        &lt;span class="n"&gt;output&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;flush&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
        &lt;span class="n"&gt;usage&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;usageMetadata&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;usage&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;&lt;code&gt;timeout=(10, 120)&lt;/code&gt; configures:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A &lt;strong&gt;10-second connection timeout&lt;/strong&gt;
&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;120-second read timeout&lt;/strong&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;In &lt;code&gt;requests&lt;/code&gt;, the read timeout is the maximum gap between received bytes, not the total request duration. A stream can continue indefinitely if it keeps sending data within that interval.&lt;/p&gt;

&lt;p&gt;The example writes each chunk to disk immediately. On Gemini 3.8 Flash, every event includes running &lt;code&gt;usageMetadata&lt;/code&gt;, so the final event contains the token counts to record for cost tracking.&lt;/p&gt;

&lt;h2&gt;
  
  
  Check timeouts at every hop
&lt;/h2&gt;

&lt;p&gt;Your HTTP client is only one part of the request path. A long stream can also cross a reverse proxy, API gateway, load balancer, and serverless runtime.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Hop&lt;/th&gt;
&lt;th&gt;What to check&lt;/th&gt;
&lt;th&gt;Symptom when misconfigured&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;HTTP client&lt;/td&gt;
&lt;td&gt;Read or idle timeout, plus any total-request timeout&lt;/td&gt;
&lt;td&gt;Exceptions mid-stream only on long responses&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Reverse proxy&lt;/td&gt;
&lt;td&gt;Read timeout and response buffering for &lt;code&gt;text/event-stream&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Events arrive in bursts, or the stream cuts off&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API gateway&lt;/td&gt;
&lt;td&gt;Maximum request duration&lt;/td&gt;
&lt;td&gt;Requests fail at the same elapsed time every run&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Load balancer&lt;/td&gt;
&lt;td&gt;Idle timeout&lt;/td&gt;
&lt;td&gt;Drops during long pauses before the first event&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Serverless function&lt;/td&gt;
&lt;td&gt;Maximum execution time&lt;/td&gt;
&lt;td&gt;The function exits while the model is still generating&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;A fixed cutoff is the key signal. If long requests always fail at the same elapsed time, a hop in the path has a hard duration limit. Streaming alone cannot solve that; move the work off the live request path.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move the longest jobs to background execution
&lt;/h2&gt;

&lt;p&gt;For very long tasks, do not hold an HTTP request open.&lt;/p&gt;

&lt;p&gt;Google’s &lt;a href="https://ai.google.dev/gemini-api/docs/interactions" rel="noopener noreferrer"&gt;Interactions API&lt;/a&gt; supports background execution with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;background=true
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Background execution depends on stored interactions. The documentation states that &lt;code&gt;store=false&lt;/code&gt; is incompatible with background execution, so leave storage enabled for these requests.&lt;/p&gt;

&lt;p&gt;For retrieving a completed background interaction, follow Google’s documented workflow rather than guessing polling endpoints. Since Google says new models launch on the Interactions API, plan for Argon’s longest jobs to run there.&lt;/p&gt;

&lt;h2&gt;
  
  
  Set output caps intentionally
&lt;/h2&gt;

&lt;p&gt;The 1M limit is a ceiling, not a target.&lt;/p&gt;

&lt;p&gt;For &lt;code&gt;generateContent&lt;/code&gt;, use &lt;code&gt;generationConfig.maxOutputTokens&lt;/code&gt; to cap each response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"generationConfig"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"maxOutputTokens"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;60000&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On Gemini 3.8 Flash, thinking counts against that cap. In one test, a cap of 2,000 produced:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;1,340 thought tokens&lt;/li&gt;
&lt;li&gt;656 visible tokens&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For the Interactions API, verify the output-cap field in Google’s current documentation before relying on it.&lt;/p&gt;

&lt;p&gt;Choose a cap based on the maximum cost you accept per request:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Output cap&lt;/th&gt;
&lt;th&gt;Worst-case standard output cost ($20/1M)&lt;/th&gt;
&lt;th&gt;Intro output cost ($10/1M)&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;64,000&lt;/td&gt;
&lt;td&gt;$1.28&lt;/td&gt;
&lt;td&gt;$0.64&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;128,000&lt;/td&gt;
&lt;td&gt;$2.56&lt;/td&gt;
&lt;td&gt;$1.28&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;500,000&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$5.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;1,000,000&lt;/td&gt;
&lt;td&gt;$20.00&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Treat a capped response as potentially incomplete. Inspect the final event’s &lt;code&gt;finishReason&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MAX_TOKENS
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;If the response ended with &lt;code&gt;MAX_TOKENS&lt;/code&gt;, either continue in a follow-up turn or increase the cap for that specific job.&lt;/p&gt;

&lt;h2&gt;
  
  
  Store and parse huge outputs without buffering
&lt;/h2&gt;

&lt;p&gt;A million tokens can produce megabytes of text. Use a streaming storage path from the start.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Write chunks as they arrive to a file or multipart object upload.&lt;/li&gt;
&lt;li&gt;Keep partial files when a connection drops. A blind retry can regenerate and rebill the same output.&lt;/li&gt;
&lt;li&gt;Request JSON Lines when you need structured output, so each line can be parsed independently.&lt;/li&gt;
&lt;li&gt;Log &lt;code&gt;usageMetadata&lt;/code&gt; and byte counts rather than full response bodies.&lt;/li&gt;
&lt;li&gt;Verify database column limits and queue message-size limits before sending an 800K-token response through them.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Test the pipeline in Apidog before access opens
&lt;/h2&gt;

&lt;p&gt;You can validate the entire workflow against a stand-in model. &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; and run these three checks.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Inspect the SSE stream
&lt;/h3&gt;

&lt;p&gt;Send the streaming request to Gemini 3.8 Flash with &lt;code&gt;GEMINI_API_KEY&lt;/code&gt; and &lt;code&gt;GEMINI_MODEL&lt;/code&gt; configured as environment variables.&lt;/p&gt;

&lt;p&gt;Apidog parses &lt;code&gt;text/event-stream&lt;/code&gt; responses and displays each event in the Timeline view as it arrives. Use this to inspect:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Chunk sizes&lt;/li&gt;
&lt;li&gt;Gaps between events&lt;/li&gt;
&lt;li&gt;Final &lt;code&gt;usageMetadata&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Whether a proxy buffers the response&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  2. Stream a much longer fake response
&lt;/h3&gt;

&lt;p&gt;Gemini 3.8 Flash currently caps output at 65,536 tokens. To test parsing, buffering, and timeout behavior beyond that limit, run a local mock that emits Gemini-shaped SSE events:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# long_stream_mock.py: Gemini-shaped SSE for parser and timeout tests (fake data)
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;json&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;time&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;http.server&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;BaseHTTPRequestHandler&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;ThreadingHTTPServer&lt;/span&gt;

&lt;span class="n"&gt;EVENTS&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20000&lt;/span&gt;
&lt;span class="n"&gt;DELAY&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mf"&gt;0.005&lt;/span&gt;
&lt;span class="n"&gt;CHUNK&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;lorem ipsum &lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;40&lt;/span&gt;


&lt;span class="k"&gt;class&lt;/span&gt; &lt;span class="nc"&gt;Handler&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;BaseHTTPRequestHandler&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
    &lt;span class="k"&gt;def&lt;/span&gt; &lt;span class="nf"&gt;do_POST&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;rfile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;read&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nf"&gt;int&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Length&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)))&lt;/span&gt;

        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send_response&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;send_header&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;Content-Type&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text/event-stream&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
        &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;end_headers&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

        &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="nf"&gt;range&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;EVENTS&lt;/span&gt;&lt;span class="p"&gt;):&lt;/span&gt;
            &lt;span class="n"&gt;event&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;candidates&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
                    &lt;span class="p"&gt;{&lt;/span&gt;
                        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;content&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                            &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;parts&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
                                &lt;span class="p"&gt;{&lt;/span&gt;
                                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;text&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;CHUNK&lt;/span&gt;
                                &lt;span class="p"&gt;}&lt;/span&gt;
                            &lt;span class="p"&gt;]&lt;/span&gt;
                        &lt;span class="p"&gt;}&lt;/span&gt;
                    &lt;span class="p"&gt;}&lt;/span&gt;
                &lt;span class="p"&gt;]&lt;/span&gt;
            &lt;span class="p"&gt;}&lt;/span&gt;

            &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;i&lt;/span&gt; &lt;span class="o"&gt;==&lt;/span&gt; &lt;span class="n"&gt;EVENTS&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
                &lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;usageMetadata&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;promptTokenCount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;candidatesTokenCount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;950000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;thoughtsTokenCount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;40000&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;totalTokenCount&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;991200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
                &lt;span class="p"&gt;}&lt;/span&gt;

            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;wfile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;write&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
                &lt;span class="sa"&gt;f&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;data: &lt;/span&gt;&lt;span class="si"&gt;{&lt;/span&gt;&lt;span class="n"&gt;json&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;dumps&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;event&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="si"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\n\n&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;encode&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="p"&gt;)&lt;/span&gt;
            &lt;span class="n"&gt;self&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;wfile&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;flush&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
            &lt;span class="n"&gt;time&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;sleep&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;DELAY&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;


&lt;span class="nc"&gt;ThreadingHTTPServer&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;127.0.0.1&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="mi"&gt;8787&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="n"&gt;Handler&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;serve_forever&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Set a mock environment base URL to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;http://127.0.0.1:8787
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then send the same streaming request through it.&lt;/p&gt;

&lt;p&gt;This mock runs for about two minutes—128 seconds in the original test—and emits 9.6 million characters of text. It is large enough to expose:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A parser that buffers the entire body&lt;/li&gt;
&lt;li&gt;A proxy that holds SSE events&lt;/li&gt;
&lt;li&gt;A timeout configured too aggressively&lt;/li&gt;
&lt;li&gt;Storage code that accumulates output in memory&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  3. Assert token and cost limits
&lt;/h3&gt;

&lt;p&gt;For a non-streaming &lt;code&gt;generateContent&lt;/code&gt; request, assert that:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;candidatesTokenCount + thoughtsTokenCount &amp;lt;= configured output cap
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Also compute the request cost and assert that it remains below your per-request ceiling at Argon pricing.&lt;/p&gt;

&lt;p&gt;The &lt;a href="http://apidog.com/blog/gemini-4-argon-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon API guide&lt;/a&gt; includes a ready-made cost script.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is 1M Gemini 4 Argon’s context window?
&lt;/h3&gt;

&lt;p&gt;No. It is the stated output limit per response, up from 64K. Google has not published Argon’s input window.&lt;/p&gt;

&lt;h3&gt;
  
  
  How much does a 1M-token Argon response cost?
&lt;/h3&gt;

&lt;p&gt;Output costs $10 at intro rates and $20 at standard rates, plus input costs. See &lt;a href="http://apidog.com/blog/gemini-4-argon-pricing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon pricing&lt;/a&gt; for more scenarios.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can I generate a 1M-token response today?
&lt;/h3&gt;

&lt;p&gt;Not unless your organization is in the Fairwind cohort with Argon access. Gemini 3.8 Flash and 3.1 Pro Preview cap output at 65,536 tokens, while Vals AI lists 262K maximum output for the Argon configuration it tested.&lt;/p&gt;

&lt;h3&gt;
  
  
  Do I have to stream long Argon responses?
&lt;/h3&gt;

&lt;p&gt;Google has not published Argon-specific streaming guidance. However, a non-streaming request that runs for minutes is exposed to every idle timeout in your stack. Stream the response or use background execution through the Interactions API.&lt;/p&gt;

&lt;h3&gt;
  
  
  How does Argon compare with GPT-6 Astra and Claude Opus 5.5?
&lt;/h3&gt;

&lt;p&gt;Both cap synchronous output at 128K. Anthropic allows 300K on Batch with a beta header. Argon’s stated 1M output limit is about 8× higher.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your next step
&lt;/h2&gt;

&lt;p&gt;Add streaming and an explicit output cap to your Gemini client now, using Gemini 3.8 Flash. Run the client against the long mock stream until no component in your stack cuts the response off.&lt;/p&gt;

&lt;p&gt;When Argon’s model ID becomes available, update &lt;code&gt;GEMINI_MODEL&lt;/code&gt; and rerun the same tests in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Gemini 4 Argon API: What's Confirmed, What It Will Cost, and How to Get Your Code Ready</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:54:14 +0000</pubDate>
      <link>https://dev.to/hassann/gemini-4-argon-api-whats-confirmed-what-it-will-cost-and-how-to-get-your-code-ready-25m3</link>
      <guid>https://dev.to/hassann/gemini-4-argon-api-whats-confirmed-what-it-will-cost-and-how-to-get-your-code-ready-25m3</guid>
      <description>&lt;p&gt;There is no public Gemini 4 Argon API yet, and Google has not published a model ID. Google announced Argon on September 30, 2026, and it is currently available only to Fairwind Program defenders: a set of Fairwind partners using it as a managed model in Gemini Enterprise. &lt;a href="https://artificialanalysis.ai/models/gemini-4-argon/providers" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt; lists Google AI Studio as Argon’s API provider, but provides no speed or latency data. That suggests allowlisted pre-release access rather than an endpoint you can sign up for. When the wider rollout begins, Google says it will start “with paid API customers and Google AI Ultra subscribers,” so a paid Gemini API key puts you first in line.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;That gap between announcement and access is useful preparation time. This guide separates confirmed Argon API details from unknowns, then shows how to build against Gemini 3.8 Flash today so moving to Argon is a one-variable change. You will:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Send Interactions API and &lt;code&gt;generateContent&lt;/code&gt; requests.&lt;/li&gt;
&lt;li&gt;Detect availability with &lt;code&gt;models.list&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Mock responses in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Add a per-request cost ceiling using Argon pricing.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For background, start with &lt;a href="http://apidog.com/blog/what-is-gemini-4-argon?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what Gemini 4 Argon is&lt;/a&gt;. For rollout timing, see the &lt;a href="http://apidog.com/blog/gemini-4-argon-release-date?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon release date guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What’s confirmed about the Gemini 4 Argon API, and what isn’t
&lt;/h2&gt;

&lt;p&gt;Google’s &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/" rel="noopener noreferrer"&gt;launch post&lt;/a&gt; confirms prices and an output limit. Most integration details are still unpublished.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;Detail&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input price&lt;/td&gt;
&lt;td&gt;Confirmed&lt;/td&gt;
&lt;td&gt;$2 per 1M tokens during the intro period; $4 after&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output price&lt;/td&gt;
&lt;td&gt;Confirmed&lt;/td&gt;
&lt;td&gt;$10 per 1M tokens during the intro period; $20 after&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cached input&lt;/td&gt;
&lt;td&gt;Confirmed&lt;/td&gt;
&lt;td&gt;95% off input: $0.10 intro, $0.20 standard&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output limit&lt;/td&gt;
&lt;td&gt;Confirmed&lt;/td&gt;
&lt;td&gt;1M tokens, up from 64K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;API surface&lt;/td&gt;
&lt;td&gt;Signaled&lt;/td&gt;
&lt;td&gt;Google docs say all new models launch on the Interactions API&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model ID&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;Absent from the models page, pricing page, and changelog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input context window&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;Google’s long-context eval used prompts up to 1M tokens, but that is not a specification&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Thinking levels&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;Evals used the “highest thinking settings”; names and defaults are unknown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rate limits&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;No tiers announced&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Batch support&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;No batch or Flex pricing announced&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Long-prompt tier&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;Gemini 3.1 Pro charges more above 200K tokens; Argon’s rule is unknown&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Intro period length&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;No end date for the $2/$10 pricing&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two details affect implementation:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The API surface expectation comes from the &lt;a href="https://ai.google.dev/gemini-api/docs/interactions" rel="noopener noreferrer"&gt;Interactions API docs&lt;/a&gt;, which state that new models “will launch on the Interactions API.” Expect Argon there first, but do not assume &lt;code&gt;generateContent&lt;/code&gt; support until Google confirms it.&lt;/li&gt;
&lt;li&gt;Google states a 1M-token output limit, but Vals AI lists a 262K maximum output for the configuration it tested. Check the API response on launch day instead of assuming every endpoint exposes the full million-token limit.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;a href="http://apidog.com/blog/gemini-4-argon-1m-output-tokens?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;1M output tokens guide&lt;/a&gt; covers the impact on streaming, timeouts, and storage.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fassets.apidog.com%2Fblog-next%2F2026%2F10%2Fimage-2.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fassets.apidog.com%2Fblog-next%2F2026%2F10%2Fimage-2.png" alt="" width="800" height="802"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;At introductory rates, a request with a 20,000-token prompt and 5,000 output tokens costs:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;20,000 × $2 / 1M = $0.04 input
 5,000 × $10 / 1M = $0.05 output
--------------------------------
Total                 = $0.09
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;At standard $4/$20 rates, that same request costs $0.18. Argon’s introductory output price ($10 per 1M tokens) is below Gemini 3.1 Pro Preview’s $12, while its standard output price matches Claude Opus 5.5. For more scenarios, including cached input and a maximum-size output, see the &lt;a href="http://apidog.com/blog/gemini-4-argon-pricing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon pricing breakdown&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Don’t hard-code a model ID you found online
&lt;/h2&gt;

&lt;p&gt;Search results may show Argon-like model IDs from benchmark sites, aggregators, or open-source pull requests. Treat them as placeholders.&lt;/p&gt;

&lt;p&gt;Google has not confirmed a public model ID. A guessed ID can cause:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A production 404 on deployment day.&lt;/li&gt;
&lt;li&gt;A silent fallback if your SDK or wrapper catches errors broadly.&lt;/li&gt;
&lt;li&gt;Incorrect assumptions about supported methods or token limits.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Keep the model name in configuration:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nb"&gt;export &lt;/span&gt;&lt;span class="nv"&gt;GEMINI_MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;gemini-3.8-flash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Every example below reads &lt;code&gt;GEMINI_MODEL&lt;/code&gt; and defaults to &lt;code&gt;gemini-3.8-flash&lt;/code&gt;, a stable model available today. Once Google publishes Argon’s real ID, update the environment variable rather than modifying application code.&lt;/p&gt;

&lt;h2&gt;
  
  
  Get your code ready on Gemini 3.8 Flash
&lt;/h2&gt;

&lt;p&gt;Gemini 3.8 Flash (&lt;code&gt;gemini-3.8-flash&lt;/code&gt;) uses the endpoints, headers, and response shapes Argon is expected to use. Its pricing is $0.75 input and $3.75 output per 1M tokens through December 31, 2026.&lt;/p&gt;

&lt;p&gt;Store your API key in &lt;code&gt;GEMINI_API_KEY&lt;/code&gt;; never commit it to source control. See the &lt;a href="http://apidog.com/blog/how-to-use-gemini-3-8-flash-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.8 Flash API guide&lt;/a&gt; for setup details.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 1: Send an Interactions API request
&lt;/h3&gt;

&lt;p&gt;The Interactions API has been generally available since June 2026 and is Google’s primary API surface for new models.&lt;/p&gt;

&lt;p&gt;Keep both the model and thinking level configurable because Argon’s thinking-level names are not published.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;GEMINI_MODEL&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;gemini&lt;/span&gt;&lt;span class="p"&gt;-3.8-flash&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;
&lt;span class="nv"&gt;THINKING&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;GEMINI_THINKING&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;medium&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="nt"&gt;-X&lt;/span&gt; POST &lt;span class="s2"&gt;"https://generativelanguage.googleapis.com/v1beta/interactions"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-goog-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$GEMINI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s2"&gt;"{
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;model&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;MODEL&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;input&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;List the retry rules a REST client should follow for HTTP 429.&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;,
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;generation_config&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: {&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;thinking_level&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;THINKING&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;}
  }"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use the Python SDK with &lt;code&gt;google-genai&lt;/code&gt; 2.3.0 or later:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="c1"&gt;# pip install "google-genai&amp;gt;=2.3.0"
&lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;from&lt;/span&gt; &lt;span class="n"&gt;google&lt;/span&gt; &lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;genai&lt;/span&gt;

&lt;span class="n"&gt;MODEL&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GEMINI_MODEL&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;gemini-3.8-flash&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;span class="n"&gt;THINKING&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GEMINI_THINKING&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;medium&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;genai&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nc"&gt;Client&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;  &lt;span class="c1"&gt;# Reads GEMINI_API_KEY from the environment
&lt;/span&gt;
&lt;span class="n"&gt;interaction&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;interactions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;MODEL&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="nb"&gt;input&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;List the retry rules a REST client should follow for HTTP 429.&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;generation_config&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;thinking_level&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;THINKING&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="n"&gt;interaction&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;output_text&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For Gemini 3.8 Flash, valid thinking levels are:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;low&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;medium&lt;/code&gt; — default&lt;/li&gt;
&lt;li&gt;&lt;code&gt;high&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Using &lt;code&gt;minimal&lt;/code&gt; returns HTTP 400. See the &lt;a href="http://apidog.com/blog/gemini-3-8-flash-thinking-levels?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;thinking levels guide&lt;/a&gt; for trade-offs.&lt;/p&gt;

&lt;p&gt;For multi-turn interactions, pass the prior response ID as &lt;code&gt;previous_interaction_id&lt;/code&gt;. To opt out of server-side storage, send &lt;code&gt;store: false&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 2: Keep the &lt;code&gt;generateContent&lt;/code&gt; path working
&lt;/h3&gt;

&lt;p&gt;Existing applications may use the legacy &lt;code&gt;generateContent&lt;/code&gt; endpoint, which Google says remains fully supported.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://generativelanguage.googleapis.com/v1beta/models/&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;MODEL&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;:generateContent"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-goog-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$GEMINI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s2"&gt;"{
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;contents&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: [{
      &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;parts&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: [{
        &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;text&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;List the retry rules a REST client should follow for HTTP 429.&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;
      }]
    }],
    &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;generationConfig&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: {
      &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;thinkingConfig&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: {
        &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;thinkingLevel&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;: &lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;THINKING&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="se"&gt;\"&lt;/span&gt;&lt;span class="s2"&gt;
      }
    }
  }"&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The thinking configuration differs between APIs:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;API&lt;/th&gt;
&lt;th&gt;Thinking field&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Interactions&lt;/td&gt;
&lt;td&gt;&lt;code&gt;generation_config.thinking_level&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;code&gt;generateContent&lt;/code&gt;&lt;/td&gt;
&lt;td&gt;&lt;code&gt;generationConfig.thinkingConfig.thinkingLevel&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If your client supports both paths, route new models through Interactions by default and retain &lt;code&gt;generateContent&lt;/code&gt; as a fallback. Apply the same care to tool definitions; see &lt;a href="http://apidog.com/blog/gemini-3-8-flash-function-calling?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.8 Flash function calling&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 3: Schedule a go-live check with &lt;code&gt;models.list&lt;/code&gt;
&lt;/h3&gt;

&lt;p&gt;Do not manually refresh changelogs. Poll the API your production application will use.&lt;/p&gt;

&lt;p&gt;The models endpoint returns available models, token limits, and supported methods:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight python"&gt;&lt;code&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;sys&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;requests&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
    &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;https://generativelanguage.googleapis.com/v1beta/models&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="n"&gt;params&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;pageSize&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;
    &lt;span class="n"&gt;headers&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;x-goog-api-key&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="n"&gt;os&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;environ&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;GEMINI_API_KEY&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;]},&lt;/span&gt;
    &lt;span class="n"&gt;timeout&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="mi"&gt;30&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;raise_for_status&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;

&lt;span class="n"&gt;hits&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;
    &lt;span class="n"&gt;model&lt;/span&gt;
    &lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;resp&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;models&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;[])&lt;/span&gt;
    &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;argon&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;].&lt;/span&gt;&lt;span class="nf"&gt;lower&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt;
&lt;span class="p"&gt;]&lt;/span&gt;

&lt;span class="k"&gt;for&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt; &lt;span class="ow"&gt;in&lt;/span&gt; &lt;span class="n"&gt;hits&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;
    &lt;span class="nf"&gt;print&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;name&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;in:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;inputTokenLimit&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;out:&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;outputTokenLimit&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
        &lt;span class="n"&gt;model&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="s"&gt;supportedGenerationMethods&lt;/span&gt;&lt;span class="sh"&gt;"&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt;
    &lt;span class="p"&gt;)&lt;/span&gt;

&lt;span class="n"&gt;sys&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;exit&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt; &lt;span class="k"&gt;if&lt;/span&gt; &lt;span class="n"&gt;hits&lt;/span&gt; &lt;span class="k"&gt;else&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;  &lt;span class="c1"&gt;# A failed run becomes the alert.
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run this hourly through cron or a CI schedule using the same API key as your production app.&lt;/p&gt;

&lt;p&gt;When this call was run on October 1, 2026 with a paid-project key, it returned 61 models, no &lt;code&gt;nextPageToken&lt;/code&gt;, and no model name containing &lt;code&gt;argon&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;When the script finds Argon, inspect:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The real model ID.&lt;/li&gt;
&lt;li&gt;Whether &lt;code&gt;outputTokenLimit&lt;/code&gt; is the full 1M tokens.&lt;/li&gt;
&lt;li&gt;Which generation methods are supported.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Step 4: Mock the response before you have access
&lt;/h3&gt;

&lt;p&gt;You can build Argon client logic before Argon access is available.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Save the Step 2 &lt;code&gt;generateContent&lt;/code&gt; request in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;Create &lt;code&gt;GEMINI_API_KEY&lt;/code&gt; and &lt;code&gt;GEMINI_MODEL&lt;/code&gt; environment variables.&lt;/li&gt;
&lt;li&gt;Send the request once against Gemini 3.8 Flash.&lt;/li&gt;
&lt;li&gt;Save the actual response as an endpoint response example.&lt;/li&gt;
&lt;li&gt;Set the project mock behavior to &lt;strong&gt;Response example first&lt;/strong&gt;:

&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;Project Settings&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Feature Settings&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Mock Settings&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Apidog’s default Smart Mock generates values from the schema. Using &lt;strong&gt;Response example first&lt;/strong&gt; makes the mock URL return your saved Gemini response instead. Your frontend, queue workers, and parsers can then run against the mock without consuming tokens.&lt;/p&gt;

&lt;p&gt;This is a current Gemini response schema mock, not an Argon-specific schema. Google has not published an Argon schema. It is still a practical target because Argon is expected to use the same API surface.&lt;/p&gt;

&lt;p&gt;To simulate a large Argon response, edit the example’s &lt;code&gt;usageMetadata&lt;/code&gt; values and verify that billing, alerting, queueing, and storage logic behave correctly.&lt;/p&gt;

&lt;h3&gt;
  
  
  Step 5: Assert on &lt;code&gt;usageMetadata&lt;/code&gt; and enforce a cost ceiling
&lt;/h3&gt;

&lt;p&gt;Every &lt;code&gt;generateContent&lt;/code&gt; response includes &lt;code&gt;usageMetadata&lt;/code&gt;. Add an Apidog post-processor that prices each response at Argon’s standard rates.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// Apidog post-processor: price this response at Gemini 4 Argon's standard rates&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;u&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;().&lt;/span&gt;&lt;span class="nx"&gt;usageMetadata&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;IN&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nx"&gt;e6&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;   &lt;span class="c1"&gt;// USD per input token after the intro period&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;OUT&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;20&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="nx"&gt;e6&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// USD per output token&lt;/span&gt;

&lt;span class="c1"&gt;// Thinking bills as output on current Gemini models.&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;out&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;candidatesTokenCount&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt;
  &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;thoughtsTokenCount&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;cost&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;u&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;promptTokenCount&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;IN&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="nx"&gt;out&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="nx"&gt;OUT&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;usageMetadata is present&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;u&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;be&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;an&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;object&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;cost under $0.10 at Argon rates&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;cost&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;be&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;below&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mf"&gt;0.10&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Choose the ceiling per request. A $0.10 limit is reasonable for the short prompt used in this example.&lt;/p&gt;

&lt;p&gt;Google has not stated how Argon bills thinking tokens. This script assumes the current Gemini billing behavior, where thinking tokens are priced as output. Keep the same saved request and run it against Gemini 3.8 Flash now, then Argon later. The token and cost differences become your regression comparison.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is the Gemini 4 Argon API available?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not publicly. Argon is currently rolling out only to Fairwind Program partners through Gemini Enterprise. Paid API customers and Google AI Ultra subscribers are next, but Google has not provided a date.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the Gemini 4 Argon model ID?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Google has not published one. Third-party strings are placeholders, so keep the model name in an environment variable and monitor &lt;code&gt;models.list&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How much will the Gemini 4 Argon API cost?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It costs $2 input and $10 output per 1M tokens during an introductory period of unknown length, then $4 input and $20 output per 1M tokens. Cached input is discounted by 95%. See &lt;a href="http://apidog.com/blog/gemini-4-argon-pricing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon pricing&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Will Argon work with &lt;code&gt;generateContent&lt;/code&gt;?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Google has not said. Its documentation says new models launch on the Interactions API, so build for Interactions first and treat &lt;code&gt;generateContent&lt;/code&gt; as a fallback.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Google AI Ultra give me API access?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. AI Ultra is a consumer subscription, not an API key. Google says API access begins with paid API customers.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your next step
&lt;/h2&gt;

&lt;p&gt;Set &lt;code&gt;GEMINI_MODEL&lt;/code&gt; in every environment, schedule the &lt;code&gt;models.list&lt;/code&gt; check, and save both Interactions and &lt;code&gt;generateContent&lt;/code&gt; requests with the cost assertion attached.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; to keep requests, mocks, and tests in one workspace. When Google publishes Argon’s model ID, update one environment variable.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>What Is Gemini 4 Argon?</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:53:05 +0000</pubDate>
      <link>https://dev.to/hassann/what-is-gemini-4-argon-18b5</link>
      <guid>https://dev.to/hassann/what-is-gemini-4-argon-18b5</guid>
      <description>&lt;p&gt;Gemini 4 Argon is Google’s new frontier model, announced on September 30, 2026. It is currently available only to Fairwind Program defenders: a vetted group of cyber defense partners. Introductory pricing is $2 per million input tokens and $10 per million output tokens, increasing to $4 and $20 respectively after the introductory period. Argon raises the output limit to 1M tokens per response. There is no public API, published model ID, or release date yet.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This guide covers what Google has and has not confirmed, where Argon fits in the Gemini lineup, expected access, pricing, benchmarks, and what you can build now. Choosing a model today? Read our &lt;a href="http://apidog.com/blog/gemini-4-argon-vs-gpt-6-astra-vs-claude-opus-5-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5 comparison&lt;/a&gt;. If you test model APIs in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;, the implementation section shows how to make Argon a one-variable model switch.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Google has confirmed vs. what it has not
&lt;/h2&gt;

&lt;p&gt;Early coverage often mixes Google statements with third-party estimates. The table below separates confirmed information from unknowns using Google’s &lt;a href="https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/" rel="noopener noreferrer"&gt;launch post&lt;/a&gt; and &lt;a href="https://deepmind.google/models/evals-methodology/gemini-4-argon" rel="noopener noreferrer"&gt;evals methodology&lt;/a&gt;.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;th&gt;What we know&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Announcement&lt;/td&gt;
&lt;td&gt;Confirmed&lt;/td&gt;
&lt;td&gt;September 30, 2026, by Koray Kavukcuoglu of Google DeepMind&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Access today&lt;/td&gt;
&lt;td&gt;Confirmed&lt;/td&gt;
&lt;td&gt;A set of Fairwind Program partners&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Next in line&lt;/td&gt;
&lt;td&gt;Confirmed, undated&lt;/td&gt;
&lt;td&gt;Paid API customers and Google AI Ultra subscribers&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Price&lt;/td&gt;
&lt;td&gt;Confirmed&lt;/td&gt;
&lt;td&gt;$2/$10 intro, then $4/$20 per 1M tokens; cached input is 95% off&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output limit&lt;/td&gt;
&lt;td&gt;Confirmed&lt;/td&gt;
&lt;td&gt;1M tokens, up from 64K&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Benchmarks&lt;/td&gt;
&lt;td&gt;Confirmed, Google-reported&lt;/td&gt;
&lt;td&gt;A 19-row table plus a methodology PDF&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model ID&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;Strings online are third-party placeholders&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Input context window&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;Google's long-context eval used prompts up to 1M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Knowledge cutoff&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;Nothing from Google&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Intro period length&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;No date for the switch to $4/$20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Rate limits, free tier&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;Nothing from Google&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Release date&lt;/td&gt;
&lt;td&gt;Not published&lt;/td&gt;
&lt;td&gt;“As soon as possible”&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;If a site lists an Argon context window, model string, or speed figure, treat it as unconfirmed until Google publishes it in official documentation.&lt;/p&gt;

&lt;h2&gt;
  
  
  What Gemini 4 Argon is and where it sits
&lt;/h2&gt;

&lt;p&gt;Argon is the only Gemini 4 model announced so far and Google’s new top-tier Gemini model. Google calls it “our new frontier model,” built “to sustain deep reasoning across complex, long-horizon workflows.” It claims frontier performance in software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense.&lt;/p&gt;

&lt;p&gt;A Google spokesperson told Reuters that Argon is larger than the company’s previous “Pro” models.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbqxf7nmshquu944yze0f.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fbqxf7nmshquu944yze0f.png" alt="Gemini 4 Argon" width="800" height="802"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;Until now, the top Gemini API model was &lt;code&gt;gemini-3.1-pro-preview&lt;/code&gt;, which remains in Preview:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Input:&lt;/strong&gt; 1,048,576 tokens&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Output:&lt;/strong&gt; 65,536 tokens&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Pricing:&lt;/strong&gt; $2 input / $12 output per million tokens for prompts up to 200K&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Argon’s introductory output price of $10 per million tokens undercuts that output rate. Google has not said whether Argon will have a separate long-prompt tier like 3.1 Pro or whether more Gemini 4 models will follow.&lt;/p&gt;

&lt;p&gt;Google says it already uses Argon internally for coding, research, and writing. It reports that Argon agents freed more than 300 TiB of fleet memory and are migrating C and C++ code to Rust, including up to 800K+ lines for the Fuchsia Zircon kernel. Those rewrites remain under audit before production use.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who can use Gemini 4 Argon today, and who is next
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Today: a subset of Fairwind partners.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Fairwind is Google’s limited-access cyber defense program with more than 650 partners. However, only “a set of Fairwind Program partners” can currently access Argon, either directly or through CodeMender, Google’s code security agent.&lt;/p&gt;

&lt;p&gt;Direct access is provided as a managed model on Gemini Enterprise with zero data retention. Partners must commit to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Phishing-resistant MFA&lt;/li&gt;
&lt;li&gt;Access restricted to internal security, incident-response, or penetration-testing teams&lt;/li&gt;
&lt;li&gt;Per-employee usage tracking&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Governments, critical infrastructure operators, and core technology platforms can apply through the &lt;a href="https://deepmind.google/fairwind-program/" rel="noopener noreferrer"&gt;Fairwind Program page&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Next: paid API customers and Google AI Ultra subscribers.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Google says it will gather feedback from early testers while iterating on guardrails. It is also participating in the U.S. government’s voluntary pre-release access process. The rollout section is titled “Rolling out soon,” but Google provides no date.&lt;/p&gt;

&lt;p&gt;Track updates in our &lt;a href="http://apidog.com/blog/gemini-4-argon-release-date?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon release date and access guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Not yet: everyone else.&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;As of October 1, Argon does not appear on the &lt;a href="https://ai.google.dev/gemini-api/docs/models" rel="noopener noreferrer"&gt;Gemini API models page&lt;/a&gt;, pricing page, Gemini app, Antigravity, or OpenRouter.&lt;/p&gt;

&lt;p&gt;Google AI Ultra, starting at $99.99 per month in the US, is a consumer subscription rather than an API key. Google has not specified which Ultra tier will receive Argon first. There is no free access path today; see &lt;a href="http://apidog.com/blog/gemini-4-argon-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Is Gemini 4 Argon free?&lt;/a&gt; for available alternatives.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gemini 4 Argon pricing
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Token type&lt;/th&gt;
&lt;th&gt;Intro price, per 1M&lt;/th&gt;
&lt;th&gt;Price after intro, per 1M&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Input&lt;/td&gt;
&lt;td&gt;$2.00&lt;/td&gt;
&lt;td&gt;$4.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cached input, 95% off&lt;/td&gt;
&lt;td&gt;$0.10&lt;/td&gt;
&lt;td&gt;$0.20&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Output&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$20.00&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;One full 1M-token response, output only&lt;/td&gt;
&lt;td&gt;$10.00&lt;/td&gt;
&lt;td&gt;$20.00&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Google has not announced how long the introductory period will last.&lt;/p&gt;

&lt;p&gt;At standard rates, Argon costs the same as Claude Opus 5.5. At introductory rates, it matches GPT-6.1 Sol and Claude Sonnet 5.5. Cached-input prices are calculated from Google’s stated 95% discount.&lt;/p&gt;

&lt;p&gt;For scenario-based estimates, see &lt;a href="http://apidog.com/blog/gemini-4-argon-pricing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon pricing&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The 1M output limit
&lt;/h2&gt;

&lt;p&gt;Google says it is “significantly expanding the model’s output token limit to an industry-leading 1M tokens, up from the previous 64K tokens.”&lt;/p&gt;

&lt;p&gt;For developers, this means a single response can theoretically contain hundreds of thousands of generated tokens. Google’s argument is that this lets the model work through complex problems in one trajectory.&lt;/p&gt;

&lt;p&gt;For comparison:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Model&lt;/th&gt;
&lt;th&gt;Maximum synchronous output&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Gemini 4 Argon&lt;/td&gt;
&lt;td&gt;1M tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6 Astra&lt;/td&gt;
&lt;td&gt;128K tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Opus 5.5&lt;/td&gt;
&lt;td&gt;128K tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Claude Fable 5.1&lt;/td&gt;
&lt;td&gt;128K tokens&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Two implementation caveats matter:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;This is an output limit, not a published context-window specification.&lt;/strong&gt; Google has not published Argon’s input window.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Use the limit reported by your endpoint.&lt;/strong&gt; &lt;a href="https://www.vals.ai/models/google_gemini-4-argon" rel="noopener noreferrer"&gt;Vals AI&lt;/a&gt;, for example, lists a 262K maximum output for the configuration it tested.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;When Argon becomes available, set explicit client-side timeout and output-cost limits rather than assuming every request can safely generate 1M tokens. See the &lt;a href="http://apidog.com/blog/gemini-4-argon-1m-output-tokens?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;guide to Argon’s 1M output tokens&lt;/a&gt; for timeout, streaming, and cost-cap guidance.&lt;/p&gt;

&lt;h2&gt;
  
  
  Gemini 4 Argon benchmarks: the headline
&lt;/h2&gt;

&lt;p&gt;All benchmark figures below come from Google’s table. Argon’s scores are Google-reported, with some self-computed and others sourced from leaderboards such as Vals AI. Competitor scores are primarily vendor-reported or public-leaderboard results, and benchmark harnesses differ.&lt;/p&gt;

&lt;p&gt;By our recount, Argon leads 13 of 19 rows outright, ties one, and trails on five.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Benchmark&lt;/th&gt;
&lt;th&gt;Gemini 4 Argon&lt;/th&gt;
&lt;th&gt;GPT-6 Astra&lt;/th&gt;
&lt;th&gt;Claude Fable 5.1&lt;/th&gt;
&lt;th&gt;Claude Opus 5.5&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Vals Index&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;68.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;63.1%&lt;/td&gt;
&lt;td&gt;65.8%&lt;/td&gt;
&lt;td&gt;67.0%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;DeepSWE v1.1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;77.9%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;74.1%&lt;/td&gt;
&lt;td&gt;67.4%&lt;/td&gt;
&lt;td&gt;74.2%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GraphWalks 256K to 1M, F1&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;84.2%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;71.8%&lt;/td&gt;
&lt;td&gt;65.0%&lt;/td&gt;
&lt;td&gt;66.8%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;FrontierSWE v2&lt;/td&gt;
&lt;td&gt;55.0%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;65.5%&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;56.3%&lt;/td&gt;
&lt;td&gt;62.3%&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Terminal-bench 4.0&lt;/td&gt;
&lt;td&gt;57.4%&lt;/td&gt;
&lt;td&gt;58.2%&lt;/td&gt;
&lt;td&gt;57.9%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;66.4%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;PostTrainBench&lt;/td&gt;
&lt;td&gt;45.3%&lt;/td&gt;
&lt;td&gt;44.3%&lt;/td&gt;
&lt;td&gt;40.2%&lt;/td&gt;
&lt;td&gt;&lt;strong&gt;49.3%&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The rows Argon loses include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;FrontierSWE v2&lt;/li&gt;
&lt;li&gt;Terminal-Bench Science 0.1&lt;/li&gt;
&lt;li&gt;OSWorld-2.0&lt;/li&gt;
&lt;li&gt;Terminal-bench 4.0&lt;/li&gt;
&lt;li&gt;PostTrainBench&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Third-party evaluations place Argon at the frontier but do not show a clear overall lead. &lt;a href="https://artificialanalysis.ai/models/gemini-4-argon" rel="noopener noreferrer"&gt;Artificial Analysis&lt;/a&gt; scores Argon at 53 on its Intelligence Index, tied with GPT-6 Astra and Claude Fable 5.1, behind Claude Opus 5.5 at 58 and Claude Sonnet 5.5 at 56 using maximum settings. Vals ranks Argon first among 41 models on the Vals Index.&lt;/p&gt;

&lt;p&gt;For the full benchmark table and methodology notes, see &lt;a href="http://apidog.com/blog/gemini-4-argon-benchmarks?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon benchmarks&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8q0dk2sp7ebli2h9htti.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2F8q0dk2sp7ebli2h9htti.png" alt="Gemini 4 Argon benchmark results" width="800" height="450"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;h2&gt;
  
  
  Cyber defense in one paragraph
&lt;/h2&gt;

&lt;p&gt;Google says Argon can “autonomously find, validate, and patch critical software vulnerabilities.” For trusted defenders and internal teams, Google says it will release Argon without cyber guardrails so they can use its cybersecurity defense capabilities.&lt;/p&gt;

&lt;p&gt;On DeepMind’s &lt;a href="https://deepmind.google/models/gemini/cyber/" rel="noopener noreferrer"&gt;cyber page&lt;/a&gt;, Argon scores:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;85.8%&lt;/strong&gt; on real-world vulnerability discovery, compared with &lt;strong&gt;71.0%&lt;/strong&gt; for Gemini 3.8 Flash Cyber&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;68%&lt;/strong&gt; on CWE-bench v1, tied with Grok 4.7 and GPT-6 Astra&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;0.7%&lt;/strong&gt; Gray Swan indirect prompt-injection attack success rate at k=15, the lowest score on that chart&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For API implications, read &lt;a href="http://apidog.com/blog/gemini-4-argon-cyber-defense?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what Argon’s cyber release means for the APIs you run&lt;/a&gt; and the overview of &lt;a href="http://apidog.com/blog/what-is-gemini-3-8-flash-cyber?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.8 Flash Cyber&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What to do now as a developer
&lt;/h2&gt;

&lt;p&gt;Do not block your implementation on Argon. Build with a model you can call today, and isolate the model name in one environment variable.&lt;/p&gt;

&lt;p&gt;Use &lt;code&gt;gemini-3.8-flash&lt;/code&gt; as the current stand-in:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Stable model&lt;/li&gt;
&lt;li&gt;1,048,576 input tokens&lt;/li&gt;
&lt;li&gt;65,536 output tokens&lt;/li&gt;
&lt;li&gt;Free tier&lt;/li&gt;
&lt;li&gt;$0.75 input / $3.75 output per million tokens through December 31, 2026&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Google’s &lt;a href="https://ai.google.dev/gemini-api/docs/interactions" rel="noopener noreferrer"&gt;Interactions API documentation&lt;/a&gt; says that all new models, multimodal capabilities, tools, and agentic features will launch on the Interactions API. Build against that API now, then replace the model variable when Google publishes Argon’s model ID.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;MODEL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;GEMINI_MODEL&lt;/span&gt;&lt;span class="k"&gt;:-&lt;/span&gt;&lt;span class="nv"&gt;gemini&lt;/span&gt;&lt;span class="p"&gt;-3.8-flash&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;

curl &lt;span class="nt"&gt;-s&lt;/span&gt; &lt;span class="s2"&gt;"https://generativelanguage.googleapis.com/v1beta/interactions"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"x-goog-api-key: &lt;/span&gt;&lt;span class="nv"&gt;$GEMINI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "'&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="nv"&gt;$MODEL&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt;&lt;span class="s1"&gt;'",
    "input": "List the breaking changes in this OpenAPI diff.",
    "generation_config": {"thinking_level": "medium"}
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Make the model switch testable
&lt;/h3&gt;

&lt;p&gt;In &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;, save this call as a reusable request and define these environment variables:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;GEMINI_API_KEY=your_api_key
GEMINI_MODEL=gemini-3.8-flash
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then add response assertions for usage fields:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;usage.total_output_tokens&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;usage.total_thought_tokens&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Thinking tokens bill as output tokens, so use both fields when enforcing a token budget. Keep the Gemini 3.8 Flash response as a baseline. Once Argon appears in the official models list:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Change &lt;code&gt;GEMINI_MODEL&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Re-run the same request.&lt;/li&gt;
&lt;li&gt;Compare latency, output quality, and token usage.&lt;/li&gt;
&lt;li&gt;Keep the token assertion enabled to catch unexpectedly large responses.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For request shapes, see the &lt;a href="http://apidog.com/blog/gemini-4-argon-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 4 Argon API guide&lt;/a&gt;. To decide whether waiting is worthwhile, read &lt;a href="http://apidog.com/blog/gemini-4-argon-vs-gemini-3-8-flash?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Argon vs. Gemini 3.8 Flash&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Can I use Gemini 4 Argon in the Gemini API?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Not yet. Only a set of Fairwind partners have access, and the public models list does not include it.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the Gemini 4 Argon model ID?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Google has not published one. Strings currently in circulation, including evaluator slugs, are third-party placeholders. Do not hard-code them.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;When is the Gemini 4 Argon release for developers?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Google has not provided a date. Paid API customers and Google AI Ultra subscribers are expected to come first, “as soon as possible.”&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Is Gemini 4 Argon better than GPT-6 Astra and Claude Opus 5.5?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;It leads most rows in Google’s table, but it ties GPT-6 Astra on Artificial Analysis and trails Claude Opus 5.5 there. For more detail, see &lt;a href="http://apidog.com/blog/what-is-claude-opus-5-5?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what is Claude Opus 5.5&lt;/a&gt; and the &lt;a href="http://apidog.com/blog/gpt-6-astra-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GPT-6 Astra API guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Gemini 4 Argon have a 1M context window?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Google has not said. The published 1M figure is an output limit. Google’s long-context evaluation used prompts up to 1M tokens, but that is not a published context-window specification.&lt;/p&gt;

&lt;h2&gt;
  
  
  Get ready before Argon opens
&lt;/h2&gt;

&lt;p&gt;Argon is announced, priced, and benchmarked, but you cannot call it through the public API yet.&lt;/p&gt;

&lt;p&gt;Build with Gemini 3.8 Flash using the &lt;a href="http://apidog.com/blog/how-to-use-gemini-3-8-flash-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Gemini 3.8 Flash API guide&lt;/a&gt;, keep the model ID in one environment variable, and add a token-usage assertion to every request.&lt;/p&gt;

&lt;p&gt;To prepare your request collection, environment variables, and baseline responses for the future model swap, &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;download Apidog&lt;/a&gt; and configure it now.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Codex after DevDay 2026: cloud tasks, a voice CLI, /agents, Code Review, and Security Cloud</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:47:33 +0000</pubDate>
      <link>https://dev.to/hassann/codex-after-devday-2026-cloud-tasks-a-voice-cli-agents-code-review-and-security-cloud-45nl</link>
      <guid>https://dev.to/hassann/codex-after-devday-2026-cloud-tasks-a-voice-cli-agents-code-review-and-security-cloud-45nl</guid>
      <description>&lt;p&gt;At OpenAI DevDay 2026 on September 29, Codex received four updates: Codex cloud with reusable environments (Plus, Pro, Business, Healthcare, Edu, Enterprise); a refreshed Codex CLI with voice, an &lt;code&gt;/agents&lt;/code&gt; view, prompt editing, resume, and worktrees (all plans); Code Review in the ChatGPT desktop app for GitHub PRs and GitLab MRs with automatic cloud reviews (all plans); and Codex Security Cloud for scheduled repository scans and fix preparation (Pro, Business, Enterprise, Edu). The same day, Codex CLI 0.159.1 made GPT-6.1 Sol the default model.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This guide covers what each change does, which commands you’ll use, and where the gaps are. For the rest of the event, see the &lt;a href="http://apidog.com/blog/openai-devday-2026?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DevDay 2026 hub for API developers&lt;/a&gt;. One gap matters for API teams: Codex writes and reviews code, but its reviews do not call your running API. That is where contract tests in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; fit in—the last section shows the CI wiring.&lt;/p&gt;

&lt;h2&gt;
  
  
  What changed in Codex at DevDay 2026
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Update&lt;/th&gt;
&lt;th&gt;What it does&lt;/th&gt;
&lt;th&gt;Where you use it&lt;/th&gt;
&lt;th&gt;Plans&lt;/th&gt;
&lt;th&gt;Status&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Codex cloud environments&lt;/td&gt;
&lt;td&gt;Prepare a setup once and reuse it in isolated workspaces&lt;/td&gt;
&lt;td&gt;Web, desktop app, &lt;code&gt;codex cloud&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Plus, Pro, Business, Healthcare, Edu, Enterprise&lt;/td&gt;
&lt;td&gt;Available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Refreshed Codex CLI&lt;/td&gt;
&lt;td&gt;Voice, &lt;code&gt;/agents&lt;/code&gt; view, prompt edit, resume, worktrees&lt;/td&gt;
&lt;td&gt;Terminal&lt;/td&gt;
&lt;td&gt;All plans&lt;/td&gt;
&lt;td&gt;Available&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Code Review&lt;/td&gt;
&lt;td&gt;PR and MR review pane, automatic cloud reviews&lt;/td&gt;
&lt;td&gt;ChatGPT desktop app, GitHub, GitLab&lt;/td&gt;
&lt;td&gt;All plans&lt;/td&gt;
&lt;td&gt;GitHub GA; GitLab MRs in preview&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Codex Security Cloud&lt;/td&gt;
&lt;td&gt;Scheduled scans, deduplicated findings, fix preparation&lt;/td&gt;
&lt;td&gt;Codex cloud, via a plugin&lt;/td&gt;
&lt;td&gt;Pro, Business, Enterprise, Edu&lt;/td&gt;
&lt;td&gt;Research preview&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;GPT-6.1 Sol as default&lt;/td&gt;
&lt;td&gt;Default in the CLI’s bundled model catalog&lt;/td&gt;
&lt;td&gt;CLI, also selectable in Codex and ChatGPT Work&lt;/td&gt;
&lt;td&gt;Plus, Pro, Business, Enterprise, Edu; admin-enabled on Enterprise and Edu&lt;/td&gt;
&lt;td&gt;Available&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sources: OpenAI’s &lt;a href="https://openai.com/index/devday-2026-recap/" rel="noopener noreferrer"&gt;DevDay 2026 recap&lt;/a&gt;, the &lt;a href="https://learn.chatgpt.com/docs/whats-new" rel="noopener noreferrer"&gt;Codex what’s new page&lt;/a&gt;, and the Codex documentation linked below.&lt;/p&gt;

&lt;h2&gt;
  
  
  Codex cloud: reusable environments
&lt;/h2&gt;

&lt;p&gt;Codex cloud runs coding tasks in the cloud so work can continue while your computer is asleep. The DevDay change is the environment model. According to the &lt;a href="https://learn.chatgpt.com/docs/environments/cloud-environments" rel="noopener noreferrer"&gt;cloud environments documentation&lt;/a&gt;, an environment is a reusable setup containing repositories, dependencies, tools, and access settings. Each task gets its own isolated workspace created from the published environment version.&lt;/p&gt;

&lt;h3&gt;
  
  
  Create a reusable environment
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;Start a new Codex task.&lt;/li&gt;
&lt;li&gt;Choose &lt;strong&gt;Work in &amp;gt; Cloud&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;Create environment&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Choose the GitHub repositories to check out.&lt;/li&gt;
&lt;li&gt;Let Codex inspect the repositories, install dependencies, and test the workflow.&lt;/li&gt;
&lt;li&gt;Provide details Codex cannot infer.&lt;/li&gt;
&lt;li&gt;Review the generated setup report.&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;Publish&lt;/strong&gt;.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Codex records the setup in two fields:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An &lt;strong&gt;install script&lt;/strong&gt; for dependencies and project setup.&lt;/li&gt;
&lt;li&gt;A &lt;strong&gt;start skill&lt;/strong&gt; that starts services and verifies they are ready.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Configure network access safely
&lt;/h3&gt;

&lt;p&gt;Three configuration details matter for API work:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Repository refresh&lt;/strong&gt; runs in the background and retains dependency caches, so new tasks do not repeat installation.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Internet access&lt;/strong&gt; is configured per environment: use a package-manager preset, custom domains, or unrestricted access.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Network secrets&lt;/strong&gt; keep credentials outside the sandbox. Programs receive a placeholder, while a proxy substitutes the real value only for approved HTTPS destinations.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;If a cloud task must call a staging API, add the staging host to the allowed domains and save its token as a network secret—not as a plain environment variable.&lt;/p&gt;

&lt;p&gt;From a terminal, use &lt;code&gt;codex cloud&lt;/code&gt; to hand work to Codex cloud and list recent cloud chats.&lt;/p&gt;

&lt;h2&gt;
  
  
  The refreshed Codex CLI
&lt;/h2&gt;

&lt;p&gt;Install the CLI with the script from the &lt;a href="https://learn.chatgpt.com/docs/codex/cli" rel="noopener noreferrer"&gt;Codex CLI documentation&lt;/a&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;-fsSL&lt;/span&gt; https://chatgpt.com/codex/install.sh | sh
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://learn.chatgpt.com/docs/changelog" rel="noopener noreferrer"&gt;Codex changelog&lt;/a&gt; also lists npm releases:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;npm &lt;span class="nb"&gt;install&lt;/span&gt; &lt;span class="nt"&gt;-g&lt;/span&gt; @openai/codex
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;CLI version 0.159.1, released September 29, added GPT-6.1 Sol as the default model in the bundled catalog. The same changelog entry also covers the Amazon Bedrock Mantle and Runtime catalogs.&lt;/p&gt;

&lt;p&gt;The CLI documentation uses &lt;code&gt;gpt-6.1-sol medium&lt;/code&gt; in its sample session. According to the &lt;a href="https://learn.chatgpt.com/docs/models" rel="noopener noreferrer"&gt;Codex models page&lt;/a&gt;, Free and Go are not included at launch, while Enterprise and Edu require an administrator to enable the model.&lt;/p&gt;

&lt;p&gt;To change models:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;/model
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Or start a session with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;codex &lt;span class="nt"&gt;--model&lt;/span&gt; gpt-6-sol
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For more context, see &lt;a href="http://apidog.com/blog/what-is-gpt-6-1-sol?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;What is GPT-6.1 Sol?&lt;/a&gt;, including the change from GPT-6 Sol and the missing &lt;code&gt;none&lt;/code&gt; effort.&lt;/p&gt;

&lt;h3&gt;
  
  
  CLI feature-to-command map
&lt;/h3&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Feature&lt;/th&gt;
&lt;th&gt;How to use it&lt;/th&gt;
&lt;th&gt;Source&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Voice&lt;/td&gt;
&lt;td&gt;Voice conversations have been on by default since 0.156.0. Press &lt;code&gt;F8&lt;/code&gt; to toggle and use &lt;code&gt;/voice settings&lt;/code&gt; to select a voice.&lt;/td&gt;
&lt;td&gt;Changelog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;
&lt;code&gt;/agents&lt;/code&gt; view&lt;/td&gt;
&lt;td&gt;View your agents, filter tasks by status, and start worktree sessions.&lt;/td&gt;
&lt;td&gt;Recap, changelog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Prompt edit&lt;/td&gt;
&lt;td&gt;Edit an earlier prompt to revert the thread to that point.&lt;/td&gt;
&lt;td&gt;Changelog&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Resume&lt;/td&gt;
&lt;td&gt;Run &lt;code&gt;codex resume&lt;/code&gt; to return to a saved chat.&lt;/td&gt;
&lt;td&gt;CLI docs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Worktrees&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;/worktree&lt;/code&gt; to run a chat in a new Git worktree. Use &lt;code&gt;/fork&lt;/code&gt; to copy a chat into a new chat or worktree.&lt;/td&gt;
&lt;td&gt;Slash command reference&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;/model&lt;/code&gt; to select a model and reasoning effort.&lt;/td&gt;
&lt;td&gt;CLI docs&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Local review&lt;/td&gt;
&lt;td&gt;Use &lt;code&gt;/review&lt;/code&gt; to check uncommitted changes, a commit, or a base branch.&lt;/td&gt;
&lt;td&gt;CLI docs&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Worktree support has been enabled by default since 0.156.0. This makes parallel agents practical because each agent gets its own checkout instead of editing the same files.&lt;/p&gt;

&lt;p&gt;Run &lt;code&gt;/review&lt;/code&gt; before committing. It reports prioritized findings without modifying your working tree.&lt;/p&gt;

&lt;h2&gt;
  
  
  Code Review: desktop app, GitHub, and GitLab
&lt;/h2&gt;

&lt;p&gt;Code Review is a plugin in the ChatGPT desktop app. The &lt;a href="https://learn.chatgpt.com/docs/code-review" rel="noopener noreferrer"&gt;Code Review documentation&lt;/a&gt; says it displays a pull request’s description, changed files, comments, and checks so you can investigate issues before approving.&lt;/p&gt;

&lt;p&gt;GitHub review is generally available. GitLab merge request support is in preview.&lt;/p&gt;

&lt;h3&gt;
  
  
  Enable automatic cloud reviews
&lt;/h3&gt;

&lt;p&gt;Automatic cloud reviews require:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;The ChatGPT Codex connector.&lt;/li&gt;
&lt;li&gt;Repository permissions.&lt;/li&gt;
&lt;li&gt;A connected GitHub repository.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Review comments appear both on GitHub and in the Code Review pane. Cloud review does not require you to create or manage a cloud environment.&lt;/p&gt;

&lt;p&gt;According to the &lt;a href="https://learn.chatgpt.com/docs/third-party/github" rel="noopener noreferrer"&gt;GitHub integration guide&lt;/a&gt;, you can manually trigger a review by commenting:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;@codex review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;On GitHub, Codex flags only P0 and P1 issues.&lt;/p&gt;

&lt;h3&gt;
  
  
  Add repository-specific review rules
&lt;/h3&gt;

&lt;p&gt;Place a &lt;code&gt;## Code Review Rules&lt;/code&gt; section in &lt;code&gt;AGENTS.md&lt;/code&gt;. Put the file closest to the code it governs so the rules apply at the right scope.&lt;/p&gt;

&lt;p&gt;Keep deterministic checks in CI. Review rules do not replace tests, branch protections, or required approvals.&lt;/p&gt;

&lt;p&gt;For a deeper workflow walkthrough, see &lt;a href="http://apidog.com/blog/codex-code-review?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Codex for code review&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Codex Security Cloud
&lt;/h2&gt;

&lt;p&gt;Codex Security Cloud scans connected GitHub repositories in Codex cloud. OpenAI’s recap describes scheduled scans, deduplicated findings, and cloud-based fix preparation. It also includes Daybreak Blue models without a separate application. See the &lt;a href="http://apidog.com/blog/openai-daybreak-blue-vs-red?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Daybreak Blue vs Red explainer&lt;/a&gt; for details on those tiers.&lt;/p&gt;

&lt;p&gt;Codex Security Cloud is in research preview for Pro, Business, Enterprise, and Edu.&lt;/p&gt;

&lt;h3&gt;
  
  
  Set up a repository scan
&lt;/h3&gt;

&lt;p&gt;Follow the &lt;a href="https://learn.chatgpt.com/docs/security/setup" rel="noopener noreferrer"&gt;Security Cloud setup guide&lt;/a&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Install the Codex Security Cloud plugin from the ChatGPT plugin marketplace.&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;Connect GitHub&lt;/strong&gt; and grant access to the repositories you want to scan.&lt;/li&gt;
&lt;li&gt;Choose a repository and cloud environment.&lt;/li&gt;
&lt;li&gt;Set the target to &lt;strong&gt;Repository&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Start a scan.&lt;/li&gt;
&lt;li&gt;For ongoing coverage, select &lt;strong&gt;Commit changes&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Adjust the environment, days of history, or pause state under &lt;strong&gt;Monitoring settings&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;For a finding, select &lt;strong&gt;Fix with Codex&lt;/strong&gt; to draft a patch.&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;Create draft pull request&lt;/strong&gt; to open the change for review.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For per-PR coverage, GitHub also supports:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;@codex security review
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This is a deeper, security-specific review currently in research preview.&lt;/p&gt;

&lt;h2&gt;
  
  
  Where Apidog fits: Codex reviews code, but your API still needs a test
&lt;/h2&gt;

&lt;p&gt;Code Review reads diffs, descriptions, comments, and checks. It does not send requests to your running service.&lt;/p&gt;

&lt;p&gt;A pull request can rename a JSON field or change a status code, look correct in review, and still break every API client. That is a contract problem, so it needs a test that calls the API.&lt;/p&gt;

&lt;p&gt;Use Codex and &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; together:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Keep your OpenAPI specification and test scenarios in Apidog.&lt;/li&gt;
&lt;li&gt;Assert status codes and validate responses against endpoint schemas.&lt;/li&gt;
&lt;li&gt;Run scenarios in CI through the Apidog CLI on every pull request.&lt;/li&gt;
&lt;li&gt;Review failing contract checks alongside Codex comments in the PR.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;See &lt;a href="http://apidog.com/blog/api-contract-testing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;API contract testing&lt;/a&gt; for the broader approach.&lt;/p&gt;

&lt;p&gt;Start with endpoints your clients depend on most. Assert:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Expected status codes&lt;/li&gt;
&lt;li&gt;Required response fields&lt;/li&gt;
&lt;li&gt;Field types&lt;/li&gt;
&lt;li&gt;Schema compatibility&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For example, if a Codex-generated pull request changes &lt;code&gt;user_id&lt;/code&gt; to &lt;code&gt;userId&lt;/code&gt;, the schema check fails. The CLI exits non-zero, the pull request check turns red, and reviewers see the break before they need to spot the rename in a long diff.&lt;/p&gt;

&lt;p&gt;This split aligns with OpenAI’s guidance: use review rules for consequential, repository-specific behavior, and use CI for deterministic checks.&lt;/p&gt;

&lt;h3&gt;
  
  
  Run Apidog contract tests in GitHub Actions
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight yaml"&gt;&lt;code&gt;&lt;span class="na"&gt;name&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;api-contract-tests&lt;/span&gt;

&lt;span class="na"&gt;on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;pull_request&lt;/span&gt;

&lt;span class="na"&gt;jobs&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
  &lt;span class="na"&gt;apidog&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
    &lt;span class="na"&gt;runs-on&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;ubuntu-latest&lt;/span&gt;
    &lt;span class="na"&gt;steps&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;npm install -g apidog-cli&lt;/span&gt;
      &lt;span class="pi"&gt;-&lt;/span&gt; &lt;span class="na"&gt;run&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;apidog run --access-token "$APIDOG_ACCESS_TOKEN" -t "$SCENARIO_ID" -e "$ENV_ID" -r cli,junit&lt;/span&gt;
        &lt;span class="na"&gt;env&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt;
          &lt;span class="na"&gt;APIDOG_ACCESS_TOKEN&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ secrets.APIDOG_ACCESS_TOKEN }}&lt;/span&gt;
          &lt;span class="na"&gt;SCENARIO_ID&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ vars.APIDOG_SCENARIO_ID }}&lt;/span&gt;
          &lt;span class="na"&gt;ENV_ID&lt;/span&gt;&lt;span class="pi"&gt;:&lt;/span&gt; &lt;span class="s"&gt;${{ vars.APIDOG_ENV_ID }}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Codex can run the same command locally before opening a pull request, allowing it to see contract-test failures while it still has task context.&lt;/p&gt;

&lt;p&gt;See &lt;a href="http://apidog.com/blog/apidog-cli-in-codex?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Using the Apidog CLI in Codex&lt;/a&gt; for the local setup, or &lt;a href="http://apidog.com/blog/build-your-own-software-factory-codex?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;building a software factory with Codex&lt;/a&gt; for the full task-to-tested-PR loop.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What did Codex get at DevDay 2026?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Reusable cloud environments, a refreshed CLI with voice, &lt;code&gt;/agents&lt;/code&gt;, prompt edit, resume, and worktrees; Code Review in the desktop app with automatic cloud reviews; and Codex Security Cloud in research preview.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;What is the default model in the Codex CLI now?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;GPT-6.1 Sol, since CLI 0.159.1 on September 29. Use &lt;code&gt;/model&lt;/code&gt; to switch models or change reasoning effort.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Can I use Codex cloud on the Free plan?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. Codex cloud is available on Plus, Pro, Business, Healthcare, Edu, and Enterprise. The CLI and Code Review are available on all plans. See the &lt;a href="http://apidog.com/blog/use-codex-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;free Codex guide&lt;/a&gt; for no-cost routes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does Codex Code Review test my API?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. It reads the pull request diff, description, comments, and checks. Add API contract tests in CI so a check fails when behavior changes.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who can use Codex Security Cloud?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;Pro, Business, Enterprise, and Edu users can access it in research preview. Install it as a plugin and connect GitHub.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next step
&lt;/h2&gt;

&lt;p&gt;Update the CLI, run &lt;code&gt;/model&lt;/code&gt; to confirm you are using &lt;code&gt;gpt-6.1-sol&lt;/code&gt;, and try &lt;code&gt;/review&lt;/code&gt; on your next branch.&lt;/p&gt;

&lt;p&gt;Then &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;download Apidog&lt;/a&gt;, turn one critical endpoint into a test scenario, and add the GitHub Actions job above. Every Codex-written pull request can then be checked against both the diff and the running API.&lt;/p&gt;

</description>
    </item>
    <item>
      <title>Sign in with ChatGPT for developers: the OAuth flow, plan usage, and what it means for your API bill</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:47:06 +0000</pubDate>
      <link>https://dev.to/hassann/sign-in-with-chatgpt-for-developers-the-oauth-flow-plan-usage-and-what-it-means-for-your-api-bill-48f7</link>
      <guid>https://dev.to/hassann/sign-in-with-chatgpt-for-developers-the-oauth-flow-plan-usage-and-what-it-means-for-your-api-bill-48f7</guid>
      <description>&lt;p&gt;Sign in with ChatGPT is OpenAI’s OAuth 2.0 and OpenID Connect login for ChatGPT users globally. Your app receives a stable account ID plus the user’s name, email address, and profile picture when available. Since DevDay on September 29, 2026, Plus and Pro users can also authorize participating apps to run eligible AI requests against their ChatGPT plan instead of your API key, subject to a weekly per-app cap they configure. Your app never receives their conversations, memories, or an API key.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This guide shows how to implement the flow, decide when to use plan usage versus your API key, handle failure states, and test the integration in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;. For the rest of the event, see the &lt;a href="http://apidog.com/blog/openai-devday-2026?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DevDay 2026 roundup&lt;/a&gt;. If identity and authorization are unclear, start with &lt;a href="http://apidog.com/blog/oauth-vs-openid?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;OAuth vs OpenID&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Sign in with ChatGPT at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;What OpenAI documents&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Identity scopes&lt;/td&gt;
&lt;td&gt;&lt;code&gt;openid profile email&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plan usage scopes (open-source flow)&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;offline_access resource.invoke chatgpt.tokens.use.direct&lt;/code&gt;, with &lt;code&gt;resource=https://api.openai.com/v1&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Your app receives&lt;/td&gt;
&lt;td&gt;An ID token; with plan usage, an access token and a refresh token&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plan usage eligibility&lt;/td&gt;
&lt;td&gt;Plus and Pro, in participating apps&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Where usage counts&lt;/td&gt;
&lt;td&gt;The plan’s ChatGPT Work and Codex usage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Per-app control&lt;/td&gt;
&lt;td&gt;Weekly cap as a share of overall weekly usage; credits after the cap are off by default&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Plan usage tokens&lt;/td&gt;
&lt;td&gt;Access token: 1 hour; refresh token: 30 days and replaced on each refresh&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Developer access&lt;/td&gt;
&lt;td&gt;Commercial apps: limited trial via an interest form. Open-source apps: self-serve&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sources: the &lt;a href="https://developers.openai.com/siwc" rel="noopener noreferrer"&gt;Sign in with ChatGPT docs&lt;/a&gt;, &lt;a href="https://developers.openai.com/siwc/token-sharing-open-source/token-reference" rel="noopener noreferrer"&gt;token reference&lt;/a&gt;, and OpenAI’s help article on &lt;a href="https://help.openai.com/en/articles/20001542-using-your-chatgpt-plan-in-other-apps-and-sites" rel="noopener noreferrer"&gt;using your ChatGPT plan in other apps&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  What your app receives—and what it does not
&lt;/h2&gt;

&lt;p&gt;Start with identity-only sign-in. Request these scopes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;openid profile email
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The &lt;a href="https://developers.openai.com/siwc/website" rel="noopener noreferrer"&gt;website guide&lt;/a&gt; documents the resulting claims:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;profile&lt;/code&gt; can include the user’s name and profile picture.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;email&lt;/code&gt; can include the email address and email-verification claims.&lt;/li&gt;
&lt;li&gt;Identity scopes do &lt;strong&gt;not&lt;/strong&gt; grant access to ChatGPT conversations or OpenAI API resources.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use the verified OIDC subject as the primary account key:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;issuer + client_id + sub
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Do not identify or automatically link users by email alone. OpenAI warns that a matching email address is not proof of account ownership. If an existing account has the same email, require the user to explicitly confirm the link.&lt;/p&gt;

&lt;p&gt;Plan usage is a separate authorization grant. When a user approves the additional scopes, the token response can include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;An access token for eligible Responses API calls&lt;/li&gt;
&lt;li&gt;A refresh token for continued access&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For identity-only sign-in, do not require an &lt;code&gt;access_token&lt;/code&gt;. The required artifact is the &lt;code&gt;id_token&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Implement the OAuth flow
&lt;/h2&gt;

&lt;p&gt;The website integration uses Authorization Code with PKCE plus OpenID Connect. See the &lt;a href="http://apidog.com/blog/oauth-2-flows?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Authorization Code grant with PKCE overview&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Load OpenAI endpoints from:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;https://auth.openai.com/.well-known/openid-configuration
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The production values documented by OpenAI are:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Issuer:                 https://auth.openai.com
Authorization endpoint: https://auth.openai.com/api/accounts/authorize
Token endpoint:         https://auth.openai.com/api/accounts/oauth/token
JWKS URI:               https://auth.openai.com/.well-known/jwks.json
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  1. Generate request-bound security values
&lt;/h3&gt;

&lt;p&gt;Before redirecting the user, generate:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;A random &lt;code&gt;state&lt;/code&gt; value&lt;/li&gt;
&lt;li&gt;A PKCE &lt;code&gt;code_verifier&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;An S256 &lt;code&gt;code_challenge&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;A random OIDC &lt;code&gt;nonce&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Store &lt;code&gt;state&lt;/code&gt;, &lt;code&gt;nonce&lt;/code&gt;, and the verifier server-side, associated with the pending browser session.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Redirect to authorization
&lt;/h3&gt;

&lt;p&gt;Build an authorization request using:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Your client ID&lt;/li&gt;
&lt;li&gt;The exact registered redirect URI&lt;/li&gt;
&lt;li&gt;Your requested scopes&lt;/li&gt;
&lt;li&gt;&lt;code&gt;state&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;nonce&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;code_challenge&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;code_challenge_method=S256&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For identity only, request:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;openid profile email
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;For eligible plan usage flows, request the documented additional scopes:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;offline_access resource.invoke chatgpt.tokens.use.direct
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Also include:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;resource=https://api.openai.com/v1
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  3. Validate the callback
&lt;/h3&gt;

&lt;p&gt;After consent, OpenAI redirects to your callback with an authorization code.&lt;/p&gt;

&lt;p&gt;Your backend must:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Validate the returned &lt;code&gt;state&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Exchange the authorization code at the token endpoint.&lt;/li&gt;
&lt;li&gt;Validate the ID token signature using OpenAI’s JWKS.&lt;/li&gt;
&lt;li&gt;Validate &lt;code&gt;iss&lt;/code&gt;, &lt;code&gt;aud&lt;/code&gt;, &lt;code&gt;exp&lt;/code&gt;, and &lt;code&gt;nonce&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Find, create, or explicitly link the local user account.&lt;/li&gt;
&lt;li&gt;Create your own application session.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Public clients do not send a client secret. A confidential client using &lt;code&gt;client_secret_basic&lt;/code&gt; sends its secret only through the HTTP Basic authentication header.&lt;/p&gt;

&lt;h3&gt;
  
  
  Open-source and locally hosted clients
&lt;/h3&gt;

&lt;p&gt;Open-source tools use a different registration path. The &lt;a href="https://developers.openai.com/siwc/token-sharing-open-source/sign-in" rel="noopener noreferrer"&gt;open-source sign-in guide&lt;/a&gt; starts with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;client_id=dynamic_agent_client
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;agent_name_hint&lt;/code&gt;: your application name&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;ext_agent_host_id&lt;/code&gt;: a persistent identifier for each host&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The callback returns an issued client ID such as &lt;code&gt;oaiapp_...&lt;/code&gt;. Save and reuse it. The redirect URI is a &lt;code&gt;127.0.0.1&lt;/code&gt; loopback URI, and this flow does not use a client secret.&lt;/p&gt;

&lt;h2&gt;
  
  
  Explain plan usage to users
&lt;/h2&gt;

&lt;p&gt;Your product UI and support documentation should set expectations before a user enables plan usage.&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Eligible requests count against the user’s plan.&lt;/strong&gt; They consume ChatGPT Work and Codex usage from a Plus or Pro subscription.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Each app has a weekly cap.&lt;/strong&gt; Users configure it as a percentage of total weekly usage. The documentation’s settings example ranges from 10% to 100%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;The cap is not reserved capacity.&lt;/strong&gt; Usage in ChatGPT or other connected apps can exhaust the plan before your app reaches its configured cap.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Credits are opt-in.&lt;/strong&gt; Continuing on credits after limits are reached is disabled by default and requires the app cap to be set to 100%.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Plus has a shared five-hour limit.&lt;/strong&gt; According to the &lt;a href="https://developers.openai.com/siwc/token-sharing-open-source/profiles-and-sessions" rel="noopener noreferrer"&gt;accounts and sessions documentation&lt;/a&gt;, it applies across every app using the plan. It does not apply to Pro.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disconnecting stops future usage.&lt;/strong&gt; It does not reverse already consumed usage. OpenAI does not notify your application directly; detect the disconnect when a request or refresh fails.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Users manage these settings in ChatGPT under Usage:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;chatgpt.com/settings/usage
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Follow OpenAI’s &lt;a href="https://developers.openai.com/siwc/ui-ux-guidelines" rel="noopener noreferrer"&gt;UI guidelines&lt;/a&gt; and provide a &lt;strong&gt;Manage usage&lt;/strong&gt; link.&lt;/p&gt;

&lt;h2&gt;
  
  
  Get a client ID
&lt;/h2&gt;

&lt;p&gt;OpenAI’s &lt;a href="https://openai.com/index/devday-2026-recap/" rel="noopener noreferrer"&gt;DevDay recap&lt;/a&gt; names 16 plan-usage partners, including Cognition’s Devin, Notion, Vercel, T3, OpenClaw, and Dactyl. &lt;a href="https://thenewstack.io/sign-in-with-chatgpt/" rel="noopener noreferrer"&gt;The New Stack&lt;/a&gt; also lists Amp, Warp, Kilo Code, and OpenCode, with Lovable marked as coming soon. If you run &lt;a href="http://apidog.com/blog/openclaw-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;OpenClaw&lt;/a&gt;, it appears in both lists.&lt;/p&gt;

&lt;p&gt;As The New Stack reported, Sam Altman summarized the value on stage: “now you don’t have to cover their token costs to get them going.”&lt;/p&gt;

&lt;p&gt;Choose the onboarding route that matches your product:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Commercial or hosted apps:&lt;/strong&gt; Sign-in is a limited trial. &lt;a href="https://developers.openai.com/siwc/request-client-id" rel="noopener noreferrer"&gt;Request a client ID&lt;/a&gt; through OpenAI’s interest form, whether you need identity only or plan usage.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Open-source and locally hosted tools:&lt;/strong&gt; Plan usage is available to open-source partners through the self-serve flow.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  Decide when to use plan usage versus your API key
&lt;/h2&gt;

&lt;p&gt;Plan usage shifts eligible model costs from your API account to the user’s ChatGPT subscription. It can reduce variable model costs, but it also means your app must operate within the user’s eligibility and plan limits.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Your API key&lt;/th&gt;
&lt;th&gt;User’s ChatGPT plan&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Who pays&lt;/td&gt;
&lt;td&gt;You, per token&lt;/td&gt;
&lt;td&gt;User’s plan; credits only if they opt in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Who can use it&lt;/td&gt;
&lt;td&gt;Every user&lt;/td&gt;
&lt;td&gt;Plus and Pro users who grant &lt;code&gt;chatgpt.tokens.use.direct&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Limits&lt;/td&gt;
&lt;td&gt;Your rate-limit tier&lt;/td&gt;
&lt;td&gt;Weekly plan usage, per-app cap, Plus five-hour window&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Request shape&lt;/td&gt;
&lt;td&gt;Full Responses API&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;store: false&lt;/code&gt; and &lt;code&gt;stream: true&lt;/code&gt; required; no &lt;code&gt;temperature&lt;/code&gt;, &lt;code&gt;max_output_tokens&lt;/code&gt;, file search, or Code Interpreter&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Typical failure&lt;/td&gt;
&lt;td&gt;429 when you exceed your tier&lt;/td&gt;
&lt;td&gt;429 &lt;code&gt;subscription_sharing_usage_limit_exceeded&lt;/code&gt;, including in a mid-stream &lt;code&gt;response.failed&lt;/code&gt; event&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Fallback&lt;/td&gt;
&lt;td&gt;Your choice&lt;/td&gt;
&lt;td&gt;None automatically; OpenAI does not switch billing&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;What to show&lt;/td&gt;
&lt;td&gt;Your pricing and usage&lt;/td&gt;
&lt;td&gt;“Using ChatGPT plan,” a Manage usage link, and supported product plans&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The &lt;a href="https://developers.openai.com/siwc/token-sharing-open-source/preview-limitations" rel="noopener noreferrer"&gt;preview limitations page&lt;/a&gt; documents the restrictions. Features requiring stored conversation state or hosted tools do not run on a user’s plan today.&lt;/p&gt;

&lt;p&gt;A practical default is a hybrid model:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Use plan usage for interactive requests from eligible Plus and Pro users.&lt;/li&gt;
&lt;li&gt;Use your API key for users without plan usage.&lt;/li&gt;
&lt;li&gt;Keep background jobs, CI, and scheduled agents on your API key.&lt;/li&gt;
&lt;li&gt;When the plan cap is reached, pause plan-backed requests.&lt;/li&gt;
&lt;li&gt;Show &lt;strong&gt;Manage usage&lt;/strong&gt; and offer your own credits or billing path as a secondary option.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;For broader architecture decisions, see &lt;a href="http://apidog.com/blog/api-key-vs-oauth?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;API key vs OAuth&lt;/a&gt; and &lt;a href="http://apidog.com/blog/ai-agent-oauth-on-behalf-of-user?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;OAuth for AI agents&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the sign-in flow and failure paths in Apidog
&lt;/h2&gt;

&lt;p&gt;Apidog does not sign users in with ChatGPT. Use it to exercise your OAuth configuration, token exchange requests, and application error handling.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; and create a dedicated environment for this integration.&lt;/p&gt;

&lt;h3&gt;
  
  
  1. Store OAuth configuration as environment variables
&lt;/h3&gt;

&lt;p&gt;Add these variables:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;SIWC_CLIENT_ID
SIWC_REDIRECT_URI
SIWC_CLIENT_SECRET
ACCESS_TOKEN
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Mark &lt;code&gt;SIWC_CLIENT_SECRET&lt;/code&gt; as sensitive for confidential clients. Reference variables in requests rather than hardcoding them:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;{{SIWC_CLIENT_ID}}
{{SIWC_REDIRECT_URI}}
{{ACCESS_TOKEN}}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This prevents secrets from being stored in saved requests or shared collections.&lt;/p&gt;

&lt;h3&gt;
  
  
  2. Run the Authorization Code with PKCE flow
&lt;/h3&gt;

&lt;p&gt;In the Auth tab:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Select &lt;strong&gt;OAuth 2.0&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Select &lt;strong&gt;Authorization Code with PKCE&lt;/strong&gt;.&lt;/li&gt;
&lt;li&gt;Enter the authorization and token endpoints.&lt;/li&gt;
&lt;li&gt;Set the scope to &lt;code&gt;openid profile email&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Use a callback URL registered for your OpenAI client.&lt;/li&gt;
&lt;li&gt;Complete the browser consent flow.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;a href="http://apidog.com/blog/test-oauth-2-apis-apidog?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog OAuth 2.0 guide&lt;/a&gt; covers the field-by-field setup.&lt;/p&gt;

&lt;h3&gt;
  
  
  3. Validate the token response
&lt;/h3&gt;

&lt;p&gt;Save the token exchange as its own POST request to the token endpoint. Include the authorization code, PKCE verifier, redirect URI, and client ID.&lt;/p&gt;

&lt;p&gt;Add a post-processing script to validate the response structure:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;json&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;token exchange returned an ID token&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;response&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;code&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;eql&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id_token&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;be&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;a&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;decode&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;require&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;atob&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;part&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id_token&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)[&lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/-/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;+&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
  &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/_/g&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;/&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;claims&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
  &lt;span class="nf"&gt;decode&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;part&lt;/span&gt; &lt;span class="o"&gt;+&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;=&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;repeat&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;part&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="o"&gt;%&lt;/span&gt; &lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt;
&lt;span class="p"&gt;);&lt;/span&gt;

&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;test&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;ID token claims match this client&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;claims&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;iss&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;eql&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;https://auth.openai.com&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;claims&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;aud&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;include&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;environment&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;get&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;SIWC_CLIENT_ID&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;));&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;claims&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sub&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;be&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;a&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;string&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;and&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;not&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;empty&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="nx"&gt;pm&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;expect&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;claims&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;exp&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nx"&gt;to&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;be&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;above&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;());&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use this only for response-level testing. Your backend must still verify the JWT signature and nonce.&lt;/p&gt;

&lt;p&gt;Log &lt;code&gt;name&lt;/code&gt;, &lt;code&gt;email&lt;/code&gt;, and &lt;code&gt;picture&lt;/code&gt; rather than asserting that they always exist, because OpenAI returns them when available.&lt;/p&gt;

&lt;p&gt;For plan usage, also verify that the response scope contains:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;chatgpt.tokens.use.direct
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  4. Mock failure paths
&lt;/h3&gt;

&lt;p&gt;Do not rely on a real Plus account as your only test fixture. Create Apidog mock responses for the cases your application must handle.&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Scenario&lt;/th&gt;
&lt;th&gt;Mock response&lt;/th&gt;
&lt;th&gt;Expected application behavior&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Plan usage declined&lt;/td&gt;
&lt;td&gt;Token response where &lt;code&gt;scope&lt;/code&gt; does not include &lt;code&gt;chatgpt.tokens.use.direct&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Keep the user signed in; offer plan usage setup or another billing path&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cap reached&lt;/td&gt;
&lt;td&gt;429 with &lt;code&gt;error.code&lt;/code&gt; set to &lt;code&gt;subscription_sharing_usage_limit_exceeded&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Pause plan-backed requests and show Manage usage&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Stream fails after starting&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;response.failed&lt;/code&gt; event with &lt;code&gt;subscription_sharing_usage_limit_exceeded&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Mark the request as failed; do not treat stream start as success&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User is not eligible&lt;/td&gt;
&lt;td&gt;403 &lt;code&gt;subscription_sharing_user_not_eligible&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Do not retry or loop through OAuth&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;User disconnected&lt;/td&gt;
&lt;td&gt;Refresh returns &lt;code&gt;invalid_grant&lt;/code&gt;, or request returns 401 &lt;code&gt;subscription_sharing_invalid_user&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Clear stored tokens and request sign-in again&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Chain these cases into a scenario and run it in CI with the Apidog CLI. OpenAI’s &lt;a href="https://developers.openai.com/siwc/token-sharing-open-source/errors-and-recovery" rel="noopener noreferrer"&gt;errors and recovery page&lt;/a&gt; lists the full error set.&lt;/p&gt;

&lt;h3&gt;
  
  
  5. Verify a live streaming request
&lt;/h3&gt;

&lt;p&gt;With a real plan token, call the Responses API using the access token:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;--no-buffer&lt;/span&gt; https://api.openai.com/v1/responses &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="k"&gt;${&lt;/span&gt;&lt;span class="nv"&gt;ACCESS_TOKEN&lt;/span&gt;&lt;span class="k"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "gpt-6.1-sol",
    "input": [{"role": "user", "content": "Say exactly: Hello, world!"}],
    "store": false,
    "stream": true
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;In Apidog, set the request authorization to:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Bearer {{ACCESS_TOKEN}}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Treat &lt;code&gt;response.completed&lt;/code&gt; as the only success signal. A stream that starts successfully can still end with &lt;code&gt;response.failed&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Can Free users sign in with ChatGPT?
&lt;/h3&gt;

&lt;p&gt;Yes. Sign in with ChatGPT is available to ChatGPT users globally. Using a ChatGPT plan inside another app requires Plus or Pro.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does my app receive the user’s OpenAI API key?
&lt;/h3&gt;

&lt;p&gt;No. Your app receives an ID token and, when plan usage is approved, an OAuth access token for eligible Responses API requests.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happens when the user hits their cap?
&lt;/h3&gt;

&lt;p&gt;Requests fail with &lt;code&gt;subscription_sharing_usage_limit_exceeded&lt;/code&gt;. This can be an HTTP 429 or a &lt;code&gt;response.failed&lt;/code&gt; event after streaming starts. Pause plan-backed requests and link the user to &lt;strong&gt;Manage usage&lt;/strong&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Can a Plus user run GPT-6.1 Sol through a partner app?
&lt;/h3&gt;

&lt;p&gt;The documentation example uses &lt;code&gt;gpt-6.1-sol&lt;/code&gt; with a plan token. List the account’s available models with that token before offering a model in your UI. See &lt;a href="http://apidog.com/blog/gpt-6-1-sol-free?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;is GPT-6.1 Sol free&lt;/a&gt; for more context.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next step
&lt;/h2&gt;

&lt;p&gt;If you run a commercial app, join the waitlist and implement cap-reached and disconnect handling with mocks while you wait. Treat plan usage as an optional billing path beside your own API billing, not an automatic replacement.&lt;/p&gt;

&lt;p&gt;Save your token assertions and failure scenarios in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;. When your client ID arrives, the real token should be the only new variable.&lt;/p&gt;

</description>
      <category>api</category>
      <category>authentication</category>
      <category>chatgpt</category>
      <category>openai</category>
    </item>
    <item>
      <title>MCP Events explained: build and test a webhook-emitting MCP server ChatGPT can subscribe to</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:45:48 +0000</pubDate>
      <link>https://dev.to/hassann/mcp-events-explained-build-and-test-a-webhook-emitting-mcp-server-chatgpt-can-subscribe-to-4od8</link>
      <guid>https://dev.to/hassann/mcp-events-explained-build-and-test-a-webhook-emitting-mcp-server-chatgpt-can-subscribe-to-4od8</guid>
      <description>&lt;p&gt;MCP Events let an MCP server push updates to ChatGPT as soon as they occur, eliminating agent-side polling. Since OpenAI DevDay on September 29, 2026, ChatGPT supports the proposed MCP Events specification at protocol version &lt;code&gt;2026-07-28&lt;/code&gt; (MCP 2.0) on all plans. To support it, implement &lt;code&gt;events/list&lt;/code&gt;, &lt;code&gt;events/subscribe&lt;/code&gt;, and &lt;code&gt;events/unsubscribe&lt;/code&gt;; advertise &lt;code&gt;events&lt;/code&gt; in &lt;code&gt;server/discover&lt;/code&gt;; verify callback endpoints; and sign every webhook delivery with Standard Webhooks HMAC. ChatGPT accepts webhook delivery only.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This guide shows the request shapes, server-side validation rules, and an end-to-end test workflow. If you need MCP fundamentals first, read &lt;a href="http://apidog.com/blog/what-is-mcp-model-context-protocol?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;what MCP is&lt;/a&gt;. You can send JSON-RPC requests and mock callback endpoints in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP Events at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;What ChatGPT expects&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Protocol&lt;/td&gt;
&lt;td&gt;MCP 2.0, version &lt;code&gt;2026-07-28&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Capability&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;"events": {}&lt;/code&gt; in &lt;code&gt;server/discover&lt;/code&gt; capabilities&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Methods&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;events/list&lt;/code&gt;, &lt;code&gt;events/subscribe&lt;/code&gt;, &lt;code&gt;events/unsubscribe&lt;/code&gt; on the same authenticated endpoint as tools&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Delivery&lt;/td&gt;
&lt;td&gt;Webhook only: no polling, streaming, &lt;code&gt;gap&lt;/code&gt;, or &lt;code&gt;terminated&lt;/code&gt; notifications&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Signing&lt;/td&gt;
&lt;td&gt;Standard Webhooks HMAC-SHA256&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Headers&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;webhook-id&lt;/code&gt;, &lt;code&gt;webhook-timestamp&lt;/code&gt;, &lt;code&gt;webhook-signature&lt;/code&gt;, &lt;code&gt;X-MCP-Subscription-Id&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Secret&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;whsec_&lt;/code&gt; plus base64 that decodes to 24–64 bytes, supplied by ChatGPT&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Payload limit&lt;/td&gt;
&lt;td&gt;256 KiB (262,144 bytes), one event per request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Subscription ID&lt;/td&gt;
&lt;td&gt;Deterministic from principal, callback URL, event name, and arguments&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Callbacks&lt;/td&gt;
&lt;td&gt;HTTPS, challenge-verified, no private addresses, no redirects&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sources: OpenAI’s &lt;a href="https://developers.openai.com/plugins/build/mcp-events" rel="noopener noreferrer"&gt;MCP Events guide&lt;/a&gt; and the draft &lt;a href="https://github.com/modelcontextprotocol/experimental-ext-triggers-events/blob/main/docs/design-sketch-proposal.md" rel="noopener noreferrer"&gt;MCP Events design sketch&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Why use events instead of polling?
&lt;/h2&gt;

&lt;p&gt;Without events, an agent repeatedly calls a tool and diffs results. That wastes requests when nothing changed and introduces delay when something does change.&lt;/p&gt;

&lt;p&gt;With MCP Events, your server sends a notification when it detects a matching change. The implementation trade-offs are the same as traditional &lt;a href="http://apidog.com/blog/webhooks-vs-polling?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;webhooks vs. polling&lt;/a&gt;, but applied to MCP.&lt;/p&gt;

&lt;p&gt;Typical use cases include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Watch a project board for new tasks, then have ChatGPT read linked documents and draft a plan.&lt;/li&gt;
&lt;li&gt;Turn channel bug reports into draft pull requests using &lt;code&gt;message.created&lt;/code&gt; filtered by &lt;code&gt;channel_id&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Apply document review comments using &lt;code&gt;comment.created&lt;/code&gt; filtered by &lt;code&gt;document_id&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The extension comes from MCP’s Triggers and Events Working Group (&lt;a href="https://github.com/modelcontextprotocol/experimental-ext-triggers-events" rel="noopener noreferrer"&gt;repository&lt;/a&gt;) and remains experimental. Pin your implementation to &lt;code&gt;2026-07-28&lt;/code&gt;. For the broader launch, see the &lt;a href="http://apidog.com/blog/openai-devday-2026?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DevDay 2026 hub&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  1. Advertise and define events
&lt;/h2&gt;

&lt;p&gt;First, add &lt;code&gt;events&lt;/code&gt; to your &lt;code&gt;server/discover&lt;/code&gt; response:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"jsonrpc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"result"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"resultType"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"complete"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"supportedVersions"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;"2026-07-28"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"capabilities"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{},&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"events"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Next, implement &lt;code&gt;events/list&lt;/code&gt;. Each event definition should include:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;name&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;description&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;delivery: ["webhook"]&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;inputSchema&lt;/code&gt; for subscription arguments, such as &lt;code&gt;document_id&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;payloadSchema&lt;/code&gt; for the &lt;code&gt;data&lt;/code&gt; field sent in each delivery&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Implementation rules:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Use stable, descriptive event names.&lt;/li&gt;
&lt;li&gt;Validate filters against &lt;code&gt;inputSchema&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Apply filters on the server, not in the client.&lt;/li&gt;
&lt;li&gt;Return only events the authenticated account is authorized to access.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  2. Implement &lt;code&gt;events/subscribe&lt;/code&gt;
&lt;/h2&gt;

&lt;p&gt;When a user asks ChatGPT to monitor something, ChatGPT calls &lt;code&gt;events/subscribe&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"jsonrpc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"method"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"events/subscribe"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"params"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"comment.created"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"arguments"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"document_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"doc_123"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"delivery"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"mode"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"webhook"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://receiver.example.com/mcp-events/callback_123"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"secret"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"whsec_&amp;lt;base64-encoded-signing-key&amp;gt;"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"cursor"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Before creating or refreshing a subscription:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Authorize the user for the event and requested arguments.&lt;/li&gt;
&lt;li&gt;Validate the event name and arguments against the event definition.&lt;/li&gt;
&lt;li&gt;Require a &lt;code&gt;whsec_&lt;/code&gt; secret whose decoded value is 24–64 bytes.&lt;/li&gt;
&lt;li&gt;Require an HTTPS callback URL.&lt;/li&gt;
&lt;li&gt;Reject private, local, and non-public callback addresses.&lt;/li&gt;
&lt;li&gt;Verify the callback endpoint before sending application data.&lt;/li&gt;
&lt;li&gt;Store the subscription owner, filters, URL, secret, and expiration.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Return a subscription record:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"jsonrpc"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2.0"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;2&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"result"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"sub_123"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"refreshBefore"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-10-02T12:00:00Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"cursor"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"truncated"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Make subscriptions deterministic and idempotent
&lt;/h3&gt;

&lt;p&gt;Use these rules to avoid duplicate subscriptions:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Deterministic IDs:&lt;/strong&gt; Derive &lt;code&gt;id&lt;/code&gt; from the authenticated principal, callback URL, event name, and canonicalized arguments. The design sketch suggests a truncated SHA-256 hash.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Idempotent upserts:&lt;/strong&gt; A repeated subscribe request with the same identity should update the existing record. Canonicalize JSON arguments before comparing them so key ordering does not create duplicates.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Refresh support:&lt;/strong&gt; ChatGPT calls &lt;code&gt;events/subscribe&lt;/code&gt; again before &lt;code&gt;refreshBefore&lt;/code&gt;, using the same identity and the last saved cursor. Return a new expiry time.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Secret rotation:&lt;/strong&gt; If a refresh includes a new secret, replace the stored secret and sign with both keys during a short transition period.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Replay support:&lt;/strong&gt; Return &lt;code&gt;cursor: null&lt;/code&gt; for event types your server cannot replay.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  3. Verify the callback before delivering data
&lt;/h2&gt;

&lt;p&gt;Before sending application events, POST a signed verification request with a fresh, single-use, short-lived challenge:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"verification"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"challenge"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"a-single-use-random-value"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use a unique &lt;code&gt;webhook-id&lt;/code&gt;, such as &lt;code&gt;msg_verification_123&lt;/code&gt;, and include the same signing headers used for event delivery.&lt;/p&gt;

&lt;p&gt;ChatGPT replies with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"challenge"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"a-single-use-random-value"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Only activate the subscription after a successful &lt;code&gt;2xx&lt;/code&gt; response and a constant-time challenge comparison.&lt;/p&gt;

&lt;p&gt;If verification fails, return JSON-RPC error &lt;code&gt;-32015&lt;/code&gt; (&lt;code&gt;CallbackEndpointError&lt;/code&gt;) with a reason such as:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"reason"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"challenge_failed"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;or:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"reason"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"timeout"&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Cache successful verification per principal and callback URL for a bounded period so normal subscription refreshes do not trigger unnecessary challenges.&lt;/p&gt;

&lt;h3&gt;
  
  
  Harden outbound webhook requests
&lt;/h3&gt;

&lt;p&gt;The callback verification exists because the subscriber provides the signing secret. Without verification, an attacker could point your server at a victim URL and cause unwanted traffic.&lt;/p&gt;

&lt;p&gt;Apply these controls to verification and normal deliveries:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Require HTTPS.&lt;/li&gt;
&lt;li&gt;Resolve and validate the target at connection time.&lt;/li&gt;
&lt;li&gt;Block loopback, private, local, and other non-public IP ranges.&lt;/li&gt;
&lt;li&gt;Do not follow redirects.&lt;/li&gt;
&lt;li&gt;Use explicit development-only allowlists for local testing.&lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;
  
  
  4. Deliver and sign events
&lt;/h2&gt;

&lt;p&gt;When a matching change occurs, POST one event object to the callback URL:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"eventId"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"evt_456"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"name"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"comment.created"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"timestamp"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"2026-10-01T12:05:00Z"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"data"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"document_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"doc_123"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"comment_id"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"comment_456"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"text"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"Can we add the rollout dates to this section?"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://docs.example.com/doc_123#comment_456"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"cursor"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;null&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Send these headers:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Content-Type: application/json
webhook-id: evt_456
webhook-timestamp: &amp;lt;unix-seconds&amp;gt;
webhook-signature: v1,&amp;lt;base64-signature&amp;gt;
X-MCP-Subscription-Id: sub_123
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Serialize the JSON body once and sign those exact bytes. Do not parse, mutate, and reserialize the payload between signing and sending.&lt;/p&gt;

&lt;p&gt;Delivery rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Send one event per request.&lt;/li&gt;
&lt;li&gt;Keep each payload below 256 KiB.&lt;/li&gt;
&lt;li&gt;For large records, send a summary and expose a read tool for full content.&lt;/li&gt;
&lt;li&gt;Treat user-authored text as data, never as model instructions.&lt;/li&gt;
&lt;li&gt;Retry transient failures with capped exponential backoff.&lt;/li&gt;
&lt;li&gt;Keep the same event ID across retries, but generate fresh signing headers for each attempt.&lt;/li&gt;
&lt;li&gt;Do not retry &lt;code&gt;410 Gone&lt;/code&gt; or &lt;code&gt;413 Payload Too Large&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Design write tools to be idempotent because events can arrive out of order.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For retry design guidance, see &lt;a href="http://apidog.com/blog/how-to-design-reliable-webhooks?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;reliable webhook design&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  5. Verify Standard Webhooks signatures
&lt;/h2&gt;

&lt;p&gt;Per the &lt;a href="https://github.com/standard-webhooks/standard-webhooks/blob/main/spec/standard-webhooks.md" rel="noopener noreferrer"&gt;Standard Webhooks specification&lt;/a&gt;, sign this exact content:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;${webhook-id}.${webhook-timestamp}.${body}
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Use these rules:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Strip &lt;code&gt;whsec_&lt;/code&gt; from the secret.&lt;/li&gt;
&lt;li&gt;Base64-decode the remaining value to get the HMAC key.&lt;/li&gt;
&lt;li&gt;Sign with HMAC-SHA256.&lt;/li&gt;
&lt;li&gt;Send one or more space-separated signatures in &lt;code&gt;webhook-signature&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Each signature uses the format &lt;code&gt;v1,&amp;lt;base64&amp;gt;&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The MCP draft requires receivers to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Reject timestamps older than five minutes.&lt;/li&gt;
&lt;li&gt;Deduplicate deliveries by &lt;code&gt;webhook-id&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;ChatGPT verifies deliveries on its side, but running a strict local receiver is useful when testing your sender.&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="c1"&gt;// receiver.mjs: strict Standard Webhooks receiver for local tests (Node 18+)&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;createServer&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;node:http&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="nx"&gt;createHmac&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;timingSafeEqual&lt;/span&gt; &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;node:crypto&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;SECRET&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;WEBHOOK_SECRET&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt; &lt;span class="c1"&gt;// whsec_...&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;MAX_BYTES&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;256&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;1024&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;TOLERANCE_S&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;5&lt;/span&gt; &lt;span class="o"&gt;*&lt;/span&gt; &lt;span class="mi"&gt;60&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;seen&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;Set&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="k"&gt;export&lt;/span&gt; &lt;span class="kd"&gt;function&lt;/span&gt; &lt;span class="nf"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;h&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;secret&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;now&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;floor&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nb"&gt;Date&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;now&lt;/span&gt;&lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;/&lt;/span&gt; &lt;span class="mi"&gt;1000&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;h&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;webhook-id&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;ts&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;h&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;webhook-timestamp&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;sigs&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;h&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;webhook-signature&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;

  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;sigs&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nc"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nb"&gt;Number&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;isInteger&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;||&lt;/span&gt; &lt;span class="nb"&gt;Math&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;abs&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;now&lt;/span&gt; &lt;span class="o"&gt;-&lt;/span&gt; &lt;span class="nx"&gt;t&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;TOLERANCE_S&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
  &lt;span class="p"&gt;}&lt;/span&gt;

  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;secret&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;replace&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="sr"&gt;/^whsec_/&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;),&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;base64&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;expected&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nf"&gt;createHmac&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;sha256&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;key&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="s2"&gt;`&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.&lt;/span&gt;&lt;span class="p"&gt;${&lt;/span&gt;&lt;span class="nx"&gt;ts&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="s2"&gt;.`&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;update&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;digest&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

  &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;sigs&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt; &lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;some&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;version&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;b64&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;signature&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;split&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;,&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;received&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="k"&gt;from&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;b64&lt;/span&gt; &lt;span class="o"&gt;??&lt;/span&gt; &lt;span class="dl"&gt;""&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;base64&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

    &lt;span class="k"&gt;return &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
      &lt;span class="nx"&gt;version&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;v1&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt;
      &lt;span class="nx"&gt;received&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="nx"&gt;expected&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt; &lt;span class="o"&gt;&amp;amp;&amp;amp;&lt;/span&gt;
      &lt;span class="nf"&gt;timingSafeEqual&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;received&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;expected&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
    &lt;span class="p"&gt;);&lt;/span&gt;
  &lt;span class="p"&gt;});&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;

&lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;SECRET&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
  &lt;span class="nf"&gt;createServer&lt;/span&gt;&lt;span class="p"&gt;((&lt;/span&gt;&lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="p"&gt;[];&lt;/span&gt;
    &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;size&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

    &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;data&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="nx"&gt;size&lt;/span&gt; &lt;span class="o"&gt;+=&lt;/span&gt; &lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;length&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;size&lt;/span&gt; &lt;span class="o"&gt;&amp;lt;=&lt;/span&gt; &lt;span class="nx"&gt;MAX_BYTES&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;push&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunk&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;

    &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;on&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;end&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;()&lt;/span&gt; &lt;span class="o"&gt;=&amp;gt;&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;size&lt;/span&gt; &lt;span class="o"&gt;&amp;gt;&lt;/span&gt; &lt;span class="nx"&gt;MAX_BYTES&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writeHead&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;413&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;end&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;

      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;raw&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;Buffer&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;concat&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;chunks&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;

      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nf"&gt;verify&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="nx"&gt;SECRET&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writeHead&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;401&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;end&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;

      &lt;span class="kd"&gt;let&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;
      &lt;span class="k"&gt;try&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;body&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;parse&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;raw&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt; &lt;span class="k"&gt;catch&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writeHead&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;400&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;end&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;

      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;type&lt;/span&gt; &lt;span class="o"&gt;===&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;verification&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writeHead&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Content-Type&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;application/json&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;});&lt;/span&gt;
        &lt;span class="k"&gt;return&lt;/span&gt; &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;end&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;JSON&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;stringify&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt; &lt;span class="na"&gt;challenge&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;challenge&lt;/span&gt; &lt;span class="p"&gt;}));&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;

      &lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;id&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;webhook-id&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;];&lt;/span&gt;

      &lt;span class="k"&gt;if &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="o"&gt;!&lt;/span&gt;&lt;span class="nx"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;has&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;))&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
        &lt;span class="nx"&gt;seen&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;add&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
        &lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;
          &lt;span class="nx"&gt;req&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;headers&lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;x-mcp-subscription-id&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
          &lt;span class="nx"&gt;body&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;name&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
          &lt;span class="nx"&gt;id&lt;/span&gt;
        &lt;span class="p"&gt;);&lt;/span&gt;
      &lt;span class="p"&gt;}&lt;/span&gt;

      &lt;span class="nx"&gt;res&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;writeHead&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;200&lt;/span&gt;&lt;span class="p"&gt;).&lt;/span&gt;&lt;span class="nf"&gt;end&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;
    &lt;span class="p"&gt;});&lt;/span&gt;
  &lt;span class="p"&gt;}).&lt;/span&gt;&lt;span class="nf"&gt;listen&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="mi"&gt;8787&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;span class="p"&gt;}&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Run the receiver:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="nv"&gt;WEBHOOK_SECRET&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;whsec_... node receiver.mjs
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;This receiver:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Returns &lt;code&gt;413&lt;/code&gt; for payloads larger than 256 KiB.&lt;/li&gt;
&lt;li&gt;Returns &lt;code&gt;401&lt;/code&gt; for invalid or stale signatures.&lt;/li&gt;
&lt;li&gt;Echoes verification challenges.&lt;/li&gt;
&lt;li&gt;Logs each &lt;code&gt;webhook-id&lt;/code&gt; once.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The &lt;code&gt;verify()&lt;/code&gt; implementation was checked against the Standard Webhooks JavaScript library signing test vector. For more background, see &lt;a href="http://apidog.com/blog/webhook-signature-verification?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;webhook signature verification&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Your server should reject &lt;code&gt;localhost&lt;/code&gt; by default because it is a private target. For development, either configure an explicit allowlist or expose the receiver through a tunnel.&lt;/p&gt;

&lt;h2&gt;
  
  
  6. Test the full subscription lifecycle
&lt;/h2&gt;

&lt;p&gt;Use &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; to test the failure modes documented by OpenAI.&lt;/p&gt;

&lt;p&gt;Create an environment with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;MCP_URL
MCP_TOKEN
CALLBACK_URL
WEBHOOK_SECRET
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Send each JSON-RPC request as a &lt;code&gt;POST&lt;/code&gt; to &lt;code&gt;{{MCP_URL}}&lt;/code&gt; with:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight plaintext"&gt;&lt;code&gt;Authorization: Bearer {{MCP_TOKEN}}
Accept: application/json, text/event-stream
MCP-Protocol-Version: 2026-07-28
Mcp-Method: &amp;lt;method-name&amp;gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;These headers follow the &lt;a href="https://modelcontextprotocol.io/specification/2026-07-28/basic/transports/streamable-http" rel="noopener noreferrer"&gt;Streamable HTTP binding&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;Every request body also needs &lt;a href="https://modelcontextprotocol.io/specification/2026-07-28/basic/index" rel="noopener noreferrer"&gt;&lt;code&gt;params._meta&lt;/code&gt;&lt;/a&gt; containing:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;io.modelcontextprotocol/protocolVersion&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;io.modelcontextprotocol/clientCapabilities&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;OpenAI’s examples omit &lt;code&gt;_meta&lt;/code&gt;, but missing metadata also returns &lt;code&gt;-32602&lt;/code&gt;. Include it so validation tests fail for the intended reason.&lt;/p&gt;

&lt;h3&gt;
  
  
  Test checklist
&lt;/h3&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Discovery&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Call &lt;code&gt;server/discover&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Assert &lt;code&gt;$.result.capabilities.events&lt;/code&gt; exists.&lt;/li&gt;
&lt;li&gt;Assert &lt;code&gt;$.result.supportedVersions&lt;/code&gt; contains &lt;code&gt;2026-07-28&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Call &lt;code&gt;events/list&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Assert every event includes &lt;code&gt;webhook&lt;/code&gt; in &lt;code&gt;delivery&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Idempotent subscription&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Subscribe using &lt;code&gt;{{CALLBACK_URL}}&lt;/code&gt; and &lt;code&gt;{{WEBHOOK_SECRET}}&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Extract &lt;code&gt;$.result.id&lt;/code&gt; into &lt;code&gt;SUB_ID&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Send the identical request again.&lt;/li&gt;
&lt;li&gt;Send it again with reordered &lt;code&gt;arguments&lt;/code&gt; keys.&lt;/li&gt;
&lt;li&gt;Assert the returned ID equals &lt;code&gt;{{SUB_ID}}&lt;/code&gt; every time.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Validation&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Send a secret that decodes to fewer than 24 bytes.&lt;/li&gt;
&lt;li&gt;Send an &lt;code&gt;http://&lt;/code&gt; callback URL.&lt;/li&gt;
&lt;li&gt;Send a private-IP callback URL.&lt;/li&gt;
&lt;li&gt;Assert each invalid request returns &lt;code&gt;-32602&lt;/code&gt; (&lt;code&gt;InvalidParams&lt;/code&gt;).&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Verification challenge&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Create an Apidog mock endpoint that responds with:
&lt;/li&gt;
&lt;/ul&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"challenge"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"wrong"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/li&gt;
&lt;/ol&gt;

&lt;ul&gt;
&lt;li&gt;Subscribe using the cloud mock URL.&lt;/li&gt;
&lt;li&gt;Assert error code &lt;code&gt;-32015&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Assert &lt;code&gt;$.error.data.reason&lt;/code&gt; is &lt;code&gt;challenge_failed&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Point &lt;code&gt;CALLBACK_URL&lt;/code&gt; to the local receiver and confirm the subscription succeeds.&lt;/li&gt;
&lt;/ul&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Oversize payload&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Trigger an event larger than 262,144 bytes.&lt;/li&gt;
&lt;li&gt;Your sender should reject it before delivery.&lt;/li&gt;
&lt;li&gt;If it reaches the receiver, expect &lt;code&gt;413&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Confirm the sender logs only one attempt and does not retry.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Replay and tampering&lt;/strong&gt;&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Copy a signed outbound delivery into a new Apidog request.&lt;/li&gt;
&lt;li&gt;Resend it immediately: expect &lt;code&gt;200&lt;/code&gt; and no duplicate receiver log entry.&lt;/li&gt;
&lt;li&gt;Change one body byte: expect &lt;code&gt;401&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Resend the original after five minutes: expect &lt;code&gt;401&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Save these requests as a scenario and run them in CI with the Apidog CLI. For additional coverage, see the &lt;a href="http://apidog.com/blog/mcp-server-testing-apidog?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;MCP server testing playbook&lt;/a&gt; and &lt;a href="http://apidog.com/blog/how-to-test-webhooks?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;how to test webhooks&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;What are MCP Events?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
An experimental MCP extension that lets a server push event notifications to a client instead of requiring polling. ChatGPT supports webhook mode at &lt;code&gt;2026-07-28&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does ChatGPT support polling or streaming for MCP Events?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
No. ChatGPT supports webhook delivery and callback verification only. Polling, streaming, and &lt;code&gt;gap&lt;/code&gt; and &lt;code&gt;terminated&lt;/code&gt; notifications are not supported.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which ChatGPT plans get MCP Events?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
OpenAI’s DevDay recap says the feature is available on all plans.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Who creates the signing secret?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The subscriber. ChatGPT sends a &lt;code&gt;whsec_&lt;/code&gt; secret in &lt;code&gt;delivery.secret&lt;/code&gt;; your server validates, stores, and signs with that secret. Your server does not generate its own replacement secret.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How is this different from the Agents API?&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
MCP Events push data from your server into ChatGPT. The &lt;a href="http://apidog.com/blog/openai-agents-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;OpenAI Agents API&lt;/a&gt; runs agents you build and reports their progress through streaming or webhooks.&lt;/p&gt;

&lt;h2&gt;
  
  
  Next step
&lt;/h2&gt;

&lt;p&gt;Start with one server, one event, and one filter:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Add &lt;code&gt;"events": {}&lt;/code&gt; to &lt;code&gt;server/discover&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Implement one &lt;code&gt;events/list&lt;/code&gt; definition.&lt;/li&gt;
&lt;li&gt;Add &lt;code&gt;events/subscribe&lt;/code&gt; with validation and callback verification.&lt;/li&gt;
&lt;li&gt;Sign a single event delivery.&lt;/li&gt;
&lt;li&gt;Run the six tests against a local receiver before connecting ChatGPT.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;&lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; to save the checks as a scenario and run them on every commit.&lt;/p&gt;

</description>
      <category>api</category>
      <category>chatgpt</category>
      <category>mcp</category>
      <category>tutorial</category>
    </item>
    <item>
      <title>OpenAI Agents API vs Responses API vs Agents SDK vs AgentKit: which one to build on</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:44:39 +0000</pubDate>
      <link>https://dev.to/hassann/openai-agents-api-vs-responses-api-vs-agents-sdk-vs-agentkit-which-one-to-build-on-46d9</link>
      <guid>https://dev.to/hassann/openai-agents-api-vs-responses-api-vs-agents-sdk-vs-agentkit-which-one-to-build-on-46d9</guid>
      <description>&lt;p&gt;These four names operate at different layers. The key decision is simple: &lt;strong&gt;who runs the agent loop?&lt;/strong&gt; The Responses API is a model endpoint, so your application runs the loop. The Agents SDK is a TypeScript and Python library whose runner executes the loop inside your application. The Agents API, in public beta since September 10, 2026, runs OpenAI’s Codex harness for you and manages the session and, optionally, the sandbox. AgentKit is the October 2025 bundle of Agent Builder, ChatKit, Connector Registry, and Evals; Agent Builder is scheduled to shut down on November 30, 2026.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;DevDay on September 29 added computer use to the Agents API, making the naming distinction more important. See the &lt;a href="http://apidog.com/blog/openai-devday-2026?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DevDay 2026 roundup&lt;/a&gt;. This guide compares the loop, compute, state, pricing, and maturity of each option, then shows how to migrate from a hand-rolled Responses loop. For sessions and approvals, read the &lt;a href="http://apidog.com/blog/openai-agents-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;OpenAI Agents API guide&lt;/a&gt;. You can test each HTTP surface in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenAI agent options side by side
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;&lt;/th&gt;
&lt;th&gt;Agents API&lt;/th&gt;
&lt;th&gt;Responses API&lt;/th&gt;
&lt;th&gt;Agents SDK&lt;/th&gt;
&lt;th&gt;AgentKit&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;What it is&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Managed agent runtime on the Codex harness&lt;/td&gt;
&lt;td&gt;Model endpoint: &lt;code&gt;POST /v1/responses&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;TypeScript and Python library&lt;/td&gt;
&lt;td&gt;Bundle: Agent Builder, ChatKit, Connector Registry, Evals&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Who runs the loop&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;OpenAI&lt;/td&gt;
&lt;td&gt;Your code&lt;/td&gt;
&lt;td&gt;SDK runner in your app&lt;/td&gt;
&lt;td&gt;Agent Builder workflows, exported to SDK code or embedded with ChatKit&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Where compute runs&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;OpenAI-hosted sandbox, your sandbox, or none&lt;/td&gt;
&lt;td&gt;Your environment, plus hosted tools&lt;/td&gt;
&lt;td&gt;Your runtime and sandbox providers&lt;/td&gt;
&lt;td&gt;Not applicable&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Where state lives&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;OpenAI session: config, turns, items&lt;/td&gt;
&lt;td&gt;Your history, &lt;code&gt;previous_response_id&lt;/code&gt;, or Conversations API&lt;/td&gt;
&lt;td&gt;Your storage, SDK sessions, or Responses state&lt;/td&gt;
&lt;td&gt;Published, versioned workflows&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;What you pay&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tokens, tools, and hosted containers; no extra fee&lt;/td&gt;
&lt;td&gt;Tokens and tools&lt;/td&gt;
&lt;td&gt;Tokens and tools, plus your hosting&lt;/td&gt;
&lt;td&gt;Underlying API usage; no separate subscription&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Integration effort (per OpenAI)&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Low&lt;/td&gt;
&lt;td&gt;High&lt;/td&gt;
&lt;td&gt;Medium&lt;/td&gt;
&lt;td&gt;Not rated&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Status&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Public beta: &lt;code&gt;OpenAI-Beta: agents=v1&lt;/code&gt;
&lt;/td&gt;
&lt;td&gt;Recommended for new projects&lt;/td&gt;
&lt;td&gt;Current&lt;/td&gt;
&lt;td&gt;Agent Builder and Evals shut down Nov. 30, 2026; ChatKit stays&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Data controls&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;US data residency only; not ZDR-eligible; state kept until deleted&lt;/td&gt;
&lt;td&gt;ZDR-eligible with limitations; regional endpoints&lt;/td&gt;
&lt;td&gt;Depends on the APIs it calls&lt;/td&gt;
&lt;td&gt;Not applicable&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sources: OpenAI’s &lt;a href="https://developers.openai.com/api/docs/guides/agents" rel="noopener noreferrer"&gt;agent runtime comparison&lt;/a&gt;, &lt;a href="https://developers.openai.com/api/docs/guides/agents-api/overview" rel="noopener noreferrer"&gt;Agents API overview&lt;/a&gt;, and &lt;a href="https://developers.openai.com/api/docs/deprecations" rel="noopener noreferrer"&gt;deprecations page&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Who runs the loop
&lt;/h2&gt;

&lt;p&gt;This is the decision that drives most of the other tradeoffs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Responses API: your code runs the loop
&lt;/h3&gt;

&lt;p&gt;Hosted tools such as web search, file search, code interpreter, and remote MCP can perform multiple internal calls in one request. Your custom functions are different: the model returns a &lt;code&gt;function_call&lt;/code&gt; item, your application executes it, and then you submit a &lt;code&gt;function_call_output&lt;/code&gt; with the same &lt;code&gt;call_id&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Your application controls:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;When the loop stops&lt;/li&gt;
&lt;li&gt;How conversation history is stored&lt;/li&gt;
&lt;li&gt;Whether responses are persisted (&lt;code&gt;store: false&lt;/code&gt; disables default storage)&lt;/li&gt;
&lt;li&gt;When long contexts are compacted with &lt;code&gt;context_management&lt;/code&gt; and &lt;code&gt;compact_threshold&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;How function calls are retried, approved, audited, or rejected&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Use the &lt;a href="http://apidog.com/blog/openai-responses-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Responses API guide&lt;/a&gt; and &lt;a href="http://apidog.com/blog/openai-function-calling?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;function calling guide&lt;/a&gt; when implementing this loop.&lt;/p&gt;

&lt;h3&gt;
  
  
  Agents SDK: the SDK runner runs the loop in your process
&lt;/h3&gt;

&lt;p&gt;The SDK runner handles the agent loop and handoffs, but your server still owns:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Deployment&lt;/li&gt;
&lt;li&gt;Tool implementations&lt;/li&gt;
&lt;li&gt;State storage&lt;/li&gt;
&lt;li&gt;Human approval flows&lt;/li&gt;
&lt;li&gt;Authentication and audit logs&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With Sandbox Agents, the harness can stay in your infrastructure while commands run in a Unix-local, Docker, or hosted-provider workspace. This keeps approval and access-control decisions outside the sandbox.&lt;/p&gt;

&lt;h3&gt;
  
  
  Agents API: OpenAI runs the loop
&lt;/h3&gt;

&lt;p&gt;The managed harness handles sessions, orchestration, context compaction, and recovery. It also adds subagents, tool search, and programmatic tool calling.&lt;/p&gt;

&lt;p&gt;Remote MCP servers are called directly by OpenAI. Your application still handles custom function tools: when a session includes a &lt;code&gt;function_call&lt;/code&gt; in &lt;code&gt;required_actions&lt;/code&gt;, submit an &lt;code&gt;agent.session.input.tool_result&lt;/code&gt; event containing the &lt;code&gt;turn_id&lt;/code&gt; and &lt;code&gt;call_id&lt;/code&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Compare the same task
&lt;/h3&gt;



&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;&lt;span class="c"&gt;# Responses API: one model call; your code owns the loop&lt;/span&gt;
curl https://api.openai.com/v1/responses &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OPENAI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "model": "gpt-6.1-sol",
    "reasoning": {"effort": "low"},
    "tools": [{"type": "web_search"}],
    "input": "Summarize the breaking changes in the latest Node.js release."
  }'&lt;/span&gt;

&lt;span class="c"&gt;# Agents API: a durable session; OpenAI owns the loop&lt;/span&gt;
curl https://api.openai.com/v1/agents/sessions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"OpenAI-Beta: agents=v1"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OPENAI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "agent": {
      "model": "gpt-6-astra",
      "tools": [{"type": "web_search"}]
    },
    "environment": {"type": "none"},
    "input": "Summarize the breaking changes in the latest Node.js release."
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The Agents API documentation uses &lt;code&gt;gpt-6-astra&lt;/code&gt; in its examples. It does not state whether other models are accepted, so verify compatibility before replacing it with &lt;code&gt;gpt-6.1-sol&lt;/code&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Compute, state, and cost
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Compute
&lt;/h3&gt;

&lt;p&gt;The Agents API can provision and manage a sandbox for the full session. Set &lt;code&gt;environment.type&lt;/code&gt; to one of:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;openai_hosted&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;self_hosted&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;none&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;With the Agents SDK, you select and pay for the sandbox provider. With Responses, code runs in your environment except for OpenAI-hosted tools.&lt;/p&gt;

&lt;h3&gt;
  
  
  State
&lt;/h3&gt;

&lt;p&gt;An Agents API session stores configuration, turns, and items on OpenAI’s side. Send a follow-up as an event on the same session ID.&lt;/p&gt;

&lt;p&gt;With Responses, chain turns with &lt;code&gt;previous_response_id&lt;/code&gt; or use the Conversations API. With the SDK, state lives in your storage, SDK sessions, or Responses state.&lt;/p&gt;

&lt;h3&gt;
  
  
  Cost
&lt;/h3&gt;

&lt;p&gt;Token pricing is identical across these options because they call the same models.&lt;/p&gt;

&lt;p&gt;The Agents API has no additional platform fee, but hosted containers cost from $0.03 for 1 GB to $0.48 for 16 GB per 20-minute session. The SDK adds the cost of your own hosting. AgentKit has no separate subscription, according to the &lt;a href="http://apidog.com/blog/openai-agentkit?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;AgentKit explainer&lt;/a&gt;.&lt;/p&gt;

&lt;h3&gt;
  
  
  Data controls
&lt;/h3&gt;

&lt;p&gt;The Agents API supports US data residency only and does not support Zero Data Retention, including when you use a self-hosted sandbox.&lt;/p&gt;

&lt;p&gt;OpenAI’s &lt;a href="https://developers.openai.com/api/docs/guides/your-data" rel="noopener noreferrer"&gt;data controls page&lt;/a&gt; lists &lt;code&gt;/v1/agents&lt;/code&gt; as not ZDR-eligible, with state retained until deletion. In comparison, &lt;code&gt;/v1/responses&lt;/code&gt; is ZDR-eligible with limitations and is available through regional endpoints such as &lt;code&gt;eu.api.openai.com&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;If ZDR or EU residency is required, the Agents API is not currently an option.&lt;/p&gt;

&lt;h2&gt;
  
  
  AgentKit in late 2026: what remains
&lt;/h2&gt;

&lt;p&gt;AgentKit launched on October 6, 2025 with four components:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agent Builder:&lt;/strong&gt; Deprecation was announced June 3, 2026. Shutdown is scheduled for November 30, 2026. OpenAI’s &lt;a href="https://developers.openai.com/api/docs/guides/agent-builder/migrate-from-agent-builder" rel="noopener noreferrer"&gt;migration guide&lt;/a&gt; exports workflows as Agents SDK code or recreates them as ChatGPT Workspace Agents for Business, Enterprise, or Edu.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Evals:&lt;/strong&gt; Existing evals become read-only on October 31, 2026. The dashboard and API are scheduled to shut down on November 30.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;ChatKit:&lt;/strong&gt; Remains available for embedded chat.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Connector Registry:&lt;/strong&gt; The administration panel for connectors and MCP servers across OpenAI products.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;For a durable, code-first AgentKit path, use the Agents SDK. See the &lt;a href="http://apidog.com/blog/openai-agentkit?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;AgentKit guide&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Which option should you build on?
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Choose&lt;/th&gt;
&lt;th&gt;Use it when&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agents API&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Tasks run for minutes, need files, commands, or a browser, and you do not want to operate the loop, sandboxes, or session storage. US residency and a beta header are acceptable.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Responses API&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;You make single calls, need complete control over each turn, require ZDR or non-US residency, or already have a working loop.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agents SDK&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Typed application code must own tools, storage, approvals, and handoffs, and the loop must run in your infrastructure.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;ChatKit&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;You need an embedded chat UI in your product.&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;&lt;strong&gt;Agent Builder&lt;/strong&gt;&lt;/td&gt;
&lt;td&gt;Do not start new work here. Export existing workflows before November 30, 2026.&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;On AWS, &lt;a href="https://aws.amazon.com/bedrock/managed-agents-openai/" rel="noopener noreferrer"&gt;Bedrock Managed Agents, powered by OpenAI&lt;/a&gt;, brings the Agents API’s core capabilities to run natively in AWS.&lt;/p&gt;

&lt;p&gt;For MCP integration in either code-first option, see &lt;a href="http://apidog.com/blog/mcp-servers-openai-agents?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;MCP servers with OpenAI agents&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Move from a Responses loop to the Agents API
&lt;/h2&gt;

&lt;p&gt;If you already run a Responses-based tool loop and want OpenAI to operate it, migrate in these steps.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Map your existing components.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Move instructions, model selection, and tools into &lt;code&gt;agent&lt;/code&gt;. Map your container configuration to &lt;code&gt;environment&lt;/code&gt;. Replace your conversation store with a session ID.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Move remote MCP servers into &lt;code&gt;agent.tools&lt;/code&gt;.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Store server tokens in a vault attached through &lt;code&gt;vault_ids&lt;/code&gt;; do not include secrets in prompts.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Rewrite custom function handling.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Replace the &lt;code&gt;function_call_output&lt;/code&gt; loop with a handler for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;agent.session.requires_action&lt;/code&gt; when streaming&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;agent.session.action_required&lt;/code&gt; when using webhooks&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Return results with &lt;code&gt;agent.session.input.tool_result&lt;/code&gt;.&lt;/p&gt;

&lt;p&gt;Subagents cannot call function tools, so keep function-tool execution on the main agent.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;&lt;p&gt;&lt;strong&gt;Remove application-side compaction logic.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
The managed harness compacts context automatically.&lt;/p&gt;&lt;/li&gt;
&lt;li&gt;
&lt;p&gt;&lt;strong&gt;Handle turn events.&lt;/strong&gt;&lt;br&gt;&lt;br&gt;
Stream or receive webhooks for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;agent.session.turn.completed&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;agent.session.turn.failed&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;agent.session.turn.cancelled&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;Do not treat an idle session as a successful turn.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;strong&gt;Validate platform constraints first.&lt;/strong&gt;
Confirm that US-only residency, no ZDR support, and the required beta header work for your application.&lt;/li&gt;
&lt;/ol&gt;

&lt;h2&gt;
  
  
  Keep both implementations in one Apidog project
&lt;/h2&gt;

&lt;p&gt;Before switching production traffic, test the old and new paths side by side.&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create a &lt;code&gt;Responses&lt;/code&gt; folder and an &lt;code&gt;Agents API&lt;/code&gt; folder in one Apidog project.&lt;/li&gt;
&lt;li&gt;Add a shared environment containing:

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;{{OPENAI_API_KEY}}&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;A model variable&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;li&gt;Send identical prompts through both implementations.&lt;/li&gt;
&lt;li&gt;Assert status codes and required output fields.&lt;/li&gt;
&lt;li&gt;Open the Agents API stream as an SSE request to inspect turn events.&lt;/li&gt;
&lt;li&gt;Save the requests as a test scenario.&lt;/li&gt;
&lt;li&gt;Run the scenario in CI with the Apidog CLI.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;This makes beta-level API changes visible as failed checks instead of production surprises. For assertion ideas, read the &lt;a href="http://apidog.com/blog/production-ai-agent-reliability?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;production AI agent reliability guide&lt;/a&gt;.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; to set up the project.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Is the Agents API replacing the Responses API?
&lt;/h3&gt;

&lt;p&gt;No deprecation has been announced. OpenAI lists the Agents API, Agents SDK, and Responses API as current options for different requirements.&lt;/p&gt;

&lt;h3&gt;
  
  
  Is OpenAI AgentKit deprecated?
&lt;/h3&gt;

&lt;p&gt;Partly. Agent Builder and Evals are scheduled to shut down on November 30, 2026. ChatKit remains available.&lt;/p&gt;

&lt;h3&gt;
  
  
  Does the Agents SDK use the Agents API?
&lt;/h3&gt;

&lt;p&gt;No. The SDK runs in your application. The Agents API runs a managed harness in OpenAI’s service.&lt;/p&gt;

&lt;h3&gt;
  
  
  What happened to the Assistants API?
&lt;/h3&gt;

&lt;p&gt;OpenAI’s deprecations page sets its removal for August 26, 2026 and directs developers to the Responses and Conversations APIs.&lt;/p&gt;

&lt;h3&gt;
  
  
  Which option is cheapest?
&lt;/h3&gt;

&lt;p&gt;Token prices are the same. The main difference is hosted-container pricing for the Agents API versus the infrastructure costs you operate with the SDK or Responses API.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pick one path this week
&lt;/h2&gt;

&lt;p&gt;Choose based on who should operate the loop, then validate the choice with real requests before building the application.&lt;/p&gt;

&lt;p&gt;If you are starting fresh, create one Agents API session and compare its output with your existing &lt;a href="http://apidog.com/blog/openai-responses-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Responses setup&lt;/a&gt; in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt;.&lt;/p&gt;

</description>
      <category>agents</category>
      <category>ai</category>
      <category>api</category>
      <category>openai</category>
    </item>
    <item>
      <title>How to use the OpenAI Agents API ?</title>
      <dc:creator>Hassann</dc:creator>
      <pubDate>Fri, 02 Oct 2026 09:44:23 +0000</pubDate>
      <link>https://dev.to/hassann/how-to-use-the-openai-agents-api--469j</link>
      <guid>https://dev.to/hassann/how-to-use-the-openai-agents-api--469j</guid>
      <description>&lt;p&gt;The OpenAI Agents API runs OpenAI’s open-source Codex harness for you. Send &lt;code&gt;POST https://api.openai.com/v1/agents/sessions&lt;/code&gt; with the &lt;code&gt;OpenAI-Beta: agents=v1&lt;/code&gt; header, an agent definition, and a task. OpenAI runs the model and tool loop, persists the session, and can provision a sandbox. There is no Agents API fee: you pay for tokens, tools, and hosted container time ($0.03 to $0.48 per 20-minute session for 1 GB to 16 GB sandbox sizes). It entered public beta on September 10, 2026, and OpenAI added computer use at DevDay on September 29.&lt;/p&gt;

&lt;p&gt;&lt;a href="https://apidog.com/?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation" class="crayons-btn crayons-btn--primary"&gt;Try Apidog today&lt;/a&gt;
&lt;/p&gt;

&lt;p&gt;This guide covers a first REST session, progress events, MCP tools, subagents, and the computer-use approval flow. For a comparison with OpenAI’s other agent surfaces, read &lt;a href="http://apidog.com/blog/openai-agents-api-vs-responses-api-vs-agents-sdk?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Agents API vs Responses API vs Agents SDK&lt;/a&gt;. For the rest of the event, see the &lt;a href="http://apidog.com/blog/openai-devday-2026?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;DevDay 2026 roundup&lt;/a&gt;. Every call is plain HTTP, so you can send it from &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; before writing application code.&lt;/p&gt;

&lt;h2&gt;
  
  
  OpenAI Agents API at a glance
&lt;/h2&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Item&lt;/th&gt;
&lt;th&gt;Value&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Status&lt;/td&gt;
&lt;td&gt;Public beta since Sep 10, 2026; computer use added Sep 29&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Create a session&lt;/td&gt;
&lt;td&gt;&lt;code&gt;POST /v1/agents/sessions&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Beta header&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;OpenAI-Beta: agents=v1&lt;/code&gt; (the OpenAI SDKs add it)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Key permissions&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;api.agents.read&lt;/code&gt;, &lt;code&gt;api.agents.write&lt;/code&gt;, &lt;code&gt;api.responses.write&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Pricing&lt;/td&gt;
&lt;td&gt;No Agents API fee; model tokens at API rates, tools at standard rates (web search $10 per 1K calls)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Hosted containers&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;$0.03&lt;/code&gt; (&lt;code&gt;small&lt;/code&gt;, 1 GB), &lt;code&gt;$0.12&lt;/code&gt; (&lt;code&gt;medium&lt;/code&gt;, 4 GB), &lt;code&gt;$0.48&lt;/code&gt; (&lt;code&gt;large&lt;/code&gt;, 16 GB) per 20-minute session&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Environments&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;none&lt;/code&gt;, &lt;code&gt;openai_hosted&lt;/code&gt;, &lt;code&gt;self_hosted&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Model in the docs’ examples&lt;/td&gt;
&lt;td&gt;&lt;code&gt;gpt-6-astra&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Data controls&lt;/td&gt;
&lt;td&gt;US data residency only; no Zero Data Retention (ZDR)&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Max request size&lt;/td&gt;
&lt;td&gt;4 MiB&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;Sources: &lt;a href="https://openai.com/index/introducing-the-agents-api/" rel="noopener noreferrer"&gt;Introducing the Agents API&lt;/a&gt;, the &lt;a href="https://developers.openai.com/api/docs/guides/agents-api/overview" rel="noopener noreferrer"&gt;Agents API overview&lt;/a&gt;, and the &lt;a href="https://developers.openai.com/api/docs/pricing" rel="noopener noreferrer"&gt;pricing page&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  The four concepts
&lt;/h2&gt;

&lt;p&gt;The API is built around four pieces:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Agent:&lt;/strong&gt; The model, instructions, tools, and MCP servers. Pass it inline, or save and reuse its &lt;code&gt;agent_id&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Environment:&lt;/strong&gt; An optional sandbox or computer where the agent reads files and runs commands.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Session:&lt;/strong&gt; A durable instance of an agent that persists configuration, conversation, and saved work.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Events and items:&lt;/strong&gt; Events report live progress; items are saved messages and tool calls.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;A message sent to an idle session starts a new turn. A message sent during a turn steers it. The harness—the hosted Codex instance that runs the model and tool loop, according to the &lt;a href="https://developers.openai.com/api/docs/guides/agents-api/architecture" rel="noopener noreferrer"&gt;architecture guide&lt;/a&gt;—also handles context compaction.&lt;/p&gt;

&lt;h2&gt;
  
  
  Pick an environment
&lt;/h2&gt;

&lt;p&gt;Use &lt;code&gt;environment.type&lt;/code&gt; to decide where commands run:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;none&lt;/code&gt;:&lt;/strong&gt; No compute. Remote MCP servers and function tools still work, but built-in Bash, apply-patch, workspace files, and executor MCPs do not.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;openai_hosted&lt;/code&gt;:&lt;/strong&gt; OpenAI manages a Linux sandbox with Python and Node.js in &lt;code&gt;/workspace&lt;/code&gt;. Set &lt;code&gt;container_size&lt;/code&gt; to &lt;code&gt;small&lt;/code&gt; (1 GB), &lt;code&gt;medium&lt;/code&gt; (default, 4 GB), or &lt;code&gt;large&lt;/code&gt; (16 GB). Configure &lt;code&gt;network.access&lt;/code&gt; as &lt;code&gt;enabled&lt;/code&gt;, &lt;code&gt;disabled&lt;/code&gt;, or &lt;code&gt;restricted&lt;/code&gt; with &lt;code&gt;allowed_domains&lt;/code&gt;. Files in &lt;code&gt;/workspace/outputs&lt;/code&gt; become artifacts when a turn completes. An idle sandbox without keep-alives can be deleted after an hour.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;self_hosted&lt;/code&gt;:&lt;/strong&gt; Your infrastructure. Run &lt;code&gt;codex exec-server&lt;/code&gt; on a laptop, container, or remote sandbox. It connects outbound using a separate environment key.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The launch post lists sandbox partners Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle, Runloop, and Vercel. The &lt;a href="https://developers.openai.com/api/docs/guides/agents-api/environments/self-hosted" rel="noopener noreferrer"&gt;self-hosted environment guide&lt;/a&gt; also adds AWS Lambda MicroVMs.&lt;/p&gt;

&lt;h2&gt;
  
  
  Your first session over REST
&lt;/h2&gt;

&lt;p&gt;Export a key with the required permissions as &lt;code&gt;OPENAI_API_KEY&lt;/code&gt;, then create a streaming session with a small container:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="nt"&gt;--no-buffer&lt;/span&gt; https://api.openai.com/v1/agents/sessions &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"OpenAI-Beta: agents=v1"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OPENAI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "agent": {
      "model": "gpt-6-astra",
      "instructions": "Write clean code, run it, and report the actual output."
    },
    "environment": {
      "type": "openai_hosted",
      "container_size": "small"
    },
    "input": "Create tree.py, a script that prints a tree of the files in the current directory. Run it and show the output.",
    "stream": true
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;With &lt;code&gt;stream: true&lt;/code&gt;, the response is the first turn’s event stream. Save the session ID from that stream, then use the same session resource for the rest of the lifecycle:&lt;/p&gt;

&lt;div class="table-wrapper-paragraph"&gt;&lt;table&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th&gt;Action&lt;/th&gt;
&lt;th&gt;Request&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td&gt;Follow up or steer&lt;/td&gt;
&lt;td&gt;
&lt;code&gt;POST /v1/agents/sessions/{id}/events&lt;/code&gt; with an &lt;code&gt;agent.session.input.message&lt;/code&gt; event&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Cancel the active turn&lt;/td&gt;
&lt;td&gt;Same endpoint with event type &lt;code&gt;agent.session.input.cancel&lt;/code&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Read saved work&lt;/td&gt;
&lt;td&gt;&lt;code&gt;GET /v1/agents/sessions/{id}/items?order=asc&amp;amp;limit=100&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td&gt;Clean up&lt;/td&gt;
&lt;td&gt;&lt;code&gt;DELETE /v1/agents/sessions/{id}&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;

&lt;p&gt;The JavaScript SDK uses the same shape. This example adds web search, subagents, and a vault:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight javascript"&gt;&lt;code&gt;&lt;span class="k"&gt;import&lt;/span&gt; &lt;span class="nx"&gt;OpenAI&lt;/span&gt; &lt;span class="k"&gt;from&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;openai&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;new&lt;/span&gt; &lt;span class="nc"&gt;OpenAI&lt;/span&gt;&lt;span class="p"&gt;();&lt;/span&gt;

&lt;span class="kd"&gt;const&lt;/span&gt; &lt;span class="nx"&gt;session&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="k"&gt;await&lt;/span&gt; &lt;span class="nx"&gt;client&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;beta&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;agents&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;sessions&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;create&lt;/span&gt;&lt;span class="p"&gt;({&lt;/span&gt;
  &lt;span class="na"&gt;agent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
    &lt;span class="na"&gt;model&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;gpt-6-astra&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="na"&gt;tools&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;web_search&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;}],&lt;/span&gt;
    &lt;span class="na"&gt;multi_agent&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt;
      &lt;span class="na"&gt;enabled&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
      &lt;span class="na"&gt;max_concurrent_subagents&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
    &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;vault_ids&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="nx"&gt;process&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;env&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;VAULT_ID&lt;/span&gt;&lt;span class="p"&gt;],&lt;/span&gt;
  &lt;span class="na"&gt;environment&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="p"&gt;{&lt;/span&gt; &lt;span class="na"&gt;type&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;openai_hosted&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt; &lt;span class="p"&gt;},&lt;/span&gt;
  &lt;span class="na"&gt;input&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt; &lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="s2"&gt;Summarize breaking changes in the latest release notes.&lt;/span&gt;&lt;span class="dl"&gt;"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;
&lt;span class="p"&gt;});&lt;/span&gt;

&lt;span class="nx"&gt;console&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nf"&gt;log&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;session&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;id&lt;/span&gt;&lt;span class="p"&gt;);&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h2&gt;
  
  
  Follow progress: stream or webhooks
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Stream events
&lt;/h3&gt;

&lt;p&gt;Open &lt;code&gt;GET /v1/agents/sessions/{id}/events?stream=true&lt;/code&gt; with &lt;code&gt;Accept: text/event-stream&lt;/code&gt; before sending input so you do not miss early events.&lt;/p&gt;

&lt;p&gt;Watch for:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;code&gt;agent.session.turn.output_text.delta&lt;/code&gt; and &lt;code&gt;agent.session.turn.output_text.done&lt;/code&gt; for text output.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;agent.session.turn.completed&lt;/code&gt;, &lt;code&gt;agent.session.turn.failed&lt;/code&gt;, or &lt;code&gt;agent.session.turn.cancelled&lt;/code&gt; for the final turn outcome.&lt;/li&gt;
&lt;li&gt;
&lt;code&gt;agent.session.requires_action&lt;/code&gt; when the agent needs a function result, environment connection, or computer-use approval.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Avoid these common mistakes:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;
&lt;code&gt;agent.session.idle&lt;/code&gt; does not mean the turn succeeded.&lt;/li&gt;
&lt;li&gt;A completed turn can still contain failed tool calls.&lt;/li&gt;
&lt;li&gt;Closing an SSE stream does not stop the running task.&lt;/li&gt;
&lt;li&gt;Streams do not replay missed events. After a disconnect, open a new stream and retrieve the session and its items.&lt;/li&gt;
&lt;/ol&gt;

&lt;h3&gt;
  
  
  Use webhooks for long-running work
&lt;/h3&gt;

&lt;p&gt;Subscribe to:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;&lt;code&gt;agent.session.created&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;agent.session.action_required&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;agent.session.in_progress&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;agent.session.idle&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;&lt;code&gt;agent.session.failed&lt;/code&gt;&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The event names differ slightly:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Stream: &lt;code&gt;requires_action&lt;/code&gt;
&lt;/li&gt;
&lt;li&gt;Webhook: &lt;code&gt;action_required&lt;/code&gt;
&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Webhook payloads omit call details. In your webhook handler, retrieve the session and inspect &lt;code&gt;required_actions&lt;/code&gt;. Verify every signature using a &lt;a href="http://apidog.com/blog/webhook-signature-verification?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;webhook signature verification&lt;/a&gt; flow. For the architecture behind this pattern, see &lt;a href="http://apidog.com/blog/ai-agent-long-running-api-operations?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;long-running API operations&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  MCP tools, tool search, programmatic tool calling, and subagents
&lt;/h2&gt;

&lt;h3&gt;
  
  
  Add an MCP server
&lt;/h3&gt;

&lt;p&gt;Add an MCP server to &lt;code&gt;agent.tools&lt;/code&gt;:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"mcp"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"server_label"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai_docs"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"transport"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"http"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"server_url"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"https://developers.openai.com/mcp"&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"required"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;By default, OpenAI creates the connection with &lt;code&gt;connection_origin: "service"&lt;/code&gt;, so the server must be reachable from OpenAI.&lt;/p&gt;

&lt;p&gt;Use these alternatives when needed:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Set &lt;code&gt;connection_origin: "environment"&lt;/code&gt; for a private-network server.&lt;/li&gt;
&lt;li&gt;Use &lt;code&gt;stdio&lt;/code&gt; to start a server inside the sandbox.&lt;/li&gt;
&lt;li&gt;Pass &lt;code&gt;transport.authorization&lt;/code&gt; for credentials scoped to one session.&lt;/li&gt;
&lt;li&gt;Attach a vault credential (&lt;code&gt;static_bearer&lt;/code&gt; or &lt;code&gt;mcp_oauth&lt;/code&gt;) through &lt;code&gt;vault_ids&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;h3&gt;
  
  
  Use tool search for large tool sets
&lt;/h3&gt;

&lt;p&gt;MCP tools are discovered automatically when the model supports tool search.&lt;/p&gt;

&lt;p&gt;For many function tools, add:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"tool_search"&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Then mark functions that should load only when needed:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"defer_loading"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Programmatic tool calling
&lt;/h3&gt;

&lt;p&gt;Programmatic tool calling is enabled by default. The agent gets an &lt;code&gt;exec&lt;/code&gt; tool that runs JavaScript in an isolated V8 runtime. It can loop over tool calls and trim large results before adding them to model context.&lt;/p&gt;

&lt;p&gt;Disable it when you need every tool call handled directly:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"programmatic_tool_calling"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;false&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;h3&gt;
  
  
  Enable subagents
&lt;/h3&gt;

&lt;p&gt;Enable subagents with a concurrency limit:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"multi_agent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"max_concurrent_subagents"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;3&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The default limit is 6. Subagents share the environment filesystem and inherit MCP tools and web search, but they cannot use function tools. A turn’s &lt;code&gt;subagent_id&lt;/code&gt; is &lt;code&gt;null&lt;/code&gt; for the main agent.&lt;/p&gt;

&lt;h2&gt;
  
  
  Computer use: the DevDay addition
&lt;/h2&gt;

&lt;p&gt;Computer use gives the agent a hosted browser. Add the tool and enable a desktop in a hosted environment:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight json"&gt;&lt;code&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"agent"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"model"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"gpt-6-astra"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"tools"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"computer_use"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
        &lt;/span&gt;&lt;span class="nl"&gt;"include_screenshots"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="nl"&gt;"environment"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"type"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"openai_hosted"&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"desktop"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"enabled"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;},&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="nl"&gt;"network"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
      &lt;/span&gt;&lt;span class="nl"&gt;"access"&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;"enabled"&lt;/span&gt;&lt;span class="w"&gt;
    &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
  &lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;The browser requires user approval before visiting every new website origin, including public sites.&lt;/p&gt;

&lt;p&gt;When you receive &lt;code&gt;agent.session.requires_action&lt;/code&gt;:&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Retrieve the session.&lt;/li&gt;
&lt;li&gt;Find &lt;code&gt;computer_use_approval_request&lt;/code&gt; entries.&lt;/li&gt;
&lt;li&gt;Inspect the nested &lt;code&gt;request.type&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Send the appropriate result through the events endpoint.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;There are two request types:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;browser_origin_access&lt;/code&gt;:&lt;/strong&gt; Show &lt;code&gt;origin&lt;/code&gt; and &lt;code&gt;reason&lt;/code&gt;, then submit &lt;code&gt;approve&lt;/code&gt;, &lt;code&gt;deny&lt;/code&gt;, or &lt;code&gt;cancel&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;&lt;code&gt;browser_authentication&lt;/code&gt;:&lt;/strong&gt; Show the sign-in &lt;code&gt;fields&lt;/code&gt;, optional login &lt;code&gt;options&lt;/code&gt;, and &lt;code&gt;credential_origin&lt;/code&gt;. Submit &lt;code&gt;action: "submit"&lt;/code&gt; with user-provided values, or &lt;code&gt;action: "cancel"&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;Submit an origin approval result like this:&lt;br&gt;
&lt;/p&gt;

&lt;div class="highlight js-code-highlight"&gt;
&lt;pre class="highlight shell"&gt;&lt;code&gt;curl &lt;span class="s2"&gt;"https://api.openai.com/v1/agents/sessions/&lt;/span&gt;&lt;span class="nv"&gt;$SESSION_ID&lt;/span&gt;&lt;span class="s2"&gt;/events"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"OpenAI-Beta: agents=v1"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Authorization: Bearer &lt;/span&gt;&lt;span class="nv"&gt;$OPENAI_API_KEY&lt;/span&gt;&lt;span class="s2"&gt;"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-H&lt;/span&gt; &lt;span class="s2"&gt;"Content-Type: application/json"&lt;/span&gt; &lt;span class="se"&gt;\&lt;/span&gt;
  &lt;span class="nt"&gt;-d&lt;/span&gt; &lt;span class="s1"&gt;'{
    "events": [
      {
        "type": "agent.session.input.computer_use_approval_request_result",
        "request_id": "REQUEST_ID",
        "response": {
          "type": "browser_origin_access",
          "decision": "approve"
        }
      }
    ]
  }'&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;

&lt;/div&gt;



&lt;p&gt;Browser work is stored as &lt;code&gt;computer_use_call&lt;/code&gt; items. Each item can include &lt;code&gt;id&lt;/code&gt;, &lt;code&gt;turn_id&lt;/code&gt;, &lt;code&gt;title&lt;/code&gt;, &lt;code&gt;status&lt;/code&gt;, and &lt;code&gt;output&lt;/code&gt;. When &lt;code&gt;include_screenshots&lt;/code&gt; is enabled and a screenshot is available, &lt;code&gt;output&lt;/code&gt; carries a base64 JPEG.&lt;/p&gt;

&lt;p&gt;Do not write screenshots to logs. They can expose account data.&lt;/p&gt;

&lt;p&gt;The &lt;a href="https://developers.openai.com/api/docs/guides/agents-api/tools/computer-use" rel="noopener noreferrer"&gt;computer use guide&lt;/a&gt; highlights these implementation constraints:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;
&lt;strong&gt;Origin approval is not action confirmation.&lt;/strong&gt; Approving a site does not make the agent ask before each purchase or delete. If you need that control, restrict the browser to resources that cannot take those actions or use a browser runtime you control.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Sign-in covers email, passwords, and verification codes.&lt;/strong&gt; Passkeys and QR-code sign-in are not supported.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Only the main agent can request authentication.&lt;/strong&gt; Subagents cannot.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Disable automatic retries on credential submissions.&lt;/strong&gt; Use &lt;code&gt;maxRetries: 0&lt;/code&gt; in the SDK or &lt;code&gt;--retry 0&lt;/code&gt; with curl.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;A &lt;code&gt;202&lt;/code&gt; means accepted, not successful.&lt;/strong&gt; Navigation or sign-in may still fail. Authentication requests expire after five minutes.&lt;/li&gt;
&lt;li&gt;
&lt;strong&gt;Origin approval does not override network policy.&lt;/strong&gt; Add the site and its redirect domains to &lt;code&gt;network&lt;/code&gt; allow rules too.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;The recap says computer use ships “through the API and in Codex and ChatGPT Work on Pro 500 and Enterprise.” For UI-driven testing with the same model, see &lt;a href="http://apidog.com/blog/gpt-6-astra-computer-use-api-testing?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;GPT-6 Astra computer use for API testing&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Give the agent your API, not your UI
&lt;/h2&gt;

&lt;p&gt;A browser is a fallback for software without an API. If you own the system, wrap it in an MCP server instead:&lt;/p&gt;

&lt;ul&gt;
&lt;li&gt;Expose typed tools.&lt;/li&gt;
&lt;li&gt;Avoid origin approval prompts.&lt;/li&gt;
&lt;li&gt;Return structured, checkable results.&lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;a href="http://apidog.com/blog/computer-use-vs-structured-apis?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Computer use vs structured APIs&lt;/a&gt; covers the trade-off. The &lt;a href="http://apidog.com/blog/apidog-mcp-server?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog MCP Server&lt;/a&gt; can feed your API specification to the coding assistant that writes the wrapper.&lt;/p&gt;

&lt;h2&gt;
  
  
  Test the Agents API in Apidog before writing code
&lt;/h2&gt;

&lt;p&gt;The API is in beta, so validate every request shape manually in &lt;a href="https://apidog.com?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Apidog&lt;/a&gt; first:&lt;/p&gt;

&lt;p&gt;&lt;a href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmanc6vo56npfm0hb7wt8.png" class="article-body-image-wrapper"&gt;&lt;img src="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fdev-to-uploads.s3.us-east-2.amazonaws.com%2Fuploads%2Farticles%2Fmanc6vo56npfm0hb7wt8.png" alt="Apidog interface" width="799" height="530"&gt;&lt;/a&gt;&lt;/p&gt;

&lt;ol&gt;
&lt;li&gt;Create an Apidog environment with &lt;code&gt;OPENAI_API_KEY&lt;/code&gt;, &lt;code&gt;VAULT_ID&lt;/code&gt;, and &lt;code&gt;SESSION_ID&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Send &lt;code&gt;Bearer {{OPENAI_API_KEY}}&lt;/code&gt; and &lt;code&gt;OpenAI-Beta: agents=v1&lt;/code&gt; on every request.&lt;/li&gt;
&lt;li&gt;Send the create-session request without &lt;code&gt;stream&lt;/code&gt;. Assert a 2xx status and a non-empty &lt;code&gt;id&lt;/code&gt;, then extract &lt;code&gt;id&lt;/code&gt; into &lt;code&gt;SESSION_ID&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;Open the event stream as an SSE request. Send input from a second request and verify that events arrive.&lt;/li&gt;
&lt;li&gt;Save approval and cancellation payloads as reusable requests for each &lt;code&gt;required_actions&lt;/code&gt; case.&lt;/li&gt;
&lt;li&gt;Chain requests into a test scenario and run it in CI with the Apidog CLI.&lt;/li&gt;
&lt;/ol&gt;

&lt;p&gt;The &lt;a href="http://apidog.com/blog/how-to-test-ai-agents-api?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;AI agent API testing guide&lt;/a&gt; includes assertion patterns for non-deterministic output. &lt;a href="https://apidog.com/download?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;Download Apidog&lt;/a&gt; to follow along.&lt;/p&gt;

&lt;h2&gt;
  
  
  FAQ
&lt;/h2&gt;

&lt;p&gt;&lt;strong&gt;Is the OpenAI Agents API free?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;There is no platform fee, but you pay for model tokens, tool calls, and hosted container time.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Which models work with the Agents API?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The docs’ examples, including every computer-use example, use &lt;code&gt;gpt-6-astra&lt;/code&gt;. The pages do not list other supported models, so test yours first.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;Does the Agents API support Zero Data Retention?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;No. It supports US data residency only and is not ZDR-eligible, even with a self-hosted sandbox.&lt;/p&gt;

&lt;p&gt;&lt;strong&gt;How is it different from the Agents SDK or the Responses API?&lt;/strong&gt;&lt;/p&gt;

&lt;p&gt;The SDK runs the loop inside your application. The Responses API is the model call that you build a loop around. See &lt;a href="http://apidog.com/blog/openai-agents-api-vs-responses-api-vs-agents-sdk?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;the full comparison&lt;/a&gt;.&lt;/p&gt;

&lt;h2&gt;
  
  
  Start with one read-only session
&lt;/h2&gt;

&lt;p&gt;Begin with a read-only session. Add one MCP server, then add computer use behind an approval handler that denies by default.&lt;/p&gt;

&lt;p&gt;When ChatGPT should react to events from your own server instead, &lt;a href="http://apidog.com/blog/mcp-events?utm_source=dev.to&amp;amp;utm_medium=wanda&amp;amp;utm_content=n8n-post-automation"&gt;MCP Events&lt;/a&gt; is the matching piece.&lt;/p&gt;

</description>
    </item>
  </channel>
</rss>
