There is no public Gemini 4 Argon API yet, and Google has not published a model ID. Google announced Argon on September 30, 2026, and it is currently available only to Fairwind Program defenders: a set of Fairwind partners using it as a managed model in Gemini Enterprise. Artificial Analysis lists Google AI Studio as Argon’s API provider, but provides no speed or latency data. That suggests allowlisted pre-release access rather than an endpoint you can sign up for. When the wider rollout begins, Google says it will start “with paid API customers and Google AI Ultra subscribers,” so a paid Gemini API key puts you first in line.
That gap between announcement and access is useful preparation time. This guide separates confirmed Argon API details from unknowns, then shows how to build against Gemini 3.8 Flash today so moving to Argon is a one-variable change. You will:
- Send Interactions API and
generateContentrequests. - Detect availability with
models.list. - Mock responses in Apidog.
- Add a per-request cost ceiling using Argon pricing.
For background, start with what Gemini 4 Argon is. For rollout timing, see the Gemini 4 Argon release date guide.
What’s confirmed about the Gemini 4 Argon API, and what isn’t
Google’s launch post confirms prices and an output limit. Most integration details are still unpublished.
| Item | Status | Detail |
|---|---|---|
| Input price | Confirmed | $2 per 1M tokens during the intro period; $4 after |
| Output price | Confirmed | $10 per 1M tokens during the intro period; $20 after |
| Cached input | Confirmed | 95% off input: $0.10 intro, $0.20 standard |
| Output limit | Confirmed | 1M tokens, up from 64K |
| API surface | Signaled | Google docs say all new models launch on the Interactions API |
| Model ID | Not published | Absent from the models page, pricing page, and changelog |
| Input context window | Not published | Google’s long-context eval used prompts up to 1M tokens, but that is not a specification |
| Thinking levels | Not published | Evals used the “highest thinking settings”; names and defaults are unknown |
| Rate limits | Not published | No tiers announced |
| Batch support | Not published | No batch or Flex pricing announced |
| Long-prompt tier | Not published | Gemini 3.1 Pro charges more above 200K tokens; Argon’s rule is unknown |
| Intro period length | Not published | No end date for the $2/$10 pricing |
Two details affect implementation:
- The API surface expectation comes from the Interactions API docs, which state that new models “will launch on the Interactions API.” Expect Argon there first, but do not assume
generateContentsupport until Google confirms it. - Google states a 1M-token output limit, but Vals AI lists a 262K maximum output for the configuration it tested. Check the API response on launch day instead of assuming every endpoint exposes the full million-token limit.
The 1M output tokens guide covers the impact on streaming, timeouts, and storage.
At introductory rates, a request with a 20,000-token prompt and 5,000 output tokens costs:
20,000 × $2 / 1M = $0.04 input
5,000 × $10 / 1M = $0.05 output
--------------------------------
Total = $0.09
At standard $4/$20 rates, that same request costs $0.18. Argon’s introductory output price ($10 per 1M tokens) is below Gemini 3.1 Pro Preview’s $12, while its standard output price matches Claude Opus 5.5. For more scenarios, including cached input and a maximum-size output, see the Gemini 4 Argon pricing breakdown.
Don’t hard-code a model ID you found online
Search results may show Argon-like model IDs from benchmark sites, aggregators, or open-source pull requests. Treat them as placeholders.
Google has not confirmed a public model ID. A guessed ID can cause:
- A production 404 on deployment day.
- A silent fallback if your SDK or wrapper catches errors broadly.
- Incorrect assumptions about supported methods or token limits.
Keep the model name in configuration:
export GEMINI_MODEL=gemini-3.8-flash
Every example below reads GEMINI_MODEL and defaults to gemini-3.8-flash, a stable model available today. Once Google publishes Argon’s real ID, update the environment variable rather than modifying application code.
Get your code ready on Gemini 3.8 Flash
Gemini 3.8 Flash (gemini-3.8-flash) uses the endpoints, headers, and response shapes Argon is expected to use. Its pricing is $0.75 input and $3.75 output per 1M tokens through December 31, 2026.
Store your API key in GEMINI_API_KEY; never commit it to source control. See the Gemini 3.8 Flash API guide for setup details.
Step 1: Send an Interactions API request
The Interactions API has been generally available since June 2026 and is Google’s primary API surface for new models.
Keep both the model and thinking level configurable because Argon’s thinking-level names are not published.
MODEL="${GEMINI_MODEL:-gemini-3.8-flash}"
THINKING="${GEMINI_THINKING:-medium}"
curl -s -X POST "https://generativelanguage.googleapis.com/v1beta/interactions" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"model\": \"${MODEL}\",
\"input\": \"List the retry rules a REST client should follow for HTTP 429.\",
\"generation_config\": {\"thinking_level\": \"${THINKING}\"}
}"
Use the Python SDK with google-genai 2.3.0 or later:
# pip install "google-genai>=2.3.0"
import os
from google import genai
MODEL = os.environ.get("GEMINI_MODEL", "gemini-3.8-flash")
THINKING = os.environ.get("GEMINI_THINKING", "medium")
client = genai.Client() # Reads GEMINI_API_KEY from the environment
interaction = client.interactions.create(
model=MODEL,
input="List the retry rules a REST client should follow for HTTP 429.",
generation_config={"thinking_level": THINKING},
)
print(interaction.output_text)
For Gemini 3.8 Flash, valid thinking levels are:
low-
medium— default high
Using minimal returns HTTP 400. See the thinking levels guide for trade-offs.
For multi-turn interactions, pass the prior response ID as previous_interaction_id. To opt out of server-side storage, send store: false.
Step 2: Keep the generateContent path working
Existing applications may use the legacy generateContent endpoint, which Google says remains fully supported.
curl -s "https://generativelanguage.googleapis.com/v1beta/models/${MODEL}:generateContent" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d "{
\"contents\": [{
\"parts\": [{
\"text\": \"List the retry rules a REST client should follow for HTTP 429.\"
}]
}],
\"generationConfig\": {
\"thinkingConfig\": {
\"thinkingLevel\": \"${THINKING}\"
}
}
}"
The thinking configuration differs between APIs:
| API | Thinking field |
|---|---|
| Interactions | generation_config.thinking_level |
generateContent |
generationConfig.thinkingConfig.thinkingLevel |
If your client supports both paths, route new models through Interactions by default and retain generateContent as a fallback. Apply the same care to tool definitions; see Gemini 3.8 Flash function calling.
Step 3: Schedule a go-live check with models.list
Do not manually refresh changelogs. Poll the API your production application will use.
The models endpoint returns available models, token limits, and supported methods:
import os
import sys
import requests
resp = requests.get(
"https://generativelanguage.googleapis.com/v1beta/models",
params={"pageSize": 1000},
headers={"x-goog-api-key": os.environ["GEMINI_API_KEY"]},
timeout=30,
)
resp.raise_for_status()
hits = [
model
for model in resp.json().get("models", [])
if "argon" in model["name"].lower()
]
for model in hits:
print(
model["name"],
"in:", model.get("inputTokenLimit"),
"out:", model.get("outputTokenLimit"),
model.get("supportedGenerationMethods"),
)
sys.exit(1 if hits else 0) # A failed run becomes the alert.
Run this hourly through cron or a CI schedule using the same API key as your production app.
When this call was run on October 1, 2026 with a paid-project key, it returned 61 models, no nextPageToken, and no model name containing argon.
When the script finds Argon, inspect:
- The real model ID.
- Whether
outputTokenLimitis the full 1M tokens. - Which generation methods are supported.
Step 4: Mock the response before you have access
You can build Argon client logic before Argon access is available.
- Save the Step 2
generateContentrequest in Apidog. - Create
GEMINI_API_KEYandGEMINI_MODELenvironment variables. - Send the request once against Gemini 3.8 Flash.
- Save the actual response as an endpoint response example.
- Set the project mock behavior to Response example first:
- Project Settings
- Feature Settings
- Mock Settings
Apidog’s default Smart Mock generates values from the schema. Using Response example first makes the mock URL return your saved Gemini response instead. Your frontend, queue workers, and parsers can then run against the mock without consuming tokens.
This is a current Gemini response schema mock, not an Argon-specific schema. Google has not published an Argon schema. It is still a practical target because Argon is expected to use the same API surface.
To simulate a large Argon response, edit the example’s usageMetadata values and verify that billing, alerting, queueing, and storage logic behave correctly.
Step 5: Assert on usageMetadata and enforce a cost ceiling
Every generateContent response includes usageMetadata. Add an Apidog post-processor that prices each response at Argon’s standard rates.
// Apidog post-processor: price this response at Gemini 4 Argon's standard rates
const u = pm.response.json().usageMetadata;
const IN = 4 / 1e6; // USD per input token after the intro period
const OUT = 20 / 1e6; // USD per output token
// Thinking bills as output on current Gemini models.
const out =
(u.candidatesTokenCount || 0) +
(u.thoughtsTokenCount || 0);
const cost = u.promptTokenCount * IN + out * OUT;
pm.test("usageMetadata is present", () => {
pm.expect(u).to.be.an("object");
});
pm.test("cost under $0.10 at Argon rates", () => {
pm.expect(cost).to.be.below(0.10);
});
Choose the ceiling per request. A $0.10 limit is reasonable for the short prompt used in this example.
Google has not stated how Argon bills thinking tokens. This script assumes the current Gemini billing behavior, where thinking tokens are priced as output. Keep the same saved request and run it against Gemini 3.8 Flash now, then Argon later. The token and cost differences become your regression comparison.
FAQ
Is the Gemini 4 Argon API available?
Not publicly. Argon is currently rolling out only to Fairwind Program partners through Gemini Enterprise. Paid API customers and Google AI Ultra subscribers are next, but Google has not provided a date.
What is the Gemini 4 Argon model ID?
Google has not published one. Third-party strings are placeholders, so keep the model name in an environment variable and monitor models.list.
How much will the Gemini 4 Argon API cost?
It costs $2 input and $10 output per 1M tokens during an introductory period of unknown length, then $4 input and $20 output per 1M tokens. Cached input is discounted by 95%. See Gemini 4 Argon pricing.
Will Argon work with generateContent?
Google has not said. Its documentation says new models launch on the Interactions API, so build for Interactions first and treat generateContent as a fallback.
Does Google AI Ultra give me API access?
No. AI Ultra is a consumer subscription, not an API key. Google says API access begins with paid API customers.
Your next step
Set GEMINI_MODEL in every environment, schedule the models.list check, and save both Interactions and generateContent requests with the cost assertion attached.
Download Apidog to keep requests, mocks, and tests in one workspace. When Google publishes Argon’s model ID, update one environment variable.

Top comments (0)