Gemini 4 Argon’s headline number—1M tokens—is its output limit, not its context window. Google says a single Argon response can run to 1 million tokens, roughly 16× the previous 64K cap, but it has not published Argon’s input window. There is also nothing to call yet: Argon is currently available only to Fairwind Program defenders, with paid API customers next when Google opens access. See the release date and access guide.
A response this large breaks common API assumptions: requests do not finish in seconds, response bodies do not safely fit in memory, and request cost is no longer small or predictable. This guide shows how to plan for cost, streaming, timeouts, output caps, storage, and testing in Apidog before access arrives. For the model overview, see what Gemini 4 Argon is. For request shapes, see the Gemini 4 Argon API guide.
1M output tokens is not a 1M context window
Do not treat Argon’s 1M figure as a context-window specification. Google’s launch post says it expanded “the model’s output token limit to an industry-leading 1M tokens.”
That 64K comparison matches the 65,536-token output cap of Gemini 3.1 Pro Preview, Google’s previous top Pro model.
Argon’s input window is a separate, unpublished value. Google’s long-context evaluation, described in its evals methodology, used prompts between 256K and 1M tokens. That is a benchmark range, not an API specification.
There is also an output caveat: Vals AI lists a 262K maximum output for the Argon configuration it tested. Google states a 1M-token model limit, while at least one third-party evaluator observed a lower cap on the endpoint it used.
| Model | Max output per response | Input or context window |
|---|---|---|
| Gemini 4 Argon | 1M (Google’s stated limit) | Not published |
| Gemini 3.1 Pro Preview | 65,536 | 1,048,576 |
| Gemini 3.8 Flash | 65,536 | 1,048,576 |
| GPT-6 Astra | 128,000 | 1,050,000 (922K max input) |
| Claude Opus 5.5 | 128K (300K on Batch with a beta header) | 1M |
Every competitor in this table caps synchronous output at 128K, making Argon’s stated 1M limit about 8× larger. Google’s stated reason is reasoning depth: the model can “generate hundreds of thousands of tokens in a single trajectory” to solve hard problems in one pass.
For other long-running API patterns, see Claude Opus 5.5’s 18-hour tasks and the GPT-6 Astra API guide.
Calculate the worst-case cost first
At Argon’s stated rates, output costs:
- Intro period: $10 per 1M output tokens
- Standard pricing: $20 per 1M output tokens
A full-length response therefore costs:
- Intro:
1,000,000 × $10 / 1,000,000 = $10.00 - Standard:
1,000,000 × $20 / 1,000,000 = $20.00
Input is additional. For example, a 200,000-token prompt adds:
- Intro:
200,000 × $2 / 1,000,000 = $0.40 - Standard:
$0.80
That makes a maxed-out call:
- $10.40 at intro rates
- $20.80 at standard rates
A nightly job that runs 100 such requests would cost $1,040 at intro pricing.
Thinking usage can make this less obvious. On current Gemini models, thinking tokens bill as output. Google has not stated whether Argon follows the same rule or whether thinking counts toward its 1M ceiling. Either way, a short visible answer can still incur substantial output usage.
For additional scenarios, including cached input at 95% off, see the Gemini 4 Argon pricing guide.
Stream every long response
A non-streaming request returns nothing until the entire response is complete. For outputs that may reach hundreds of thousands of tokens, that creates a long silent connection. Any client, proxy, gateway, load balancer, or serverless timeout can terminate it before the first byte arrives.
Use server-sent events (SSE) instead.
For generateContent, use:
:streamGenerateContent?alt=sse
Read each event as it arrives and write each text chunk directly to storage. Do not accumulate the full response body in memory.
This example runs against Gemini 3.8 Flash today. Keep the model name in an environment variable because Google has not published Argon’s model ID. See the Gemini 3.8 Flash API guide for setup details.
import json
import os
import requests
MODEL = os.environ.get("GEMINI_MODEL", "gemini-3.8-flash")
URL = (
"https://generativelanguage.googleapis.com/v1beta/models/"
f"{MODEL}:streamGenerateContent?alt=sse"
)
body = {
"contents": [
{
"parts": [
{
"text": "Write a test plan for every endpoint in a payments API."
}
]
}
],
"generationConfig": {
"maxOutputTokens": 60000
},
}
usage = None
with requests.post(
URL,
json=body,
stream=True,
timeout=(10, 120),
headers={"x-goog-api-key": os.environ["GEMINI_API_KEY"]},
) as response, open("response.txt", "a", encoding="utf-8") as output:
response.raise_for_status()
for line in response.iter_lines(decode_unicode=True):
if not line or not line.startswith("data:"):
continue
event = json.loads(line[5:])
for candidate in event.get("candidates", []):
for part in candidate.get("content", {}).get("parts", []):
output.write(part.get("text", ""))
output.flush()
usage = event.get("usageMetadata", usage)
print(usage)
timeout=(10, 120) configures:
- A 10-second connection timeout
- A 120-second read timeout
In requests, the read timeout is the maximum gap between received bytes, not the total request duration. A stream can continue indefinitely if it keeps sending data within that interval.
The example writes each chunk to disk immediately. On Gemini 3.8 Flash, every event includes running usageMetadata, so the final event contains the token counts to record for cost tracking.
Check timeouts at every hop
Your HTTP client is only one part of the request path. A long stream can also cross a reverse proxy, API gateway, load balancer, and serverless runtime.
| Hop | What to check | Symptom when misconfigured |
|---|---|---|
| HTTP client | Read or idle timeout, plus any total-request timeout | Exceptions mid-stream only on long responses |
| Reverse proxy | Read timeout and response buffering for text/event-stream
|
Events arrive in bursts, or the stream cuts off |
| API gateway | Maximum request duration | Requests fail at the same elapsed time every run |
| Load balancer | Idle timeout | Drops during long pauses before the first event |
| Serverless function | Maximum execution time | The function exits while the model is still generating |
A fixed cutoff is the key signal. If long requests always fail at the same elapsed time, a hop in the path has a hard duration limit. Streaming alone cannot solve that; move the work off the live request path.
Move the longest jobs to background execution
For very long tasks, do not hold an HTTP request open.
Google’s Interactions API supports background execution with:
background=true
Background execution depends on stored interactions. The documentation states that store=false is incompatible with background execution, so leave storage enabled for these requests.
For retrieving a completed background interaction, follow Google’s documented workflow rather than guessing polling endpoints. Since Google says new models launch on the Interactions API, plan for Argon’s longest jobs to run there.
Set output caps intentionally
The 1M limit is a ceiling, not a target.
For generateContent, use generationConfig.maxOutputTokens to cap each response:
{
"generationConfig": {
"maxOutputTokens": 60000
}
}
On Gemini 3.8 Flash, thinking counts against that cap. In one test, a cap of 2,000 produced:
- 1,340 thought tokens
- 656 visible tokens
For the Interactions API, verify the output-cap field in Google’s current documentation before relying on it.
Choose a cap based on the maximum cost you accept per request:
| Output cap | Worst-case standard output cost ($20/1M) | Intro output cost ($10/1M) |
|---|---|---|
| 64,000 | $1.28 | $0.64 |
| 128,000 | $2.56 | $1.28 |
| 500,000 | $10.00 | $5.00 |
| 1,000,000 | $20.00 | $10.00 |
Treat a capped response as potentially incomplete. Inspect the final event’s finishReason:
MAX_TOKENS
If the response ended with MAX_TOKENS, either continue in a follow-up turn or increase the cap for that specific job.
Store and parse huge outputs without buffering
A million tokens can produce megabytes of text. Use a streaming storage path from the start.
- Write chunks as they arrive to a file or multipart object upload.
- Keep partial files when a connection drops. A blind retry can regenerate and rebill the same output.
- Request JSON Lines when you need structured output, so each line can be parsed independently.
- Log
usageMetadataand byte counts rather than full response bodies. - Verify database column limits and queue message-size limits before sending an 800K-token response through them.
Test the pipeline in Apidog before access opens
You can validate the entire workflow against a stand-in model. Download Apidog and run these three checks.
1. Inspect the SSE stream
Send the streaming request to Gemini 3.8 Flash with GEMINI_API_KEY and GEMINI_MODEL configured as environment variables.
Apidog parses text/event-stream responses and displays each event in the Timeline view as it arrives. Use this to inspect:
- Chunk sizes
- Gaps between events
- Final
usageMetadata - Whether a proxy buffers the response
2. Stream a much longer fake response
Gemini 3.8 Flash currently caps output at 65,536 tokens. To test parsing, buffering, and timeout behavior beyond that limit, run a local mock that emits Gemini-shaped SSE events:
# long_stream_mock.py: Gemini-shaped SSE for parser and timeout tests (fake data)
import json
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer
EVENTS = 20000
DELAY = 0.005
CHUNK = "lorem ipsum " * 40
class Handler(BaseHTTPRequestHandler):
def do_POST(self):
self.rfile.read(int(self.headers.get("Content-Length", 0)))
self.send_response(200)
self.send_header("Content-Type", "text/event-stream")
self.end_headers()
for i in range(EVENTS):
event = {
"candidates": [
{
"content": {
"parts": [
{
"text": CHUNK
}
]
}
}
]
}
if i == EVENTS - 1:
event["usageMetadata"] = {
"promptTokenCount": 1200,
"candidatesTokenCount": 950000,
"thoughtsTokenCount": 40000,
"totalTokenCount": 991200,
}
self.wfile.write(
f"data: {json.dumps(event)}\n\n".encode()
)
self.wfile.flush()
time.sleep(DELAY)
ThreadingHTTPServer(("127.0.0.1", 8787), Handler).serve_forever()
Set a mock environment base URL to:
http://127.0.0.1:8787
Then send the same streaming request through it.
This mock runs for about two minutes—128 seconds in the original test—and emits 9.6 million characters of text. It is large enough to expose:
- A parser that buffers the entire body
- A proxy that holds SSE events
- A timeout configured too aggressively
- Storage code that accumulates output in memory
3. Assert token and cost limits
For a non-streaming generateContent request, assert that:
candidatesTokenCount + thoughtsTokenCount <= configured output cap
Also compute the request cost and assert that it remains below your per-request ceiling at Argon pricing.
The Argon API guide includes a ready-made cost script.
FAQ
Is 1M Gemini 4 Argon’s context window?
No. It is the stated output limit per response, up from 64K. Google has not published Argon’s input window.
How much does a 1M-token Argon response cost?
Output costs $10 at intro rates and $20 at standard rates, plus input costs. See Gemini 4 Argon pricing for more scenarios.
Can I generate a 1M-token response today?
Not unless your organization is in the Fairwind cohort with Argon access. Gemini 3.8 Flash and 3.1 Pro Preview cap output at 65,536 tokens, while Vals AI lists 262K maximum output for the Argon configuration it tested.
Do I have to stream long Argon responses?
Google has not published Argon-specific streaming guidance. However, a non-streaming request that runs for minutes is exposed to every idle timeout in your stack. Stream the response or use background execution through the Interactions API.
How does Argon compare with GPT-6 Astra and Claude Opus 5.5?
Both cap synchronous output at 128K. Anthropic allows 300K on Batch with a beta header. Argon’s stated 1M output limit is about 8× higher.
Your next step
Add streaming and an explicit output cap to your Gemini client now, using Gemini 3.8 Flash. Run the client against the long mock stream until no component in your stack cuts the response off.
When Argon’s model ID becomes available, update GEMINI_MODEL and rerun the same tests in Apidog.

Top comments (0)