DEV Community

Cover image for Gemini 4 Argon's 1M Output Tokens: What a Million-Token Response Does to Your API Stack
Hassann
Hassann

Posted on Originally published at apidog.com

Gemini 4 Argon's 1M Output Tokens: What a Million-Token Response Does to Your API Stack

Gemini 4 Argon’s headline number—1M tokens—is its output limit, not its context window. Google says a single Argon response can run to 1 million tokens, roughly 16× the previous 64K cap, but it has not published Argon’s input window. There is also nothing to call yet: Argon is currently available only to Fairwind Program defenders, with paid API customers next when Google opens access. See the release date and access guide.

Try Apidog today

A response this large breaks common API assumptions: requests do not finish in seconds, response bodies do not safely fit in memory, and request cost is no longer small or predictable. This guide shows how to plan for cost, streaming, timeouts, output caps, storage, and testing in Apidog before access arrives. For the model overview, see what Gemini 4 Argon is. For request shapes, see the Gemini 4 Argon API guide.

1M output tokens is not a 1M context window

Do not treat Argon’s 1M figure as a context-window specification. Google’s launch post says it expanded “the model’s output token limit to an industry-leading 1M tokens.”

That 64K comparison matches the 65,536-token output cap of Gemini 3.1 Pro Preview, Google’s previous top Pro model.

Argon’s input window is a separate, unpublished value. Google’s long-context evaluation, described in its evals methodology, used prompts between 256K and 1M tokens. That is a benchmark range, not an API specification.

There is also an output caveat: Vals AI lists a 262K maximum output for the Argon configuration it tested. Google states a 1M-token model limit, while at least one third-party evaluator observed a lower cap on the endpoint it used.

Model Max output per response Input or context window
Gemini 4 Argon 1M (Google’s stated limit) Not published
Gemini 3.1 Pro Preview 65,536 1,048,576
Gemini 3.8 Flash 65,536 1,048,576
GPT-6 Astra 128,000 1,050,000 (922K max input)
Claude Opus 5.5 128K (300K on Batch with a beta header) 1M

Every competitor in this table caps synchronous output at 128K, making Argon’s stated 1M limit about 8× larger. Google’s stated reason is reasoning depth: the model can “generate hundreds of thousands of tokens in a single trajectory” to solve hard problems in one pass.

For other long-running API patterns, see Claude Opus 5.5’s 18-hour tasks and the GPT-6 Astra API guide.

Gemini 4 Argon output token illustration

Calculate the worst-case cost first

At Argon’s stated rates, output costs:

  • Intro period: $10 per 1M output tokens
  • Standard pricing: $20 per 1M output tokens

A full-length response therefore costs:

  • Intro: 1,000,000 × $10 / 1,000,000 = $10.00
  • Standard: 1,000,000 × $20 / 1,000,000 = $20.00

Input is additional. For example, a 200,000-token prompt adds:

  • Intro: 200,000 × $2 / 1,000,000 = $0.40
  • Standard: $0.80

That makes a maxed-out call:

  • $10.40 at intro rates
  • $20.80 at standard rates

A nightly job that runs 100 such requests would cost $1,040 at intro pricing.

Thinking usage can make this less obvious. On current Gemini models, thinking tokens bill as output. Google has not stated whether Argon follows the same rule or whether thinking counts toward its 1M ceiling. Either way, a short visible answer can still incur substantial output usage.

For additional scenarios, including cached input at 95% off, see the Gemini 4 Argon pricing guide.

Stream every long response

A non-streaming request returns nothing until the entire response is complete. For outputs that may reach hundreds of thousands of tokens, that creates a long silent connection. Any client, proxy, gateway, load balancer, or serverless timeout can terminate it before the first byte arrives.

Use server-sent events (SSE) instead.

For generateContent, use:

:streamGenerateContent?alt=sse
Enter fullscreen mode Exit fullscreen mode

Read each event as it arrives and write each text chunk directly to storage. Do not accumulate the full response body in memory.

This example runs against Gemini 3.8 Flash today. Keep the model name in an environment variable because Google has not published Argon’s model ID. See the Gemini 3.8 Flash API guide for setup details.

import json
import os
import requests

MODEL = os.environ.get("GEMINI_MODEL", "gemini-3.8-flash")
URL = (
    "https://generativelanguage.googleapis.com/v1beta/models/"
    f"{MODEL}:streamGenerateContent?alt=sse"
)

body = {
    "contents": [
        {
            "parts": [
                {
                    "text": "Write a test plan for every endpoint in a payments API."
                }
            ]
        }
    ],
    "generationConfig": {
        "maxOutputTokens": 60000
    },
}

usage = None

with requests.post(
    URL,
    json=body,
    stream=True,
    timeout=(10, 120),
    headers={"x-goog-api-key": os.environ["GEMINI_API_KEY"]},
) as response, open("response.txt", "a", encoding="utf-8") as output:
    response.raise_for_status()

    for line in response.iter_lines(decode_unicode=True):
        if not line or not line.startswith("data:"):
            continue

        event = json.loads(line[5:])

        for candidate in event.get("candidates", []):
            for part in candidate.get("content", {}).get("parts", []):
                output.write(part.get("text", ""))

        output.flush()
        usage = event.get("usageMetadata", usage)

print(usage)
Enter fullscreen mode Exit fullscreen mode

timeout=(10, 120) configures:

  • A 10-second connection timeout
  • A 120-second read timeout

In requests, the read timeout is the maximum gap between received bytes, not the total request duration. A stream can continue indefinitely if it keeps sending data within that interval.

The example writes each chunk to disk immediately. On Gemini 3.8 Flash, every event includes running usageMetadata, so the final event contains the token counts to record for cost tracking.

Check timeouts at every hop

Your HTTP client is only one part of the request path. A long stream can also cross a reverse proxy, API gateway, load balancer, and serverless runtime.

Hop What to check Symptom when misconfigured
HTTP client Read or idle timeout, plus any total-request timeout Exceptions mid-stream only on long responses
Reverse proxy Read timeout and response buffering for text/event-stream Events arrive in bursts, or the stream cuts off
API gateway Maximum request duration Requests fail at the same elapsed time every run
Load balancer Idle timeout Drops during long pauses before the first event
Serverless function Maximum execution time The function exits while the model is still generating

A fixed cutoff is the key signal. If long requests always fail at the same elapsed time, a hop in the path has a hard duration limit. Streaming alone cannot solve that; move the work off the live request path.

Move the longest jobs to background execution

For very long tasks, do not hold an HTTP request open.

Google’s Interactions API supports background execution with:

background=true
Enter fullscreen mode Exit fullscreen mode

Background execution depends on stored interactions. The documentation states that store=false is incompatible with background execution, so leave storage enabled for these requests.

For retrieving a completed background interaction, follow Google’s documented workflow rather than guessing polling endpoints. Since Google says new models launch on the Interactions API, plan for Argon’s longest jobs to run there.

Set output caps intentionally

The 1M limit is a ceiling, not a target.

For generateContent, use generationConfig.maxOutputTokens to cap each response:

{
  "generationConfig": {
    "maxOutputTokens": 60000
  }
}
Enter fullscreen mode Exit fullscreen mode

On Gemini 3.8 Flash, thinking counts against that cap. In one test, a cap of 2,000 produced:

  • 1,340 thought tokens
  • 656 visible tokens

For the Interactions API, verify the output-cap field in Google’s current documentation before relying on it.

Choose a cap based on the maximum cost you accept per request:

Output cap Worst-case standard output cost ($20/1M) Intro output cost ($10/1M)
64,000 $1.28 $0.64
128,000 $2.56 $1.28
500,000 $10.00 $5.00
1,000,000 $20.00 $10.00

Treat a capped response as potentially incomplete. Inspect the final event’s finishReason:

MAX_TOKENS
Enter fullscreen mode Exit fullscreen mode

If the response ended with MAX_TOKENS, either continue in a follow-up turn or increase the cap for that specific job.

Store and parse huge outputs without buffering

A million tokens can produce megabytes of text. Use a streaming storage path from the start.

  • Write chunks as they arrive to a file or multipart object upload.
  • Keep partial files when a connection drops. A blind retry can regenerate and rebill the same output.
  • Request JSON Lines when you need structured output, so each line can be parsed independently.
  • Log usageMetadata and byte counts rather than full response bodies.
  • Verify database column limits and queue message-size limits before sending an 800K-token response through them.

Test the pipeline in Apidog before access opens

You can validate the entire workflow against a stand-in model. Download Apidog and run these three checks.

1. Inspect the SSE stream

Send the streaming request to Gemini 3.8 Flash with GEMINI_API_KEY and GEMINI_MODEL configured as environment variables.

Apidog parses text/event-stream responses and displays each event in the Timeline view as it arrives. Use this to inspect:

  • Chunk sizes
  • Gaps between events
  • Final usageMetadata
  • Whether a proxy buffers the response

2. Stream a much longer fake response

Gemini 3.8 Flash currently caps output at 65,536 tokens. To test parsing, buffering, and timeout behavior beyond that limit, run a local mock that emits Gemini-shaped SSE events:

# long_stream_mock.py: Gemini-shaped SSE for parser and timeout tests (fake data)
import json
import time
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer

EVENTS = 20000
DELAY = 0.005
CHUNK = "lorem ipsum " * 40


class Handler(BaseHTTPRequestHandler):
    def do_POST(self):
        self.rfile.read(int(self.headers.get("Content-Length", 0)))

        self.send_response(200)
        self.send_header("Content-Type", "text/event-stream")
        self.end_headers()

        for i in range(EVENTS):
            event = {
                "candidates": [
                    {
                        "content": {
                            "parts": [
                                {
                                    "text": CHUNK
                                }
                            ]
                        }
                    }
                ]
            }

            if i == EVENTS - 1:
                event["usageMetadata"] = {
                    "promptTokenCount": 1200,
                    "candidatesTokenCount": 950000,
                    "thoughtsTokenCount": 40000,
                    "totalTokenCount": 991200,
                }

            self.wfile.write(
                f"data: {json.dumps(event)}\n\n".encode()
            )
            self.wfile.flush()
            time.sleep(DELAY)


ThreadingHTTPServer(("127.0.0.1", 8787), Handler).serve_forever()
Enter fullscreen mode Exit fullscreen mode

Set a mock environment base URL to:

http://127.0.0.1:8787
Enter fullscreen mode Exit fullscreen mode

Then send the same streaming request through it.

This mock runs for about two minutes—128 seconds in the original test—and emits 9.6 million characters of text. It is large enough to expose:

  • A parser that buffers the entire body
  • A proxy that holds SSE events
  • A timeout configured too aggressively
  • Storage code that accumulates output in memory

3. Assert token and cost limits

For a non-streaming generateContent request, assert that:

candidatesTokenCount + thoughtsTokenCount <= configured output cap
Enter fullscreen mode Exit fullscreen mode

Also compute the request cost and assert that it remains below your per-request ceiling at Argon pricing.

The Argon API guide includes a ready-made cost script.

FAQ

Is 1M Gemini 4 Argon’s context window?

No. It is the stated output limit per response, up from 64K. Google has not published Argon’s input window.

How much does a 1M-token Argon response cost?

Output costs $10 at intro rates and $20 at standard rates, plus input costs. See Gemini 4 Argon pricing for more scenarios.

Can I generate a 1M-token response today?

Not unless your organization is in the Fairwind cohort with Argon access. Gemini 3.8 Flash and 3.1 Pro Preview cap output at 65,536 tokens, while Vals AI lists 262K maximum output for the Argon configuration it tested.

Do I have to stream long Argon responses?

Google has not published Argon-specific streaming guidance. However, a non-streaming request that runs for minutes is exposed to every idle timeout in your stack. Stream the response or use background execution through the Interactions API.

How does Argon compare with GPT-6 Astra and Claude Opus 5.5?

Both cap synchronous output at 128K. Anthropic allows 300K on Batch with a beta header. Argon’s stated 1M output limit is about 8× higher.

Your next step

Add streaming and an explicit output cap to your Gemini client now, using Gemini 3.8 Flash. Run the client against the long mock stream until no component in your stack cuts the response off.

When Argon’s model ID becomes available, update GEMINI_MODEL and rerun the same tests in Apidog.

Top comments (0)