Gemini 4 Argon costs $2 per 1M input tokens and $10 per 1M output tokens during an introductory period, then $4 and $20. Cached input is 95% off the input price: $0.10 per 1M tokens at the intro rate and $0.20 afterward. Google has not said how long the intro period lasts, and you cannot buy Argon yet: it is announced but not generally available, and is currently rolling out only to Fairwind Program cyber defenders.
This post turns the rate card into implementation-ready budget numbers: comparisons with current models, four pricing scenarios, third-party per-task costs, and cost controls you can configure now. For the model overview, start with what Gemini 4 Argon is; for availability, see is Gemini 4 Argon free. You can build cost assertions in Apidog against a model you can call today, then swap the model name when Argon becomes available.
The price sheet
Google published these prices in the Gemini 4 Argon launch post. Footnote 1 specifies the post-intro rate: “After the introductory period expires, the price of $4 per 1M input tokens and $20 per 1M output tokens will apply.”
| Gemini 4 Argon, per 1M tokens | Intro | Standard (after intro) |
|---|---|---|
| Input | $2.00 | $4.00 |
| Cached input (95% off input) | $0.10 | $0.20 |
| Output | $10.00 | $20.00 |
| Max output per response (Google) | 1M tokens | 1M tokens |
| Cost of one full 1M-token output | $10.00 | $20.00 |
The cached-input rates are derived from Google’s 95% discount rule:
intro cached input = $2 × 0.05 = $0.10 per 1M tokens
standard cached input = $4 × 0.05 = $0.20 per 1M tokens
Google has not published an input context window or a model ID. It says the output limit is 1M tokens, “up from the previous 64K tokens.” At least one third-party evaluator, Vals AI, lists a 262K max output for the configuration it tested.
Argon vs. current frontier and Gemini models
All prices are per 1M tokens.
| Model | Input | Cached input | Output | Max output |
|---|---|---|---|---|
| Gemini 4 Argon (intro) | $2 | $0.10 | $10 | 1M |
| Gemini 4 Argon (standard) | $4 | $0.20 | $20 | 1M |
| GPT-6 Astra | $10 | $1 | $50 | 128,000 |
| Claude Opus 5.5 | $4 | $0.20 | $20 | 128K (300K on Batch, beta) |
| Claude Sonnet 5.5 | $2 | $0.20 | $10 | 128K (300K on Batch, beta) |
| GPT-6.1 Sol | $2 | $0.10 | $10 | 128,000 |
| Claude Fable 5.1 | $10 | $0.25 | $50 | 128K |
| gemini-3.1-pro-preview (prompt <=200K / >200K) | $2 / $4 | $0.20 / $0.40 | $12 / $18 | 65,536 |
| gemini-3.8-flash (intro / from 2027-01-01) | $0.75 / $1.50 | $0.075 / $0.15 | $3.75 / $7.50 | 65,536 |
Competitor prices come from vendor documentation: OpenAI’s GPT-6 Astra page, Anthropic’s pricing page, and Google’s Gemini API pricing page.
Key comparisons:
- Standard Argon matches Claude Opus 5.5. Both are $4 input, $20 output, and $0.20 cached input.
- Intro Argon matches Claude Sonnet 5.5 and GPT-6.1 Sol at $2 input and $10 output. Its $0.10 cached-input rate also matches GPT-6.1 Sol.
- GPT-6 Astra and Claude Fable 5.1 cost 5x intro Argon and 2.5x standard Argon on a per-token basis. See Claude Fable 5.1 pricing.
- Argon’s intro output price ($10) is below gemini-3.1-pro-preview’s ($12), Google’s previous top Pro model.
- Long prompts can change the calculation. OpenAI bills prompts over 272K input tokens at 2x input and cache prices and 1.5x output price for the full request. Gemini 3.1 Pro costs more above 200K input tokens; Anthropic does not use that tier. Google has not said whether Argon will have a long-prompt tier.
For capability comparisons alongside price, see Gemini 4 Argon vs GPT-6 Astra vs Claude Opus 5.5.
Treat thinking tokens as output cost
On current Gemini models, thinking tokens are billed as output. Google’s thinking documentation states that “response pricing is the sum of output tokens and thinking tokens.”
Google has not documented Argon-specific billing behavior, but budget as though the same rule applies. Google ran Argon evaluations with its highest thinking settings, so the $10 per 1M output-token intro price should be treated as the cost of both visible answer tokens and reasoning tokens.
Use this cost model:
cost =
(uncached_input_tokens / 1,000,000 × input_rate)
+ (cached_input_tokens / 1,000,000 × cached_input_rate)
+ ((output_tokens + thinking_tokens) / 1,000,000 × output_rate)
Four worked scenarios
Every example below uses:
tokens / 1,000,000 × per-1M rate
Scenarios 1–3 assume no caching.
1. Typical API call: 20K input, 5K output
Intro pricing
input = 20,000 / 1M × $2 = $0.04
output = 5,000 / 1M × $10 = $0.05
total = $0.09
Standard pricing
input = 20,000 / 1M × $4 = $0.08
output = 5,000 / 1M × $20 = $0.10
total = $0.18
At 1,000 calls per day:
| Pricing period | Daily cost |
|---|---|
| Intro | $90 |
| Standard | $180 |
For the same 20K-input, 5K-output call:
- Claude Opus 5.5: $0.18
- GPT-6 Astra: $0.45 ($0.20 input + $0.25 output)
- gemini-3.1-pro-preview: $0.10 ($0.04 input + $0.06 output)
- gemini-3.8-flash at its intro rate: about $0.034 ($0.015 input + $0.019 output)
2. Long agent turn: 200K input, 50K output
Intro pricing
input = 200,000 / 1M × $2 = $0.40
output = 50,000 / 1M × $10 = $0.50
total = $0.90
Standard pricing
input = 200,000 / 1M × $4 = $0.80
output = 50,000 / 1M × $20 = $1.00
total = $1.80
For comparison:
- GPT-6 Astra: $4.50 ($2.00 input + $2.50 output)
- gemini-3.1-pro-preview: $1.00 ($0.40 input + $0.60 output)
The 200K input size is Gemini 3.1 Pro’s pricing threshold. Prompts beyond that threshold use $4 input and $18 output pricing. If Argon introduces a similar tier, longer agent turns will cost more than the flat-rate examples above.
3. Maximum-size response: 1M output tokens
Output alone costs:
1,000,000 / 1M × $10 = $10.00 intro
1,000,000 / 1M × $20 = $20.00 standard
Add a 100K-token prompt:
| Pricing period | Output | Input | Total |
|---|---|---|---|
| Intro | $10.00 | $0.20 | $10.20 |
| Standard | $20.00 | $0.40 | $20.40 |
No other model in the comparison table produces 1M tokens in one synchronous response. Competitors cap output at 128K, and the previous Gemini Pro model caps output at 65,536 tokens.
At Vals AI’s listed 262K max output, a full response would cost approximately:
262,000 / 1M × $10 = $2.62 intro
262,000 / 1M × $20 = $5.24 standard
For engineering considerations such as streaming, timeouts, and gateway limits, see Gemini 4 Argon’s 1M output tokens.
4. Cached 500K-token prompt reused 10 times
Assume you send the same 500K-token codebase or contract set 10 times and receive 5K output tokens per call. The first call pays the full input price; the next nine calls use cached input.
| 10 calls: 500K prompt, 5K output each | Intro | Standard |
|---|---|---|
| Input without caching (10 × 500K) | $10.00 | $20.00 |
| Input with caching (1 full + 9 cached) | $1.00 + 9 × $0.05 = $1.45 | $2.00 + 9 × $0.10 = $2.90 |
| Output (10 × 5K) | $0.50 | $1.00 |
| Total with caching | $1.95 | $3.90 |
| Total without caching | $10.50 | $21.00 |
Caching cuts input cost by 85.5% in this example.
However, account for two unknowns:
- Google has not said whether Argon charges cache-storage fees. Gemini 3.1 Pro charges $4.50 per 1M tokens per hour, which would add $2.25 to hold 500K tokens for one hour.
- A 500K-token prompt is above Gemini 3.1 Pro’s 200K threshold. Argon may introduce a similar tier.
Argon’s standard cached-input rate matches Claude Opus 5.5’s $0.20 rate, so the break-even approach in Claude Opus 5.5 prompt caching cost math also applies here.
Per-token price is not per-task cost
Do not budget from rate cards alone. Measure how many tokens your workload actually uses.
Artificial Analysis lists $1.99 per task to run its Intelligence Index on Gemini 4 Argon (High) at intro prices. The Decoder puts the same run at $3.98 at standard prices. Artificial Analysis scores Argon at 53, the same as GPT-6 Astra (max), which it lists at $3.26 per task.
The important variable is output volume:
- Artificial Analysis counts roughly 62K output tokens per Argon task at High.
- It counts roughly 27K output tokens per GPT-6 Astra task at its maximum setting.
- That is about 2.3x more output tokens for Argon.
At intro pricing:
Argon output cost = 62,000 / 1M × $10 = $0.62
Astra output cost = 27,000 / 1M × $50 = $1.35
At standard Argon prices, The Decoder’s full-task estimate for Argon ($3.98) exceeds Artificial Analysis’s listed $3.26 per task for GPT-6 Astra (max), despite Argon’s standard per-token rates being 60% lower than Astra’s.
Vals AI shows the same pattern. It lists $15.68 per test for Argon on the Vals Index, using the $4/$20 standard rate, compared with:
- Claude Opus 5.5: $32.14 per test
- Claude Sonnet 5.5: $21.34 per test
- GPT-6.1 Sol: $3.24 per test
Sonnet 5.5 and GPT-6.1 Sol share Argon’s intro price on paper, but their per-test costs differ by more than 6x.
Use your own production token counts to budget per-task cost.
Cost controls to implement before Argon ships
1. Cap output tokens
On generateContent, set generationConfig.maxOutputTokens to limit output cost.
{
"generationConfig": {
"maxOutputTokens": 8192
}
}
Google says new models launch on the Interactions API, so configure the equivalent cap there once Argon-specific documentation is available.
A 1M-token output limit is also a maximum output charge of:
$10 per request at intro pricing
$20 per request at standard pricing
Set a lower limit for normal production paths. Reserve large outputs for endpoints that explicitly require them.
2. Cache stable prompt content
Cache content that repeats across requests:
- System instructions
- Tool schemas
- API specifications
- Reference documents
- Codebase context
- Contract sets
Avoid caching user-specific or frequently changing content. The 95% cached-input discount is most useful when a large, stable prompt is reused across multiple calls.
3. Select thinking levels intentionally
Google has not published Argon’s thinking levels. Current defaults include:
-
gemini-3.1-pro-preview:high -
gemini-3.8-flash:medium
When Argon’s settings are documented, benchmark lower thinking levels against your task-quality requirements. Since thinking tokens are expected to be billed as output, reducing unnecessary reasoning can reduce cost.
4. Add usage and cost assertions
Every generateContent response ends with usageMetadata. Save the model name in a GEMINI_MODEL environment variable, then calculate cost from the returned token counts.
In Apidog, you can:
- Save a Gemini request in a collection.
- Use
{{GEMINI_MODEL}}in the request URL or body. - Run the request against
gemini-3.8-flashtoday. - Add a post-response script that reads
usageMetadata. - Fail the request when calculated cost exceeds your per-call budget.
- Replace only
GEMINI_MODELwhen Argon’s model ID ships.
Use this calculation:
cost = uncached_input / 1,000,000 × input_rate
+ cached_input / 1,000,000 × cached_rate
+ (output + thinking) / 1,000,000 × output_rate
What Google has not priced yet
Before committing a production budget, track these unresolved items:
- The intro period length and the date that $4/$20 pricing starts
- Whether prompts above 200K tokens have a higher price tier
- Batch or Flex pricing
- Cache-storage fees
- Rate limits
- Whether a free tier exists
For reference, Gemini 3.1 Pro has Batch and Flex pricing at $1/$2 input and $6/$9 output.
FAQ
How much does Gemini 4 Argon cost per million tokens?
$2 input and $10 output during the intro period, then $4 input and $20 output. Cached input is 95% off: $0.10 during intro pricing and $0.20 afterward.
How long does the intro price last?
Google has not said. The launch post gives no duration or end date.
Are thinking tokens billed as output?
For current Gemini models, yes. Google has not documented Argon’s rules specifically.
How much does a 1M-token output response cost?
$10 at intro pricing and $20 at standard pricing, before input-token charges.
Is Argon cheaper than Claude Opus 5.5 or GPT-6 Astra?
Per token, standard Argon matches Claude Opus 5.5 and is 60% cheaper than GPT-6 Astra. Per task, the result depends on token volume: Artificial Analysis measured Argon using about 2.3x Astra’s output tokens. For a cheaper Gemini model available today, see Gemini 3.8 Flash pricing.
Can I pay for Argon today?
No. It is rolling out only to Fairwind Program partners. Paid API customers and Google AI Ultra subscribers are next, but Google has not provided a date.
Budget it now, measure it later
Use your current traffic data to estimate an Argon range:
intro estimate = input_tokens × $2/1M + output_tokens × $10/1M
standard estimate = input_tokens × $4/1M + output_tokens × $20/1M
Then implement measurement now:
- Download Apidog.
- Save a request using
gemini-3.8-flash. - Store the model name in
GEMINI_MODEL. - Add a cost assertion based on
usageMetadata. - Set an output cap and per-request cost ceiling.
- When Google publishes Argon’s model ID, update one variable and rerun the suite.

Top comments (0)