Gemini 3.8 Flash is available today: a stable model with a free tier, priced at $0.75 input and $3.75 output per 1M tokens through 2026-12-31, then $1.50 and $7.50, with a 64K output cap. Gemini 4 Argon is not generally available. Google’s launch post lists Argon at $2 input and $10 output during an intro period of unstated length, then $4 and $20. Google states a 1M-token output limit, but availability is currently limited to Fairwind Program defenders. If you need to ship this quarter, build on Flash and keep Argon behind a configuration switch.
This guide compares the models’ published specs, shared benchmark rows, cyber scores, and the cost of an identical call. It also shows how to make the eventual model switch a single environment-variable change. For background, see what Gemini 4 Argon is and what Gemini 3.8 Flash is.
Side by side: what you can call today
| Gemini 3.8 Flash | Gemini 4 Argon | |
|---|---|---|
| Availability | Stable on the Gemini API, AI Studio, Antigravity, and Gemini Enterprise | A subset of Fairwind Program partners, as a managed model on Gemini Enterprise |
| Model ID | gemini-3.8-flash |
Not published |
| Price per 1M tokens (input / output) | $0.75 / $3.75 through 2026-12-31, then $1.50 / $7.50 | $2 / $10 intro pricing, then $4 / $20 |
| Cached input per 1M | $0.075, then $0.15 | $0.10 intro, then $0.20 |
| Output limit | 65,536 tokens | 1M tokens, according to Google |
| Input window | 1,048,576 tokens | Not published |
| Thinking levels |
low, medium (default), high; minimal returns HTTP 400 |
Not documented |
| Free API tier | Yes, rate-limited | None announced |
| Consumer access | Gemini app on AI Pro and Ultra; AI Mode in Search | Not yet; AI Ultra subscribers are in the first wave, with no date announced |
A few details need qualification:
- Google has not published Argon’s input window, although its long-context evaluation used prompts up to 1M tokens.
- Google states a 1M-token Argon output limit. Third-party evaluator Vals AI lists a 262K maximum output for the configuration it tested.
- Argon evaluations used “the highest thinking settings,” which suggests a high setting exists, but Google has not documented the level names or default.
- Flash thinking levels are documented in Google’s thinking documentation and in this Gemini 3.8 Flash thinking-levels guide.
Flash’s free tier is rate-limited, and Google says free-tier data is “used to improve our products.” Google has not announced a free way to use Argon. See the Argon free-access explainer for currently available alternatives.
Shared benchmark rows
Google’s launch tables share two benchmark rows. These are Google-reported results, and Argon’s methodology attributes both rows to Vals AI.
| Benchmark | Gemini 3.8 Flash | Gemini 4 Argon |
|---|---|---|
| Vals Finance Agent v2 | 61.4% | 65.4% |
| Harvey Legal Agent Benchmark | 10.0% | 19.6% |
Treat these results as directional, not as a controlled head-to-head test: Google published the tables a month apart and compared them against different model sets.
Still, the published gap is notable:
- Finance: Argon leads by 4.0 percentage points.
- Legal: Argon scores 19.6% versus Flash’s 10.0%, nearly doubling the result.
Other rows do not have a direct Flash counterpart in the launch-post text. For example, Flash reported HLE-Verified while Argon did not, and Flash’s DeepSWE score appeared only in model-card images. On Arena’s Text leaderboard, a third-party ranking, Argon (High) is #1 on preliminary votes while Gemini 3.8 Flash (High) is #11.
For the full table and measurement sources, see this Argon benchmarks breakdown.
Cyber: Argon vs. 3.8 Flash Cyber
The DeepMind cyber page compares Argon with Gemini 3.8 Flash Cyber, a Fairwind-gated cyber-focused Flash variant.
| Cyber evaluation | Gemini 4 Argon | Gemini 3.8 Flash Cyber |
|---|---|---|
| Real-world vulnerability discovery | 85.8% | 71.0% |
| Wiz Penetration Test Benchmark | 70.9% | 58.2% |
Both models are gated:
- Gemini 3.8 Flash Cyber is generally available behind an allowlist on Gemini Enterprise Agent Platform.
- Gemini 4 Argon is Fairwind-only.
If you are not a Fairwind partner, these scores do not change the model you can use today. The generally available Gemini 3.8 Flash model is the practical implementation choice.
Cost of the same call on both
Use the same workload to compare pricing:
- Input: 100,000 tokens
- Output: 8,000 tokens
- Assumption: For current Gemini models, including Flash, thinking tokens are billed as output tokens. Google has not documented Argon’s thinking-token billing, so the Argon calculation assumes it follows current Gemini billing.
| Model and period | Input cost | Output cost | Total per call |
|---|---|---|---|
| Gemini 3.8 Flash, intro | 0.1M × $0.75 = $0.075 |
0.008M × $3.75 = $0.030 |
$0.105 |
| Gemini 3.8 Flash, from 2027-01-01 | 0.1M × $1.50 = $0.150 |
0.008M × $7.50 = $0.060 |
$0.210 |
| Gemini 4 Argon, intro | 0.1M × $2 = $0.200 |
0.008M × $10 = $0.080 |
$0.280 |
| Gemini 4 Argon, standard | 0.1M × $4 = $0.400 |
0.008M × $20 = $0.160 |
$0.560 |
Argon’s per-token rates are 8 / 3, or about 2.67× Flash pricing, during both intro and standard pricing periods.
At 1,000 calls per day using this workload:
Gemini 3.8 Flash intro pricing: 1,000 × $0.105 = $105/day
Gemini 4 Argon intro pricing: 1,000 × $0.280 = $280/day
Three implementation details can change the real ratio:
Token counts will differ.
The same prompt can produce different answer lengths and thinking-token counts. Argon’s published evaluations used its highest thinking settings, while Google warns that higher Flash effort levels can consume more tokens.Long outputs may require request splitting.
A 200,000-token answer exceeds Flash’s 65,536-token limit, so Flash would require multiple calls. Under Argon’s stated limit, it fits in one response.
Argon, 200K output tokens:
0.2M × $10 = $2.00 at intro pricing
0.2M × $20 = $4.00 at standard pricing
Argon, 1M output tokens:
1M × $10 = $10.00 at intro pricing
1M × $20 = $20.00 at standard pricing
- Only Flash has a published intro-pricing end date. Flash intro pricing ends on 2026-12-31. Google has not stated when Argon intro pricing ends.
Flash also supports Batch at a 50% discount: $0.375 input and $1.875 output per 1M tokens through December 31, according to the Gemini API pricing page. Google has not published Batch pricing for Argon.
For more detail, see the Argon pricing guide and Gemini 3.8 Flash pricing guide.
Decision rule: ship on Flash or plan for Argon?
Ship on Gemini 3.8 Flash now when
- You are cost-bound. Flash costs 3/8 of Argon per token, and Batch can halve Flash costs again.
- You run high-volume workloads. Classification, extraction, routing, and chat costs compound quickly at $10 per 1M output tokens.
- You need a model today. Flash has a stable model ID, free-tier prototyping access, and a published price schedule.
- Your responses fit within 65,536 tokens. Most API workloads do.
Plan for Gemini 4 Argon when
- Your workflows are long-horizon. Google describes Argon as built for deep reasoning across complex, long-running workflows.
- You need outputs beyond 64K tokens. Large migrations, full-module rewrites, and long reports are the clearest use cases for the stated 1M-token output limit.
- You build legal or finance agents. These are the two benchmark rows Google published for both models, and Argon leads on both.
- You are a Fairwind-eligible defender. Argon is the model currently led by that program.
The practical approach is not to choose permanently. Route normal production traffic to Flash, then test and route only the hard tail to Argon when it becomes available.
One saved request, two models
Keep the model ID in an environment variable. Run production-like requests against Flash now, then update the variable when Google publishes Argon’s model ID.
Do not guess an Argon model ID.
MODEL="${GEMINI_MODEL:-gemini-3.8-flash}"
curl -s -X POST \
"https://generativelanguage.googleapis.com/v1beta/models/${MODEL}:generateContent" \
-H "x-goog-api-key: $GEMINI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"contents": [
{
"parts": [
{
"text": "Summarize this contract clause in 3 bullets."
}
]
}
],
"generationConfig": {
"thinkingConfig": {
"thinkingLevel": "medium"
},
"maxOutputTokens": 8192
}
}'
Set the environment values:
export GEMINI_API_KEY="your-api-key"
export GEMINI_MODEL="gemini-3.8-flash"
When Argon receives a published model ID, the intended switch is:
export GEMINI_MODEL="published-argon-model-id"
Add cost and response checks
Every generateContent response includes usageMetadata. Use it to enforce a cost ceiling and compare model behavior.
In Apidog:
- Create an environment with
GEMINI_API_KEYandGEMINI_MODEL. - Save the request above.
- Add assertions for:
- HTTP status
200 - The response fields your application consumes
- A maximum
usageMetadata.thoughtsTokenCountbased on your token budget
- HTTP status
- Keep
maxOutputTokensset so an unexpectedly long response cannot exceed your output budget. - Add the request to a test scenario and run it against Flash.
- Save the report as a baseline for the Argon comparison.
For example, set an explicit output cap per request:
{
"generationConfig": {
"thinkingConfig": {
"thinkingLevel": "medium"
},
"maxOutputTokens": 8192
}
}
Prepare for the Interactions API
There are two launch-day caveats:
- Google says “all new models” will launch on the Interactions API, but has not confirmed whether Argon supports
generateContent. - Argon’s thinking-level names are not documented, so Flash’s
mediumsetting may not transfer directly.
Keep an Interactions API version of your request beside the generateContent version:
POST https://generativelanguage.googleapis.com/v1beta/interactions
Use generation_config.thinking_level for the Interactions API request.
When Argon is available:
- Update only
GEMINI_MODEL. - Run the same prompt suite.
- Compare output quality, latency, and
usageMetadata. - Calculate cost from actual input, output, and thinking-token usage.
- Route only workloads where Argon’s result justifies its higher cost.
FAQ
Is Gemini 4 Argon better than Gemini 3.8 Flash?
On the two benchmark rows shared by both launch tables, yes:
- Vals Finance Agent v2: 65.4% vs. 61.4%
- Harvey Legal Agent Benchmark: 19.6% vs. 10.0%
However, Argon costs about 2.67× more per token and is not generally callable yet.
Should I wait for Gemini 4 Argon?
Not if you need to ship. Google has not provided a release date, only “as soon as possible,” starting with paid API customers and AI Ultra subscribers.
Gemini 3.8 Flash or Argon for a new project?
Start with Gemini 3.8 Flash and place the model ID in an environment variable. Move requests requiring long outputs or stronger legal and finance performance to Argon only after testing it against your own prompts.
Is there a free way to use Argon?
No. Google has not announced a free tier. See the Argon free-access explainer for free Gemini models available today.
Does Argon replace Gemini 3.8 Flash?
Google has not said so. Google calls Argon its new frontier model, while Reuters reports it is larger than the previous Pro line. It appears to be a different tier from Flash.
Next step
Set up one saved request today, run it on Gemini 3.8 Flash, and store the output and usage report. When Argon ships, you can evaluate its quality and cost against your real traffic within minutes.
Download Apidog to build and save the comparison.
Top comments (0)