Gemini 4 Argon: a million tokens of output, and what that actually buys you
Google's new frontier model lifts the output ceiling from 64,000 tokens to a million. That is a bigger deal than another million-token context window, and a smaller one than it sounds. Here is what changes, what it costs, and when you should still be chunking.
On 30 September 2026 Google announced Gemini 4 Argon, and the number everyone led with was a million. Not the context window — a million tokens of output, up from the 64,000-token ceiling on the Gemini 3 line. Roughly a fifteenfold jump in how much a single response is allowed to contain.
Worth being precise about why that is the headline, because "1M tokens" has been a marketing line for two years and it has always meant input. Gemini models have read a million tokens since 1.5. What they could not do was write more than a long chapter before the response was cut off.
- output tokens in one responsewas 64K on Gemini 3
- 1M
- the old output ceiling≈ 750,000 words
- 15.6×
- per million input / output tokensintroductory rate
- $2/$10
- off for cached input tokens$0.10 per million
- 95%
Output ceilings were the quiet bottleneck
If you have built anything that generates long artefacts, you have hit this wall and probably worked around it without ever filing it as a model limitation. A 64K output cap means a codebase migration, a book-length document, a full test suite or a long agent transcript has to come out in pieces. Pieces mean orchestration: splitting the job, holding the plan outside the model, stitching the parts back together, and reconciling the seams where part four forgot a decision made in part one.
That stitching layer is where a surprising amount of my agent code lives, and almost none of it is interesting. It exists because of a ceiling, not because the work was ever really parallel.
What Google says it did with it
The launch leans on internal engineering results rather than chat demos, which is the more useful kind of evidence:
- A migration of the Fuchsia Zircon kernel — 800,000+ lines — handled as one long-horizon job.
- A video decoder rewritten to run 2.7× faster than an existing Rust port while keeping memory safety.
- Memory optimisations that freed over 300 TiB across Google infrastructure.
- A quantum algorithm optimisation that beat the published baseline by 40%.
Those are claims about a lab with unusual engineers and unusual access, not a promise about your Tuesday. But they describe the right shape of task: long, single-threaded, and dependent on holding one set of decisions consistent across an enormous amount of generated output. Precisely the thing a 64K cap made awkward.
The benchmarks, read honestly
Argon does not win everything, and the pattern in where it wins is more informative than the top-line scores.
- Argon
- Opus 5.5
- Astra
Vals Index
enterprise knowledge work
AutomationBench
multi-step office automation
DeepSWE v1.1
real-world software engineering
GraphWalks F1
long-context reasoning, 256K–1M tokens
FrontierSWE v2
harder agentic coding
Argon leads on knowledge work, long-context reasoning and everyday software engineering, and trails both rivals on the harder agentic coding set. Google also reports 91.7% on LVBench for long video and a first-place tie with Astra at 68% on CWE-bench v1. Terminal-bench 4.0 is the other loss: 57.4% against 66.4% for Opus 5.5.
GraphWalks is the row I would stare at. 84.2% against 71.8% and 66.8% is a wide margin, and it measures exactly the regime a million-token output is meant to serve: reasoning that stays coherent across hundreds of thousands of tokens. A large output ceiling sitting on top of weak long-range coherence would be a trap. These two numbers suggest they were built together.
The coding story is more mixed than the announcement's framing implies. Argon tops DeepSWE but loses FrontierSWE v2 by ten points and Terminal-bench by nine. If your workload is an agent driving a terminal, the leaderboard does not say Argon.
The price is the aggressive part
| Per 1M tokens | Introductory | After introduction | GPT-6 Astra |
|---|---|---|---|
| Input | $2 | $4 | $10 |
| Output | $10 | $20 | $50 |
| Cached input | $0.10 | 95% off input | — |
At the introductory rate Argon undercuts Astra by five times on both sides of the meter. Even at the standard rate it is half the input price and well under half the output price. Google is buying the enterprise workload, and the cache discount is aimed squarely at agents that re-read the same codebase or document set on every single turn.
The rollout is the genuinely unusual bit
Argon did not ship to everyone. It went first to vetted cybersecurity defenders through a new Fairwind Program, and it went to them with the cyber guardrails removed. Google's named partner is Wiz, which used it in a "Scan for Good" effort to find critical vulnerabilities in healthcare infrastructure.
Deliberately handing an unrestricted, offensively capable model to one side of the security balance is a strong position to take, and it is the same bet OpenAI made differently with Astra's gated rollout: capability at this level is dual-use, so decide who gets it first rather than pretending the question away. Everyone else gets Argon after further safety testing, beginning with paid API customers and Google AI Ultra subscribers.
The four safeguards Google describes:
- Refusing cyber and CBRN misuse while preserving legitimate dual-use research.
- Resistance to indirect prompt injection — their most resilient model yet, they claim, with leading Gray Swan IPI results.
- Chain-of-thought and action monitoring for misalignment.
- Sandboxed execution with isolation.
Two things are conspicuously absent from the announcement: the input context window and the API model identifier. Neither is a scandal this early in a staged rollout, but it does mean you cannot size a workload properly yet. Argon was evaluated through a million tokens of input. What the commercial limit will be is unstated.
Whether to use the ceiling
The reflex when an output limit goes up fifteenfold is to delete the chunking code. Resist that for a while. A single enormous response is one unit of work that either succeeds or fails as a whole, and when it fails at token 900,000 you have paid for 900,000 tokens and have nothing reviewable to show for it.
A job that needs more than 64K of output
Do the later parts depend on decisions made in the earlier ones?
↳ No — the sections are independent. Keep chunking: cheaper, parallel, retryable.
Can you verify the result mechanically?
Tests, a build, a schema, a diff review — something that is not a person reading prose.
↳ No — a human has to read all of it. A million tokens is about 1,500 pages. Nobody is reading that.
Is $10–20 per attempt acceptable?
↳ No — the budget has decided for you. Chunk it.
Run it long, with checkpoints you can resume from
One unit of work, but not one unit of risk.
The honest answer for most pipelines is still chunking. The cases where one long run wins are the ones where the seams were the problem all along: migrations, repository-wide refactors, and documents whose later sections must obey earlier ones.
Which is to say the workloads I would reach for this on look like the Zircon migration: a change applied consistently across a very large body of code, where the expensive failure mode is inconsistency between chunks rather than any individual chunk being wrong. That class of problem has been badly served until now.
The ceiling moved. The engineering question did not: what is this allowed to touch, how do you find out when it was wrong, and what happens next when it is.
Frequently asked questions
What is Gemini 4 Argon?
Google's frontier model, announced on 30 September 2026, aimed at long-horizon professional work: software engineering, legal and financial knowledge work, and cybersecurity. Its headline feature is a one-million-token output limit.
Is the million tokens input or output?
Output, and this is the distinction most of the coverage blurs. Gemini models have accepted a million tokens of input for generations. Argon is the first that can produce a million tokens in a single response, up from 64,000.
What does Gemini 4 Argon cost?
$2 per million input tokens and $10 per million output tokens at the introductory rate, doubling to $4 and $20 afterwards. Cached input tokens are 95% cheaper, which is $0.10 per million at the introductory rate.
Can I use it today?
Only if you are a vetted cybersecurity defender in the Fairwind Program. Google says wider access follows after further safety testing, beginning with paid API customers and Google AI Ultra subscribers. The API model identifier has not been published.
Is it better than GPT-6 Astra or Opus 5.5?
It depends on the workload, which is the only honest answer. Argon leads on enterprise knowledge work, long-context reasoning and the DeepSWE software engineering set. Astra leads FrontierSWE v2 and Opus 5.5 leads Terminal-bench 4.0, so for an agent working against a terminal it is not the top pick.
What is the Fairwind Program?
Google's initiative giving trusted cyber defenders access to Argon without its cyber guardrails, on the argument that defenders need the same capability attackers will eventually have. Wiz's "Scan for Good" vulnerability hunt across healthcare infrastructure is the named example.
Does a bigger output limit mean I can stop chunking?
Not yet, and for a lot of pipelines not ever. One long response is a single unit that fails as a whole, costs $10–20 per attempt at full length, and produces more text than a person can review. Chunking stays right for work with independent sections. One long run wins when later output has to stay consistent with earlier decisions.
Building something like this?
I design and ship these systems for clients: retrieval over private data, agents that complete real tasks, and the Laravel platforms underneath them.