On 30 September, Google DeepMind announced Gemini 4 Argon, its new frontier model, built "to sustain deep reasoning across complex, long-horizon workflows" (Google's announcement). It is aimed at long, demanding jobs: real-world software engineering, legal and financial work, and cybersecurity. You can't try it yet. Google is giving it to a vetted group of cyber defenders first.
What Google announced
- A new generation and a new name. Argon is the first Gemini 4 model, and it breaks from the old "Pro" and "Flash" labels. It follows Gemini 3.8 Flash; a planned Gemini 3.5 Pro was dropped, 9to5Google reports.
- Built for long work. Google says it is strong at coding, knowledge work such as finance and law, cybersecurity and creative writing, and that it can keep going on multi-step tasks.
- Staged release. It goes first to cyber defenders in Google's Fairwind Program, then to paid API customers and Google AI Ultra subscribers, then to developers, businesses and everyone else. No date has been given for the public release.
- Price. An introductory $2 per million input tokens and $10 per million output tokens, rising later to $4 and $20. Repeated (cached) input costs 95% less.
What it has already done
Google says Argon already powers its internal work, with "thousands of Googlers" using it. Its examples:
- Quantum computing: it improved a subroutine that slows down important quantum programs, beating the published baseline by 40% "in a matter of minutes".
- Data centres: teams of Argon agents studied memory use across Google's computers and applied fixes that free more than 300 TiB of memory once rolled out, with 500 TiB to 1 PiB expected in total.
- Safer code: Argon agents are helping move Google code from C and C++ to Rust, a language that prevents many memory bugs. The work ranges from small libraries to the 800,000+ lines of the Fuchsia Zircon kernel, with heavy testing and human review before anything ships.
- Faster video: in libgav1, Google's open-source video decoder, Argon replaced 32,000 lines of hand-tuned speed-up code (SIMD) with safe Rust. The result runs 2.7 times faster than the earlier Rust version, with identical video output.
How it scores
| Test | What it measures | Argon |
|---|---|---|
| DeepSWE v1.1 | Long, real-world software engineering tasks | 77.9%, ahead of Claude Opus 5.5 (74.2%) and GPT-6 Astra (74.1%), per 9to5Google |
| AutomationBench (Zapier) | Doing business tasks from start to finish | 51.3%, first place |
| LVBench | Understanding long videos | 91.7%, the best so far |
| CWE-bench v1 | Fixing security bugs | 68%, tied first |
| Vals Index | Finance, coding, legal and tax enterprise tasks | First place, per Google |
| Artificial Analysis Intelligence Index | An independent mix of 10 tests | 53, against a median of 26 for comparable models |
Most of these numbers come from Google. The independent Artificial Analysis places Argon among the leading models. According to Engadget's report of its findings, Argon matches OpenAI's GPT-6 Astra on that index at about 60% of the cost per task at the introductory prices, and has the lowest hallucination rate of the leading models, at 15%. Artificial Analysis also notes that Argon is wordy: it used 110 million tokens to complete their tests, against a median of 81 million.
Why Argon is different
1. It can write a million tokens in one answer
AI models read and write in tokens: words or pieces of words. Two limits matter. The context window is how much a model can read at once. The output limit is how much it can write in a single reply. Argon's output limit jumps from Google's previous 64,000 tokens to 1 million. Engadget reports that GPT-6 Astra's limit is 128,000.
Why would anyone need such a long answer? Reasoning models "think" by writing out their steps before they reply, and long jobs such as rewriting a codebase produce huge amounts of text. In Google's words, when the model "has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning". Long answers are expensive to produce, though: everything the model writes stays in its working memory, the KV cache, until the answer is done.
2. Cyber defenders get it first, without the usual limits
Google trained Argon to "autonomously find, validate, and patch critical software vulnerabilities". The same skill that fixes security holes could help attackers find them, so Google is giving defenders a head start. Argon goes first to a set of partners in its Fairwind Program, which works with more than 650 vetted organisations: governments, operators of hospital, telecom, energy and banking networks, and big software platforms. Only their security teams may use it, behind strict logins, and access can't be shared or resold.
For these partners and Google's own teams, Argon comes "without cyber guardrails", so they can use its full security abilities. Wiz is running its "Scan for Good" initiative with Argon — a programme to find and fix vulnerabilities in critical public infrastructure. In one early test, Argon found a critical flaw in healthcare software used by hospitals worldwide that exposed sensitive personal information, a risk that earlier models had missed. Google also says Argon's safety systems cover CBRN misuse (chemical, biological, radiological and nuclear), and that it is designed to refuse harmful requests while continuing to support legitimate scientific work. Google is also taking part in the US government's voluntary programme for testing models before release.
Starting with a small, vetted group instead of the public is a cautious way to launch a flagship model.
3. Google watches how it thinks
When a reasoning model works through a problem, it writes its thinking down first: its chain of thought. Google says it watches that chain of thought, and Argon's actions, and stops the model if it starts going beyond what the user asked for. It also monitors the model's internal signals to catch misuse, and says Argon is its most resistant model yet to prompt injection: instructions hidden in a web page or document to hijack the AI.
One detail stands out. Google was careful not to feed what its monitors found back into training, so that Argon wouldn't learn to hide its reasoning. It urged other AI companies to "preserve reasoning transparency" too.
4. It is cheaper than its closest rival, for now
At the introductory $2 and $10 per million tokens, Argon costs a fifth of what Engadget reports for GPT-6 Astra ($10 and $50). After the introductory period, Argon's price doubles to $4 and $20.
What to watch
- When everyone gets it. Google hasn't given a date for developers or the public.
- Independent checks. Most results are Google's own. Early independent tests agree that it is among the best models, and more will follow.
- The guardrail-free version. Giving outside organisations a model without cyber limits puts a lot of weight on how well Google checks who gets in.
- The cost of thinking. A model that writes more tokens can cost more and take longer, even at a low price per token.
For how Gemini fits with Google's other AI products, read our guide to Google's AI stack.
