Gemini 4 Argon is Google's frontier model for long coding, knowledge work, and cyber defense, opened first to trusted defenders.
Who can run it today
Koray Kavukcuoglu, SVP of Google DeepMind and Google's chief AI architect, announced Argon on 30 September 2026. It is not on general sale. The first people outside Google are trusted cyber defenders in the Fairwind Program. Google says it is also in the U.S. government's voluntary process for pre-release model access. Paid Gemini API customers and Google AI Ultra subscribers are named next. No date is given for either.
Fairwind opened on 2 September 2026 around the smaller Gemini 3.8 Flash Cyber model. The program page now says Google works with over 650 partners worldwide. Only a set of those partners receive Argon. Google reviews applications and says it answers eligible ones as soon as it can. It does not publish a waiting time.
- Priority goes to governments and national cyber authorities, critical infrastructure (healthcare, telecommunications, energy, and financial networks), and core technology platforms.
- Academic labs that benchmark defenses can apply. Students are pointed at CodeMender on Google Cloud, using models that are already public.
- Applicants get a background check on security history and how they have operated.
- A partner may give Argon only to its own cybersecurity, incident-response, or penetration-testing staff, and must track who uses it. Phishing-resistant multi-factor authentication is required.
- Partners may not share, resell, or redistribute the model. Permitted dual-use work is authorized threat simulation, reverse engineering, and malware analysis for defense or academic research. Creating malware is not allowed.
- When Argon is used as a managed model on Gemini Enterprise, Google says it supports zero data retention.
- An organization that does not qualify can still run CodeMender with public models, plus Google's other AI Threat Defense products.
Jobs Google says it is already doing
Google says thousands of its own people are using Argon for specialized coding, deeper research, and writing. Three of the examples on the announcement are specific enough to repeat. These are Google's accounts of its own work.
- Quantum researchers used it on subroutines that were eating qubits and gates. In one case Google says it beat a published baseline by 40 percent, in minutes.
- A team of Argon agents read fleet-wide memory profiles and applied optimizations across Google's data centers. Google says the rollout freed more than 300 TiB, and estimates the total at 500 TiB to 1 PiB.
- Agents are moving C and C++ code to Rust. The span Google names runs from tens of thousands of lines in libraries such as re2 and libgav1, up to more than 800,000 lines in the Fuchsia Zircon kernel. Those rewrites are still in automated and manual audit, emulation testing, and review before they go to production.
- On libgav1, Google's open-source video decoder, the agents started from an existing Rust port and replaced 32,000 lines of hand-written SIMD. They ran rounds of profile-guided experiments, read the compiler's output, and wrote safe Rust so the compiler would vectorize it. Google says the pictures match and the decoder runs 2.7 times faster than that Rust port, closer to the optimized C++.
A million tokens of output
The output limit is 1 million tokens, up from 64,000. Google's reason is a single trajectory: the model can keep thinking and writing for hundreds of thousands of tokens instead of stopping and being called again. The announcement also lists creative writing among the jobs Argon is built for. The published comparison table does not include a writing score.
The table on the model page
The DeepMind model page compares Argon with GPT-6 Astra, Claude Fable 5.1, and Claude Opus 5.5. Every figure below is from that table. Google links a methodology page for how the runs were done. The table does not include every model on the market, so a lead here is a lead against these three, on Google's own runs.
Where Argon is ahead, or tied
- Vals Index, knowledge work weighted by each sector's share of U.S. GDP: 68.9 percent. Opus 5.5 is 67.0, Fable 5.1 is 65.8, Astra is 63.1.
- AutomationBench, Zapier's end-to-end business tasks: 51.3 percent. Opus 5.5 is 42.5, Astra is 41.4, Fable 5.1 is 31.4.
- Vals Finance Agent v2: 65.4 percent. Fable 5.1 is 58.9, Opus 5.5 is 58.6, Astra is 53.5.
- Harvey's Legal Agent Benchmark: 19.6 percent. Fable 5.1 is 6.7, Astra is 5.4, Opus 5.5 is 3.8.
- DeepSWE v1.1, long software-engineering tasks: 77.9 percent. The announcement calls this a new state of the art. On the table, Opus 5.5 is 74.2, Astra is 74.1, Fable 5.1 is 67.4.
- Vibe Code Bench: 91.9 percent. Fable 5.1 and Opus 5.5 are 90.3, Astra is 89.6.
- LABBench 2: 88.8 percent. Astra is 85.4, Opus 5.5 is 73.1, Fable 5.1 is 68.6.
- RiemannBench: 76.0 percent. Astra is 72.0, Opus 5.5 is 69.6, Fable 5.1 is 65.6.
- GraphWalks, breadth-first search, F1. At or under 128,000 tokens: 99.7 percent, against Astra 98.7, Fable 5.1 91.4, Opus 5.5 90.6. From 256,000 tokens to 1 million: 84.2 percent, against Astra 71.8, Opus 5.5 66.8, Fable 5.1 65.0.
- LVBench, long video: 91.7 percent. Google calls this state of the art. Astra is 87.5, Opus 5.5 is 83.7, Fable 5.1 is 79.7.
- Chartography: 71.6 percent. Astra is 71.0, Opus 5.5 is 66.3, Fable 5.1 is 46.2.
- Agent's Last Exam, computer-use pass rate: 39.5 percent. Opus 5.5 is 38.2, Astra is 34.2. Fable 5.1 has no score on that row.
- CWE-bench v1, repairing known classes of security weakness: 68 percent, tied with Astra. Opus 5.5 is 67, Fable 5.1 is 58.
Where the same table has Argon behind
- FrontierSWE v2: 55.0 percent, behind Astra at 65.5, Opus 5.5 at 62.3, and Fable 5.1 at 56.3.
- Terminal-bench 4.0: 57.4 percent, behind Opus 5.5 at 66.4, Astra at 58.2, and Fable 5.1 at 57.9.
- PostTrainBench, machine-learning engineering: 45.3 percent, behind Opus 5.5 at 49.3. Argon is ahead of Astra at 44.3 and Fable 5.1 at 40.2 on this row.
- Terminal-Bench Science 0.1: 57.6 percent, behind Astra at 68.1 and Opus 5.5 at 63.3, ahead of Fable 5.1 at 52.6.
- OSWorld-2.0, offline subset, partial score: 69.2 percent, behind Astra at 72.6. Fable 5.1 and Opus 5.5 are blank on that row.
Cyber defense stays inside Fairwind
Google says it trained Argon to find a vulnerability, check that it is real, and patch it. Fairwind partners and Google's own teams get a copy with the cyber guardrails removed. The copy everyone else would get, later, keeps those guardrails. Partners can also run Argon inside CodeMender, Google's agent for proposing software fixes, so a defender does not have to build that harness.
Wiz is an early user, through its Scan for Good program, which looks for serious exposures in public infrastructure without charging for that work. Google says one early run found a critical flaw that exposed personal information in healthcare software used by hospitals worldwide, and that earlier frontier models had missed it. The post does not name the software or the flaw.
- On Google's internal vulnerability set, Argon found exposures in codebases spanning 20 programming languages.
- On Wiz's internal test of live websites, with no source code, Google says Argon beat Gemini 3.8 Flash Cyber at mapping what is exposed, naming weaknesses, and producing evidence that they are real.
Four checks before a wider release
- Misuse. Argon is designed to refuse harmful cyber requests and harmful chemical, biological, radiological, and nuclear requests, while still allowing legitimate research under Google's Frontier Safety Framework. Google says it is also watching the model's internal activations, and that internal and external red teams tested those refusals.
- Prompt injection. Google calls Argon its most resistant model so far to indirect prompt injection, where hidden instructions in a page or a file try to hijack the run. It says Argon leads Gray Swan's Indirect Prompt Injection benchmark. The announcement does not print the score.
- Going past the ask. A monitor reads the chain of thought and the actions, and can stop the run. Google used a similar monitor during training. Alerts went to an incident team, and those findings were kept out of later training so the model would not learn to hide from the monitor.
- The box around the test. High-risk training and evaluations start only after the sandbox is isolated and sealed. Google says it plans to share that practice with partners.
The price, once it is for sale
The launch price on the announcement is $2 per million input tokens and $10 per million output tokens. Cached input is 95 percent off the input price, which is $0.10 per million. Footnote 1 says that after the introductory period the price becomes $4 per million input tokens and $20 per million output tokens. Google does not say when that period ends. Until paid API access actually opens, there is no public invoice to check the footnote against.