NEWS
Google’s Gemini 4 Argon Goes to Cyber Defenders First
Gemini 4 Argon writes 1 million tokens and hunts software bugs on its own, so Google is giving it to cyber defenders first, with no public date.
Google DeepMind opened Gemini 4 Argon on September 30 to vetted cyber defenders, with a 1 million token output limit. Paid API customers and Google AI Ultra subscribers are next. Free Gemini plans are not in the rollout.
Koray Kavukcuoglu, Google’s chief AI architect, said a model at this level needs a staged release and U.S. government pre-release review. Inside Google the same system is already moving C and C++ into Rust and hunting bugs that earlier frontier models missed.
Introducing Gemini 4 Argon – our new frontier model.
It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program. pic.twitter.com/X8acOWJOSF
— Google DeepMind (@GoogleDeepMind) September 30, 2026
A Million Tokens Changes the Size of One Job
Google is raising Gemini’s 1 million token output limit from 64,000 on earlier models. Output is how much the model can write and think through in one response, so a 64,000 cap forced long audits and migrations into fragments. At the new ceiling, a job can stay in one run.
Kavukcuoglu wrote that when the model has room to generate hundreds of thousands of tokens in a single pass, it can finish hard problems without a handoff every few pages. That is the shape of an agent loop: plan, try, fail, and close the task. It is also why a wrong turn in the middle of a long run gets expensive to catch, which is why Google says it will watch chain-of-thought and actions and stop Argon when it steps past what the user asked.
ARGON AT A GLANCE
- Output ceiling: 1 million tokens, up from 64,000 on prior Gemini models.
- DeepSWE v1.1: 77.9 percent, Google’s published score on long-horizon software engineering.
- AutomationBench: 51.3 percent, first place on Zapier’s test of end-to-end business tasks.
- LVBench: 91.7 percent, Google’s published lead on long video understanding.
OpenAI published 68.8 percent for GPT-6 Sol at max effort on DeepSWE v1.1, and each lab reports its own run. Google also says Argon leads the Vals Index, a GDP-weighted mix of finance, coding, legal, and tax work, and Harvey’s legal-agent test. Input context for Argon was not disclosed in the launch post.
Inside Google, Argon Is Already Rewriting Production Code
Kavukcuoglu said thousands of Googlers are already using Argon on specialist coding, research, and writing. The company named live engineering jobs, not demo prompts, and it still sends the large rewrites through automated checks, emulation, and human review before production.
WHAT ARGON ALREADY DID INSIDE GOOGLE
| Job | What the agents did | Published result |
|---|---|---|
| Quantum subroutines | Cut spacetime cost (qubits × gates) on a published baseline | 40 percent better, in minutes |
| Data-center memory | Read fleet telemetry and applied memory fixes | Over 300 TiB freed on rollout; 500 TiB to 1 PiB estimated in total |
| Fuchsia Zircon kernel | Migrate C and C++ to Rust, after smaller ports of re2 and libgav1 | 800,000+ lines in the kernel |
| libgav1 video decoder | Replace 32,000 lines of SIMD with safe Rust the compiler can vectorize | 2.7 times faster than the prior Rust port, identical video output |
The libgav1 pass is the clearest specimen. Argon took an existing Rust port, studied compiler output across profile-guided rounds, and produced memory-safe code that still matches the optimized C++ on video. Google is doing the same class of work on Zircon, which is why those patches wait on audit rather than shipping on the agent’s say-so.
The Hospital Flaw Previous Frontier Models Missed
Google trained Argon to find, check, and fix critical software bugs on its own. Wiz is already running it through the Scan for Good initiative, a free hunt for high-risk holes in public services and critical systems. Google said the model found a critical flaw that exposed sensitive personal information in healthcare software used by hospitals worldwide, a miss for earlier frontier models. The company did not name the product or say whether the hole is closed.
That finding is separate from the cases Wiz published on September 24, when Scan for Good still ran on Gemini 3.8 Flash Cyber. Those earlier hits included a public hospital whose missing access controls left a campus-wide mobile alert channel open to anyone online, and a private hospital whose appointment site had an unsafe upload that could have exposed a server, patient identifiers, clinical information, and consent signatures. Wiz also described an administrator key on a public server that opened 8.8 million files in a national archive, a municipal data service that exposed health and financial records for about 5,000 older residents, and a public rail operator whose leaked production database exposed live administrator sessions.
On the CWE-bench v1 remediation test, Argon tied for first at 68 percent. Google also said it beat 3.8 Flash Cyber on an internal discovery set covering 20 languages and on Wiz’s black-box pentest, which scores live web systems without source code. Human researchers still validate impact and handle disclosure.
For trusted defenders and our own internal teams at Google, we’ll be releasing Argon without cyber guardrails so they can leverage its full frontier-level cybersecurity defense capabilities.
Koray Kavukcuoglu, SVP, Google DeepMind, Gemini 4 Argon launch post
No Cyber Guardrails Inside the Fairwind Program
Fairwind is Google DeepMind’s defender-first access path. The program page is blunt about why it exists: frontier models can find and fix serious bugs fast, and the same skills are a threat in the wrong hands. High-priority groups such as governments, healthcare providers, and telecoms get the model before new threats show up. A set of those partners can run Argon on its own or inside CodeMender, DeepMind’s code-security agent for research and patching.
A million-token unguarded worker is a full agent loop, not a longer chat reply, so Google is not opening a front door and iterating later. The people who already write long agent workflows will be ready when the API appears. Everyone waiting on a Gemini chat box will keep waiting.
FAIRWIND ACCESS RULES
- Partner count: DeepMind lists over 650 partners globally.
- Who is prioritized: Governments and national cyber authorities, operators of healthcare, telecom, energy, and financial networks, and core technology platforms.
- Who may click it: Only internal cybersecurity, incident response, or pentest teams, with user-level login, phishing-resistant MFA, and an access log.
- Allowed dual-use work: Authorized threat simulation, reverse engineering, and malware analysis for defense and academic research; creating malware is banned.
- Sharing: Partners may not share, redistribute, or sell access, and DeepMind runs background checks. Direct Gemini Enterprise access can use zero data retention.
Groups that do not get in can still run CodeMender on public Gemini models and the rest of Google’s threat-defense stack. Academic labs that do defensive benchmarks may apply; DeepMind points students at CodeMender on Google Cloud instead.
Launch Pricing Tracks GPT-6 Sol, Then Doubles
Argon’s launch rate matches GPT-6 Sol. After an intro period with no end date, the rate doubles. Cached input is 95 percent off the input rate, which is $0.10 per million tokens at the launch input price. A full 1 million token reply at the launch output rate costs $10, because that is the window times the $10-per-million output fee.
PRICE AND OUTPUT CAPS
| Model | Input per 1 million tokens | Output per 1 million tokens | Max output |
|---|---|---|---|
| Gemini 4 Argon (intro) | $2 | $10 | 1 million |
| Gemini 4 Argon (after intro) | $4 | $20 | 1 million |
| GPT-6 Sol | $2 | $10 | 128,000 |
OpenAI’s GPT-6 Sol API pricing landed on September 22 with Luna at $0.10 input and $0.50 output per million tokens. GPT-6 Astra, the top GPT-6 model, lists at $10 and $50. Sol’s context window is 1,050,000 tokens with a 128,000 output cap, so Argon’s published edge is how much it can write in one go, not a disclosed input window. The cheap launch rate is the invite; the later $4 and $20 line is the list price Google expects once the gate opens.
Who Can Use Gemini 4 Argon Right Now?
As of the September 30 launch, Gemini 4 Argon is in the hands of Google’s own teams and approved Fairwind cyber defenders. The next groups named are paid Gemini API customers and Google AI Ultra subscribers. Google has not given dates for those steps, and free Gemini plans are not part of the announced rollout.
THE ROLLOUT GOOGLE HAS ANNOUNCED
- September 22, 2026: OpenAI releases GPT-6 Sol and Luna at $2 and $10 per million tokens, with a 128,000 output cap.
- September 24, 2026: Wiz publishes Scan for Good results running on Gemini 3.8 Flash Cyber, including the hospital, archive, and rail cases.
- September 30, 2026: Google announces Gemini 4 Argon, opens it to Fairwind defenders without cyber guardrails, and Wiz adds Argon to Scan for Good.
- No date set: Paid Gemini API access, Google AI Ultra, consumer Gemini, and the end of the $2 and $10 intro rate.
Google says it is in the U.S. government’s voluntary process for pre-release model access while it widens the circle. Early testers are there to stress the guardrails, not to stand in for a public beta. Fairwind applications stay on DeepMind’s form, with no published wait time.
The Safeguards Google Wants Before a Wider Release
Before a broad launch, Google listed four controls it is still hardening. They read as the bill for a model that can run for 1 million tokens and poke live systems.
FOUR CHECKS ON THE UNGUARDED WORKER
- Misuse refusals: The model is built to refuse cyber and CBRN attack requests while still allowing dual-use scientific work under DeepMind’s Frontier Safety Framework, with monitors on internal activations and red-team tests.
- Prompt injection: Google calls Argon its toughest model yet on indirect prompt injection and says it leads Gray Swan’s IPI benchmark after automated red-teaming and adversarial training.
- Misalignment stops: Separate monitors watch chain-of-thought and actions and can halt a run that goes beyond the user’s intent, with training-run alerts that Google says it does not feed back into training, so the model cannot learn to hide from the watchers.
- Sealed sandboxes: High-risk training and evals run in isolated, sealed environments, in line with DeepMind’s agent-control roadmap, and Google says it will share those practices with partners.
Google still has not named a day when paid API customers or Google AI Ultra subscribers will get Argon, or when the $2 and $10 intro rate expires. Fairwind applications remain open on DeepMind’s program page, and Wiz is still taking Scan for Good assessment requests from groups that run exposed public systems.
-
AUTO3 years agoBMW’s Heated Seat Retreat Taught Automakers Which Fees Survive
-
NEWS4 weeks agoJohn Ternus Debuts a $2,099 Foldable and Holds iPhone 18
-
NEWS4 weeks agoTesla Burns AI Cash While SpaceX Sends the Invoices
-
ENTERTAINMENT1 month agoDolly Parton Laid to Rest as the Public Funeral Began
-
NEWS4 weeks agoThe Book on China’s Car Brain Still Stars Infineon
-
NEWS3 years agoCopilot’s Election Problem Shifted From Errors to Silence
-
TECHNOLOGY3 years agoTwitter Search Not Showing All Results: How to Fix it?
-
NEWS4 weeks agoSAP Puts Token Spend on the Same Line as Hiring
