Inside Gemini 4 Argon: A 2.7x Faster Video Decoder and a Million-Token Output

Google has announced Gemini 4 Argon, a frontier model that is rolling out first to a set of trusted cyber defenders through what the company calls its Fairwind Program. In a post dated September 30, 2026, Koray Kavukcuoglu, SVP of Google DeepMind and Google's Chief AI Architect, described a model built to sustain deep reasoning across long, multi-step work in software engineering, enterprise knowledge work such as legal and finance, and cybersecurity defense. Developers, enterprises and consumers will get access later, starting with paid API customers and Google AI Ultra subscribers.
Access and pricing
Google is taking a phased approach. The company says it is engaged in the U.S. government's voluntary process for pre-release model access while it gradually expands availability, and that it will keep gathering feedback from early testers while it adjusts guardrails. No general release date was given beyond "as soon as possible."
Pricing is already set for launch. Argon will carry an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95 percent off the input rate. For a workload that reuses a large shared context, such as a long codebase or a document set queried repeatedly, the cached rate would be about ten cents per million tokens.
What Google says it has done with Argon internally
Google says thousands of its employees already use Argon for specialized coding, deeper research and writing. The post lists three examples.
The first comes from quantum computing research. Argon was asked to optimize the spacetime resources, measured as qubits multiplied by gates, of subroutines that bottleneck important applications. In one case it beat the published baseline by 40 percent in a matter of minutes.
The second is about data center memory. A group of Argon agents analyzed fleet-wide profiling telemetry and applied memory optimizations on their own. Google says this will free more than 300 TiB of memory once rolled out, with estimated total savings of 500 TiB to 1 PiB.
The third is code migration. Argon agents are moving C and C++ codebases to Rust across Google, from tens of thousands of lines in core libraries such as re2 and libgav1 up to more than 800,000 lines for the Zircon kernel of Fuchsia OS. Google says these rewrites are going through automated and manual auditing, emulation testing and review before reaching production, given how many of the systems are critical.
The libgav1 project is the most detailed. libgav1 is Google's open source video decoder. Argon agents started from an existing Rust port and replaced 32,000 lines of SIMD code. They ran many rounds of profile-guided experiments and studied the compiler's output, then wrote safe Rust that the compiler could vectorize automatically. The result, Google reports, is a memory-safe decoder that runs 2.7 times faster than the Rust port and produces identical video output, which brings it closer to the speed of the optimized C++ version.
A larger output window
Argon raises the output token limit to 1 million, up from 64,000. Google says that when a model has room to think and generate hundreds of thousands of tokens in a single trajectory, it can work through hard problems in one pass. For agent workflows that currently chain many shorter calls, a larger single output changes how tasks can be structured, though the post does not include data on how the extra headroom affects accuracy or cost.
Benchmark results
Google reports the following results, all self-published:
DeepSWE v1.1: 77.9 percent, described as a new state of the art on real-world, long-horizon software engineering tasks.
Vals Index: first place. The index measures economic impact across finance, coding, legal and tax work, with each sector weighted by its contribution to U.S. GDP.
Vals Finance Agent v2 and Harvey's Legal Agent Benchmark: described as leading performance on multi-step financial research and on legal research and drafting. The text gives no scores for these.
AutomationBench: first place at 51.3 percent. Zapier's benchmark measures end-to-end execution across core business functions.
LVBench: 91.7 percent, state of the art for long video understanding.
The post also says Argon is strong where knowledge work depends on visual understanding, including professional chart analysis, finding details in long videos and acting on a series of documents. The announcement includes charts for several of these benchmarks. The comparison models and margins are in those images, and the text only states rankings and a few scores, so readers should check the charts directly before citing relative performance.
Cybersecurity defense
Google says it trained Argon specifically to be capable at defending against cyberattacks, including autonomously finding, validating and patching critical vulnerabilities. For trusted defenders and Google's internal teams, the company will release Argon without cyber guardrails so they can use its full defensive capability. This is the reason the Fairwind Program is first in line.
Wiz is already using the model through its Scan for Good initiative, a program that finds and remediates high-risk exposures in critical public infrastructure at no charge. According to the post, Argon uncovered a critical vulnerability exposing sensitive personal information across healthcare software used by hospitals worldwide, a risk that previous frontier models had missed. Google does not name the software or say whether the flaw has been fixed.
Other cyber results cited:
On CWE-bench v1, which tests the remediation of security vulnerabilities, Argon ties for first place with 68 percent. Google notes this builds on the performance of its earlier 3.8 Flash Cyber on CWE-bench v0.
On Google's internal vulnerability benchmark, Argon found a wide range of exposures across complex codebases in 20 programming languages.
On Wiz's internal black-box penetration testing benchmark, which asks a model to analyze live web systems without source code, Argon outperformed 3.8 Flash Cyber at discovering attack surface, identifying vulnerabilities and producing proof-of-concept evidence.
The internal benchmarks cannot be checked from outside, and the post does not give scores for them.
Safeguards before broad release
Google lists four areas of work it is continuing before wider rollout.
Misuse. The model is designed to refuse harmful requests for cyber and chemical, biological, radiological and nuclear attacks while still supporting legitimate dual-use scientific research, in line with Google's Frontier Safety Framework. The company says it is strengthening these safeguards, including techniques that monitor the model's internal activations to spot misuse. Internal and external red teams tested them with manual and automated attacks.
Prompt injection. Google calls Argon its most resilient model yet against indirect prompt injection, where malicious instructions hidden in content hijack a model's behavior. It reports leading results on Gray Swan's Indirect Prompt Injection benchmark, reached through automated red teaming and adversarial training.
Misalignment. Google is deploying mitigations that monitor Argon's chain-of-thought and actions and halt execution when the model appears to step beyond the user's intentions. A similar system watched the training runs and sent alerts to a dedicated incident response team. Google says it took precautions against feeding those findings back into training, to avoid shaping the model's reasoning to evade monitoring. The company also urged the rest of the industry to preserve reasoning transparency so that model thoughts remain useful for diagnosing misalignment.
Hardening systems. Following its agent control roadmap, Google is isolating and sealing its sandboxed environments before high-risk training or evaluations begin, and says it will share these security practices with partners.
What to watch
Several questions remain open. The release date for developers and consumers is undefined. The benchmark charts need to be read against the comparison models, and independent evaluations of DeepSWE, AutomationBench and CWE-bench results will show how the numbers hold up. The Rust migrations are still under audit, so the production results are not yet in. The memory savings are stated as what Google expects once rollout is complete.
For teams planning around the model, the published price and the 1 million token output limit give enough to begin modeling costs. Security teams that want early access can look at the Fairwind Program, since the no-guardrail cyber version is going to that group first.

