Google has announced Gemini 4 Argon, a frontier model designed for complex, long-running workflows in software engineering, enterprise knowledge work and cybersecurity defense. In a post on X, @GoogleDeepMind says the model is initially rolling out to trusted testers through its Fairwind Program rather than launching broadly.

Image credit: @GoogleDeepMind on X
Google says broader access will begin with paid API customers and Google AI Ultra subscribers after the early testing phase. It has not announced a launch date or detailed eligibility requirements for either route.
What Gemini 4 Argon is designed to do
Google positions Argon around three main areas: software engineering, enterprise knowledge work and cybersecurity defense. The company also describes capabilities in reasoning, multimodal understanding and creative writing, but its examples focus most heavily on coding and security.
The announcement says thousands of Google employees are already using Argon for specialized coding, research and writing tasks. Google presents these as internal examples of the model’s potential, not as typical results for external users.
What the 1 million-token output limit changes
A token is a unit of text or other model input and output; the amount represented by a token varies by language and content. Google says Gemini 4 Argon can produce up to 1 million output tokens in a single trajectory, up from a previous limit of 64,000 tokens.
That limit is relevant to workflows in which the model must sustain a long sequence of analysis and actions rather than produce a short answer. In principle, a larger output allowance can give an agent more room to inspect a codebase, run or interpret experiments, revise an approach and document the result in one extended process. It does not by itself show that the model will solve every long task reliably.

Image credit: @GoogleDeepMind on X
Google’s description of the capability is tied to “long-horizon” work: problems that involve multiple steps and require the model to maintain a coherent plan over time. The practical value will depend on factors such as tool access, safeguards, context quality and how well the model’s intermediate work can be checked.
Coding and large codebase migrations
Google says Argon agents are being used for everyday debugging, algorithm design and large migrations from C and C++ to Rust. The company says those migrations range from tens of thousands of lines in libraries such as re2 and libgav1 to more than 800,000 lines in the Fuchsia OS Zircon kernel.
The announcement gives a more specific example involving libgav1, Google’s open-source video decoder. Google says Argon agents replaced 32,000 lines of SIMD code in an existing Rust port through repeated profile-guided experiments and compiler analysis. The resulting decoder reportedly ran 2.7 times faster than that Rust port while producing identical video output.
That is an internal Google example, not a promise about the result a typical developer will obtain. Google also says large migrations undergo automated and manual auditing, emulation testing and review before production deployment, underscoring that the model’s output still requires engineering controls.
The company additionally says Argon helped quantum-computing researchers optimize a subroutine’s spacetime resources—qubits multiplied by gates—and beat a published baseline by 40% in one example. Google reports that Argon agents also analyzed fleet-wide profiling data and identified memory optimizations that freed more than 300 tebibytes of memory once deployed, with an estimated 500 tebibytes to 1 pebibyte in total savings.
Enterprise knowledge work and multimodal tasks
Google says Argon is intended for knowledge work in areas including finance, legal research and drafting, tax and business automation. It reports leading results on the Vals Index, which measures work across finance, coding, legal and tax tasks, as well as on Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark.
The company also reports a first-place result on AutomationBench, a Zapier benchmark for end-to-end execution across core business functions, with a score of 51.3%.
For work involving visual material, Google says Argon can analyze professional charts, identify details in long videos and act on a sequence of documents. It reports a 91.7% score on LVBench, an evaluation of long-video understanding, and a 71.6% score on Chartography.
These figures are Google-reported evaluation results. They do not establish how Argon would perform on a particular company’s documents, software or business processes.
Cybersecurity defense
Google says it trained Argon to find, validate and patch software vulnerabilities. The company plans to release the model without cyber guardrails to trusted defenders and its own internal teams so they can use its full cybersecurity capabilities.
Google also says Wiz is using Argon through its Scan for Good initiative, which aims to find and remediate high-risk exposures in critical public infrastructure. In one early demonstration, Google reports that the model found a critical vulnerability exposing sensitive personal information in healthcare software used by hospitals worldwide. The announcement does not provide detailed technical information about the vulnerability or the remediation process.
On CWE-bench v1, which evaluates vulnerability remediation, Google reports that Argon tied for first place with a score of 68%. The company also describes improvements over its earlier 3.8 Flash Cyber model on internal vulnerability-discovery and black-box penetration-testing evaluations.
Because these capabilities can be dual-use, the announcement’s cybersecurity access restrictions are part of the product story. Argon’s initial rollout is aimed at trusted testers and cyber defenders, not unrestricted public use.
How to read the reported benchmarks
Google’s benchmark comparison shows Argon alongside GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 across selected evaluations. The table below reproduces the figures shown in that comparison.
Evaluation | Area | Gemini 4 Argon | GPT-6 Astra | Claude Fable 5.1 | Claude Opus 5.5 |
|---|---|---|---|---|---|
Vals Index | Knowledge work | 68.9% | 63.1% | 65.8% | 67.0% |
AutomationBench | Knowledge work | 51.3% | 41.4% | 31.4% | 42.5% |
DeepSWE v1.1 | Agentic coding | 77.9% | 74.1% | 67.4% | 74.2% |
Chartography | Multimodal understanding | 71.6% | 71.0% | 46.2% | 66.3% |
LVBench | Multimodal understanding | 91.7% | 87.5% | 79.7% | 83.7% |
CWE-bench v1 | Cybersecurity | 68.0% | 68.0% | 58.0% | 67.0% |
Google DeepMind reports these figures, and they have not been independently verified here. The announcement does not provide the complete prompts, model configurations, test conditions or validation process needed to determine how directly the scores can be compared across every evaluation. A higher score on one benchmark also does not establish broader superiority across all enterprise, coding or security tasks.
The reported results are most useful as indicators of the tasks Google is prioritizing: multi-step coding, business automation, visual understanding and vulnerability remediation. They are not substitutes for testing a model against an organization’s own data, tools and review requirements.
Pricing and access
Google says Gemini 4 Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens. Cached input tokens are priced at 95% below the input-token price, according to the announcement.
The announcement references a pricing footnote but does not spell out additional conditions, limits or product-specific details. It also does not say when the introductory prices will change.
Access is currently staged. Google says Argon is rolling out to a set of trusted testers through the Fairwind Program, with the initial announcement emphasizing trusted cyber defenders. Google has not published the program’s eligibility rules or application process in the supplied announcement.
Google says it plans to expand access to developers, enterprises and consumers after gathering feedback and strengthening its safeguards. The company says that broader release will start with paid API customers and Google AI Ultra subscribers, but no specific date or detailed access requirements are provided.
Safeguards before broader release
Google says it is continuing work on four areas of frontier safeguards before making Argon more widely available:
Misuse prevention: The model is designed to refuse harmful cyber and chemical, biological, radiological and nuclear requests while supporting legitimate dual-use research. Google says internal and external red teams tested these safeguards.
Prompt-injection resistance: Google describes protections against indirect prompt injection, in which malicious instructions hidden in documents or other context attempt to redirect the model.
Misalignment monitoring: Google says it monitors the model’s reasoning and actions and can stop execution when behavior moves beyond the user’s intent.
Hardened environments: The company says it is isolating and sealing sandboxed environments before high-risk training or evaluations and plans to share related agent-security practices with partners.
These measures help explain why Google is describing Argon as a phased announcement rather than a general launch. The model’s reported capabilities are aimed at workflows where mistakes, unauthorized actions or successful misuse could have significant consequences. For now, Gemini 4 Argon is best understood as a limited-access system undergoing evaluation, with broader availability and real-world performance still to be established.





0 comments
No approved comments yet. You can start the conversation.
Leave a comment