Is Gemini 4 Argon available now, and what do Google's benchmark claims show?
Google has announced a frontier model for coding, professional work and cyber defence, but access is initially limited to trusted defenders. Its 19-row comparison is broad and often strong; it is still a vendor evaluation, not an independent test or a public release.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
What Google's limited Gemini 4 Argon launch establishes about capability, access and evidence quality
At a glance
- 1Gemini 4 Argon is not generally available: Google says it is first going to trusted cyber defenders through its Fairwind programme while pre-release review and safety work continue.
- 2Google publishes 19 benchmark rows and Argon leads outright on 13, ties one and trails on five; the mixed result is more informative than a claim that it is simply the best model.
- 3The introductory API price is $2 per million input tokens and $10 per million output tokens, later doubling to $4 and $20, but paid API customers and Google AI Ultra subscribers do not yet have a firm access date.
Living evidence record
Impact record IAI-0CIE7ED
Evidence stage
Announced
Confidence
Supported
Reporting basis
Multi-source analysis
Independent support
Not yet
Record status
Monitoring
Last checked
1 October 2026
Source trail
3 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
The announcement is a limited deployment, not a public launch
Google announced Gemini 4 Argon on 30 September as a frontier model for long, multi-step work in software engineering, finance, legal research and cybersecurity. The first users are a set of trusted cyber defenders in Google's Fairwind programme and Google's own teams. The company says it is also taking part in the United States government's voluntary pre-release access process. Developers, enterprises and consumers are promised later access, beginning with paid API customers and Google AI Ultra subscribers, but no date is given.
That distinction matters for anyone making a purchasing or workflow decision. A product page and benchmark table show what Google intends Argon to do; they do not show that an ordinary customer can test it today. The announced introductory price is $2 per million input tokens and $10 per million output tokens, with cached input at 95% below the input rate. Google says the regular price after the introductory period will be $4 and $20 respectively. Until access widens, teams cannot verify total cost per completed task in their own environment.[1][2]
The benchmark table is broad, but the result is not a clean sweep
Google's model page publishes 19 comparison rows spanning professional knowledge work, agentic coding, machine-learning engineering, science and mathematics, long context, computer use, multimodal understanding and vulnerability repair. Counting those rows, Argon leads outright on 13, ties GPT-6 Astra on CWE-bench v1 and trails a rival on five. It scores 77.9% on DeepSWE v1.1, 68.9% on the Vals Index and 51.3% on AutomationBench. Those figures support a claim of competitive breadth, not universal superiority.
The losses help readers understand the boundary. Argon scores 55.0% on FrontierSWE v2 against 65.5% for GPT-6 Astra, 57.4% on Terminal-bench 4.0 against 66.4% for Claude Opus 5.5, and 45.3% on PostTrainBench against 49.3% for Opus. On the published offline subset of OSWorld-2.0, Argon's 69.2% trails Astra's 72.6%. Benchmark names, harnesses, subsets and scoring rules differ, so small gaps should not be read as stable real-world rankings without repeated independent runs.[2][3]
A million output tokens expands the design space and the failure surface
Google says Argon's output limit rises from 64,000 to one million tokens. A larger output allowance can let one model trajectory inspect more material, use tools over more steps and produce a substantial migration or research product without being cut off. It also creates practical questions that a token ceiling cannot answer: whether the model stays on task, recovers from tool errors, preserves instructions, exposes sensitive data or produces a reviewable audit trail across an extremely long run.
Google gives internal examples rather than controlled workplace studies. It says Argon helped optimise a quantum-computing subroutine, identified memory savings across data centres and assisted C or C++ to Rust migrations, including work on an 800,000-line kernel. The company also says thousands of employees have used the model. These examples show plausible applications and scale, but no denominator is provided for attempted projects, failed runs, reviewer hours or comparison with existing engineering methods.[1][2]
Cyber capability explains both the urgency and the narrow access
Google says Argon can find, validate and patch vulnerabilities and will give approved defenders access without cyber guardrails so they can use its full capability. It reports a 68% score on CWE-bench v1, tied with GPT-6 Astra in Google's table. It also cites internal testing across codebases in 20 programming languages and a Wiz black-box penetration test. Google says an early deployment found a critical flaw exposing personal information in healthcare software used by hospitals, but it does not identify the vendor, affected denominator, disclosure timeline or independent verifier.
The safety case is therefore consequential but incomplete. Google says internal and external red teams tested safeguards, that the model is more resilient to indirect prompt injection, and that monitoring can stop actions when they appear misaligned. It does not publish the number of red-team attempts, the false-positive rate of monitoring or the performance difference between the guarded public configuration and the unguarded defender version. Limited release is a sensible source of additional evidence, but it is not itself proof that every control works.[1][2]
What this changes for people and organisations
Security teams may gain a faster way to locate and repair difficult flaws, while software developers may eventually automate larger migrations and investigations. The same capabilities can raise the cost of poor permissions, weak logging or automatic execution. Organisations should keep least-privilege tool access, sandboxing, human approval for consequential changes and their own acceptance tests outside the model. Procurement should compare solved tasks, reviewer time, incident rates and total cost rather than importing a vendor's benchmark rank into a business case.
For workers in legal, finance and engineering roles, the announcement is evidence that frontier models are being aimed at longer professional workflows, not proof that whole occupations have been automated. Google's strongest knowledge-work result is still a score on a constructed evaluation. Actual work includes accountability, confidential context, client communication, regulation and repair when a tool is wrong. Those requirements determine whether a model saves time, moves risk to reviewers or changes staffing.[1][2][3]
What would change our assessment
Confidence would rise when independent evaluators can run the same model and configuration, when Google reports sample sizes and variance for every benchmark, and when users publish measured results on their own workloads. A model or safety report should separate internal evaluations, third-party leaderboards and competitor-reported numbers and document the exact harness, tools, thinking settings and number of attempts used for each row.
The assessment would improve further if the Fairwind deployment produces responsibly disclosed vulnerabilities and evidence that fixes reached affected systems without exposing new risks. It would worsen if public access slips without explanation, if independent tests fail to reproduce the largest gains, or if the unguarded cyber model causes an incident. For now, Argon is a material technology announcement with promising vendor evidence and unusually narrow availability. It is not Breaking news and it is not yet a generally testable product.[1][2][3]
What this means for people
- Developers, analysts and security teams may eventually run longer workflows with fewer hand-offs, but need stronger review, access control and cost measurement for long autonomous runs.
- People whose data sit in vulnerable systems could benefit from faster patching, while also facing greater harm if comparable capability is misused or deployed with weak controls.
Global context
Google is a US-headquartered company and the pre-release process described is US-based, but its customers, cloud infrastructure and the software it may inspect are global. Availability, data residency, cyber law, incident disclosure and professional-accountability rules differ by country. Benchmark performance should therefore be separated from the legal and organisational conditions that determine whether Argon can be used safely in a particular jurisdiction.
What the evidence does not yet show
- The principal evidence comes from Google, the model developer; general access is not yet available for independent reproduction.
- Google's 19-row table combines different benchmarks, harnesses, subsets and result sources rather than one neutral head-to-head experiment.
- Internal productivity and cybersecurity examples lack denominators for attempts, failures, reviewer time and affected systems.
What to watch next
- A firm release date and documented configuration for paid API customers and Google AI Ultra subscribers.
- Independent replications of DeepSWE v1.1, FrontierSWE v2, AutomationBench and CWE-bench v1 under comparable harnesses.
- A detailed model or safety report separating guarded and unguarded cyber capability and quantifying red-team coverage.
Evidence trail
Sources used for this report
Links checked 1 October 2026
This report is labelled multi-source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Technology
Does GPT-6.1 Sol make AI work cheaper? Count successful tasks, not just tokens
OpenAI's 29 September release reports stronger performance at lower task costs in selected evaluations. The useful comparison for buyers is the cost of checked, completed work in their own setting.
4 min · 2 sources
Technology
What does beating a Stratego champion prove about AI under hidden information?
Ataraxos won 15 of 20 games against one of Stratego's most decorated players and transferred its methods to three other games. The peer-reviewed Nature paper shows a real advance in efficient game strategy—not that the system can already run negotiations, markets or military decisions.
6 min · 3 sources
Technology
Does self-hosting keep an AI coding agent away from sensitive source code?
IBM has made a customer-managed version of its Bob development agent generally available, including supported air-gapped deployments. Local inference can keep code inside an approved environment, but the announcement does not prove that every integration, model or generated change is secure.
6 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.