AIAI News Online

Gemini 4 Argon explained: why Google's best model launched behind a gate

Google's Gemini 4 Argon leads most of its benchmark table at an introductory $2/$10 per million tokens, yet only vetted cyber defenders can use it. Here's why.

14 min at full depth15 sources

In 60 seconds

  • Google DeepMind announced Gemini 4 Argon on September 30, 2026, but released it only to vetted cyber defenders in its Fairwind Program; developers and consumers get it later, with no date given.
  • One model, two doors: trusted defenders get Argon with cyber guardrails off, while the public version will sit behind refusals, activation probes, reasoning monitors and prompt-injection defences.
  • Argon leads 13 of 19 rows in Google's own table, but Google computed Argon's score itself on 10 of the 19 rows, independent testers place it near the top rather than clearly first, and no model card has been published.

Google DeepMind's new flagship model, Gemini 4 Argon, leads most of the benchmark table Google published for it, and almost nobody is allowed to use it. Outside Google, it went on launch day only to vetted cybersecurity defenders, because the skill that lets it patch holes in software is the same skill that would let it find those holes for an attacker. Google joins Anthropic and OpenAI in keeping its strongest cyber capability behind a gate, and the way that gate works tells you more about where AI is heading than any score does.

For: Everyone

The plain-English version

Imagine a locksmiths' guild invents a machine that can study any lock and, within minutes, work out how to open it and how to repair that weakness. Give it to a building manager and it is a gift: every door in the city gets checked and re-keyed. Give it to a burglar and it is the opposite. Same machine, same skill.

That is Google's position. On September 30, Google DeepMind announced Gemini 4 Argon, an AI model that Google says can find serious security flaws in software, confirm they are real and write the fix, without a person guiding each step. Software flaws are how hospitals, banks and power companies get hacked, so fixing them faster is valuable. But finding a flaw is also the first step of breaking in.

So Google did what a sensible guild would do. It lent the machine first to people it has checked out: governments, operators of critical infrastructure such as hospitals and energy networks, and core technology platforms in its Fairwind Program, which Google says has more than 650 partners. Those defenders get Argon with its hacking-related safety limits switched off, so they can use its full ability to repair things. Google's stated aim for the programme is "a vital adaptation window": time to harden systems before bad actors can exploit the new capabilities.

Everyone else waits. Google says developers, businesses and consumers will get Argon "as soon as possible", with limits fitted: it is designed to refuse harmful requests, and automated monitors watch what it is doing and can stop it. Google gave no date.

Two more things are worth knowing. First, Argon is also Google's bid to compete with OpenAI and Anthropic at everyday professional work such as programming, finance and legal research, and it will be priced low to start. Second, the impressive scores come from a report card the student partly filled in: Google ran about half of the tests on its own model itself. Independent testers who have tried Argon place it near the top of the class, not clearly at the top.

For: Curious

How it actually works

The problem. A software vulnerability is a mistake in code that lets someone do what they should not, such as read private records or take over a machine. Until recently, finding these took scarce human experts. A model that does it autonomously changes the economics for both sides at once, and the two sides cannot be separated. One of Google's own internal tests for Argon measures the "ability to write exploits against real web application vulnerabilities, without access to the codebase", according to its evaluation document. That is a penetration test: an attack carried out with permission.

The old way. One model, one launch day, one rulebook. The model is trained to refuse anything that looks like hacking. This fails in both directions. Defenders are blocked, because proving a bug is real usually means writing an exploit for it. And determined attackers rephrase until something gets through.

The new idea: one model, two doors.

  1. Train one highly capable model. Google says it "trained Gemini 4 Argon to be highly capable at cybersecurity defense".
  2. Door A: vetted defenders, guardrails off. Fairwind applicants face background checks, must limit access to internal security, incident-response or penetration-testing teams, must use phishing-resistant multi-factor authentication, and may not share or resell access. In exchange, Google is "releasing Argon without cyber guardrails" to them, and to its own internal teams.
  3. Door B: everyone else, later, guardrails on. The public version waits until Google has finished iterating on its safeguards.
  4. Use the gap. Defenders patch. Google collects feedback. The US government gets a look through what Google calls a "voluntary process for pre-release model access".

Google's announcement describes four layers of safeguards. The first three wrap the model people will use; the fourth protects Google's own training and testing:

Layer What it guards against How
Misuse defences A person asking for help with a cyber, chemical, biological, radiological or nuclear attack Trained refusals, plus monitors that read the model's internal activations; tested by internal and external red teams
Prompt-injection hardening Malicious instructions hidden in a web page, email or file the model reads while working Automated red teaming and adversarial training
Misalignment monitoring The model itself going beyond what the user authorised Monitors read its chain-of-thought and actions and "stop execution when necessary"
Sealed sandboxes A model reaching real systems during risky training or testing Environments are isolated and sealed before high-risk runs begin

Why it can work. Patches last. If defenders fix a flaw during the head start, an attacker who later gets the same capability finds one less door open. The approach buys time; it does not buy permanence. Comparable skills are already spreading: on the independent CWE-bench vulnerability-fixing test, Argon is tied with two rival models, and software teams write new bugs every day.

For: Practitioner

The deep dive

What Google published, and what it did not

The public record is a blog post by Koray Kavukcuoglu, a five-page evaluation methodology PDF, and product pages. There is no architecture description, parameter count, training-data statement, knowledge cutoff or API model ID. We could not find a model card, and an independent briefing that looked for one reports the same. The post also does not say whether Argon reached a cyber Critical Capability Level under Google's Frontier Safety Framework, which calls for safety case reviews before external launches when relevant levels are reached.

What is stated: the output limit rises to 1M tokens from 64K. The input context window is not given in the post; Google's own long-context tests run to 1M tokens, and Artificial Analysis lists 1M.

Reading the benchmark table: who ran what

Google's evaluation document has 19 rows covering 18 benchmarks (GraphWalks appears at two context lengths). By our count Argon leads 13 rows outright, ties one and trails on five. Google computed Argon's own score on 10 of the 19 rows; the other nine come from third-party leaderboards. Six of the 13 leads are on rows Google ran itself. A selection, all figures as published by Google:

Benchmark Argon GPT-6 Astra Claude Fable 5.1 Claude Opus 5.5 Source of Argon's score
Vals Index 68.9% 63.1% 65.8% 67.0% Vals AI
AutomationBench 51.3% 41.4% 31.4% 42.5% Zapier leaderboard
DeepSWE v1.1 77.9% 74.1% 67.4% 74.2% Google, own harness
FrontierSWE v2 55.0% 65.5% 56.3% 62.3% Proximal leaderboard
Terminal-Bench 4.0 57.4% 58.2% 57.9% 66.4% Google
Terminal-Bench Science 0.1 57.6% 68.1% 52.6% 63.3% Google, 6x verifier timeout
GraphWalks 256k–1M 84.2% 71.8% 65.0% 66.8% Google ran all models
OSWorld-2.0 (offline subset) 69.2% 72.6% — — Google, max over 3 runs
LVBench 91.7% 87.5% 79.7% 83.7% Google ran all models
CWE-bench v1 68.0% 68.0% 58.0% 67.0% Public leaderboard

Three details in the footnotes matter. The headline DeepSWE figure is "self computed, using a mini-swe agent harness", while Astra's number comes from the public leaderboard and the Claude numbers from Anthropic's system cards, so the scores in that row were not produced under one setup. On LVBench, a long-video test, Google fed Gemini one frame per second but gave Astra 800 frames, Fable 300 and Opus 600 "due to API limitations", so the models did not see the same footage. And Argon's Terminal-Bench Science run got a six-times-longer verifier timeout and still trails.

What independent testers found

Evaluator Metric Argon Nearest comparison
Vals AI Vals Index 68.90%, 1st of 41 Claude Sonnet 5.5, 67.04%
Artificial Analysis Intelligence Index ("High" setting) 53, 8th of 223 Claude Opus 5.5, about 58 (as reported)
CWE-bench v1 Programmatic pass@1 68% Grok 4.7 and GPT-6 Astra, 68%
CWE-bench v1 Judge-panel pass@1 62% Claude Opus 5.5, 67%
Proximal (via Google's table) FrontierSWE v2 55.0% GPT-6 Astra, 65.5%

CWE-bench is the most relevant here: 120 held-out audit-and-patch tasks across 73 weakness types, built by Collinear AI and evaluated independently by Artificial Analysis. A patch passes the programmatic gate if the exploit stops working and existing tests still pass; a panel of three AI judges from different model families then grades fix quality. Argon's six-point drop from programmatic to judged score is the largest among the top four models.

Cost: cheap per token, not per task

Google lists $2 per million input tokens and $10 per million output during an introductory period of unstated length, then $4 and $20, with cached input discounted 95%. Artificial Analysis measured $1.99 per task on its index at the introductory price, but also found that Argon generated 110M output tokens across the index against a median of 82M for comparable reasoning models. On CWE-bench, Argon cost $6.63 per rollout against $2.85 for Astra and $0.79 for Opus 5.5 at similar pass rates. On the Vals Index it cost $15.68 per test, while GPT-6.1 Sol scored 61.15% at $3.24.

The misuse layer: activation probes

Google says it improved techniques that monitor Argon's "internal activations" and links to a DeepMind paper, Building Production-Ready Probes For Gemini. A probe is a tiny classifier that reads the model's hidden state instead of its text. The simplest version averages a learned linear readout over all nin_i tokens of an input SiS_i:

fLinear(Si)=1ni∑j=1niw⊤xi,jf_{\text{Linear}}(S_i) = \frac{1}{n_i}\sum_{j=1}^{n_i} w^\top x_{i,j}

Here xi,jx_{i,j} is the activation at token jj and ww is a learned direction meaning, roughly, "this looks like an attack request". The average is the weakness: bury one harmful paragraph in a 900k-token prompt and the mean washes it out. The paper's fix replaces the mean with a maximum, as in its MultiMax probe:

fMultiMax(Si)=∑h=1Hmax⁡j∈[ni]vh⊤yi,jf_{\text{MultiMax}}(S_i) = \sum_{h=1}^{H} \max_{j \in [n_i]} v_h^\top y_{i,j}

with HH heads vhv_h over MLP-transformed activations yi,jy_{i,j}. Now one strongly firing position is enough. On long-context attacks, the paper reports false-negative rates (the share of attacks missed) of 99.13% for a mean-pooled linear probe and 87.88% for a standard attention probe, against 3.03% for the design the authors selected, a "Max of Rolling Means" attention probe that takes the maximum over short sliding windows. Probes are then cascaded: confident scores are decided by the probe, and only an uncertain band τ1<fp(Si)<τ2\tau_1 < f_p(S_i) < \tau_2 is escalated to a full LLM classifier. The authors report that such a cascade achieves lower test loss than Gemini 2.5 Flash alone at roughly 1/50th of the inference cost.

Two caveats. The paper studies probes on Gemini 2.5 Flash and predates Argon, and Google does not say which probe design Argon uses. And the paper's own adaptive red-teaming set was much harder: the selected probe missed 42.46% of those attacks, and the five best probes missed between 22.97% and 50.58%. The authors say their techniques "do not significantly reduce the success rate of adaptive adversarial attacks".

The misalignment layer: reading the reasoning

DeepMind's June agent control roadmap treats untrusted AI agents deployed inside Google as potential insider threats. The full roadmap defines four detection tiers, D1 to D4, keyed to how well a model can evade oversight: chain-of-thought monitoring is enough at D1, while D4 calls for monitoring of internal activations. It also defines three prevention-and-response tiers, R1 to R3: after-the-fact review of logs at R1, real-time blocking at R2, and infrastructure to shut agents down at R3. In that vocabulary, a monitor that reads Argon's chain-of-thought and halts execution is D1-style detection with a real-time ("synchronous") response. That mapping is our reading; Google does not place Argon on these tiers. Google adds that a similar system watched training runs and alerted an incident-response team, and that it took precautions against feeding those findings back into training so as not to risk "shaping Argon's reasoning to evade our monitoring". That choice matters: train against your monitor and you risk teaching the model to hide from it.

Closest prior work

Anthropic's Fable and Mythos are one underlying model with two access tiers; on the public tier, requests flagged as cybersecurity, biology or chemistry are answered by an older model, Claude Opus 4.8, instead. OpenAI's system card for GPT-6 Astra says that model reaches the "Critical" level of cybersecurity capability under its Preparedness Framework and describes a Trusted Access for Cyber programme; Astra went first to a limited set of organizations on September 3, with OpenAI saying paying ChatGPT users and the API would follow within days, Unite.AI reported. Google's version differs in two ways: the public tier has not shipped at all, and there is no model card yet.

For: Everyone

Why it matters

Everyday users. Nothing changes today. When Argon reaches the Gemini app, Google AI Ultra subscribers come first. The benefit most people will feel is indirect: Google says Argon agents are working on migrating C and C++ code to memory-safe Rust inside the company, with projects ranging up to 800K+ lines for the Fuchsia Zircon kernel. If that work lands, fewer memory bugs in widely used software means fewer breaches, whoever you are.

Developers and builders. The introductory $2/$10 price matches the rates at which OpenAI launched GPT-6.1 Sol a day earlier, Forkast reports, and VentureBeat notes that it is a fifth of GPT-6 Astra's. A 1M-token output limit opens up single-response jobs such as large code migrations. Budget by task, though, not by token, given the token counts independent testers measured. And if you build security tooling, expect the public model to refuse or halt on some legitimate work.

Companies. Access to the strongest models is becoming a function of who you are. Security teams at infrastructure operators and software platforms now have a reason to go through verification programmes. Organizations outside those categories will get frontier defence capability later than their better-connected peers.

The field. Three labs have now converged on giving vetted defenders earlier or less restricted access to cyber-capable models than the general public gets. It happened in the same week that OpenAI held back GPT-6.1 Astra because, in the words of its head of safety systems Saachi Jain, it "didn't quite meet the bar in terms of staying within scope and authorization", and that six tech leaders signed a voluntary White House AI safety accord.

The second-order effect is a two-tier ecosystem in which a few companies decide who counts as a trusted defender. That arrangement holds only while the capability is scarce. On CWE-bench, DeepSeek-V4.1-Flash already scores 55% at $0.09 per rollout. The head start defenders are being given is real, and it is short.

For: Critical

What to be skeptical of

The safety paperwork is missing. Google cites its Frontier Safety Framework but has published no model card, no capability-level determination and no safety case for Argon. OpenAI's system card for Astra opens with a safety overview that states its cyber risk level. Until Google publishes the equivalent, outsiders cannot audit the claim that this rollout is proportionate.

The table flatters. The announcement text names benchmarks Argon wins or ties; the five it trails appear only in the table. Rivals that would complicate the picture are absent from that table: Grok 4.7 ties Argon on CWE-bench and beats it on the pass@4 tie-break (81% to 75%) that Google's own methodology cites, and Claude Sonnet 5.5 sits within two points on the Vals Index, where Vals puts the standard error for the top Claude models at about ±0.9.

Insiders are reportedly unconvinced. Bloomberg reported, as summarized by Analytics India Magazine, that some Google employees with direct access say Argon's benchmark strength does not fully carry over to real coding work. The sources are unnamed and we have not seen the original report. Google disputed that characterisation, and another employee cited in the same report described a "large consensus" internally that the model is at the frontier.

The showcase results are unverified. A quantum-computing subroutine that "beat the published baseline by 40%", data-centre memory savings of more than 300 TiB, a memory-safe video decoder running 2.7x faster than an earlier Rust port and a critical vulnerability in hospital software uncovered in work with Wiz are all Google's own accounts, without papers, code or advisories attached so far.

The safeguard claims lack numbers. The "most resilient model yet" claim against prompt injection rests on a 0.7% attack success rate at 15 attempts on Gray Swan's benchmark, per Google's cyber page, with no sample size disclosed. No recall or false-alarm figures are given for the reasoning monitors, which are only as good as the reasoning is honest: OpenAI's card says Astra is "more capable of controlling its own CoT" than GPT-5.6 Sol, and DeepMind researchers have themselves urged the industry to keep reasoning readable.

Sandboxes are not a full answer. Cryptographer Matthew Green argued this week that useful agents need access to real information, and that the deeper problem, which he illustrates with OpenAI's agents, is that they "will do what they're told by whoever manages to get text in front of them".

Finally, nobody has yet shown that a defenders-first window measurably reduces harm. That is an open empirical question.

For: Everyone

What to watch next

  • A general-availability date and API model ID. Google has promised paid API customers and AI Ultra subscribers first. Watch whether the introductory price gets an end date.
  • A model card and a capability-level statement. Specifically, whether Google says Argon did or did not reach a cyber Critical Capability Level, and whether a safety case is published.
  • Third-party reruns of the self-computed rows. DeepSWE v1.1 on Datacurve's own leaderboard and Terminal-Bench 4.0 on the official one would settle whether Google's harness numbers hold. Artificial Analysis has reportedly put Argon's Terminal-Bench score at 57%, in line with Google's.
  • Evidence from the head start. Public advisories or CVEs credited to Fairwind partners, including the healthcare flaw Google says Argon uncovered in work with Wiz, would show whether the window produces patches.
  • Washington's follow-through. SiliconANGLE reports that President Trump said the White House is considering a 10-person AI safety committee and that he plans to name an "AI czar". Whether voluntary pre-release access becomes a standing process is the thing to track.
  • What OpenAI does with GPT-6.1 Astra. Whether it ships, and how it then handles "scope and authorization", will show whether the industry's monitoring tools are improving as fast as its models.

Check your understanding

Pick an answer — you'll see why right away.

1. Why did Google give Gemini 4 Argon to vetted cyber defenders before anyone else?

2. Google reports 77.9% for Argon on DeepSWE v1.1 against about 74% for its closest rivals. What is the main reason to treat that gap with caution?

3. A simple activation probe averages a signal over every token in the prompt. Why does that fail on very long inputs, and what fixes it?

4. Google says monitors read Argon's chain-of-thought and can stop it. What does that safeguard depend on?

Glossary

Staged (phased) release
Shipping a model to a small vetted group first and widening access step by step as safeguards are tested.
Dual-use
A capability that serves both protective and harmful purposes, such as finding software vulnerabilities.
Fairwind Program
Google's limited-access programme that gives vetted governments, critical infrastructure operators and core technology platforms early access to its most capable cyber-defence models.
Cyber guardrails
The filters and trained refusals that stop a model from helping with requests that look like hacking.
Activation probe
A small classifier that reads a model's internal activations, rather than its text output, to flag things like attack requests.
Indirect prompt injection
An attack that hides malicious instructions in content an AI agent reads, such as a web page or email, to hijack its behaviour.
Chain-of-thought monitoring
Automatically reading a model's written intermediate reasoning to catch it going off-task or acting deceptively.
Critical Capability Level (CCL)
A threshold in Google DeepMind's Frontier Safety Framework at which a model may pose heightened risk of severe harm without mitigations.
Harness
The scaffolding of tools, prompts and loops wrapped around a model so it can act as an agent; it strongly affects benchmark scores.
pass@1
The share of benchmark tasks a model solves on its first and only attempt.

Questions people ask

What is Gemini 4 Argon?

Gemini 4 Argon is Google DeepMind's new frontier AI model, announced on September 30, 2026. Google positions it for coding, knowledge work such as finance and legal tasks, creative writing and cybersecurity defence, including finding and patching software vulnerabilities autonomously.

When will Gemini 4 Argon be available to the public?

Google has not given a date. It says Argon will reach developers, enterprises and consumers "as soon as possible", starting with paid API customers and Google AI Ultra subscribers, after it finishes iterating on safeguards with early testers.

How much does Gemini 4 Argon cost?

Google lists introductory API pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the introductory period, with a 95% discount on cached input. Independent testers note the model uses many tokens per task, so per-task cost can be higher than the per-token price suggests.

Is Gemini 4 Argon better than GPT-6 Astra and Claude Opus 5.5?

It depends on the test. In Google's own table Argon leads 13 of 19 rows, and it ranks first on the independent Vals Index, but it trails both rivals on FrontierSWE v2 and scores about 53 on the Artificial Analysis Intelligence Index against roughly 58 for Claude Opus 5.5. There is no clear overall winner.

What is Google's Fairwind Program?

Fairwind is a limited-access programme Google introduced on September 2, 2026 for governments, critical infrastructure operators and core technology platforms. Google says it has more than 650 partners, runs background checks on applicants, and restricts use to internal security teams doing defensive work.

Why is Gemini 4 Argon restricted to cyber defenders?

Google says the model can autonomously find, validate and patch critical vulnerabilities, a skill attackers could also use. Releasing it to vetted defenders first is meant to give them time to harden systems while Google strengthens safeguards for the public version.

Sources

  1. Gemini 4 Argon: our next era of frontier intelligence — Google · official announcement
  2. Gemini 4 Argon Model evaluation: Approach, methodology & results — Google · docs
  3. Fairwind Program — Google DeepMind · docs
  4. Proactive cyber defense for governments and enterprises — Google · official announcement
  5. Strengthening our Frontier Safety Framework — Google DeepMind · official announcement
  6. Building Production-Ready Probes For Gemini — arXiv (Google DeepMind authors) · paper
  7. Securing the Future of AI Agents — Google DeepMind · official announcement
  8. GDM AI Control Roadmap — Google DeepMind · paper
  9. Vals Index — Vals AI · analysis
  10. CWE-bench v1 — Collinear AI · analysis
  11. Gemini 4 Argon (High) model page — Artificial Analysis · analysis
  12. Gemini 4 Argon leads 13 of 18 rows; Google ran 9 of them itself — pk-sharma.com · analysis
  13. Google Gemini 4 Struggles With Coding Despite Strong Benchmark Numbers: Report — Analytics India Magazine (reporting on Bloomberg) · news
  14. GPT-6 Astra System Card — OpenAI · docs
  15. OpenAI holds off on releasing new model over safety concerns, saying it "didn't quite meet the bar" — CBS News · news

How this was made: researched and written by an AI model (Claude) from the primary sources listed above, then checked claim-by-claim against those sources in a separate AI fact-check pass. Spotted an error? Email [email protected] and we correct it publicly. Our process.