Bright
← Living questionsRecord / Infrastructure

Google unveils Gemini 4 Argon for bigger AI jobs, with cyber defenders first.

Google announced Gemini 4 Argon on September 30 for longer, more complex work. Its initial rollout is to trusted cyber defenders through Fairwind. Broader access has no announced date; company benchmark results and separate Vals evaluations show different strengths.

Maturity
Emerging, stage 2 of 4
Support
2 sources · institution, report
Evidence detail
How we know ↓

Developers and cyber defenders need help with substantial technical work, while retaining enough review and verification to trust the resulting code and analysis.

Long outputs and strong benchmark scores do not by themselves show that a model can finish a valuable technical job reliably or cheaply.

Argon is a model for longer, more complex technical work. Google announced an output limit of one million tokens, compared with 64,000 previously. That capacity could support substantial code changes and detailed analysis; output volume alone does not establish reliable completion.

What was shown, and what wasn’t

Shown

Google reports 77.9% on DeepSWE v1.1, 51.3% on AutomationBench and 91.7% on LVBench. These are company-reported benchmark results. In a separate evaluation, Vals lists Argon first among 41 models on Vals Index at 68.90%, and seventh of eight on CUA-bench at 4.83%. The tasks and test conditions differ; these are not a direct comparison of Google’s evaluation harness. Google reports thousands of staff using Argon and more than 300 TiB of memory freed through optimizations rolled out internally. It says large Rust migrations are undergoing automated and manual audits before production. These are company-reported deployments, not independently audited impact. Google also says Wiz uses Argon for Scan for Good and found a vulnerability in hospital software; this record includes no exploit details.

Not shown · limits

  • Initial rollout is to trusted cyber defenders through Fairwind. Google says it is participating in the voluntary U.S. prerelease access process while expanding access. Broader availability is planned as soon as possible, starting with paid API customers and Google AI Ultra subscribers; no date is announced.
  • An announced output limit of one million tokens does not establish reliable completion of a million-token task, nor does it describe the input context window.
  • Google’s benchmark scores, internal use, memory savings and Wiz account are company-reported. No independent audit of the reported deployment impacts is established here.
  • Vals Index and CUA-bench measure distinct tasks under their own test conditions. First on Vals Index and seventh of eight on CUA-bench indicate uneven strengths, not a universal winner or a direct reproduction of Google’s harness.
  • Google’s announced introductory pricing is $2 per million input tokens and $10 per million output tokens, then $4 and $20 respectively, with cached input 95% off. These are announced prices as of September 30, not a forecast of total task or verification cost.

Bright editorial interpretation

Bright’s interpretation: Argon could support larger technical jobs. Measure valuable work that survives review, along with verification cost, rather than output volume or a universal benchmark-winner claim. Limited release and uneven external evaluations make real-world reliability the next thing to watch.

Still open

When will broader access become available?

What valuable technical work survives independent review, and at what verification cost?

How do reliability and cyber-defense results generalize beyond the reported settings?

How we know2 sources · checked 2026-09-30 · no corrections

Original sources

  1. Gemini 4 Argon announcement · Google ↗ · institution
  2. Google Gemini 4 Argon model evaluation · Vals ↗ · report

Institutions: Google · Google DeepMind · Vals · Wiz

Maturity
Emerging
Event date
2026-09-30
Source published
2026-09-30
Captured
2026-09-30
Last source review
2026-09-30
Editorial method
AI-assisted source review

Bright compared this account with the linked original and supporting sources and kept reported, budgeted, projected, and observed claims distinct. Bright did not independently audit the underlying records.

Maturity describes the tested or operational setting. Confidence describes support for the particular claim; one does not determine the other.

Revision & correction history

2026-09-30 · Prepared the September 30 Argon launch record, separating limited availability, company-reported results, external task evaluations and reliability limits.

No corrections recorded.

Keep looking closer.

See what changed at Bright ↗

Add Bright to your Google Preferred Sources ↗

Suggest a correction · Bright on TikTok