# Google unveils Gemini 4 Argon for bigger AI jobs, with cyber defenders first.

Agent contract: 1.2.0

Google announced Gemini 4 Argon on September 30 for longer, more complex work. Its initial rollout is to trusted cyber defenders through Fairwind. Broader access has no announced date; company benchmark results and separate Vals evaluations show different strengths.

Canonical: https://brightaifuture.com/discoveries/gemini-4-argon
Format: discovery
Source publication: 2026-09-30
Bright publication: 2026-09-30
Substantive update: None recorded
Evidence and review: Emerging; confidence: unassessed; approved; ai-assisted. AI-assisted source review of Google’s September 30 announcement and the separate Vals evaluation. Company-reported results, availability and deployment accounts remain attributed; the external tests use distinct tasks and conditions. No independent deployment audit or hands-on model test is claimed.

## The human problem

Developers and cyber defenders need help with substantial technical work, while retaining enough review and verification to trust the resulting code and analysis.

## The prior constraint

Long outputs and strong benchmark scores do not by themselves show that a model can finish a valuable technical job reliably or cheaply.

## AI’s actual role

Argon is a model for longer, more complex technical work. Google announced an output limit of one million tokens, compared with 64,000 previously. That capacity could support substantial code changes and detailed analysis; output volume alone does not establish reliable completion.

## The documented result

Google reports 77.9% on DeepSWE v1.1, 51.3% on AutomationBench and 91.7% on LVBench. These are company-reported benchmark results. In a separate evaluation, Vals lists Argon first among 41 models on Vals Index at 68.90%, and seventh of eight on CUA-bench at 4.83%. The tasks and test conditions differ; these are not a direct comparison of Google’s evaluation harness.

Google reports thousands of staff using Argon and more than 300 TiB of memory freed through optimizations rolled out internally. It says large Rust migrations are undergoing automated and manual audits before production. These are company-reported deployments, not independently audited impact. Google also says Wiz uses Argon for Scan for Good and found a vulnerability in hospital software; this record includes no exploit details.

## Why it may matter

Bright’s interpretation: Argon could support larger technical jobs. Measure valuable work that survives review, along with verification cost, rather than output volume or a universal benchmark-winner claim. Limited release and uneven external evaluations make real-world reliability the next thing to watch.

## Limitations

Initial rollout is to trusted cyber defenders through Fairwind. Google says it is participating in the voluntary U.S. prerelease access process while expanding access. Broader availability is planned as soon as possible, starting with paid API customers and Google AI Ultra subscribers; no date is announced.

An announced output limit of one million tokens does not establish reliable completion of a million-token task, nor does it describe the input context window.

Google’s benchmark scores, internal use, memory savings and Wiz account are company-reported. No independent audit of the reported deployment impacts is established here.

Vals Index and CUA-bench measure distinct tasks under their own test conditions. First on Vals Index and seventh of eight on CUA-bench indicate uneven strengths, not a universal winner or a direct reproduction of Google’s harness.

Google’s announced introductory pricing is $2 per million input tokens and $10 per million output tokens, then $4 and $20 respectively, with cached input 95% off. These are announced prices as of September 30, not a forecast of total task or verification cost.

## Unresolved questions

When will broader access become available?

What valuable technical work survives independent review, and at what verification cost?

How do reliability and cyber-defense results generalize beyond the reported settings?

## Provenance and history

{
  "dates": {
    "eventDate": "2026-09-30",
    "publicationDate": "2026-09-30",
    "captureDate": "2026-09-30",
    "lastReviewedDate": "2026-09-30"
  },
  "provenance": {
    "origin": "editorial",
    "externalId": "https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/"
  },
  "revisions": [
    {
      "id": "revision:cedbb2d890d45483828e",
      "recordedAt": "2026-09-30",
      "summary": "Prepared the September 30 Argon launch record, separating limited availability, company-reported results, external task evaluations and reliability limits.",
      "sourceIds": [
        "source:gemini-4-argon-google",
        "source:gemini-4-argon-vals"
      ]
    }
  ],
  "corrections": []
}

## Original sources

- [Gemini 4 Argon announcement · Google](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/)
- [Google Gemini 4 Argon model evaluation · Vals](https://www.vals.ai/models/google_gemini-4-argon)

## Continue exploring


