{"schemaVersion":"1.0","generatedFrom":"https://brightaifuture.com/discoveries/gemini-4-argon","record":{"id":"gemini-4-argon","headline":"Google unveils Gemini 4 Argon for bigger AI jobs, with cyber defenders first.","canonicalUrl":"https://brightaifuture.com/discoveries/gemini-4-argon","datePublished":"2026-09-30","dateModified":null,"sourcePublicationDate":"2026-09-30","author":null,"publisher":{"name":"Bright AI Future","url":"https://brightaifuture.com/"},"topics":[],"summary":"Google announced Gemini 4 Argon on September 30 for longer, more complex work. Its initial rollout is to trusted cyber defenders through Fairwind. Broader access has no announced date; company benchmark results and separate Vals evaluations show different strengths.","evidenceState":"Emerging","keyFacts":[{"label":"AI’s role","value":"Argon is a model for longer, more complex technical work. Google announced an output limit of one million tokens, compared with 64,000 previously. That capacity could support substantial code changes and detailed analysis; output volume alone does not establish reliable completion."},{"label":"Documented result","value":"Google reports 77.9% on DeepSWE v1.1, 51.3% on AutomationBench and 91.7% on LVBench. These are company-reported benchmark results. In a separate evaluation, Vals lists Argon first among 41 models on Vals Index at 68.90%, and seventh of eight on CUA-bench at 4.83%. The tasks and test conditions differ; these are not a direct comparison of Google’s evaluation harness.\n\nGoogle reports thousands of staff using Argon and more than 300 TiB of memory freed through optimizations rolled out internally. It says large Rust migrations are undergoing automated and manual audits before production. These are company-reported deployments, not independently audited impact. Google also says Wiz uses Argon for Scan for Good and found a vulnerability in hospital software; this record includes no exploit details."},{"label":"Important limitation","value":"Initial rollout is to trusted cyber defenders through Fairwind. Google says it is participating in the voluntary U.S. prerelease access process while expanding access. Broader availability is planned as soon as possible, starting with paid API customers and Google AI Ultra subscribers; no date is announced."}],"limitations":["Initial rollout is to trusted cyber defenders through Fairwind. Google says it is participating in the voluntary U.S. prerelease access process while expanding access. Broader availability is planned as soon as possible, starting with paid API customers and Google AI Ultra subscribers; no date is announced.","An announced output limit of one million tokens does not establish reliable completion of a million-token task, nor does it describe the input context window.","Google’s benchmark scores, internal use, memory savings and Wiz account are company-reported. No independent audit of the reported deployment impacts is established here.","Vals Index and CUA-bench measure distinct tasks under their own test conditions. First on Vals Index and seventh of eight on CUA-bench indicate uneven strengths, not a universal winner or a direct reproduction of Google’s harness.","Google’s announced introductory pricing is $2 per million input tokens and $10 per million output tokens, then $4 and $20 respectively, with cached input 95% off. These are announced prices as of September 30, not a forecast of total task or verification cost."],"evidenceLinks":[{"title":"Gemini 4 Argon announcement · Google","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/","type":"institution"},{"title":"Google Gemini 4 Argon model evaluation · Vals","url":"https://www.vals.ai/models/google_gemini-4-argon","type":"report"}],"evidencePackUrl":"https://brightaifuture.com/evidence-pack/gemini-4-argon","embedUrl":"https://brightaifuture.com/embed/story/gemini-4-argon","attribution":{"credit":"Bright AI Future","requirements":["Link to the canonical Bright record.","Keep material limitations with the claim they qualify.","Link to the original evidence when repeating a substantive claim.","Do not describe a source check or organization-reported result as independent verification."],"sourceRights":"Linked source material, quotations, trademarks and media remain subject to their owners’ terms. No reuse right is granted for third-party media."}},"claim":{"humanProblem":"Developers and cyber defenders need help with substantial technical work, while retaining enough review and verification to trust the resulting code and analysis.","priorConstraint":"Long outputs and strong benchmark scores do not by themselves show that a model can finish a valuable technical job reliably or cheaply.","aiRole":"Argon is a model for longer, more complex technical work. Google announced an output limit of one million tokens, compared with 64,000 previously. That capacity could support substantial code changes and detailed analysis; output volume alone does not establish reliable completion.","documentedResult":"Google reports 77.9% on DeepSWE v1.1, 51.3% on AutomationBench and 91.7% on LVBench. These are company-reported benchmark results. In a separate evaluation, Vals lists Argon first among 41 models on Vals Index at 68.90%, and seventh of eight on CUA-bench at 4.83%. The tasks and test conditions differ; these are not a direct comparison of Google’s evaluation harness.\n\nGoogle reports thousands of staff using Argon and more than 300 TiB of memory freed through optimizations rolled out internally. It says large Rust migrations are undergoing automated and manual audits before production. These are company-reported deployments, not independently audited impact. Google also says Wiz uses Argon for Scan for Good and found a vulnerability in hospital software; this record includes no exploit details.","whyItMayMatter":"Bright’s interpretation: Argon could support larger technical jobs. Measure valuable work that survives review, along with verification cost, rather than output volume or a universal benchmark-winner claim. Limited release and uneven external evaluations make real-world reliability the next thing to watch.","unresolvedQuestions":["When will broader access become available?","What valuable technical work survives independent review, and at what verification cost?","How do reliability and cyber-defense results generalize beyond the reported settings?"]},"evidenceAssessment":{"state":"Emerging","claimConfidence":"unassessed","reviewState":"approved","reviewMethod":"ai-assisted","reviewNote":"AI-assisted source review of Google’s September 30 announcement and the separate Vals evaluation. Company-reported results, availability and deployment accounts remain attributed; the external tests use distinct tasks and conditions. No independent deployment audit or hands-on model test is claimed.","lastSourceReview":"2026-09-30","independentVerification":"not-established-by-this-source-review"},"sources":[{"id":"source:gemini-4-argon-google","title":"Gemini 4 Argon announcement · Google","url":"https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-4-argon/","type":"institution"},{"id":"source:gemini-4-argon-vals","title":"Google Gemini 4 Argon model evaluation · Vals","url":"https://www.vals.ai/models/google_gemini-4-argon","type":"report"}],"revisions":[{"id":"revision:cedbb2d890d45483828e","recordedAt":"2026-09-30","summary":"Prepared the September 30 Argon launch record, separating limited availability, company-reported results, external task evaluations and reliability limits.","sourceIds":["source:gemini-4-argon-google","source:gemini-4-argon-vals"]}],"corrections":[]}