Google Releases Gemini 4 Argon: Leads 12 of 18 Official Benchmarks, but the First Wave Goes Only to Cyber Defenders

On September 30 Google released its new flagship model, Gemini 4 Argon, raising the maximum output from 64,000 tokens to 1 million. The first wave is available only to vetted cybersecurity defenders through the Fairwind Program; developers, enterprises and consumers will get it later, with paid API customers and Google AI Ultra subscribers first. Across the 18 benchmarks Google published, Argon leads outright on 12 and ties for first on one, scoring 77.9% on DeepSWE v1.1, but it still trails Claude Opus 5.5 or GPT-6 Astra on Terminal-Bench 4.0, FrontierSWE v2 and a few others. Introductory pricing is $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 after the introductory period.

Defenders first, everyone else later

Gemini 4 Argon is the new flagship Google DeepMind released on September 30, positioned for real-world software engineering, legal and financial knowledge work, and cybersecurity defense. Its most visible change is the per-response output limit, raised from 64,000 tokens to 1 million, enough to produce a large codebase migration or a long report in one pass.

It is not broadly available, though. The first users come only from the Fairwind Program: defensive organizations such as critical infrastructure operators and core technology platforms get access after vetting, and must accept terms including defensive use only, mandatory phishing-resistant multi-factor authentication and no resale. For developers, enterprises and consumers, Google says only that it is "rolling out soon," with no date.

Read the scorecard line by line

On the 18 benchmarks Google published, Argon leads outright on 12 and ties on one: 77.9% on DeepSWE v1.1 (Opus 5.5: 74.2%), 19.6% on Harvey's legal agent benchmark, and 91.7% on the long-video benchmark LVBench. But it trails Opus 5.5 by about 9 points on Terminal-Bench 4.0 (57.4% vs. 66.4%) and GPT-6 Astra by 10.5 points on FrontierSWE v2. Third-party Artificial Analysis gives it 53 on its Intelligence Index, eighth on that leaderboard, which suggests the lead in Google's own numbers narrows under neutral testing.

What it means

Most developers can't use it yet. The price is the signal worth remembering: the introductory $2/$10 per million tokens is about one-fifth of GPT-6 Astra's list price, and the post-introductory $4/$20 matches Opus 5.5. Once the API opens, terminal-heavy work should still be tested against Opus 5.5 first, while long legal and financial documents and long-video analysis are the places to try it early.

via: Google blog, Fairwind Program, VentureBeat report, Artificial Analysis model page