Google Unveils Gemini 4 Argon, Returning to the AI Frontier but with Limited Access

Deep News
1 hour ago

Google has launched its flagship AI model Gemini 4 Argon, surpassing the top products from OpenAI and Anthropic across multiple benchmarks, signaling that the search giant has fought its way back to the competitive frontier after months of lagging behind.

However, this highly anticipated model is currently only available to select cybersecurity partners, with ordinary developers and enterprise users still waiting.

On September 30, Google officially launched its latest flagship AI model Gemini 4 Argon, which ranked first or tied for first in 13 out of 18 benchmarks, excelling particularly in knowledge work and long-context reasoning.

Given that AI safety issues have sparked widespread discussion in the tech community in recent weeks, Google stated that its new model will initially be available only to specific companies and organizations focused on cybersecurity defense. The model will be released more broadly once defenders have had the opportunity to patch system vulnerabilities and fix flaws discovered by the model.

Google added that Gemini 4 Argon includes new safety guardrails to prevent misuse and unintended behavior. These measures include monitoring capabilities that can catch the model when it oversteps into tasks beyond user expectations and can terminate its activities if it goes out of bounds.

Tulsee Doshi, Senior Director and Product Lead for Gemini, said:

We see these guardrails working.

Leading in Knowledge Work, Mixed Performance in Coding

Gemini 4 Argon demonstrates overwhelming advantages in knowledge work benchmarks, but its coding performance is more complex.

In comparison data published by Google, Argon scored 68.9% on the Vals Index, ahead of Opus 5.5's 67.0%; scored 65.4% on Vals Finance Agent v2, leading Fable 5.1 by approximately 6.5 percentage points; and scored 19.6% on Harvey's Legal Agent Benchmark, roughly triple Fable 5.1's score.

In coding, Argon topped the DeepSWE v1.1 leaderboard with 77.9%, leading Astra and Opus by about 3 to 4 percentage points, which Google defined as a new state-of-the-art level for this test.

However, on FrontierSWE v2, Argon's 55.0% trailed Astra's 65.5% by 10.5 percentage points; on Terminal-bench 4.0, Argon's 57.4% also lagged Opus 5.5's 66.4% by nearly 9 percentage points.

It is worth noting that the aforementioned benchmarks were not conducted under completely uniform experimental conditions. For DeepSWE, for example, Google calculated Argon's score using its own mini-swe agent framework, while competitor data came from public leaderboards and each model developer's own reports. Methodological differences mean that direct numerical comparisons require careful interpretation.

One Million Token Output Cap Opens Space for Complex Tasks

A significant technical specification breakthrough for Gemini 4 Argon is the dramatic increase of the output token limit from the previous 64,000 to 1 million. Google directly links this design to deep reasoning capabilities, believing that a larger generation space enables the model to complete more complex task chains in a single run.

Google's disclosed internal deployment cases confirm this potential. In data center memory optimization, Argon agents analyzed cluster performance telemetry data and implemented optimization solutions, freeing over 300 TiB of memory, with potential savings estimated at 500 TiB to 1 PiB.

In the field of code migration, Argon has participated in engineering projects migrating C/C++ code to Rust, involving core libraries such as re2, as well as over 800,000 lines of code in the Fuchsia operating system's Zircon kernel.

In the migration of the video decoding library libgav1, Argon agents replaced 32,000 lines of SIMD code, ultimately achieving a decoder that runs 2.7 times faster than the previous Rust port.

Google also reported that Argon assisted in optimizing the spatiotemporal resources of a subroutine in the quantum computing field, achieving a 40% improvement.

Cybersecurity as the First-Release Core, with Strict Access Controls

Security defense is the primary focus of Argon's first batch of application scenarios.

Google stated that the model has been specially trained to autonomously discover, verify, and patch software vulnerabilities. Trusted participants in the Fairwind Program and Google's internal teams will receive access to the raw model without cybersecurity guardrails.

Google's subsidiary Wiz has already used Argon for its "Scan for Good" vulnerability scanning project. Google said the model discovered a critical vulnerability in medical software used by multiple hospitals worldwide, a risk that several previous frontier models failed to identify.

In internal security assessments, Argon scored 85.8% on source code vulnerability discovery tests, higher than Gemini 3.8 Flash Cyber's 71.0%; and scored 70.9% on Wiz's black-box penetration testing evaluation, higher than the latter's 58.2%.

On the public CWE-bench v1 vulnerability remediation test, Argon tied with GPT-6 Astra at 68.0%, with Opus 5.5 closely following at 67.0%.

Pricing Strategy: Competitive During Promotional Period, Set to Double Later

Google announced Argon's API pricing: during the promotional period, $2 per million input tokens, $10 per million output tokens, with a 95% discount on cached input, translating to approximately $0.10 per million tokens.

After the promotional period ends, prices will increase to $4 per million input tokens and $20 per million output tokens, both doubling, but Google did not disclose the specific end date of the promotional period.

Compared with major competitors, Argon's promotional pricing offers significant advantages:

GPT-6 Astra and Claude Fable 5.1's standard API prices are $10 per million input tokens and $50 per million output tokens respectively, both five times Argon's promotional rates.

Claude Opus 5.5 is priced at $4 per million input and $20 per million output, matching Argon's post-promotional pricing.

However, for users seeking value for money, the pricing landscape is not simply either-or.

OpenAI's GPT-6.1 Sol is priced the same as Argon's promotional rates, at $2 input and $10 output; Google's own Gemini 3.8 Flash is currently priced at $0.75 input and $3.75 output, which is lower.

Google's benchmark comparisons did not include GPT-6.1 Sol, so the actual quality difference between the two models at the same promotional price point remains to be tested by the market.

Skipping 3.5 Pro, Moving Directly into the Gemini 4 Era

Argon's release also means Google has abandoned its previously planned Gemini 3.5 Pro.

Google had announced plans at its I/O developer conference in May this year to launch Gemini 3.5 Pro in June, but the plan never materialized, and it was subsequently replaced by a series of Flash models.

According to analysis by AI evaluation firm Vals AI, this shift reflects that Gemini 3.5 Pro's performance was no longer sufficient to catch up with competitors, leading Google to decide to skip directly to releasing Gemini 4 to re-contest the AI frontier.

In July this year, Pichai acknowledged Google's lagging position on an investor call, but stated that "we remain at the frontier in many dimensions," and viewed Gemini 4 as key to catching up.

In November last year, Google briefly topped the AI leaderboard with a new model, once triggering a "code red" emergency response within OpenAI; whether Argon can hold onto the gains of this counteroffensive remains to be verified by the market after full public release.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10