AI

Google Unveils 'Gemini 4 Argon,' Will First Roll It Out to Cybersecurity Experts

IT DAILY ·

Gemini 4 Argo. [Photo: Google Korea]

✦ AI Summary

Google unveiled its next-generation AI model, 'Gemini 4 Argon.'

The model increases maximum output from 64,000 tokens to 1 million tokens and is focused on enterprise knowledge work such as software engineering, legal and financial work, as well as cybersecurity.

General release has been put on hold, and it will first be offered to cybersecurity experts and a limited group of users before access is expanded.

Google unveiled its next-generation AI model, 'Gemini 4 Argon.' According to Google's announcement as of 1 day after the announcement, the model is designed to carry out multiple stages of complex tasks over long periods. To that end, the maximum output was expanded from 64,000 tokens to 1 million tokens, more than 15 times the previous limit. It also supports long-duration reasoning.

Google said 'Gemini 4 Argon' is focused on software engineering, enterprise knowledge work such as legal and financial services, and cybersecurity. As examples of use, it cited handling large-scale code tasks and solving complex, multi-step problems along a single execution path. Google said Argon can be used to process large-scale code tasks and complex multi-step problems along a single execution path.

General release to users has been deferred for now. Instead, it will first be provided to cybersecurity experts. Google plans to expand access after verifying it in real-world environments.

Meanwhile, the source of the photo of 'Gemini 4 Argon' is Google Korea.

Google is already applying Argon to internal operations and development. One internal use case cited by Google is analysis of profiling data for data center server groups used by multiple Argon agents.

Based on this analysis, Google said it derived and applied memory optimization measures. As a result, it said it secured more than 300 TiB of memory.

Google also deployed Argon in its effort to convert C and C++ codebases to Rust. The target applications included a library of tens of thousands of lines and the Zircon kernel of the Fuchsia OS, which exceeds 800,000 lines. The photo of the Gemini 4 Argon benchmark was provided by Google Korea.

Google's main focus area for Argon is cybersecurity. Google said it trained Argon to autonomously detect, verify, and patch software vulnerabilities. Argon ranked tied for first place with 68% in 'CWE-bench v1,' an evaluation of vulnerability-fixing capabilities.

Through the 'Fairwind Program,' Google will first provide Argon to trusted cybersecurity experts. Google plans to provide a version without cyber guardrails to external cybersecurity defense experts and Google's internal security team. The goal is to unconstrain the model's cyber capabilities in defense work such as vulnerability discovery and verification.

Google is further verifying the safety measures before broad release. The model is designed to refuse harmful requests that could be used for cyberattacks and chemical, biological, radiological, and nuclear (CBRN) attacks. Internal and external red teams are also testing whether the safeguards can be bypassed, and Google is strengthening defenses against indirect prompt injection attacks. Indirect prompt injection is a method that attempts to change AI behavior through commands hidden in external documents and elsewhere.

The model also monitors attempts to carry out tasks that deviate from user intent and permitted scope, and applies execution stoppage when necessary. While carrying out these steps, Google is also participating in the U.S. government's voluntary pre-release model access process.

Argon is an evaluation target released by Google DeepMind. Argon recorded 77.9% in 'DeepSWE v1.1,' an evaluation of long-term software engineering tasks, 51.3% in 'AutomationBench,' an enterprise task automation evaluation, and 91.7% in 'LVBench,' an evaluation of long-duration video understanding.

In 'LVBench,' it scored the highest among the compared models. It also showed some advantages in evaluations of specialized work such as finance and legal services. In 'Vals Index,' an evaluation of finance, coding, legal, and tax tasks, it scored 68.9%, outperforming 67.0% for Claude Opus 5.5 and 63.1% for GPT-6 Astra in 'Vals Index.'

In a separate evaluation, it scored 65.4% on a separate finance research benchmark and 19.6% on a separate legal benchmark. In this way, Argon posted higher scores than comparison models in some specialized task evaluations.

However, it lagged behind competing models in some other coding-related benchmarks. In 'FrontierSWE v2,' an evaluation of long-term software engineering, it scored 55.0%, below GPT-6 Astra's 65.5%. In 'Terminal-bench 4.0,' an agent coding evaluation, it scored 57.4%, below Claude Opus 5.5's 66.4%, and it trailed competing models in some coding benchmarks.

Google plans to first offer it to a limited user base and improve the safety measures by incorporating feedback from early users. Access rights will also be expanded gradually, beginning with paid API users and subscribers to Google AI Ultra.

The rollout will then expand to developers, enterprises, and general users, in that order. Google also announced launch pricing of USD 2 per 1 million input tokens and USD 10 per 1 million output tokens. A 95% discount will apply to cached input tokens compared with the input token price.

Source: IT DAILY · Kim Byeong-ju
Original: https://www.itdaily.kr/news/articleView.html?idxno=241960

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.