AI

Google Unveils “Gemini 4 Argon,” Boosting Long-Horizon Reasoning, Coding, and Cybersecurity

TECHWORLD ·

Gemini 4 Argon (Gemini 4 Argon). [Photo: Google]

✦ AI Summary

Google unveiled its frontier model, "Gemini 4 Argon (Gemini 4 Argon)," and focused on long-horizon reasoning, software engineering, enterprise work, and cybersecurity.

It raised the output token limit from 64,000 to as many as 1 million, and before the full public release, it provided the model to trusted cybersecurity experts in the Fairwind Program.

Google applied Argon to internal development, data center operations, and code migration, and highlighted benchmark performance and security evaluation results while signaling a phased rollout and stronger safeguards.

Google unveiled its frontier model, "Gemini 4 Argon," on the 30th (local time) last month. The model is focused on agentic workflows that sustain complex tasks over long periods. The enhanced areas are long-horizon reasoning, software engineering, enterprise work, and cybersecurity.

To that end, Google expanded the output token limit from 64,000 to as many as 1 million. As a result, it is now possible to generate hundreds of thousands of tokens in a single run. Google is aiming to handle long-duration tasks such as large-scale code analysis, research, and document writing without interruption.

The model is being made available before its full public release to trusted cybersecurity experts in the Fairwind Program. Google plans to refine its safeguards based on their early usage results.

Google is expanding Argon across a range of internal development tasks, including research, infrastructure operations, and code migration. Google said it has applied Argon to its internal development environment and used it in quantum computing research to optimize bottlenecks in key applications. According to Google, spatial and temporal resource efficiency improved by more than 40% compared with the existing standard.

It is also being applied in the form of agents to data center operations. This agent analyzes telemetry data from server fleet profiling to derive and apply memory optimization measures. Google said it has secured more than 300 TiB of memory since deployment.

Google has also deployed Argon for large-scale codebase migration work. Targets include major libraries such as "re2" and "libgav1," as well as the Zircon kernel of Fuchsia OS, which spans more than 800,000 lines of code. The migration involves moving C and C++ code to Rust.

In the "libgav1" case, iterative performance analysis and code improvements were carried out based on an existing Rust port. As a result, it replaced 32,000 lines of SIMD code. According to the company, performance improved by 2.7 times compared with the existing Rust version.

Google highlighted coding and enterprise work performance based on benchmark results. In "DeepSWE v1.1," which evaluates long-horizon software engineering capability, it scored 77.9%. Google also said it delivered top-tier performance in "Vals Index," which evaluates finance, legal, tax, and coding tasks.

Looking at a broader evaluation range, it received high scores in "Vals Finance Agent v2," a multi-step financial research benchmark, and also recorded high marks in "Harvey's Legal Agent Benchmark," which evaluates legal research and document drafting. It posted 51.3% in "AutomationBench," an evaluation of work automation, and 91.7% in "LVBench," which measures understanding of long-form video.

Google also explained its strengthened cybersecurity capabilities. Argon has been trained to detect, verify, and patch software vulnerabilities, and Google said it tied for first place with 68% in "CWE-bench v1," which evaluates vulnerability-fixing ability. Google cited its top-tier performance in "Vals Index" and its shared first place in "CWE-bench v1."

Wiz used Argon in the public infrastructure vulnerability detection program "Scan for Good." In early proof-of-concept testing, it detected security vulnerabilities that an existing frontier model had missed. The company also disclosed a case in which it found a vulnerability in healthcare software that could expose personal information.

The model rollout will proceed in stages. Google is considering the potential for misuse in cyberattacks, as well as chemical, biological, radiological, and nuclear, or CBRN, domains. Accordingly, Google plans to apply its frontier safety framework and detect anomalous behavior based on internal activation-state monitoring.

Google is also strengthening its response to indirect prompt injection attacks. To that end, it will conduct automated red-team testing and adversarial training. It will also apply model inference-process and behavior monitoring, and it will add a function that stops execution if a task begins moving in a direction different from the user's intent. Alignment-disagreement mitigation features will also be applied.

Google said it has put in place an isolated sandbox security system for high-risk model training and evaluation processes. Google added that it plans to share its related approach to agent security operations with external partners in the future.

Google also disclosed pricing for Gemini 4 Argon. As a launch promotional price, Gemini 4 Argon costs USD 2 per 1 million input tokens and USD 10 per 1 million output tokens. The price for cached input tokens has been cut by 95% from the standard input price.

Google announced that it will initially prioritize cybersecurity experts and testers among users of the model, and then gradually expand access to paid API users, Google AI Ultra subscribers, developers and enterprises, and general users.

Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407642

References

This article was produced with the help of an automated content generation algorithm.


Source: TECHWORLD

View original

This article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.