AI

Google Unveils Gemini 3.8 Flash, Boosting Coding and Agent Performance

TECHWORLD ·

Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. [Photo: Google]

✦ AI Summary

Google unveiled 'Gemini 3.8 Flash' and 'Gemini 3.8 Flash Cyber' on the 2nd (local time).

'Gemini 3.8 Flash' has improved coding, reasoning, and autonomous agent task performance, while 'Gemini 3.8 Flash Cyber' is optimized for cybersecurity use cases.

Google will not immediately release the cyber model to general users and will instead offer it only in a limited form through the 'Fairwind Program.'

Google unveiled 'Gemini 3.8 Flash' and 'Gemini 3.8 Flash Cyber' on the 2nd (local time). The new releases come three weeks after Google announced 'Gemini 3.7 Flash'.

Both models use the same base model. Of the two, 'Gemini 3.8 Flash' is optimized for general coding and reasoning environments, with improved coding performance, multi-step reasoning performance, and autonomous agent task performance. 'Gemini 3.8 Flash Cyber' is a cybersecurity-focused model optimized for cybersecurity use cases.

Google's flagship 'Gemini 3.8 Flash' delivers improved software engineering performance, agent task performance, and domain-specific multi-step reasoning performance compared with 3.7 Flash. The model is designed to handle complex tasks and addresses long-running agent work through additional reasoning execution methods and repeated tool-calling.

Gemini 3.8 Flash outperformed many large frontier models on 'DeepSWE v1.1 (DeepSWE v1.1),' an evaluation that measures real-world long-term software engineering performance. On 'HLE-Verified,' a domain reasoning benchmark, Gemini 3.8 Flash scored 54.9%. Its performance also improved over 3.7 Flash on financial and legal agent benchmarks.

The newly announced model is Gemini 3.8 Flash Cyber. The model focuses on vulnerability detection and automated patching, and it outperformed previous cyber models and large frontier models on 'CyberGym,' a benchmark for security vulnerability detection. In internal Google evaluations, it achieved a vulnerability detection success rate above 70% across 20 programming languages, while patching performance also improved.

Patch performance was also shown on the external benchmark 'CWE-Bench (CWE-Bench).' Gemini 3.8 Flash Cyber posted a pass@1 of 47.2% on 'CWE-Bench (CWE-Bench).' Pass@1 refers to the rate at which a correct patch is generated in a single attempt, and the top model score on the benchmark was 47.8%, putting Gemini 3.8 Flash Cyber's 47.2% close to the leader.

Google first applied the cyber model to internal cyber use cases, including code security. In tests by the Chrome security team, it generated correct patch code 2.6 times more often than existing large commercial models. Google's cloud vulnerability research team used the model to find major infrastructure vulnerabilities within 2 hours.

Google will not immediately release the cyber model to general users. Instead, it will be made available in a limited form through the 'Fairwind Program,' with initial access granted to trusted government agencies, major infrastructure operators, and software maintainers. Google plans to apply cybersecurity-related safety restrictions more flexibly than with general models, while limiting access to verified security experts.

Meanwhile, Gemini 3.8 Flash will apply the same launch pricing as 3.7 Flash through the end of the year. Input pricing is USD 0.75 per 1 million input tokens, and output pricing is USD 3.75 per 1 million output tokens. Availability includes Google AI Studio, Android Studio, and Gemini Enterprise, while Google AI Pro and Ultra subscribers can also use it in the Gemini app, Google Search AI Mode, and Google Sheets.

Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=406459

References

This article was produced with the help of an automated content generation algorithm.


Source: TECHWORLD

View original

This article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.