Google Unveils Gemini 3.8 Flash, Lifting Performance Again After 3 Weeks
IT DAILY ·
✦ AI Summary
Google on the 2nd unveiled its AI model, "Gemini 3.8 Flash," and its cybersecurity-focused model, "Gemini 3.8 Flash Cyber." "Gemini 3.8 Flash" has the same token-based pricing as 3.7 Flash, but is designed to use more tokens for complex tasks, boosting reasoning, coding and agent performance. "Gemini 3.8 Flash Cyber" is tuned for vulnerability detection and patch generation, and will be offered first through the "Fairwind Program" rather than to general users.
Google on the 2nd unveiled its AI model, "Gemini 3.8 Flash." On the same day, it also introduced "Gemini 3.8 Flash Cyber," a model specialized for cybersecurity. The announcement focused on expanding performance gains while maintaining the Flash line's emphasis on cost efficiency.
"Gemini 3.8 Flash" keeps the same price per token as the previous model, while being designed to use more tokens for complex tasks. Based on that, its reasoning, coding and agent performance have been improved. The move reflects an effort to raise performance under the same pricing conditions while preserving the Flash line's long-standing focus on cost efficiency.
The performance comparison is based on its predecessor, "3.7 Flash." Compared with "3.7 Flash," "Gemini 3.8 Flash" delivered better software engineering performance and improved agent task performance. Its multi-step reasoning performance also improved over "3.7 Flash." In terms of the article's headline, the interval between Google's performance upgrade announcements is 3 weeks.
3.8 Flash is designed to perform additional reasoning and repeatedly call tools on complex tasks. Its high reasoning level allows for expanded token usage to improve performance. Developers can lower the reasoning level by task, and they can also adjust token usage using the existing 3.7 Flash.
Launch pricing is USD 0.75 per 1 million input tokens and USD 3.75 per output token. The pricing for 3.8 Flash is the same as 3.7 Flash, but the same token rates do not mean the actual cost of a task is identical. Artificial Analysis estimated the per-task cost of high-reasoning-level 3.8 Flash at USD 0.58, compared with USD 0.40 for 3.7 Flash. That put high-reasoning-level 3.8 Flash at about 40% more expensive than 3.7 Flash. The increase stemmed from about 30% higher average output token usage and more task steps in agent evaluation processes.
According to Artificial Analysis, high-reasoning-level 3.8 Flash has a higher per-task cost than 3.7 Flash, but also a higher performance score. Performance improved, and high-reasoning-level 3.8 Flash scored 59 on the Artificial Analysis Intelligence Index. That is 3 points ahead of 3.7 Flash. The 3.6 Flash released in July scored 50 in the same evaluation at the time, highlighting the rapid performance gains across the Flash line.
In its own evaluation criteria, Google said "Gemini 3.8 Flash" outperformed 3.7 Flash across multiple benchmarks. For long-duration software engineering performance, it surpassed 3.7 Flash on "DeepSWE v1.1," and it also exceeded 3.7 Flash on agent benchmarks in finance and legal domains.
Google also said that "Gemini 3.8 Flash" scored 54.9% on HLE-Verified. HLE-Verified evaluates specialized knowledge and reasoning ability.
Google then unveiled "Gemini 3.8 Flash Cyber," a cybersecurity-focused version of the same base model. The model is aimed at defensive work such as vulnerability detection and patch generation.
However, "Gemini 3.8 Flash Cyber" was not made immediately available to general users. Instead, Google set up a new early-access channel called the "Fairwind Program."
Priority access under the "Fairwind Program" is for trusted government agencies, critical infrastructure operators and software maintainers. This access model combines relaxed safety restrictions in cybersecurity use cases with limits on who can use the model, compared with general-purpose models.
Google said its 3.8 Flash Cyber recorded a pass@1 score of 47.2% in the external patch-generation benchmark "CWE-Bench." That figure is compared with a top model score of 47.8% in the evaluation. The performance figure also reflects Google internal evaluations and deployment cases, and the model showed a vulnerability detection success rate above 70% across 20 programming-language codebases used in Google's own assessments.
Google also said it applied the model to internal security work. The Chrome security team conducted vulnerability patch tests, and the results showed 2.6 times as much accurate patch-code generation as a conventional large commercial model. Google Cloud's vulnerability research team also shared use cases, and according to its announcement, the model helped uncover a critical vulnerability in core infrastructure technology within 2 hours.
At the same time, Google also noted that despite the Flash performance gains, a gap with top-tier models remains. In the external benchmark "CWE-Bench," 3.8 Flash Cyber recorded a pass@1 score of 47.2%, falling short of the top model's 47.8%.
Meanwhile, Google continued to sharply shorten the Flash line's release cycle. After launching 3.6 Flash in July, it rolled out 3.7 Flash in August and 3.8 Flash in September in quick succession. In particular, 3.8 Flash was presented as a follow-up model released 3 weeks after the unveiling of 3.7 Flash.
Gemini 3.8 Flash achieved high scores in independent evaluations compared with previous Flash models, and its performance improved. However, at higher reasoning levels, token usage rises and per-task costs also increase. To that end, Google offers a lower reasoning-level option for cost-sensitive tasks and allows continued use of 3.7 Flash for cost-sensitive workloads.
At the same time, while the Flash line is advancing rapidly, the absence of a top-tier "Pro" line continues. This announcement did not include a new Pro model, and the timing for a Pro line release to counter top-tier model competition also remains unconfirmed. In July, Google said the next Pro model was under partner testing, and it also said it would release the next Pro model once preparations were complete.
Gemini 3.8 Flash is available in Google AI Studio, Android Studio, the Gemini API and Gemini Enterprise. Google AI Pro and Ultra subscribers can also use it in the Gemini app, Google Search AI Mode and Google Sheets. Launch pricing will apply through December 31, and starting in January 2027, the input price will be USD 1.50 per 1 million tokens and the output price will be USD 7.50 per 1 million tokens.
Source: IT DAILY · Kim Byung-joo
Original: https://www.itdaily.kr/news/articleView.html?idxno=241377
References
This article was produced with the help of an automated content generation algorithm.
Source: IT DAILY
View originalThis article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.