Frontier AI New Models Debut in Succession as Performance, Cost, and Agent Competition Heat Up
TECHWORLD ·
✦ AI Summary
The center of competition for new models among global AI companies is shifting from high benchmark scores to cost and real-world task execution.
Meta launched 'Muse,' which can perform real tasks such as sending emails, booking travel, and buying products, but it also ran into conflict with Amazon.
XAI, Anthropic, and OpenAI each unveiled 'Grok 4.7,' 'Claude Opus 5.5,' and 'GPT-6 Sol' and 'GPT-6 Luna,' emphasizing long-duration coding and specialized work as well as cost efficiency.
The center of competition among global AI companies’ new models is shifting from performance to cost and real-world task execution. As the frontier AI market heats up again, the direction of competition is also moving away from simply scoring high on benchmarks and toward spending longer periods coding and handling specialized work while reducing total task costs.
Major AI companies rolled out new models on three consecutive days. XAI unveiled 'Grok 4.7' on the 21st local time, Anthropic released 'Claude Opus 5.5' on the 22nd, and OpenAI expanded the GPT-6 lineup by adding 'GPT-6 Sol' and 'GPT-6 Luna' on the same day.
A common thread in the models released recently is that they are focused more on multi-step real-world task execution than on one-off question-and-answer interactions. Agents are expanding into code modification, external tool calls, and computer control, and their task duration is also getting longer. As a result, the factors determining cost efficiency are broadening beyond token prices to include task steps, cache utilization, and processing speed.
Meta launched its personal AI agent, 'Muse,' on the 8th. 'Muse' is designed not merely for simple question-and-answer exchanges, but to perform real tasks on users’ behalf in external services. Examples of tasks it can handle include sending emails, booking travel, and purchasing products.
To do this, Meta built 'Muse' on 'Muse Spark.' 'Muse' also runs in a dedicated virtual environment called 'Muse Secure VM.' Users can configure which applications to connect, as well as the scope of access.
For sensitive actions such as sending emails and making purchases, Meta required an approval process before execution. However, conflicts arose with platform operators during the process of integrating with real services.
Amazon blocked 'Muse' from product search and purchasing functions, citing access to its shopping mall without prior approval. Amazon raised concerns that 'Muse' may have had insufficient identification as an AI agent in the process. It also cited the possibility that account data could be processed while accessing the site as another concern.
Agent competitiveness is not determined by the model’s reasoning performance alone. In conditions involving real task execution, there is the issue of whether external services allow agent access, and the way user authority is delegated also needs to be resolved.
In this context, XAI unveiled 'Grok 4.7' on the 21st. 'Grok 4.7' is aimed at coding and specialized knowledge work. It uses a larger base model than 'Grok 4.6.'
In terms of training, it is based on reinforcement learning that increased the share of difficult tasks taking several hours. Through this, 'Grok 4.7' strengthened its ability to maintain context during long tasks and also improved its ability to recheck its own results. XAI also trained the model to directly understand the execution environment of 'Grok Bot' for conversation and general knowledge tasks.
Performance was presented through CursorBench 4.0, a long-duration coding benchmark, and AA Briefcase, a long-duration office work evaluation. 'Grok 4.7' scored 46.3% on CursorBench 4.0, while 'Grok 4.6' scored 40.4%. As a result, 'Grok 4.7' improved its performance on CursorBench 4.0 compared with 'Grok 4.6.'
'Grok 4.7' also showed higher performance than the previous model on AA Briefcase. However, detailed evaluation results varied in terms of which competing model had the stronger points.
XAI emphasized cost performance for long-duration work rather than the highest score on individual benchmarks. XAI said it places importance on cost performance for long-duration work.
Prices for the Grok lineup were set starting at USD 2 per 1 million input tokens and USD 6 per 1 million output tokens. This price level is similar to that of Grok 4.6. The fast model has twice the output speed, and its price is twice that of the base model.
Grok 4.7 can be used in Cursor, Grok Build, the Grok API, external coding tools, model routers, and cloud platforms.
Anthropic released 'Claude Opus 5.5' on the 22nd. 'Claude Opus 5.5' is the first model in the Claude 5.5 lineup. The model is intended to improve complex coding, computer use, and specialized knowledge work that require multi-step reasoning and tool use.
Anthropic said 'Claude Opus 5.5' improved both performance and efficiency compared with the existing Opus 5. External evaluations also confirmed the gains. In the Artificial Analysis Intelligence Index (AAII), Opus 5.5 at maximum reasoning settings scored 58 points. It also showed higher performance than the existing Opus 5 in specialized work evaluations and agentic coding evaluations.
Anthropic stressed that the gap in benchmark scores between models has limits in representing the actual user experience. As an important factor in real work, Anthropic pointed to completing tasks with fewer tokens and fewer steps rather than higher scores.
In line with this standard, the API price for Opus 5.5 was set at USD 4 per 1 million input tokens and USD 20 per 1 million output tokens. That represents a 20% reduction each for input and output compared with the existing Opus 5. The cache read price was cut 60% to USD 0.20 per 1 million tokens. Cache read pricing is widely used in long-duration agentic and coding work.
Anthropic said that when token usage reductions are reflected in typical workloads under the default setting, total task costs are about 40% lower than with the existing Opus 5. Output speed also improved by more than 30%.
Anthropic also strengthened safeguards for long-duration work. Compared with the previous model, Opus 5.5 reduced hard-to-reverse actions and actions that drift outside the given scope. Prompt injection defense performance also improved.
Anthropic plans to release Sonnet 5.5 and Haiku 5.5 within weeks. Anthropic’s plan is to extend the performance and cost-efficiency improvements seen in Opus across the entire Claude lineup.
On the same day, OpenAI released GPT-6 Sol and GPT-6 Luna. The release of GPT-6 Astra came earlier this month.
OpenAI segmented the GPT-6 lineup so users can choose models based on task difficulty, speed, and budget. GPT-6 Sol and Luna were trained in a manner similar to Astra.
The development focus for GPT-6 Sol and Luna was to extend Astra’s specialized work, factual accuracy, coding, computer use, and alignment performance into faster, lower-cost models. In its positioning for enterprise work, cost performance was emphasized.
In enterprise work evaluations, a comparison using AutomationBench figures was presented. AutomationBench is characterized by tasks that use 47 tools across sales, marketing, operations, customer support, finance, and human resources. GPT-6 Sol scored 33.2% at the AutomationBench xhigh reasoning level, while GPT-6 Astra low, the comparison target, scored 30.3%. The comparison showed that GPT-6 Sol posted a higher figure than GPT-6 Astra low.
GPT-6 Sol delivered 56.4% at max reasoning level in 'Agents' Last Exam' (ALE), an example of a long-duration specialized work performance evaluation. ALE evaluates long-duration performance in computer-based specialized work across 55 detailed industries.
In coding and computer use, cost efficiency improved compared with the previous generation. In DeepSWE v1.1, which measures long-duration software development ability based on real codebases, and OSWorld 2.0, which evaluates computer manipulation ability, OpenAI said GPT-6 Sol achieved performance similar to competing frontier models while reducing task cost.
Factual accuracy also improved. The evaluation was based on de-identified ChatGPT conversations in which users pointed out factual errors in previous models, and GPT-6 Sol’s error level was presented as about half that of GPT-5.6 Sol.
OpenAI said the result came from selecting only conversations with a high likelihood of errors. It added that the figure does not represent the factual error rate in typical ChatGPT usage, drawing a line under the point that it should not be taken as the error rate for all general ChatGPT use.
OpenAI then lowered the token prices for GPT-6 Sol and Luna. GPT-6 Sol is priced at USD 2 per 1 million input tokens and USD 10 per 1 million output tokens, while GPT-6 Luna is priced at USD 0.10 for input and USD 0.50 for output. The prices for GPT-6 Sol and Luna were each cut by 50% compared with the promotional prices for GPT-5.6 Sol and Luna.
OpenAI also improved prompt caching as a factor related to the cost of long-duration agent work. In GPT-6, the base cache hit rate for repeated context increased, and a 90% discount applies to cached input tokens. In addition, previous context can be reused even when the reasoning level or tools used differ.
GitHub said it has applied related caching improvements over the past several months. The changes affected billions of OpenAI model requests, and as a result, the share of prompt tokens that needed to be newly processed fell by more than 50%.
OpenAI’s GPT-6 Sol and Luna are available to Plus, Pro, Business, Enterprise, and Edu users. They are available in ChatGPT Work and Codex. Free and Go users can use GPT-6 Luna in the desktop app.
In the latest model competition, the importance of the cost and time required to complete real tasks is rising alongside the highest benchmark scores for individual models. The thrust of the remarks is that the standard for frontier competition is shifting from a focus on top scores to a focus on task completion costs.
This change is tied to the characteristics of agent environments. Agent environments involve multiple tool calls and repeated reading of long context. Because of this, it is difficult to judge total costs using only input and output token prices.
Even when the same result is produced, actual operating costs can vary depending on the number of task steps, token usage, cache utilization, and processing time. Accordingly, model evaluation methods are also changing.
Model evaluation is expanding its emphasis on the assessment of long-duration real work execution rather than the ability to produce correct answers. Examples of long-duration task evaluations included reading and modifying real codebases and moving among several applications to complete work end to end.
AI companies are moving toward expanding product lineups that separate performance, speed, and cost by task type rather than centering on a single top-performing model. OpenAI has Astra, Sol, and Luna as the segmented models in GPT-6. Anthropic has signaled Sonnet 5.5 and Haiku 5.5 after Opus.
Meta’s Muse is presented as a real service expansion case of internal model changes. The model has begun accessing external sites, has begun carrying out purchases, and has begun making reservations. This shows that frontier AI competition is shifting away from a model centered on making one top-performing model.
Accordingly, the standard for competition is moving toward real task execution ability, lower cost, long-duration performance, and whether tasks are completed to the end. Going forward, the factors used to judge agent competitiveness are presented as reasoning performance, the degree of stable integration with other services, and the way required permissions are managed.
Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407176
References
This article was produced with the help of an automated content generation algorithm.
Source: TECHWORLD
View originalThis article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.