Hardware

Apple Expands Enterprise AI to On-Device Computing, Targets 'Token Costs'

IT DAILY ·

Apple Mac Studio equipped with the M5 Ultra. Apple is expanding on-device processing for enterprise AI. [Photo: Apple]

✦ AI Summary

Apple emphasized expanding on-device computing in the enterprise AI market for reasons of performance, security, and cost.

Using examples from the MacBook Pro and Mac Studio, it pointed to cloud API costs and the possibility of offsetting initial device purchase costs.

Apple described a model in which cloud and device processing are separated, with sensitive data processing and repetitive inference handled on devices.

Apple is bringing on-device computing to the forefront of the enterprise AI market. It says that as generative AI becomes more agentic, cloud calls are rising, and with them the burden of higher cloud costs and greater data security concerns. Apple is expanding on-device processing for enterprise AI.

On the 3rd, Apple identified performance, security, and cost as the key elements of enterprise on-device AI. It also pointed to cloud API costs that accumulate in proportion to AI usage as a major reason for expanding on-device computing. It is therefore promoting a strategy to ease cost and security burdens by processing AI on devices such as Mac.

As supporting evidence, Apple cited a MacBook Pro use case involving a 26 billion-parameter model. For its cost calculation, Apple assumed a workload capable of handling tasks such as financial simulations, with the assumed usage conditions set at 4 hours a day and 150 million tokens processed per month.

Apple estimated the cloud API processing cost for the same task at USD 450–850 per user per month. It also presented a photo of the Apple Mac Studio equipped with M5 Ultra.

Based on estimates for the Mac Studio, Apple said that if a certain scale of AI workloads is continuously run, cloud usage fees can grow large enough to recoup the initial device purchase cost in a relatively short period. According to Apple’s calculations, for an agentic coding workload using an 80 billion-parameter model, 12 hours of use per day, and 300 million tokens processed per month, continued AI use above a certain level could offset the initial Mac purchase cost within months.

At that point, the estimated cloud cost was presented at USD 900–1,500 per user per month. However, that figure is not the result of directly applying the pricing list of a specific cloud provider, nor is it the result of directly applying the pricing list of a specific AI model.

Apple said the range was calculated based on cloud API costs for commonly used frontier models and actual software usage patterns. However, the specific model names were not disclosed, and the calculation formula by input and output token was also not disclosed. Additional materials on enterprise AI cost calculations are expected to be released within months.

Under these assumptions, Apple focused on running enterprise AI separately in cloud and on-device environments. Large-scale training tasks and workloads requiring elastic computing resources were described as cloud use cases, while sensitive data processing, repetitive inference, and field work in environments with limited network connectivity were described as on-device use cases.

Data sovereignty and network-constrained environments were cited together as reasons for expanding on-device AI. Finance and healthcare were said to face strong data-processing regulations, making it necessary to process information without sending it to external infrastructure.

Transportation and manufacturing were highlighted for their potential to use local computing because network connectivity is not always guaranteed. Omdia conducted a survey commissioned by Apple of more than 1,500 corporate technology leaders and developers.

In the survey, the No. 1 cost item in AI infrastructure TCO was cloud computing and storage at 39%, followed by security and compliance management at 36%, and software licenses and tools at 32%. Also, 57% of the AI models used by companies were under 10 billion parameters, and Omdia said that memory capacity and architecture, rather than model size alone, are the factors that determine whether a model will run locally.

On-device AI is expanding into real-world settings such as aviation and healthcare. Cathay Pacific applied a 7 billion-parameter model to pilots' iPads, and the model is used to search flight manuals. Nabla implemented on-device voice processing for clinicians' consultations.

In these cases, processing is done on the device without a network connection, or by reducing transmission to external servers. Cathay Pacific's processing method involves handling information on the device without an internet connection. Nabla processes sensitive medical data without sending it to external servers.

Apple cited the unified memory architecture of Apple silicon as a technology supporting the expansion of on-device AI. The unified memory architecture gives the CPU, GPU, and Neural Engine access to a single memory pool. Apple said this reduces data-movement overhead and secures memory capacity for AI model use.

M5 Ultra was also recently unveiled. It is characterized by support for up to 512GB of unified memory and 1.2TB of memory bandwidth per second. Apple said that, based on M5 Ultra, it is possible to run a large language model (LLM) on a Mac directly at the scale of hundreds of billions of parameters.

Apple is providing a software-side framework for enterprise use of AI models and deployment, and both its own foundation models and external models can be used. It offers Core AI to support developers in deploying their own models on Apple devices, and it also supports use of Apple silicon computing resources. MLX provides model training support and fine-tuning support.

Apple is also expanding device management features for enterprise environments. Its iPad 'authenticated Guest Mode' can be used in workplaces where multiple employees take turns using the device, and in this setup an employee logs in with a company account, performs their work, and the user’s data is deleted from the device when they log out. Enterprise features include customizing the work environment on the login screen, as well as authentication-based customization of the work environment.

Against this backdrop, Apple's policy is to run cloud and on-device systems in parallel. Its approach is to expand device-distribution strategies for enterprise AI workloads depending on cost, security, and network conditions.

Source: IT DAILY · Kim Byung-joo
Original: https://www.itdaily.kr/news/articleView.html?idxno=241375

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.