Cloudera Partners With Nvidia to Accelerate Apache Spark and Cut Cloud Costs
IT DAILY ·
✦ AI Summary
Cloudera announced performance-enhancing features for Spark processing in response to enterprises' AI data-preparation speed challenges.
Cloudera Data Engineering will add native GPU acceleration for Apache Spark 4.1 based on Nvidia cuDF, and Cloudera Anywhere Cloud is also scheduled to be supported.
The feature accelerates Spark workloads without code changes, reducing processing time and cloud infrastructure costs while maintaining security and governance.
Cloudera is pushing to advance data engineering, with the goal of helping enterprises speed up AI data preparation. On the 20th, Cloudera announced performance-enhancing features for Spark processing in response to the data-preparation speed challenges created by the expanding adoption of AI in enterprises.
According to the announcement, Cloudera Data Engineering will add native GPU acceleration for Apache Spark 4.1 based on Nvidia CUDA-X library cuDF. An Nvidia cuDF plugin for Apache Spark is also scheduled to be supported in Cloudera Anywhere Cloud.
Cloudera said that for large-scale Spark workloads, it is common for completion to take several hours. It explained that when such workloads take a long time, the rollout of analytics and AI applications is delayed and cloud computing costs rise.
Cloudera said that directly integrating GPU acceleration can shorten processing time without changing existing Spark applications or operational workflows. As a result, it expects to accelerate Spark workloads without modifying PySpark or SQL code, reduce cloud infrastructure costs across hybrid environments, and improve the speed of AI-ready data preparation.
Apache Spark is used in data pipelines at many companies. Cloudera said it has built native GPU acceleration into Cloudera Data Engineering so enterprises can improve performance while maintaining the security and governance of operational workloads.
Cloudera plans to use Nvidia cuDF for Spark workloads, and it said Nvidia GPUs will enable up to 4 times faster workloads than existing CPU infrastructure.
The feature provides zero-code GPU acceleration for Apache Spark 4.1 workloads, improving ETL and data preparation speeds for analytics and AI, shortening compute execution time, and reducing cloud infrastructure costs. It also supports built-in deployment that does not require manual driver configuration, and provides enterprise-grade security and governance based on Cloudera's Unified Data Fabric. Consistent performance is also supported across public cloud, private cloud, sovereign cloud, and on-premises environments.
Cloudera said that for many enterprises, the AI bottleneck is not the model itself, but whether raw data can be quickly turned into trustworthy and usable insights. Leo Brunnick, Cloudera's CPO, said that Spark acceleration inside Cloudera Data Engineering addresses that bottleneck and speeds the move from the data-preparation stage to the analytics and AI stage, while keeping governance, security, and operational consistency at the center of the strategy.
Pat Lee, Nvidia's vice president of strategic enterprise partnerships, said that Nvidia AI infrastructure and the CUDA-X library have been natively integrated into Cloudera Data Engineering. He said this will allow enterprises to cut costs and improve the speed of Apache Spark pipelines without modifying PySpark or SQL code, turning business data into AI-driven data.
Source: IT DAILY · Yang Seung-gap
Original: https://www.itdaily.kr/news/articleView.html?idxno=241120
References
This article was produced with the help of an automated content generation algorithm.
Source: IT DAILY
View originalThis article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.