Software

E-papyrus Says PyMuPDF Surpasses 1 Billion Global Cumulative Downloads

IT DAILY ·

[Photo: Epapyrus]

✦ AI Summary

E-papyrus said PyMuPDF surpassed 1 billion global cumulative downloads to mark the 10th anniversary of its launch.

PyMuPDF is a PDF processing library for the Python environment, offering high-speed rendering and text and image extraction, conversion, and editing functions.

E-papyrus said the recent increase in downloads was driven by demand for LLM- and AI-based document data preprocessing, as well as growing demand for building generative AI and RAG pipelines.

E-papyrus announced on the 2nd that global cumulative downloads of PyMuPDF have surpassed 1 billion, marking the 10th anniversary of the product's launch.

PyMuPDF is a PDF processing library for the Python environment. It supports high-speed rendering of PDF documents and provides text extraction, conversion, and editing functions, as well as image extraction, conversion, and editing functions. E-papyrus said the library, based on fast processing speeds and a lightweight engine, is being used by developers and companies around the world as a tool for processing document data.

Downloads have risen recently. E-papyrus attributed this to growing demand for document data preprocessing based on large language models (LLM) and AI. Of the 1 billion cumulative downloads, 83% came in the most recent 12 months.

The company said the spread of generative AI and LLMs, along with increased demand for document data parsing and preprocessing in the process of building retrieval-augmented generation (RAG) pipelines, led to a surge in downloads in the global market.

E-papyrus and the development team unveiled PyMuPDF4LLM and released graph neural network (GNN)-based layout analysis technology as part of efforts to respond to market changes in the AI era. The company said the technology delivers up to 10 times faster performance than vision AI models, making infrastructure cost savings possible. It also said the technology can accurately identify the internal structure of PDFs and the meaning of tables, and can convert content into formats used by LLMs, such as Markdown and JSON.

Along with that, the company also unveiled PyMuPDF Pro, an extensible document preprocessing library. E-papyrus said PyMuPDF Pro is a library that expands the scope of document preprocessing beyond existing PDFs, enabling not only PDF parsing and preprocessing but also Hangul Word Processor (HWP and HWPX) parsing and preprocessing and Microsoft Office (Word, Excel, and PowerPoint) parsing and preprocessing. E-papyrus introduced PyMuPDF Pro as an engine for processing various document formats used by companies and public institutions.

The company said PyMuPDF and PyMuPDF Pro are being used by Goldman Sachs, Oracle, Snowflake, Palantir, and Mistral AI in global finance, cloud, and AI sectors. In Korea, it said the products are being supplied to LG AI Research, Lotte Innovate, HD Hyundai Samho, Korea National Railway, the Korea Institute of Maritime and Fisheries Technology, and the Korea Institute of Machinery and Materials for use by companies, public institutions, and research institutes.

Kim Jeong-a E-papyrus vice president said the reason PyMuPDF reached 1 billion downloads is that PyMuPDF has been recognized as a standard for data preprocessing in the global AI ecosystem and has become indispensable. Kim added that the company will actively support successful AX and LLM implementations for domestic and overseas customers by leveraging PyMuPDF Pro, a leading product that supports parsing of various document formats including PDFs, Hangul, and MS Office.

Source: IT DAILY · Yang Seung-gap
Original: https://www.itdaily.kr/news/articleView.html?idxno=241385

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.