Software

Epapyrus Proves Its Competitiveness in AI Document Parsing, Outpacing IBM and Other Global Technologies

IT DAILY ·

[Photo: Epapyrus]

✦ AI Summary

Epapyrus said on the 2nd that PyMuPDF4LLM demonstrated the competitiveness of its document parsing technology.

In the ParseBench evaluation hosted by LlamaIndex, PyMuPDF4LLM recorded better results than IBM's open source technology Docling.

Epapyrus said PyMuPDF4LLM showed superiority across overall metrics for RAG build and AI training data preprocessing.

Epapyrus, a document AI specialist, said on the 2nd that its own "PyMuPDF4LLM" has demonstrated the competitiveness of its document parsing technology. Epapyrus said that "PyMuPDF4LLM" has proven its performance in converting documents into data that can be used for generative AI. "PyMuPDF4LLM" is a version of the open source library "PyMuPDF" specialized for LLM preprocessing. "PyMuPDF" has surpassed 1.1 billion cumulative global downloads.

Epapyrus cited the results of the benchmark "ParseBench," hosted by LlamaIndex, as evidence. In the evaluation, "PyMuPDF4LLM" recorded better results than IBM's open source technology "Docling."

Epapyrus emphasized that "PyMuPDF4LLM" showed relative superiority across overall metrics related to RAG build and AI training data preprocessing. The detailed metrics in which it led were Tables, Content Faithfulness, Semantic Formatting, and Visual Grounding.

The company explained that PyMuPDF4LLM requires about one-hundredth the hardware resources of Docling and processes documents more than 10 times faster. It also said that while dozens of global big tech companies are competing in the global parsing market, and many domestic companies have failed to enter the rankings, PyMuPDF4LLM is the only Korean technology to break into the global top tier in that market.

It added that PyMuPDF4LLM has proven its standing in the domestic market and, compared with other major AI document parsing technologies that also ranked domestically, is also more than about 2 times better on most core metrics.

Based on this, Epapyrus is actively strengthening its marketing for the commercial product "PyMuPDF Pro." "PyMuPDF Pro" is tailored for enterprise and institutional environments, and supports PDF, Hangul Word Processor (HWP·HWPX), and Microsoft Office (Word·Excel·PowerPoint). It also provides the same high-performance layout analysis and data extraction functions across a wide range of business documents without restrictions.

PyMuPDF Pro is offered as an ultralight Python package, and was introduced as a method that does not require a separate server build and can be directly embedded into application code. As a result, it can operate in a variety of secure environments, including on-premises, closed networks, and in-house PCs, and it was emphasized that there are no CPU or GPU constraints.

The product analyzes tables, images, lists, coordinate information, and Reading Order within documents, and is said to help reduce hallucinations when building RAG services and document automation. Kim Jeong-a, vice president of Epapyrus, said that to improve the performance of generative AI, it is important not only to use AI models but also to accurately structure and preserve source documents, adding that the data extraction technology proven by PyMuPDF's benchmark evaluation demonstrates Epapyrus's technological edge in the global market.

The speaker added that, based on PyMuPDF Pro's integrated support for all document formats, the company will help enterprises build AI solutions and support businesses in the process of building RAG, thereby helping them experience the best document preprocessing efficiency.

Source: IT DAILY · Yang Seung-gap
Original: https://www.itdaily.kr/news/articleView.html?idxno=241980

References

This article was produced with the help of an automated content generation algorithm.


Source: IT DAILY

View original

This article was summarized and organized by BizCrush based on the original article from IT DAILY. For exact quotations and full details, please refer to the original article.