AI

Twelve Labs Releases “Pegasus 1.6” for Automated Video Cleanup in Physical AI Training

TECHWORLD ·

[Photo: Twelve Labs]

✦ AI Summary

Twelve Labs announced on the 7th the release of VLM “Pegasus 1.6.”

“Pegasus 1.6” is designed to structure robot training data from first-person video and is notable for strengthening egocentric data understanding performance.

The system supports action segmentation and labeling, detailed caption labeling, quality evaluation, search and curation, and privacy and compliance detection.

Twelve Labs announced on the 7th the release of VLM “Pegasus 1.6.” “Pegasus 1.6” is designed to structure robot training data from first-person video and is notable for strengthening egocentric data understanding performance.

Egocentric data refers to first-person video captured with action cameras, smart glasses, and other devices worn by workers. This data is used to reflect real human work processes in robot training, but securing the original footage alone does not complete the training data. “Pegasus 1.6” was introduced in the context of supporting the automatic cleanup of such training videos.

To use this data for training, preprocessing is required, including selecting necessary scenes and segmenting actions by interval. It is also necessary to distinguish interactions among hands, tools, and objects, as well as to separate work outcomes.

This preprocessing is highly manual. Twelve Labs said that more than 200,000 hours have been cumulatively spent on action and workflow annotation for the public video dataset “Ego-Exo4D.” “Ego-Exo4D” runs 1,286 hours in length, and the amount of work required to process 1 hour of raw video is about 155 hours.

Pegasus 1.6 is a system focused on automating the video cleanup process. It structures workflows by analyzing human actions, surrounding objects, and preceding and following context. Based on functions for separating work stages, organizing the timing of actions, cataloging appearing objects, and arranging execution order, it can turn raw footage into data for robot training. As an example, it was presented as analyzing a worker picking up parts, processing tools, and moving finished objects, while identifying the start and end points of each action and the tools and objects used.

Five key functions were presented: action segmentation and labeling, detailed caption labeling, quality evaluation, search and curation, and privacy and compliance detection. These functions support dividing work stages, filtering out low-quality videos, finding needed scenes, and extending to the scope of dataset construction.

The system also provides time-based metadata (TBM). Under the TBM approach, the user specifies metadata items and structure. As a result, it can extract the timing of specific actions and the surrounding context.

It can also process recorded video that includes narrated speech. In this case, it can connect audio and action timing to organize the work process in greater detail.

Pegasus was improved with enhanced object recognition for consistent identification of hands, objects, and tools across the full video span, faster processing of large-scale video, and better cost efficiency, with the aim of reducing the burden on robot and physical AI companies of building data for repetitive, large-volume video analysis.

These capabilities are already being used in Pegasus for real robot training data construction. Twelve Labs said a robotics training data company in the U.S. is using Pegasus. The videos targeted by the company are egocentric videos collected at factory worksites.

The usage process involves preliminary quality evaluation of videos followed by attaching detailed descriptions to the selected clips. The company processes hundreds of thousands of hours of video cleanup per day. In this process, it replaces full human review and performs repetitive tasks subject to automation, such as quality evaluation and caption generation, helping cut the time needed to build training data.

Twelve Labs' existing application areas were media, entertainment, sports, and the public sector. Its existing technology analyzes actions, events, and surrounding context in video. With Pegasus 1.6, Twelve Labs plans to expand its applications into physical AI, including robot and manufacturing settings.

Twelve Labs CEO Lee Jae-sung said that the company has long aimed for machines to understand the real world through video, and that, in that context, Pegasus 1.6 functions to convert video into structured knowledge, helping physical AI and robotics companies use real human experience for training.

Source: TECHWORLD · Kim Seung-gi
Original: https://www.epnc.co.kr/news/articleView.html?idxno=407889

References

This article was produced with the help of an automated content generation algorithm.


Source: TECHWORLD

View original

This article was summarized and organized by BizCrush based on the original article from TECHWORLD. For exact quotations and full details, please refer to the original article.