Abstract
This research presents a framework for real-time task detection and digital twin modeling based on human posture estimation from 360° videos. The system integrates markerless human posture estimation with task classification and digital twin visualization. Posture estimation using MediaPipe provided accurate skeletal tracking, while kinematic feature extraction enabled detailed motion analysis. Gaussian Mixture Models effectively segmented task transitions, distinguishing between different phases of ladder use. Gaussian Splatting helped realistic and adaptive visualizations, for a digital twin that accurately represented human-environment interactions. Using these techniques, the framework achieves a non-intrusive and scalable approach to task detection and digital twin modeling. The system captured human movements from 360° videos and classified them into task-specific segments. The results demonstrated that task detection based on posture estimation could improve workplace safety by identifying inefficient or hazardous postures. The digital twin representation can be analyzed for movement patterns and ergonomic risk assessment.
Keywords
Introduction
Digital twin technology has been extensively explored in industrial applications, where virtual models of physical entities enable real-time simulation and analysis (Javaid et al., 2023; Jiang et al., 2021). Its use in human-centered applications is growing, particularly in safety and ergonomics (Caputo et al., 2019). A key motivation behind this research is to improve the monitoring and analysis of human movements in complex environments, where real-time task detection can help mitigate safety risks. Workplace safety remains a critical concern, especially for physically demanding tasks such as ladder use (Iyer et al., 2024; Ogunseiju et al., 2021). Workers engaged in climbing, maintenance, and construction activities are at risk for falls and other injuries due to awkward posture and unstable movements (Tanvi Newaz et al., 2022). A system capable of tracking and analyzing these movements in real time–identifying hazardous postures and task inefficiencies–could improve safety training and accident prevention. This study aims to develop a real-time framework that utilizes 360° video recordings to estimate human posture and dynamically detect task transitions. Unlike traditional motion tracking systems that rely on wearable sensors or single-camera setups, the proposed method integrates markerless human posture estimation with task classification and digital twin visualization (Iyer & Jeong, 2024). This framework represents an early-stage integration of key modules and provides a scalable, non-intrusive foundation for real-time monitoring and modeling of human-environment interactions.
Background
Digital Twin technology has its roots in engineering and optimization of industrial processes, where it has been applied to monitor and simulate machine performance (Liu et al., 2021). More recently, researchers have tried to extend these capabilities to human-centric applications, particularly in occupational safety, healthcare, and ergonomics (Erol et al., 2020; Greco et al., 2020; Park et al., 2023). Human posture estimation techniques have been developed to analyze skeletal movements, with methods such as OpenPose and MediaPipe offering accurate tracking of key body landmarks (Cao et al., 2017; Lugaresi et al., 2019). These models use deep learning algorithms to process image and video data, estimate joint positions, and capture motion patterns. The advantage of markerless tracking over traditional sensor-based methods lies in its ability to collect movement data without physical attachments, reducing setup complexity and minimizing interference with natural movement. However, a challenge in human motion analysis is the accurate classification of task transitions, particularly in dynamic environments where multiple postures can overlap or change rapidly (Yadav et al., 2021). Gaussian Mixture Models (GMMs) have been used for pattern recognition in motion data to segment movement sequences into distinct task phases (Reynolds, 2009). Although previous research has demonstrated the feasibility of using HPE for motion tracking, integrating task classification and digital twin modeling into a single framework remains an open challenge. This study seeks to bridge this gap by developing a system that can process 360° video data, detect human tasks in real time, and generate interactive digital twins.
Approach
The framework developed in this study follows a structured approach for processing 360° video data, estimating posture, classifying tasks, and generating digital twin visualizations (see Figure 1). First, raw video frames are extracted and processed using MediaPipe, which detects and tracks 33 skeletal landmarks per frame (Gyamenah et al., 2025; Iyer & Jeong, 2024). These landmarks are represented as three-dimensional coordinates and form the basis for motion analysis. Once posture data have been collected, kinematic features such as velocity, acceleration, and joint displacement are calculated to quantify movement dynamics. The next stage involves task classification using GMMs, which analyze the extracted motion features to segment activities into predefined categories (Reynolds, 2009). Task transitions such as climbing a ladder, working on a ladder, and descending a ladder are identified based on variations in joint angles and movement patterns. Finally, classified task data and posture information are integrated into a digital twin framework using Gaussian Splatting, which enables real-time rendering of human movement. This visualization method represents skeletal motion as a series of Gaussian functions, producing a smooth and continuous depiction of posture transitions.

A framework for posture estimation from 360° videos for task detection and digital twin modeling.
Outcome
The system developed in this study successfully captured human movements from 360° videos and classified them into task-specific segments. Posture estimation using MediaPipe provided accurate skeletal tracking, while kinematic feature extraction enabled detailed motion analysis. GMMs effectively segmented task transitions, distinguishing between different phases of ladder use. Gaussian Splatting facilitated realistic and adaptive visualizations, creating a digital twin that accurately represented human-environment interactions. The results demonstrated that task detection based on posture estimation could improve workplace safety by identifying inefficient or hazardous postures. The digital twin representation enabled a more intuitive analysis of movement patterns, making it easier to assess ergonomic risks and develop targeted interventions.
Conclusion
This research presents a framework for real-time task detection and digital twin modeling based on human posture estimation from 360° videos. By integrating posture estimation, motion analysis, task classification, and neural rendering, the study offers a scalable and non-intrusive approach to workplace safety monitoring. Future work will focus on optimizing computational efficiency, expanding the framework to support a wider range of physical tasks, and incorporating additional sensor data for improving accuracy. Validation in real-world industrial environments will further refine the system’s capabilities and assess its practical impact on worker safety and training. While this study confirms visual consistency between pose estimates and digital twin output, quantitative validation against ground-truth joint angles is planned for future work.
Footnotes
Acknowledgements
The authors gratefully acknowledge Ms. Debra Bradley from Impacto Protective Products Inc. for their generous donation of personal protective equipment used in this research project. Their support significantly contributed to the successful execution of the study.
Author’s Note
*Current affiliation: Arizona State University, Tempe, USA.
Declaration of Conflicting Interests
The author(s) declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Funding
The author(s) received no financial support for the research, authorship, and/or publication of this article.
