Wednesday 05 March 2025
Researchers have made significant progress in developing Large Vision Language Models (LVLMs) that can understand and analyze daily living activities, such as cooking or cleaning, by leveraging the complementary nature of egocentric views. These models are capable of learning to extract fine-grained interactions and spatial relationships from exocentric videos, which is crucial for applications like elderly monitoring and cognitive assessment.
To achieve this, researchers have proposed an online ego2exo distillation approach that learns ego-augmented exo representations in LVLMs. However, collecting paired ego-exo training data for real-world daily living activities is impractical. To address this limitation, the team developed EgoMimic, a skeleton-guided method that can generate mimicked ego views from exocentric videos.
EgoMimic uses skeleton-motion to dynamically determine the most active joints for each video, adaptively focusing on the skeleton joints most relevant to the action. This approach outperforms existing methods that focus exclusively on cropping around hands under the assumption that hand regions capture all meaningful interactions. The researchers also ablated the number of skeleton joints selected by EgoMimic and found that using 6 joints performs best.
The team has also created a pipeline to automatically generate high-quality instruction-tuning data from EgoExo4D, a large-scale dataset containing synchronized ego-exo video pairs. This dataset is used to train LVLMs on exocentric daily living activities. The researchers have also developed a benchmark called EgoPerceptionMCQ, which consists of multiple-choice questions that test understanding of the ego perspective.
The categorization of these questions into four categories – active hand identification, relevant object identification, active hand and relevant object identification, and other – allows for analysis of the strengths and weaknesses of LVLMs across different types of ego cues. This benchmark provides a comprehensive evaluation of LVLMs’ ability to understand daily living activities from exocentric videos.
The proposed approach has several advantages over existing methods. First, it can learn to extract fine-grained interactions and spatial relationships from exocentric videos, which is crucial for understanding daily living activities. Second, it can generate high-quality instruction-tuning data automatically, reducing the need for manual annotation. Third, it provides a comprehensive evaluation of LVLMs’ ability to understand daily living activities from exocentric videos.
Cite this article: “Large Vision Language Models for Understanding Daily Living Activities”, The Science Archive, 2025.
Large Vision Language Models, Egocentric Views, Exocentric Videos, Daily Living Activities, Skeleton-Guided Method, Egomimic, Instruction-Tuning Data, Egoexo4D, Benchmark, Egoperceptionmc







