LongViTU: A Large-Scale Dataset for Understanding Long-Form Video Content

Tuesday 04 March 2025


The pursuit of understanding long-form video content has been a challenge for artificial intelligence (AI) researchers and developers in recent years. The complexity of processing extended sequences of images, paired with the need to extract relevant information, has proven to be a difficult task. However, a new dataset and approach have been developed to tackle this issue.


LongViTU, a large-scale dataset, has been created to facilitate the development of AI models capable of understanding long-form video content. This dataset consists of over 121,000 question-answer pairs, paired with nearly 900 hours of video footage. The videos are organized into a hierarchical structure, allowing for efficient processing and analysis.


One of the key features of LongViTU is its ability to incorporate self-revision mechanisms. These mechanisms allow AI models to refine their answers based on feedback from humans. This approach helps to improve the accuracy and relevance of the AI’s responses, making it more effective in understanding long-form video content.


The dataset has been evaluated using a range of metrics, including GPT-4 scores, which measure the model’s ability to accurately answer questions about the video content. The results show that LongViTU is capable of achieving high levels of accuracy, even when compared to state-of-the-art models such as LLaMA-VID and Video LLAVA.


LongViTU has also been used to fine-tune existing AI models, resulting in significant improvements in their performance. For example, a model called LongVU was able to achieve a GPT-4 score of 49.9% on the dataset, compared to its initial score of around 20%. This improvement demonstrates the potential of LongViTU to enhance the capabilities of AI models.


The development of LongViTU has also highlighted the importance of considering the temporal and spatial aspects of video content when designing AI models. The hierarchical structure of the dataset allows for the extraction of relevant information from both short-term and long-term memory, enabling AI models to better understand complex video sequences.


In addition to its use in fine-tuning existing models, LongViTU has also been used as a benchmark for evaluating the performance of new AI models. This allows researchers to compare the capabilities of different models and identify areas where improvements can be made.


The creation of LongViTU represents an important step forward in the development of AI models capable of understanding long-form video content.


Cite this article: “LongViTU: A Large-Scale Dataset for Understanding Long-Form Video Content”, The Science Archive, 2025.


Artificial Intelligence, Long-Form Video Content, Dataset, Question-Answer Pairs, Video Footage, Self-Revision Mechanisms, Gpt-4 Scores, State-Of-The-Art Models, Fine-Tuning, Temporal And Spatial As


Reference: Rujie Wu, Xiaojian Ma, Hai Ci, Yue Fan, Yuxuan Wang, Haozhe Zhao, Qing Li, Yizhou Wang, “LongViTU: Instruction Tuning for Long-Form Video Understanding” (2025).


Leave a Reply