Aligning Video Generation with User Prompts through Human Feedback

Friday 14 March 2025


The quest for perfect video generation has been a long-standing challenge in the field of computer vision and machine learning. Researchers have made significant progress in recent years, but there is still much to be desired when it comes to generating high-quality videos that meet human expectations.


One major obstacle has been the difficulty in aligning generated videos with user prompts. This problem arises from the fact that video generation models often produce outputs that are not perfectly synchronized with the input text or image prompt. As a result, the generated videos may contain mismatched scenes, incorrect objects, or even nonsensical content.


To address this issue, a team of researchers has developed a novel approach that leverages human feedback to improve the alignment between generated videos and user prompts. Their method, known as VideoGen-RewardBench, relies on a large-scale dataset of human-annotated video pairs to train a reward model that can predict the quality of generated videos based on their alignment with the input prompt.


The researchers began by constructing a dataset consisting of 26.5k video pairs, each annotated with four dimensions: visual quality, motion quality, text alignment, and overall quality. The dataset was designed to be representative of modern T2V models, featuring higher resolutions (480×720 – 576×1024) and longer durations (4s – 6s).


To train the reward model, the researchers employed a reinforcement learning framework that maximizes the cumulative reward obtained from the human annotations. The model was trained using a combination of pairwise comparisons and pointwise scores to predict the quality of generated videos.


The results were impressive: the reward model demonstrated significant improvements in its ability to align generated videos with user prompts, outperforming existing state-of-the-art models on both GenAI-Bench and VideoGen-RewardBench. The researchers also developed a range of alignment algorithms that extend the diffusion-based models used for video generation, including Flow-DPO, Flow-RRW, and Flow-NRG.


Flow-DPO, in particular, demonstrated superior performance compared to other baselines, with human evaluators rating its generated videos as significantly more aligned with user prompts. The algorithm’s success can be attributed to its ability to incorporate human feedback into the training process, allowing it to learn from mistakes and adapt to changing conditions.


The implications of this work are far-reaching, enabling applications such as video editing, animation, and even AI-generated content creation.


Cite this article: “Aligning Video Generation with User Prompts through Human Feedback”, The Science Archive, 2025.


Video Generation, Machine Learning, Computer Vision, Video Editing, Animation, Ai-Generated Content, Text Alignment, Visual Quality, Motion Quality, Reinforcement Learning.


Reference: Jie Liu, Gongye Liu, Jiajun Liang, Ziyang Yuan, Xiaokun Liu, Mingwu Zheng, Xiele Wu, Qiulin Wang, Wenyu Qin, Menghan Xia, et al., “Improving Video Generation with Human Feedback” (2025).


Leave a Reply