Saturday 22 March 2025
Artificially intelligent video generators have made tremendous progress in recent years, capable of producing realistic and engaging videos from text prompts. But a crucial question remains: are these generated videos physically plausible? Researchers have now developed a benchmark to evaluate the physical coherence of such videos, ensuring that they adhere to real-world laws of physics.
The new benchmark, called PhyCoBench, consists of 120 carefully designed prompts covering seven categories of physical principles, including gravity, collision, vibration, friction, fluid dynamics, projectile motion, and rotation. For each prompt, four state-of-the-art video generation models were asked to produce videos that demonstrate these physical principles in action.
To evaluate the generated videos, human evaluators were instructed to focus on the object motions depicted in the videos, rather than details such as human faces or camera movements. They ranked the videos based on how well they adhered to real-world physical laws, taking into account factors like object mass, volume, and forces acting upon them.
The results showed that the automated evaluation model, called PhyCoPredictor, closely aligned with the human rankings in 80% of the cases. This suggests that PhyCoPredictor is a reliable tool for evaluating the physical coherence of generated videos.
But why is this important? Physically plausible videos can have significant implications for fields such as robotics, animation, and virtual reality. For instance, robots designed to interact with their environment need to understand how objects behave under different conditions. Similarly, animators require realistic simulations to create convincing special effects. Virtual reality experiences can also benefit from more accurate physics-based simulations.
To develop PhyCoPredictor, researchers combined two key components: a latent flow diffusion module and a frame prediction model. The former generates a sequence of optical flows that capture the motion trajectories of objects in the video, while the latter uses these flows to predict the next frame in the sequence.
The team trained their models using a combination of datasets, including OpenVid and Motion Data. OpenVid provides a large-scale collection of videos with text captions, which allowed the researchers to filter out static scenes and focus on dynamic scenarios. Motion Data, on the other hand, consists of action recognition datasets, which provided a wealth of information about object movements.
In summary, PhyCoBench offers a new benchmark for evaluating the physical coherence of generated videos, ensuring that these videos adhere to real-world laws of physics.
Cite this article: “Physically Plausible Videos: A New Benchmark for Evaluating AI-Generated Content”, The Science Archive, 2025.
Artificial Intelligence, Video Generation, Physical Coherence, Benchmark, Phycobench, Physics, Robotics, Animation, Virtual Reality, Simulation.







