Multimodal Models Fall Short: Uncovering Audio-Visual Synchronization Challenges in Human-Like Reasoning

Thursday 10 April 2025


In a recent study, researchers have developed a novel benchmark dataset designed to evaluate audio-visual models in their ability to integrate and interpret information from both visual and auditory modalities. This dataset, called DAVE (Diagnostic Audio Visual Evaluation), is specifically tailored to assess the multimodal integration capabilities of large language models (LLMs) and other audio-visual learning machines.


The study reveals that existing benchmarks for evaluating LLMs often suffer from a strong visual bias, where answers can be inferred solely from visual data. This limitation compromises the validity of performance metrics on these benchmarks, as high scores may reflect proficiency in unimodal reasoning rather than true multimodal integration.


To address this issue, DAVE is designed to ensure that both modalities are necessary to answer questions correctly. The dataset consists of 40 questions that require models to integrate audio and visual information to identify actions, recognize sounds, or synchronize temporal events. For example, one question asks participants to identify the person doing a specific action when a certain sound is heard in the background.


Researchers evaluated the performance of several open-source LLMs on DAVE, including PandaGPT, video-SALMONN, and Video-LLama-2. The results show that these models often struggle with audio-visual synchronization, indicating that they are not yet capable of effectively integrating multimodal information.


To further understand model limitations, researchers also evaluated human performance on the same task using a web-based interface. Participants were asked to answer multiple-choice questions without prior task-specific information, providing a baseline for comparison with model performance. The results suggest that humans are able to accurately identify actions and recognize sounds in audio-visual scenarios, highlighting the potential benefits of developing more sophisticated multimodal models.


The study also explores the impact of modality availability on DAVE performance. Researchers found that when models have access to all modalities (audio, video, and text), they achieve higher accuracy compared to situations where one or more modalities are absent. This suggests that DAVE is effective in requiring genuine cross-modal reasoning from models.


Finally, researchers developed a pipeline-based approach to evaluate the performance of LLMs on DAVE. The approach involves an audio model extracting timestamps of overlaid sounds, which are then used to select corresponding video frames for input to a video model. The results show that this approach achieves higher accuracy than individual models, highlighting the potential benefits of combining multiple modalities and processing streams.


Cite this article: “Multimodal Models Fall Short: Uncovering Audio-Visual Synchronization Challenges in Human-Like Reasoning”, The Science Archive, 2025.


Audio-Visual Integration, Multimodal Learning, Large Language Models, Benchmark Dataset, Dave, Visual Bias, Audio-Visual Synchronization, Human Performance, Modality Availability, Pipeline Approach


Reference: Gorjan Radevski, Teodora Popordanoska, Matthew B. Blaschko, Tinne Tuytelaars, “DAVE: Diagnostic benchmark for Audio Visual Evaluation” (2025).


Leave a Reply