Saturday 22 March 2025
The quest for better audiovisual active speaker detection has led researchers to develop innovative techniques that leverage speaker-specific information from both reference speech and candidate audio signals. A recent study proposes a novel approach, dubbed SCAN (Speaker Comparison Auxiliary Network), which integrates this information via cross-attention mechanisms to disambiguate challenging scenarios.
Audiovisual active speaker detection is a crucial task in various applications, including video conferencing, public speaking events, and even surveillance systems. The goal is to identify the active speakers in a mixed-audio signal, where multiple people are talking simultaneously. Conventional approaches rely on modeling the temporal correspondence of audiovisual cues, such as lip movement and audible speech, but these methods struggle with noisy or occluded visual signals.
To address this limitation, researchers have turned to speaker recognition models, which can extract robust speaker embeddings from reference speech. These embeddings can then be used to inform active speaker detection systems. However, previous attempts have relied solely on reference speech, neglecting the valuable information present in the candidate audio signal.
SCAN changes this by introducing a cross-attention mechanism that compares the speaker embeddings extracted from both the reference speech and candidate audio signals. This allows the system to identify similarities and distinctions between the two sources of speaker-specific information, effectively disambiguating challenging scenarios where visual cues are absent or unreliable.
To evaluate SCAN’s effectiveness, researchers tested it on the Ego4D-AVD dataset, which consists of 572 video clips recorded from the egocentric perspective. The dataset is particularly challenging due to its high levels of noise and occlusion. Compared to state-of-the-art systems, SCAN demonstrated a significant improvement in performance, with an increase in mean average precision (mAP) of up to 14.5% on the Ego4D-AVD validation set.
The study also proposed a novel method for generating identity-speech libraries, which are crucial for speaker recognition models. By fine-tuning a pre-trained face recognition model using self-supervised learning and transformer encoder layers, researchers were able to create more robust identity-speech libraries that can better handle the challenges of egocentric recordings.
The implications of this research are far-reaching, with potential applications in various fields, including video conferencing, public speaking events, and surveillance systems. By leveraging speaker-specific information from both reference speech and candidate audio signals, SCAN provides a powerful tool for disambiguating challenging scenarios and improving overall performance in audiovisual active speaker detection tasks.
Cite this article: “Advancing Audiovisual Active Speaker Detection with SCAN”, The Science Archive, 2025.
Speaker Comparison Auxiliary Network, Audiovisual Active Speaker Detection, Ego4D-Avd Dataset, Cross-Attention Mechanism, Reference Speech, Candidate Audio Signals, Speaker Embeddings, Identity-Speech Libraries, Face Recognition Model, Transformer Encoder Layers.







