Saturday 22 March 2025
The quest for relevance in online advertising has led researchers to develop a new approach that could significantly improve the effectiveness of search ads on short video platforms. By combining visual and textual data, a team of scientists has created a model that can better match users with relevant advertisements.
The problem of irrelevant ads is a common one – users often find themselves bombarded with promotions for products they don’t need or want. This not only frustrates the user but also wastes the advertiser’s budget. The key to solving this issue lies in developing a relevance model that can accurately predict which ads are most likely to resonate with individual users.
Traditionally, relevance models have relied on language-based approaches, using natural language processing techniques to analyze text and identify relevant matches. However, these methods often struggle when it comes to multimedia content like videos. To address this limitation, the researchers turned to a multimodal approach, incorporating both visual and textual data into their model.
The new model, known as HCMRM (High-Consistency Multimodal Relevance Model), uses a combination of image recognition and natural language processing techniques to analyze query-video pairs. The system first extracts visual features from the video using a pre-trained convolutional neural network, and then generates text embeddings based on the video’s description.
The researchers also introduced a novel triplet relevance modeling approach, where pseudo-queries are generated from the video text and used to fine-tune the model. This technique enhances the consistency between pre-training and relevance tasks, allowing the model to better learn the relationship between queries and videos.
To evaluate the effectiveness of HCMRM, the team deployed it in a real-world search advertising system for over a year, observing significant improvements in ad relevance. The results showed that the model reduced the proportion of irrelevant ads by 6.1% and increased ad revenue by 1.4%.
The implications of this research are significant – with the ability to accurately match users with relevant advertisements, online advertisers can optimize their campaigns for better ROI, while users benefit from more targeted and engaging content. As the team continues to refine their approach, it’s likely that we’ll see even more innovative applications of multimodal learning in the years to come.
The success of HCMRM also highlights the potential for AI-driven solutions to tackle complex problems in online advertising. By combining visual and textual data, the model is able to capture nuances in user behavior and preferences that might be missed by traditional language-based approaches.
Cite this article: “Multimodal Approach Boosts Relevance of Search Ads on Short Video Platforms”, The Science Archive, 2025.
Online Advertising, Relevance Model, Multimodal Approach, Visual Data, Textual Data, Natural Language Processing, Convolutional Neural Network, Triplet Relevance Modeling, Ad Revenue, Roi, Ai-Driven Solutions







